Grading group work you can defend
A fair individual mark needs more than one signal, and it needs to survive an appeal. Here is how to build one.
The goal of assessing group work is not to catch people. It is to give each student a mark that reflects what they did, and to be able to explain that mark if they ask. Both halves matter, and they need different things.
Peer assessment is useful, and partial
Student peer marks can track teacher marks reasonably well, especially when peers make global judgements against criteria they understand rather than rating many fine-grained dimensions. 1 Peer assessment shows adequate reliability and validity across many settings. 2 That is agreement, not equivalence, and it is not a stable quantity either: the mean correlation across fifty-six comparisons is 0.69, which leaves less than half the variation in peer marks tracking variation in teacher marks, and the individual studies disagree with each other far more than chance can explain. 1 Reliability was adequate in 18 of 25 studies comparing teacher and peer marks and unacceptably low in the rest, and one of the four applications named in that minority is peer assessment of individual contributions to a group project. 2 Two fears are smaller than assumed. Peers do not systematically mark higher or lower than teachers: once a single outlying study is set aside, the average gap between peer and teacher marks is effectively zero. 1 And reciprocity, meaning whether a student rated highly by a teammate returns the favour, accounted for about one percent of the variance in the one study that measured it directly. That measurement captured tit-for-tat only and explicitly excluded friendship, so it is narrower than it sounds. 3 A separate study found the marks students awarded bore no relationship to the marks they received. 2 That reassurance has a hard limit, and it is the most useful thing we have learned about our own design. Where students rated each other repeatedly and saw their feedback in between, reciprocity appeared plainly. Students who received positive ratings liked their teammates more and rated them higher next time; students who received negative ratings rated them lower. It happened even though the feedback they saw was aggregated across several raters. 4 So reciprocity is not a property of peer assessment. It is a property of showing people their scores and asking again.
But peer marking has real limits. Scores inflate once a grade is attached, most of all for underperforming students, 5 and a rubric improves how closely peer marks track an expert’s while leaving friendship bias larger, not smaller, than without it. 6 The lesson is not to discard peer assessment. It is to keep the grade weight low, ask for written justification, and never make it the only thing the assessment can see. 5 On anonymity Dwixel had been placing itself on the wrong side of a debate it is not in. It records who wrote each rating and withholds it only from the person rated, which is precisely the partial anonymity Yang recommends — names attached on collection, blinded before the person evaluated sees anything — and the mechanism WebPA has shipped for fifteen years. 5 7 One gap under all of this is worth naming rather than hiding: in the broadest review of the field, peer assessment of group and project work is the category with the thinnest outcome evidence of any, reported largely as student perceptions rather than as measured gains. 2 Peer assessment of writing and of tests has been tested far harder than peer assessment of teammates has.
Established practice: weight the group mark by contribution
The mature approach in the literature is not to grade everyone the same and is not to grade everyone separately. It is to individualise a group mark using peer-assessed contribution. Open tools have done exactly this for decades. WebPA began at Loughborough University in 1998, moved from paper to a spreadsheet to the web, and was released open source under JISC funding; by 2008 it was running in 56 per cent of departments across the university’s four faculties and being adopted by eight more UK institutions. 7 Its authors found afterwards that the method they had each arrived at independently already existed in the literature, published in 1990. 7 Instruments such as CATME formalise the dimensions of team-member effectiveness with validated scales. 4 This is settled practice, not novelty. It is also not settled in one respect, and it would be dishonest to skip past it: the meta-analytic evidence above favours a single overall judgement made against clear criteria over scoring several dimensions separately, 1 and the one study to test both in a group-project context found the same thing. 2 Dimension-based instruments are the norm and the evidence prefers the alternative. There is a resolution, and it comes from CATME’s own validation rather than from us. Across its dimensions the average correlation was 0.76, high enough that the authors combined all five into a single overall score to run their analysis, and when they tested which dimensions predicted anything independently, two of the five contributed nothing. 4 Peer ratings of separate dimensions are inflated by halo, by 63 per cent in one meta-analysis. 4 So the dimensions are best read as one judgement expressed five ways, which is what the meta-analysis recommended in the first place. Score them if the wording helps students understand what good contribution looks like, and use the overall figure for the mark. The tension is not ours to invent, either: WebPA’s own developers weighed the same trade-off and wrote that CATME “fixes the criteria so that academics cannot change it, which is one solution, but also is very restrictive”, while admitting their own flexible alternative left so many choices that it was daunting for new staff. 7 That question is open, and it is one we would rather put to a coordinator than answer for them.
Where objective traces come in
Peer judgement answers "how did this person work with us". It does not, by itself, show how much of the artifact each person actually produced. That is what an attributed contribution record adds: a second, independent signal grounded in the work itself rather than in opinion. Used together, the two cover each other’s weaknesses. Peer ratings catch the things traces miss (coordination, ideas, leadership); traces catch the things ratings distort (who actually wrote it). Dwixel shows both separately, and offers a suggested weighting drawn from the two, which the instructor accepts or overrides. The suggestion is an average of them; neither signal disappears into it, and the instructor sets the final figure.
Defensible means it survives an appeal
A mark is only as strong as the record behind it on the day a student challenges it. When work is handed in through Dwixel, every member signs off and the artifacts are frozen into a locked, timestamped snapshot, alongside the contribution picture at that moment. If the grade is later questioned, you are not reconstructing events from memory. You are pointing at a fixed record of what was submitted and who did it.
References
- 1.Falchikov, N., & Goldfinch, J. (2000). Student peer assessment in higher education: A meta-analysis comparing peer and teacher marks. Review of Educational Research, 70(3), 287–322. Link ↗
- 2.Topping, K. J. (1998). Peer assessment between students in colleges and universities. Review of Educational Research, 68(3), 249–276. Link ↗
- 3.Magin, D. J. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63. Link ↗
- 4.Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630. Link ↗
- 5.Yang, A., Brown, A., Gilmore, R., & Persky, A. M. (2022). A practical review for implementing peer assessments within teams. American Journal of Pharmaceutical Education, 86(7), 8795. Link ↗
- 6.Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on construct validity of peer assessment, perceived fairness and comfort, and performance. Studies in Educational Evaluation, 39(4), 195–203. Link ↗
- 7.Loddington, S., Pond, K., Wilkinson, N., & Willmot, P. (2009). A case study of the development of WebPA: An online peer-moderated marking tool. British Journal of Educational Technology, 40(2), 329–341. Link ↗