← All posts
Research

Peer assessment, done properly

Students rating each other is a useful signal and a dangerous one. The difference is entirely in the design.

The Dwixel team · June 2026 · 10 min read

Ask students to rate each other and you get something genuinely valuable: a view of contribution that no edit log can capture, like who led, who organised, who carried the meetings. You also get a set of well-documented biases. Whether peer assessment helps or harms depends almost entirely on how you set it up.

What it is good at

Peer marks are more reliable than their reputation suggests. Across many settings, peer assessment shows adequate reliability and validity, with learning outcomes at least equivalent to teacher assessment. 1 The meta-analysis comparing peer and teacher marks is more specific, and its finding is the most useful sentence in this literature: agreement is highest when peers make one overall judgement in the knowledge of clear criteria (r = 0.85), lower when they judge globally with no criteria at all (0.72), and lowest when they score several dimensions separately (0.53). 2 Of every factor tested in that analysis — subject, task, study quality, level of course — whether the judgement was global or dimensional was the single strongest predictor of whether peer and teacher marks agreed. 2 Keep the judgement holistic and the criteria explicit.

It is also worth knowing what peer assessment is best at, because it happens to be the thing group work needs. Split by what is being judged, agreement is strongest for academic processes such as presentations and participation in group activities (r = 0.83), weaker for academic products like essays and exams (0.75), and weakest for professional practice such as clinical skills (0.54). 2 Peers are better at judging how someone worked than at judging what they produced.

Where the evidence is thinner than it sounds

Two caveats belong here rather than in a footnote. The first is that "adequate reliability" describes marks, not feedback: almost the entire literature measures the numbers peers award and not the quality of what they write. 1 The second is more awkward for a page like this one. In the broadest review of the field, reliability was acceptable in 18 of 25 studies comparing peer and teacher marks, and peer assessment of individual contributions to a group project is one of the four applications named in the minority where it was not. 1 The same review rates the outcome evidence for each application separately, and group work comes last: positive in the few studies to date, but measured almost entirely as student perceptions rather than as demonstrated gains. 1 Peer assessment of writing and of tests has been tested far harder than peer assessment of teammates has.

One more thing the field does not say out loud: how much students trust peer assessment has no relationship to whether it is working. The review reports twice that student acceptance appeared unrelated to actual reliability. 1 Students were confident in unreliable systems and sceptical of reliable ones, so satisfaction surveys are evidence about acceptability and none at all about accuracy.

Where it goes wrong

The failure modes are real and specific. When scores count toward a grade, students become lenient, and the bias is more pronounced among underperforming students. 3 The size of it is worth seeing. In CATME’s three validation studies, on a five-point scale where 3 was defined as satisfactory and 5 as excellent, the average rating on every dimension was about 4.2. Students barely used the bottom half of the scale, and the authors name the consequence themselves, which is that it reduces the accuracy of the ratings. 4 Friendship distorts scores too: students over-score their peers in general, and closer friends are over-scored more, an effect a rubric did not remove — in that study the friendship effect was larger with the rubric than without it. 5 That result comes from students rating one peer’s individual work rather than a teammate, so read it as a warning that transfers rather than a measurement of group projects.

A quieter one is rarely mentioned and matters most for how much weight you put on the result. When CATME’s developers decomposed where the variance in ratings actually came from, the rater explained more of it than the person being rated on three of the five dimensions, by a factor of four to six. 4 They are careful about it, and note the alternative reading — a teammate genuinely does not interact identically with everyone — and that their design cannot separate the two. 4 Either way, on some dimensions the score describes the rater at least as much as the rated, which is an argument for treating peer ratings as one input rather than as a measurement.

And the formative claim does not hold evenly. Fifteen years of WebPA use found that only the more mature students, postgraduates in particular, valued peer assessment as a development tool, while undergraduates approached it instrumentally, as a way to maximise a mark. 6 Confidentiality carries a cost of its own worth stating plainly: under anonymity students may tolerate a teammate’s behaviour in the moment and settle up in the final rating rather than raising it while it can still be fixed. 3 The answer to that is to ask more than once, not to remove confidentiality.

Reciprocity: not what you think, but not nothing

The first objection anyone raises is vote-trading, and the evidence on it is genuinely split until you notice what separates the studies. Where reciprocity, meaning whether a student rated highly by a teammate returns the favour, accounted for about one percent of the variance in the one study that measured it directly. That measurement captured tit-for-tat only and explicitly excluded friendship, so it is narrower than it sounds. 7 A separate study found the marks students awarded bore no relationship to the marks they received. 1 But where students rated each other repeatedly and saw their feedback in between, it appeared plainly: those who received positive ratings rated their teammates higher next time, and those who received negative ratings rated them lower — even though the feedback they saw was averaged across several raters. 4

So reciprocity is not a property of peer assessment. It is a property of showing people their scores and then asking again. If you assess once, it is close to a non-issue. If you assess repeatedly, what students see between rounds is the design decision that matters.

The reassurances that do hold

Three, all from the same meta-analysis, and none of them the one usually offered. Peers do not systematically mark higher or lower than teachers: setting aside a single outlying study, the average gap between peer and teacher marks is effectively zero. 2 Better-run studies find better agreement, not worse — high-quality designs average r = 0.78 against 0.50 for low-quality ones — so the pessimistic numbers come disproportionately from the weakest research. 2 And peer marks agree with teacher marks best on exactly the thing group work needs judged, which is participation and process. 2

One reassurance we used to offer here was wrong, and it is worth correcting rather than quietly deleting. We had said that groups of two to seven raters track teacher marks better than a single rater or a large cohort. That holds only while one discredited study is left in the data. With it removed, the two-to-seven band falls to r = 0.59, below both single raters (0.72) and groups of eight to nineteen (0.77), and the meta-analysis reports that correlations got significantly smaller as the number of raters rose, with ratings by single raters no less reliable than the rest. 2 More peers is not a fix.

The design rules that follow

  • 1.Keep the grade weight of peer scores modest, not decisive, to blunt leniency. 3
  • 2.Ask for a short written justification, not just a number. 3
  • 3.Use partial anonymity: collect ratings with the rater’s name attached, and blind them before the person being rated sees anything. Yang recommends this over guaranteed anonymity, because it keeps honest feedback while attaching responsibility to the act of rating. 3
  • 4.Use clear shared criteria and favour one overall judgement over many fine dimensions. 2
  • 5.Rate what people did rather than what they are like: task-oriented behaviours have proved easier to rate reliably than group-maintenance ones. 1
  • 6.Assess more than once, early and mid-project, not only at the end 3 — but not in every module of every year, because heavy repeated use encourages students to play the instrument rather than answer it. 6
  • 7.If you assess repeatedly, decide deliberately when students see their feedback. Releasing scores between rounds is the condition under which peers start returning the ratings they were given. 4
  • 8.Do not read a peer factor computed from ratings that are all within half a point of each other. Flat ratings are the normal case, not the exception. 4

Turning ratings into marks

The mature way to use peer scores is not as the grade, but as a weighting on the group mark. Open tools have done this for decades: WebPA has each student rate teammates and themselves, then moderates the overall group mark into individual marks, 6 and the formula it uses was already published in 1990, eight years before WebPA’s authors independently arrived at it. 6 CATME computes the same thing, as the ratio of a student’s average rating to their team’s average, and reports it twice, once including the student’s self-rating and once without. 4 None of this is experimental; it is established practice.

There is an apparent contradiction here worth resolving rather than skipping. CATME scores five dimensions, and the meta-analysis above prefers a single overall judgement. CATME’s own validation answers it: the five dimensions intercorrelate at 0.76 on average, high enough that the authors combined them into one composite score to run their analysis, and when they tested which dimensions predicted anything independently, two of the five contributed nothing. 4 Peer ratings of separate dimensions are also inflated by halo, by 63 per cent in one meta-analysis. 4 The dimensions are best understood as one judgement expressed five ways. Use the wording to teach students what good contribution looks like, and use the overall figure for the mark.

How Dwixel handles it

Dwixel records who wrote each rating and never shows it to the person being rated. That is not merely confidentiality, it is the specific arrangement Yang recommends: turn evaluations in with your name attached, and blind them before they reach the person being evaluated. He prefers it to guaranteed anonymity because it keeps the honesty and adds responsibility for the rating you gave. 3 It is also what the established tools ship. In WebPA the ratings are recorded and stored against the student who gave them, and only the tutor can see who said what, 6 and CATME made the same design decision. 4 In the Loughborough student survey, 78 per cent were comfortable assessing their peers, and the authors are explicit that the finding assumes ratings stay hidden from other students. 6

Everything above is a reason to treat a peer score as one input rather than an answer. So Dwixel puts it beside the objective contribution record and your own academic judgement, and never asks it to carry the whole grade on its own. The instrument is good at what data cannot see. It is not good enough to be the only thing you look at.

The takeaway
Peer assessment is a good instrument pointed at the things data cannot see. Keep it confidential, keep the judgement global and the criteria clear, keep the grade weight modest, and never make it the only thing the assessment can see.

References

  1. 1.Topping, K. J. (1998). Peer assessment between students in colleges and universities. Review of Educational Research, 68(3), 249–276. Link ↗
  2. 2.Falchikov, N., & Goldfinch, J. (2000). Student peer assessment in higher education: A meta-analysis comparing peer and teacher marks. Review of Educational Research, 70(3), 287–322. Link ↗
  3. 3.Yang, A., Brown, A., Gilmore, R., & Persky, A. M. (2022). A practical review for implementing peer assessments within teams. American Journal of Pharmaceutical Education, 86(7), 8795. Link ↗
  4. 4.Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630. Link ↗
  5. 5.Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on construct validity of peer assessment, perceived fairness and comfort, and performance. Studies in Educational Evaluation, 39(4), 195–203. Link ↗
  6. 6.Loddington, S., Pond, K., Wilkinson, N., & Willmot, P. (2009). A case study of the development of WebPA: An online peer-moderated marking tool. British Journal of Educational Technology, 40(2), 329–341. Link ↗
  7. 7.Magin, D. J. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63. Link ↗