← All guides
Guide

Running peer assessment in your course

A short, evidence-based setup that gets the value of peer assessment while avoiding its known biases.

3 min read

Peer assessment works when it is designed around what the research actually shows. This is a practical setup you can apply to any group assignment, with the reasoning behind each choice.

1. Write clear criteria, and keep judgements global

Peer marks agree with teacher marks most closely when students make a global judgement against criteria they understand, rather than rating many separate dimensions. 1 Give students three or four plain statements of what a good contribution looks like, and ask for an overall judgement against them, not a long grid.

2. Keep the grade weight modest

When peer scores count heavily toward a grade, students become lenient, and the bias is more pronounced among underperforming students. 2 Let peer assessment shape the mark, not decide it. A modest weight preserves the signal without inviting inflation.

3. Require a sentence of justification

Asking for a short written reason alongside each rating makes the rater accountable for it, and gives you something to read when scores disagree. 2 A number on its own is hard to act on; a number with a reason is evidence.

4. Use partial anonymity, and control the timing

Collect ratings with the rater’s name attached and blind them before the person being rated sees anything. Yang recommends this over guaranteed anonymity: it keeps the honesty that anonymity buys while making a student answerable for the rating they gave. 2 And do not let fear of vote-trading drive the design: reciprocity, meaning whether a student rated highly by a teammate returns the favour, accounted for about one percent of the variance in the one study that measured it directly. That measurement captured tit-for-tat only and explicitly excluded friendship, so it is narrower than it sounds. 3 A separate study found the marks students awarded bore no relationship to the marks they received. 4 It does appear if you assess repeatedly and let students see their scores in between, which is a reason to control the timing of feedback rather than to abandon the method. 5 Do not reach for more raters as the fix, either: agreement got weaker as rater numbers rose, and single raters were no less reliable than groups. 1 A rubric helps validity, but it does not remove friendship bias and in the study that tested it the bias was larger with the rubric than without. Read that as a warning rather than a measurement of group work: the students there each scored one stranger-or-friend’s concept map in a single sitting, not a teammate they had worked with for ten weeks. 6

5. Weight the group mark by contribution

The established way to use the results is to moderate the group mark into individual marks, the model open tools such as WebPA have used for years. 7 This is what turns peer input into a fair, individualised grade rather than a popularity score.

How Dwixel supports this

Dwixel records who wrote each rating and never shows it to the person being rated, which is the partial anonymity Yang recommends rather than blanket confidentiality, and treats the result as one input alongside the objective contribution record and your judgement. You get the signal peer assessment is good at, without asking it to carry the whole grade.

In one line
Clear global criteria, confidential ratings with a written reason, modest grade weight, deliberate timing, and use the result to weight the group mark, not to be it.

References

  1. 1.Falchikov, N., & Goldfinch, J. (2000). Student peer assessment in higher education: A meta-analysis comparing peer and teacher marks. Review of Educational Research, 70(3), 287–322. Link ↗
  2. 2.Yang, A., Brown, A., Gilmore, R., & Persky, A. M. (2022). A practical review for implementing peer assessments within teams. American Journal of Pharmaceutical Education, 86(7), 8795. Link ↗
  3. 3.Magin, D. J. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63. Link ↗
  4. 4.Topping, K. J. (1998). Peer assessment between students in colleges and universities. Review of Educational Research, 68(3), 249–276. Link ↗
  5. 5.Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630. Link ↗
  6. 6.Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on construct validity of peer assessment, perceived fairness and comfort, and performance. Studies in Educational Evaluation, 39(4), 195–203. Link ↗
  7. 7.Loddington, S., Pond, K., Wilkinson, N., & Willmot, P. (2009). A case study of the development of WebPA: An online peer-moderated marking tool. British Journal of Educational Technology, 40(2), 329–341. Link ↗