Your contribution score is not your grade
A contribution number is evidence, not a verdict. Reading it as a grade is the fastest way to make it unfair.
When you can measure contribution, the temptation is to turn the number straight into a mark. Resist it. A contribution figure is a strong piece of evidence about participation, and a weak basis for a grade on its own. Understanding the difference is what keeps the measurement honest.
What the number is, and is not
The researcher who built the closest thing to this measurement put it more sharply than we would have dared. Trentin scored individual contribution to a group wiki from four components, and deliberately gave the least weight to volume of text produced, because it is "a quantitative and not qualitative evaluation of each student’s written contributions". 1 His reason for the ranking is the part worth sitting with: word and page counts should be read — his words — "as an element of product quality and not as an indicator of students’ contribution level". 1 Volume tells you about the document. It does not tell you about the person.
His data shows why. Comparing an average contributor with the strongest one in the group, the gap came almost entirely from dialogue and peer review (135.8 against 184.6) and barely at all from what they produced (34.6 against 37.8). 1 Volume separated almost nobody. The conversation separated everybody.
We should be straight about where that leaves Dwixel. Our contribution figure weights words written far more heavily than comments, which is the opposite ordering to the one Trentin argues for. We think there is a defensible reason — a shared document is not a discussion forum, and comment volume is easier to inflate than prose is — but it is a live disagreement with the closest thing this field has to a specification, and it is a large part of why this page exists. The number is not the finding. It is the thing you interrogate.
Why survival is not the answer either
The obvious repair is to count what lasts rather than what was typed. That is the right instinct and it is not sufficient, because survival has a bias of its own. The study of collaborative authorship on Wikipedia found a first-mover advantage: initial text tends to survive longer and be modified less than later contributions, because the person who starts a page sets its shape. 2 Survival rewards being early as much as being right.
The same work found that most text does not last long at all — across more than 600,000 revisions, the median piece of content survived about ninety minutes 2 — and that roughly a fifth of edits shrank the page, cuts the authors describe as often beneficial, tightening prose or removing what did not belong. 2 A person whose work is mostly deletion can be doing the most valuable job in the group and will register as barely present on any measure built around what remains.
Why Dwixel shows a profile, not a verdict
For this reason the contribution view is a profile to inspect, with the attributed history behind it, rather than a single score handed down. That mirrors the tools researchers built to surface collaborative effort, which reconstructed the full revision history so people could see who did what and when, rather than collapsing it to one figure. 3 The number is a way in, not the last word.
The ethics of inference
Drawing conclusions from trace data carries a specific risk, and it is not the one people expect. In the broadest review of the ethics of educational trace data, the validity and integrity of the data and the algorithms was the second most discussed concern of four, behind only privacy and consent. 4 Treating a contribution score as a grade is exactly the invalid inference that literature warns about. The data tells you where to look, not what to conclude.
One of the review’s four themes is uncomfortable for anyone who builds this kind of tool, and we would rather raise it than wait to be asked. It calls it the obligation to act: where data flags an outcome that might result in student harm, that raises a duty to respond, on moral or legal grounds, and the review poses the question directly — what obligation does an institution have to intervene when there is evidence a student could benefit from support? 4 A tool that shows you in week three that someone has contributed nothing has not only given you evidence for a grade in week twelve. It has arguably given you a reason to do something now.
Where the grade actually comes from
Students do not judge the fairness of group assessment on the mark alone. A UK study of 190 students found their perceptions organised into three distinct things: grade congruence, whether the mark matched the contribution; performative group dynamics, whether tasks were allocated and communicated properly; and interpersonal treatment, whether people were dealt with respectfully. 5
The ordering is the useful part, and it is not the one you would guess. The authors conclude that the procedural dimension is the strongest contributor to a sense of fairness, followed by interpersonal treatment, with the grade itself last. 5 Their own summary of what that means: a student may get the best mark available and still feel the process was unfair. 5 The correlations point the same way. Grade congruence was only weakly related to the other two dimensions (r = 0.34 and 0.25), while those two were tightly bound to each other (r = 0.69), 5 so getting the number right does not buy you a fair process and running a fair process does not excuse the wrong number. A student can hold both against you at once.
You meet all three by combining signals: the objective contribution record, confidential peer assessment of how each person worked, and your own academic judgement of the product. The contribution score informs the mark. It does not become it. And there is a cost to individualising that the same review names: assessing people separately inside a group can itself produce the behaviour you were trying to prevent, including withholding work from teammates and intragroup rivalry. 5 That is an argument for keeping the group mark as the basis and moderating it, rather than grading everyone alone. Worth noting on the evidence itself, since this page is about not over-reading numbers: that study covered 190 students at three UK institutions and used exploratory analysis only, its sample was 78 per cent White and 84 per cent domestic, and the authors call both for confirmatory work and for more research on the relative weight of the three dimensions before the ordering is treated as settled. 5 One figure from it is worth carrying regardless of what the factor analysis settles: 53 per cent of those students said they would rather work alone than in a group. 5
References
- 1.Trentin, G. (2009). Using a wiki to evaluate individual contribution to a collaborative learning project. Journal of Computer Assisted Learning, 25(1), 43–55. Link ↗
- 2.Viégas, F. B., Wattenberg, M., & Dave, K. (2004). Studying cooperation and conflict between authors with history flow visualizations. Proc. ACM CHI Conference on Human Factors in Computing Systems (CHI ’04), 575–582. Link ↗
- 3.Wang, D., Olson, J. S., Zhang, J., Nguyen, T., & Olson, G. M. (2015). DocuViz: Visualizing collaborative writing. Proc. ACM CHI Conference on Human Factors in Computing Systems (CHI ’15), 1865–1874. Link ↗
- 4.Hakimi, L., Eynon, R., & Murphy, V. A. (2021). The ethics of using digital trace data in education: A thematic review of the research landscape. Review of Educational Research, 91(5), 671–717. Link ↗
- 5.Rasooli, A., Turner, J., Varga-Atkins, T., Pitt, E., Asgari, S., & Moindrot, W. (2024). Students' perceptions of fairness in groupwork assessment: Validity evidence for peer assessment fairness instrument. Assessment & Evaluation in Higher Education, 50(1), 111–126. Link ↗