← All features
Feature

Before anyone writes a word

Most of what makes a group mark defensible is decided in week one. This is the part of Dwixel that happens before there is anything to measure.

5 min read

Group work goes wrong in week seven for reasons that were settled in week one. Measuring contribution is the last thing Dwixel does, not the first.

Week one, before anyone writes

The first thing a group does is answer five questions, together, in their own words: how often they will meet, where they will talk and how fast someone should reply, what their own deadlines are before the real one, what finished work looks like, and what they will do if someone stops contributing. Then each person writes one line saying what they are taking on, and everybody signs it.

Editing the wording afterwards clears every signature, because those were given to a different version. Reopening a signed agreement is recorded. Plans change, and hiding that they changed is what causes the argument at the end.

How often will you meet, and where?
Where will you talk, and how fast should someone reply?
What are your own deadlines, before the real one?
What does finished work look like here?
What will you do if someone stops contributing?
0 of 4 signed
Five questions, in their own words, signed by everyone.

A contract of contribution agreed at the start is the first anti-freeloading strategy in the published guidance for academics, 1 and the framework has to be scaffolded while the rules stay the group’s own rather than conforming to rules enforced by the academic. 1 That is why Dwixel asks questions instead of supplying a template.

What counts here, decided by them

Alongside the standard criteria, the group adds its own. Turns up to the lab sessions, not just the write-up. Those lines then print on the peer review form six weeks later, under the heading “And what your group agreed”. That second part is the whole mechanism: criteria sitting in a document nobody reads at rating time do nothing.

Criteria that students helped derive and agree produced peer marks agreeing with the tutor far better than teacher-supplied criteria did. 2 And the same meta-analysis finds a single overall judgement reaches its highest agreement only where the rater knows the criteria, dropping sharply without them. 2 Student involvement in developing criteria is highly desirable. 3

More than one assessment point

Peer assessment runs more than once, and only the last round carries any grade weight. A student rated badly in week five who then changed is marked on what they did afterwards, which is the only thing that makes a warning worth responding to.

Formative · carries no marks
Omar
4.6
Grace
4.2
Finn
3.9
Ruby
2.4

Ruby sits well below the group. Round one stays on the record and carries no grade weight.

Does the student see round one?

The default. Nobody can respond to a score they have not seen, so round two rates the work rather than the last round.

There is no setting that gets both. Dwixel picks the safer default and says which one it picked.

The mid-point round exists so that free-riding is identified early enough that a student has an opportunity to begin contributing or receive help, 4 which requires telling the student. Doing that invites the next round to become a reply to the last one, and the answer to that is one line: instructors control whether and when they release this feedback to their students. 5 There is no setting that gets both, so Dwixel defaults to the safer one and says which it picked.

The form itself

One overall rating per teammate, not five. The five criteria are printed above it as what to weigh up, and each rater is asked about one of them in detail rather than all five. A written sentence is required, and the prompt asks for something specific, because “they were great” is the inflation this exists to interrupt.

What to weigh up
  • Contributing to the team's work
  • Interacting with teammates
  • Keeping the team on track
  • Expecting quality
  • Bringing relevant knowledge and skills
And what your group agreed
  • Turns up to the lab sessions, not just the write-up
  • Answers in the group chat within a day
FDFinn Doherty
Overall, how did Finn contribute compared with the rest of the group?
You have been asked to look at keeping the team on track for Finn. What did he do there?
A sentence is enough. Naming something concrete is what keeps this useful.
The scale compares each person with their own group, not against a fixed standard. Try submitting without writing anything.

One global judgement made in the knowledge of the criteria agrees with tutor marks far better than scoring separate dimensions does. 2 CATME’s own validation explains why: the dimensions intercorrelate at 0.76, the authors combined them into a composite to run their analysis, and only two of the five predicted anything independently. 5 Five boxes never produced five judgements. They produced one, written down five times.

The scale is relative to the group rather than absolute, which matters because ratings cluster near the top: on an instrument designed with three as satisfactory, the mean recorded in practice was about 4.2. 5

Five minutes of practice first

Before rating anyone, a student can rate four made-up teammates and see how experienced markers scored the same people. One of them did their job properly, on time, and nothing more, and the expert answer is three. Most people give that person a four, and that is exactly the habit worth interrupting before it reaches a real group.

It never touches a real rating. On three of five dimensions the rater explained more of the variance than the person being rated, by a factor of four to six, and the remedy the same authors shipped is a practice-rating exercise compared against expert raters. 5 Correcting somebody’s actual marks against a four-question exercise would claim far more than that supports.

When it still goes wrong

The group can ask staff to look at a member who has stopped contributing. What happens next is fixed: the academic reads the record and speaks to the student before anything counts. Only then can a card be issued, and it can be lifted the moment contribution resumes. The person who raised it is visible to staff and to nobody else.

1
The group asks

A member raises it, in their own words. Nothing happens automatically and no number can trigger this.

2
You read it and speak to them

The record, then the student. This step is not optional and the system enforces it.

3
Yellow card

Marks at risk, and liftable the moment contribution resumes. It exists to change what happens next.

4
Red card

After a written investigation. Removal and a mark of zero, with the rest of the group not disadvantaged.

Rescinding is a first-class action, not an undo. A card exists to change what happens next.
Why the meeting is not optional
Illness, caring responsibilities and undisclosed disability all look identical to a low word count. Dwixel refuses to issue a card until the academic confirms they have spoken to the student. And where a card is open, the marker is told the contribution record and the peer ratings are no longer independent: a group that pushes someone out tends to rate them badly too, so averaging the two looks like corroboration when it is one dispute counted twice.

What this does not fix

A shared mark makes students share an outcome. It does not make them need each other, and the second is what the research on team learning finds more consistently effective. Education mostly operationalises interdependence as the shared result, while task interdependence — the degree to which members have to rely on one another to accomplish the task — shows the better effects. 6

That is uncomfortable, because everything Dwixel measures sits on the result side. If four people can split a brief into four private sections and staple it together on the last night, no measurement added afterwards rescues it. So when a brief is written, Dwixel asks three questions that are not about the mark scheme at all.

  • ·Does each person hold something the others cannot get on their own — a different reading, dataset, interview or site visit?
  • ·Does the task have genuinely different jobs in it, or is it one job done four times and stapled together?
  • ·Is there a point where one person cannot start until another has finished?

If the answer to all three is no, the group is an administrative arrangement rather than a team, and no amount of contribution tracking downstream will fix that.

References

  1. 1.Francis, N., Allen, M., & Thomas, J. (Advance HE) (2022). Using group work for assessment — an academic’s perspective. Advance HE, UK. Link ↗
  2. 2.Falchikov, N., & Goldfinch, J. (2000). Student peer assessment in higher education: A meta-analysis comparing peer and teacher marks. Review of Educational Research, 70(3), 287–322. Link ↗
  3. 3.Topping, K. J. (1998). Peer assessment between students in colleges and universities. Review of Educational Research, 68(3), 249–276. Link ↗
  4. 4.Hall, D., & Buzwell, S. (2013). The problem of free-riding in group projects: Looking beyond social loafing as reason for non-contribution. Active Learning in Higher Education, 14(1), 37–49. Link ↗
  5. 5.Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630. Link ↗
  6. 6.Kyndt, E., Raes, E., Lismont, B., Timmers, F., Cascallar, E., & Dochy, F. (2013). A meta-analysis of the effects of face-to-face cooperative learning. Do recent studies falsify or verify earlier findings?. Educational Research Review, 10, 133–149. Link ↗