1,339 teachers were each shown a student’s exercise with a wrong mark already on it, and the ones told an algorithm had set the mark finished further from the fair grade than the ones told a colleague had — but only where the mark was too harsh, with no detectable difference where it was too generous

Date:

A Teacher’s Encounter with Algorithmic Grading in Greece

A teacher in Greece opens a survey link and finds a student’s exercise displayed on screen. The exercise is divided into five parts, each already labelled as correct or incorrect. Beneath the exercise sits a mark someone else has assigned: five out of ten. The teacher’s only remaining task is to enter their own mark.

Four of the five parts are ticked correct, which, by the exercise’s own arithmetic, should amount to a score of eight. The displayed mark of five understates the work by three points, and all the information required to determine this is clearly visible on the screen.

The Promise of Human Oversight in Algorithmic Grading

The prevailing reassurance regarding automated decisions is that a human remains in the loop. The machine drafts, and a professional reviews and corrects any errors the machine might make. This principle underpins much policy on algorithmic grading, triage, and screening.

Grading offers an ideal testing ground for this promise: it combines evaluative judgment with objective arithmetic, bears significant real-world consequences, and involves professionals who ultimately shape outcomes. To explore this, researchers Sofoklis Goulas, Rigissa Megalokonomou, and Panagiotis Sotirakopoulos conducted a large-scale study involving 1,339 in-service teachers across Greece. Their findings were published in PNAS Nexus in June 2026 (read the study here).

The Experimental Setup: Who Made the Mistake?

The study’s vignette introduced two key variables. First, the source of the incorrect mark: either a colleague or an algorithmic grading system. Second, the nature of the error: whether the mark was harsher or more generous than the fair grade. Each teacher saw an exercise tailored to their subject area, and within each subject group, the same exercise was shown to all participants.

In the harsh condition, four of five answers were correct, making the fair mark eight, but the displayed mark was five, which is unreasonably low. In the generous condition, only one of five answers was correct, making the fair mark two, but the displayed mark was still five, representing an overly generous score. The recommended mark never changed from five.

The study measured the “grading fairness gap”—the absolute difference between the teacher’s own mark and the fair mark. Each participant contributed one data point, as they graded one exercise. Importantly, the trial was pre-registered with the AEA registry, enhancing its methodological transparency.

Notably, the study’s scope was deliberately narrow: it focused on one country, presented only one exercise per teacher, and was conducted via a survey completed in a median time of 7.2 minutes rather than a full marking session. It excluded evening, vocational, and special-education settings, focusing solely on day schools. The study recorded exact mark placements before querying teachers about their beliefs.

Findings: Human vs. Algorithmic Labels and the Grading Fairness Gap

In the harsh marking condition, when the incorrect mark was labelled as coming from a human, the average fairness gap was 1.384 points. Teachers still deviated by about 1.4 points from the correct mark, even when the correct answer was a simple calculation away. This shows a strong anchoring effect on the mark presented, regardless of its source.

When the same harsh mark was labelled as algorithmically assigned, the fairness gap increased to 1.584 points—an increase of 0.300 points or roughly 22% compared to the human baseline (p = 0.003 with controls, p = 0.022 without). This suggests that teachers were more reluctant to deviate from algorithmic marks when the grade was harsh.

In contrast, the generous condition showed no significant differences. The average fairness gap was 1.930 under the human label and 1.813 under the algorithmic label, with an insignificant estimate of −0.108 (p = 0.685). Thus, algorithmic labelling did not influence teachers’ willingness to adjust overly generous marks.

Pooling data across conditions confirmed this asymmetric pattern: only harsh algorithmic marks led to increased deference. Interestingly, the study’s prose and Table 2 present slightly different mean values, but the direction and magnitude of effects remain consistent.

Teachers’ Grading Patterns: Softening Harsh Marks and Inflating Generous Ones

Under the human label, teachers’ deviations were asymmetric. They tended to soften harsh grades—moving the mark closer to the fair grade of eight—and inflate generous marks, pushing the score further from the fair grade of two. The label effect was an addition to this existing bias rather than a replacement.

This pattern reflects a natural tendency among educators to protect students from unfairly harsh evaluations and, conversely, to adjust generous marks upwards—perhaps reflecting a complex interplay of fairness and leniency.

Why Do Teachers Defer More to Harsh Algorithmic Marks?

After entering their marks, teachers rated the grader on five dimensions: ability, comprehension, fairness, intent, and responsibility. Teachers consistently rated algorithmic graders lower than human graders across all dimensions, especially in the generous condition, where perceived ability, intent, fairness, and responsibility showed the widest gaps.

However, when the mark was harsh, the perceived ability gap narrowed. The authors interpret this as teachers perceiving severity as a sign of competence—harsh algorithmic grading was seen as a marker of the grader’s knowledge and seriousness.

Mediation analysis supported this: on the harsh version, perceived ability and responsibility significantly mediated deference to the algorithmic grader, accounting for about 73% and 47% of the effect respectively (noting these percentages overlap due to shared variance). Other dimensions like comprehension, fairness, and intent did not significantly mediate the effect. Conversely, in the generous condition, all five perception paths ran negatively and significantly, nullifying any net effect.

The study acknowledges that perception measures were taken after teachers saw the grader’s identity and entered their marks, limiting causal claims. The possibility remains that teachers who accepted the mark felt motivated to rate the grader positively afterward, a dynamic the authors note but do not fully explore.

Deference Patterns Among Different Teacher Subgroups

Deference to algorithmic graders under the harsh condition was not uniform across all teachers. Significant effects appeared among younger teachers (under 51 years), those holding a master’s or doctorate degree, humanities specialists, and teachers with high self-rated technological literacy. For older teachers, those with only a bachelor’s degree, STEM specialists, and low-tech-literacy teachers, the effect was small and statistically indistinguishable from zero.

The authors caution that confidence intervals for these subgroup comparisons overlap, so differences should be interpreted carefully. The fact that deference was stronger among technologically literate and highly educated teachers challenges narratives that portray teachers as uniformly wary of algorithmic grading.

The sample itself was somewhat unrepresentative: 8% of respondents held doctorates compared to roughly 2% in the Greek K-12 teaching population, and the average respondent age was 49, higher than the profession’s average of 40. Consequently, the findings may primarily reflect the perspectives of relatively senior educators in Greece.

Teacher Attitudes Toward AI Grading and Ethical Considerations

At the survey’s conclusion, teachers expressed ambivalent to negative attitudes toward AI grading. On a scale from −5 to +5, the average belief in AI’s fairness in grading was near neutral at 0.03. Willingness to allow AI to grade was negative at −1.03, and perceptions of ethicality were even lower at −1.33.

Nearly half of the teachers used generative AI tools weekly or more for lesson preparation, but more than half rarely or never encouraged colleagues to adopt such tools. Open-ended comments revealed substantial reservations, with about 75% of coded concerns focusing on moral or ethical blind spots—such as neglecting students’ unique contexts and challenges—and 25% addressing technical limitations. This skepticism was real and openly stated.

Interestingly, this skepticism appeared to influence responses predominantly in the generous condition, where all perception paths ran negative and the label left no detectable effect, contrasting with the harsher condition where deference prevailed.

Limitations and Broader Implications for Algorithmic Oversight

The task’s simplicity was intentional: by making the correct mark deducible from a straightforward list of ticks, the study isolated the effect of the grader’s label. However, this also means the results measure reluctance to overrule a source rather than oversight under real-world conditions involving ambiguity, fatigue, large workloads, and tight deadlines.

The authors acknowledge that real-world algorithmic grading systems often provide additional supports such as explanations, confidence scores, and accuracy records—features absent in this vignette-based study.

Ultimately, this experiment offers a controlled glimpse into how the label of an algorithmic grader can influence teacher judgments. It shows that even when teachers have all the information needed to identify the fair grade, they tend to anchor on the given mark, and that this anchoring intensifies when the mark is harsh and labelled as algorithmic.

The central question remains: if this is what oversight looks like when the error is obvious and no real student is affected, what can it achieve in the complex, high-stakes environment where errors are subtle and affect a student’s academic future?

For more detailed insights and the full study, visit Here.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Share post:

Popular

More like this
Related