What it means

A study published August 27, 2026 found that AI large language models tend to mark students' essays higher than human graders do. Researchers concluded that LLMs cannot be relied on to give an accurate indication of student performance — but that reliability gap is the central obstacle standing between pilot and policy.

What to do

If this is the problem on your desk, [talk to us](/contact).

Quick Answer

A study published August 27, 2026 found that AI large language models mark student essays higher than human graders do, and researchers concluded LLMs cannot reliably indicate student performance. Universities are piloting AI grading to ease workload pressure, but the reliability gap remains the central obstacle to wider adoption.

A study published August 27, 2026 found that AI large language models tend to mark students' essays higher than human graders do. Researchers concluded that LLMs cannot be relied on to give an accurate indication of student performance — but that reliability gap is the central obstacle standing between pilot and policy.

What the New Study Actually Found

A study published August 27, 2026 found that AI large language models tend to mark student essays higher than human graders do, and that LLMs cannot be relied on to give an accurate indication of student performance. arXiv cs.AI research from 2024–2025 documents a systematic positivity bias, with LLMs rating text quality 10–20% higher on average than calibrated human raters when no explicit rubric anchoring is applied.

A study published August 27, 2026 found that AI large language models tend to mark students' essays higher than human graders do. Researchers concluded that LLMs cannot be relied on to give an accurate indication of student performance.

ArXiv cs.AI preprints from 2024–2025 document a systematic positivity bias: LLMs rate text 10–20% higher than calibrated human raters without rubric anchoring. A 2025 study found GPT-4-class models agreed with human expert scores on only 58% of essay pairs, versus 74% human-to-human.

How AI Grading Student Essays Differs From Human Scoring

AI graders score essays 10–20% higher than calibrated human raters on average, per arXiv research (2025), while human inter-rater reliability benchmarks sit at Cohen's kappa 0.60–0.80. AI is faster and cheaper but less explainable and less consistent with expert judgment than trained human markers.

Human inter-rater reliability for holistic essay scoring typically produces Cohen's kappa values of 0.60–0.80 — the benchmark automated systems must meet for high-stakes use. Without structured rubric anchoring, LLMs rate text 10–20% higher on average than calibrated human raters.

Why Do LLMs Inflate Scores?

LLMs trained with reinforcement learning from human feedback (RLHF) tend to favor fluent, confident-sounding text regardless of accuracy. ArXiv cs.AI preprints from 2024–2025 show LLMs rate text quality 10–20% higher than calibrated human raters when no rubric anchoring is applied. OWASP GenAI guidance flags this pattern as a primary overreliance risk.

Models trained with RLHF favor outputs that appear confident and fluent, independent of factual accuracy — producing a positivity bias of 10–20% higher scores than calibrated human raters.

Automated scoring systems trained on historical grades can also inherit and amplify leniency bias already present in that data, with no built-in mechanism to correct for it.

OWASP GenAI guidance from 2024 flags overreliance as a primary risk and recommends mandatory human-in-the-loop review wherever outputs directly affect grades.

What Does This Mean for Institutions Exploring AI Grading?

Universities running AI grading pilots face real accreditation risk. Score inflation of even 0.3–0.5 points on a 4-point rubric can distort grade distributions across a cohort (Assessment in Education research). At least a dozen R1 universities launched AI grading pilots by 2025, most driven by cost rather than pedagogical evidence.

Universities are exploring AI-assisted grading to relieve pressure on overstretched graders. At least a dozen R1 universities launched internal pilots, with administrators citing cost and scalability rather than pedagogical evidence — creating risk when inflated scores distort grade distributions.

MIT Sloan Management Review research from 2024 found that organizations skipping validation pilots were three times more likely to face rollbacks. OWASP GenAI guidance recommends mandatory human-in-the-loop review for any AI application affecting grades.

Is Any AI Grading Tool Accurate Enough to Trust for High-Stakes Assessment?

Research sets the bar at Cohen's kappa 0.60–0.80 for high-stakes use, and current LLMs fall short without structured rubric prompts and human review. GPT-4-class models agreed with expert scores on only 58% of essay pairs in a 2025 study when rubric anchoring was absent. Formative feedback and human-in-the-loop workflows offer the safest near-term use case.

Human inter-rater reliability sets the benchmark at Cohen's kappa 0.60–0.80 for high-stakes scoring. GPT-4-class models reached only 58% agreement with expert scores without rubric prompts, versus 74% human-to-human — ruling out unsupervised AI grading for consequential assessments.

OWASP recommends human-in-the-loop oversight wherever grades are affected. EdSurge reported in 2025 that more than 20 vendors market these tools, with procurement driven by provost offices rather than faculty — raising accuracy-oversight concerns.

What Administrators and Edtech Buyers Should Do Now

Before buying any AI grading tool, require vendors to show accuracy benchmarks against human raters. Research shows LLMs rate essays 10–20% higher than calibrated human raters without rubric anchoring (arXiv cs.AI, 2025). Start with low-stakes formative work, build human-review checkpoints, and insist on structured validation pilots before any summative use.

Require vendors to provide accuracy benchmarks against human raters. LLMs rate essays 10–20% higher on average than calibrated human raters without rubric anchoring.

Pilot on low-stakes work first. OWASP recommends human-in-the-loop review wherever grades are affected.

More than 20 vendors market AI grading tools. Build human-review checkpoints for every summative grade.

Comparison of AI grading, human grading, and hybrid approaches across key decision dimensions. Cost and timeline figures are not available from verified public benchmarks and are omitted. Accuracy and bias figures from arXiv cs.AI preprints (2024–2025); inter-rater reliability benchmarks from Assessment in Education and Journal of Educational Measurement; workload context from Chronicle of Higher Education (2025); risk guidance from OWASP GenAI (2024) and MIT Sloan Management Review (2024); market context from EdSurge (2025) and Inside Higher Ed (2026).
DimensionAI GradingHuman GradingHybrid (AI + Human Review)
Typical cost rangevaries — no reliable public benchmarkvaries — no reliable public benchmarkvaries — no reliable public benchmark
Typical timeline
Best fitHigh-volume, low-stakes courses where speed and scalability matter (EdSurge, 2025)High-stakes assessments where accuracy and defensibility are essential (Assessment in Education)Pilots at R1 universities seeking cost relief without fully removing human oversight (Chronicle of Higher Education, 2025)
Key riskSystematic leniency bias: LLMs rate text 10–20% higher than calibrated human raters without explicit rubric anchoring; GPT-4-class models agreed with human experts on only 58% of essay pairs unaided (arXiv cs.AI, 2025)Inter-rater reliability produces Cohen's kappa values of 0.60–0.80; workload pressure means large courses can generate hundreds of essays per week a single instructor cannot review (Assessment in Education; Chronicle, 2025)OWASP GenAI guidance warns that skipping structured validation before full deployment raises risk of accuracy-related rollbacks; human-in-the-loop review is recommended for any AI application directly affecting grades (OWASP GenAI, 2024; MIT SMR, 2024)
SourcesarXiv cs.AI (2025); EdSurge (2025); Inside Higher Ed (2026)Assessment in Education; Chronicle of Higher Education (2025)OWASP GenAI (2024); MIT Sloan Management Review (2024); Chronicle of Higher Education (2025)