AI Grading Accuracy Center

How accurate is our AI Essay Grader?

For every accuracy claim we make - the study behind it, the methodology we adopted and the sample size is all listed here.
Last updated August 2026
94%
Agreement with experienced teacher scores
0.84
QWK on our top state-standard alignment
22K+
Human-scored essays in our benchmark set
99%
Essays scored consistently across three graded passes
Agreement means the AI score matched the teacher's score exactly or within one point on the same rubric criterion.
Methodology

How we measure grading accuracy

Every number in our Grading Accuracy Center comes from the same method. Here's exactly what we test against, how we score it, and what each metric does and doesn't tell you. Terms like agreement and kappa are defined in our accuracy glossary.
Blind scoring, criterion by criterionQuadratic weighted kappa reported alongside agreementConsistency: the same essay graded three times
Research

Studies that were run

Each study lists its sample, its protocol and its result including the ones where we found a weakness and fixed it. Open any study to read the full report.

Table of Contents

Teacher agreement

State-assessment alignment

Grading consistency

EssayGrader vs. human raters

Feedback at the right level

Actionable feedback

Study 01· JULY 2026

Teacher agreement benchmark, 22,000 essays

Every essay carries an official human score published by the program that released it: state anchor sets (Texas STAAR, Florida B.E.S.T., MCAS, PSSA/Keystone, NY Regents, Smarter Balanced), the ASAP 2.0 corpus, and College Board AP scoring guides. Our engine scored each essay blind against that program's own rubric, with no access to the human score. Results were compared criterion by criterion rather than on the overall grade, and scored with quadratic weighted kappa on each rubric's native scale.

Read the full report
94%
criterion-level agreement
Study 02· AUGUST 2026

State-assessment alignment, six programs

Released state writing prompts with published anchor papers and official scores, across Texas STAAR, Florida B.E.S.T., MCAS, PSSA/Keystone, NY Regents, Smarter Balanced, graded on each state's own rubric to see whether our scores track the real ones.

Read the full report
0.75
average QWK vs. official anchor scores
Study 03 · JANUARY 2026

Grading consistency: same essay, three times

If a teacher regraded the same essay, would the marks hold? We asked that of our own grader, using the same essays, the same rubrics and the same settings, graded three separate times, and measured how far the scores moved.

Read the full report
99%
scored within 10% across three passes
Study 04· July 2026

How we compare to human raters

The bar for essay scoring isn't perfection; it's a trained human rater, and two qualified teachers don't fully agree either. We set our agreement with official scores beside the human-to-human figures each state publishes for its own raters.

Read the full report
0.79 vs 0.70
our agreement vs. reported human agreement
Study 05· AUGUST 2025

Feedback at the student's reading level

Feedback that's accurate but written above a student's reading level doesn't get read. We scored every piece of feedback blind against a grade-level-appropriateness evaluator, across elementary through college.

Read the full report
95%
of feedback within one grade band of the student
Study 06· June 2026

Actionable feedback, not just corrections

Education research separates feedback that names a fault from feedback that names a next step. We measured, criterion by criterion, which kind ours produces, against Hattie & Timperley's model of effective feedback.

Read the full report
83%
of suggestions give a strategy, not just a correction
Glossary

Accuracy terms, defined

What our accuracy terms mean

Every figure on this page uses these definitions. Vague accuracy claims usually hide a loose one.

Agreement

The AI score matched the human score exactly, or within one point on the same rubric criterion. This is the figure we lead with, state programs use the same "within one point" bar to qualify their own raters, where it's called adjacent agreement.

Exact Match

The AI score and the human score were identical on that criterion, with no tolerance. A stricter measure, reported alongside agreement so the looser number is never read on its own.

Quadratic weighted kappa

A measure of how closely two raters agree, corrected for the agreement you'd get by chance. It penalises being badly wrong far more than being slightly wrong. Below 0.60 is weak, 0.60-0.70 is acceptable, 0.70-0.80 is solid working agreement, and above 0.80 is comparable to a well-trained human rater.

Criterion

A single line of a rubric, scored on its own scale. We compare criterion by criterion rather than on an overall grade, because a total can match while the underlying judgements disagree.

Consistency

Whether the same essay, graded again with identical inputs, gets the same score. Measured separately from accuracy, because a grader can be repeatable and repeatably wrong, or accurate on average and too noisy to defend a single grade.

Human score

The score a study measures against: for state assessments, the official score the program published with its released papers; for research corpora, the score that dataset's trained raters assigned. Each study names its source.

Exemplars

Scored example papers attached to a rubric as reference points. Each study states whether a run used them, and where used they were held out of the scored set, so no paper was graded against itself.

Academic integrity

How we handle AI-writing and plagiarism detection

No AI detector on the market is accurate enough to justify an academic integrity decision on its own, and any vendor claiming otherwise is selling you risk. Paraphrased AI text defeats every detector we have tested, including ours, and non-native English writing is disproportionately misflagged industry-wide.

So our reports go to the teacher, never automatically to the student or a parent, and they show evidence rather than a verdict. Use them to decide whether a conversation is worth having  that conversation, not the score, is what establishes what happened.

Safeguards

How we keep grading accurate

Accuracy is not a one-time benchmark. These are the checks that run continuously, and the design choices that make a wrong score visible rather than silent.

Every score carries its reasoning

The report cites the rubric language and the passage that earned each score, so a teacher can check the logic in seconds instead of trusting a number.

The teacher approves everything

No grade or comment reaches a student until a teacher releases it. Human review is not an optional setting it is how the product works.

We report the honest metric

We publish quadratic weighted kappa, the measure states use to qualify their own scorers, not just the flattering agreement scores.

Grading is tested blind

When we benchmark, the engine never sees the human score, so agreement is earned, not fed to it.

No model ships without a re-run

The full benchmark suite runs against every model change. If agreement or kappa drops on any rubric type or student group, the change does not ship.

Calibration support for districts

Send us the essays where scoring felt wrong. We investigate with your team and tune the rubric alignment for your district, not just our benchmark.

Where we are honest

What our AI grader is not good at

Any vendor claiming their grader has no weaknesses has not measured carefully. These are the cases where we tell teachers to look closely.

Creative and highly personal writing

Rubrics reward structure. Deliberate rule-breaking that a human reads as voice can score lower than it deserves, so poetry and experimental narrative need your eyes.

Factual claims outside the source material

If you have not attached the reading, the engine cannot always tell a confident wrong claim from a correct one. Reference materials close most of this gap.

Very short responses

Under roughly 75 words there is not enough evidence for criterion-level scoring to be reliable. These essays are flagged rather than scored with false confidence.

AI-writing detection is a signal, not a verdict

No detector, ours included, is accurate enough to accuse a student. We show the evidence and leave the judgment where it belongs - you, the teacher.

Verify it yourself

Do not take our word for it — audit us

Our numbers are only useful if they hold on your students' writing. Here is how to check in an afternoon.
Pick 20 essays you have already graded ideally a spread from strongest to weakest
Upload your rubric and grade them without looking at your original scores
Compare criterion by criterion, not total by total that is where you learn something
Send us the ones that missed we will tell you why and tune the alignment with you

Frequently asked questions?

Our answers to the most common questions teachers ask us.

How accurate is EssayGrader?

Our score matched the official human score exactly or within one point on 94% of rubric criteria, across eight programs. Each essay was already scored by the program that released it: state test papers and public research corpora. That is the same bar states use to qualify their own human scorers. Benchmarks are re-run on every model release.

How is EssayGrader accuracy measured?

Each essay is graded blind on the benchmark's own rubric, with no access to the human score. We then compare criterion by criterion rather than on the overall grade, since a total can match while the individual criterion scores disagree. Consistency is measured separately, by grading the same essay three times with identical inputs.

How closely do AI essay scores match teacher scores?

Four state programs publish how well their own trained scorers agree with each other. On all four, our agreement with the official score was equal or higher, averaging 0.79 where the programs report an average of 0.70. Two qualified humans do not fully agree on an essay either, and this is the realistic bar for comparisons.

Will the same essay get the same score twice?

Yes, almost every time. We graded 940 essays three times each, holding every input constant. 99% stayed within 10% of themselves. Average movement was a fraction of a point on a rubric's scale. Consistency held across elementary through college work.

Does using my own rubric improve the accuracy?

Your rubric is what we grade against by default, not a rubric of ours. Every benchmark we publish was run on the assessment program's own rubric for the same reason. Attaching two or three scored examples improves agreement further.

Is EssayGrader accurate enough for state writing assessments like STAAR or Florida B.E.S.T.?

Yes, for practice and formative scoring. We benchmark against released papers from eight assessment programs, graded on each program's own published rubric. EssayGrader's agreement scores are equal or higher than among trained teachers. EssayGrader is not an official scoring engine, and its scores do not replace the state's own.

How accurate is the feedback, not just the score?

We measure the feedback separately from the score, against Hattie and Timperley's model of effective feedback. 83% of EssayGrader's suggestions name a method the student can reuse on the next assignment. We also check readability: 95% of comments are written at or below the reading level of the essay they respond to.

Does a new AI model change how my essays are graded?

Not without being checked first. A new model reaches classrooms only after the full benchmark has been re-run, so an upgrade cannot quietly shift how essays are scored. If agreement drops on any rubric type, the change does not ship.

Can an AI essay grader make mistakes?

Yes. Agreement with official scores runs about 94% on individual rubric lines, so some judgments differ from what a human rater would give. When our scores disagree with a teacher it is almost always by a single point on one criterion. Teachers can change any score before it reaches the student.

How often does an AI essay grader get the score wrong?

Across our benchmark, about 6 rubric criteria in every 100 landed more than a point from the human score. Misses cluster in predictable places: very short responses, highly personal writing, and anything the rubric does not cover. The rate varies by program, and it improves when the rubric includes released exemplars.

What are the limitations of an AI essay grader?

An essay grader can only be as good as the rubric it grades against. It scores what the criteria describe, so very short responses, highly personal writing, and anything outside the rubric need a human read. It also cannot say whether a student's claims are factually true.

Is AI essay grading accurate enough to use for real grades?

Yes, when a teacher issues the final grade. Agreement with official scorers on individual rubric lines runs 94%, and every score arrives attached to the rubric line and the reasoning behind it.

No grade is final until it is approved by the teacher.

Can an AI essay grader replace a teacher or professor?

No, and it is not built to. An AI grader applies a rubric and drafts feedback, which is one task inside teaching. Deciding what the class needs next, sitting with a student who has stopped trying, and setting the standards in the rubric all remain with the teacher, who also decides which scores to keep.

Can I see why it gave a score, and can I override it?

Yes to both. Every score shows the rubric criterion it maps to and the reasoning behind it, so what gets reviewed is an argument rather than a number. A single click changes any score, which usually settles a borderline line faster than re-reading the essay. The grade a student sees is the one the teacher released.

How can I check the accuracy on my own students' essays?

Run a small audit on a single class. Take 20 essays already graded by hand, upload the same rubric, and grade them again. Comparing the two sets shows where the scores diverge and by how much. Our team will review any that missed and explain the cause. The first 100 essays are free.