How accurate is our AI Essay Grader?
How we measure grading accuracy
Every benchmark essay already carries an official score published by the state program or research corpus that released it. We grade against that authoritative score, not one we commissioned. It's the fairest available target on a rubric, and not one we chose.
The engine never sees the human scores. We compare at the rubric-criterion level rather than the overall grade, because a right total built from two wrong criteria is not accuracy.
Exact-agreement percentages flatter any grader on a four-point rubric. We publish quadratic weighted kappa alongside them, the measure state assessment programs use, which penalises being badly wrong more than slightly wrong.
A grader can be accurate on average but too noisy to trust on any single essay. So we grade the same essay three times with identical inputs and measure how far the score drifts. A grade you can't reproduce is a grade you can't defend.

.avif)
.avif)
.avif)
Studies that were run
Each study lists its sample, its protocol and its result including the ones where we found a weakness and fixed it. Open any study to read the full report.
Teacher agreement
State-assessment alignment
Grading consistency
EssayGrader vs. human raters
Feedback at the right level
Actionable feedback
Teacher agreement benchmark, 22,000 essays
Every essay carries an official human score published by the program that released it: state anchor sets (Texas STAAR, Florida B.E.S.T., MCAS, PSSA/Keystone, NY Regents, Smarter Balanced), the ASAP 2.0 corpus, and College Board AP scoring guides. Our engine scored each essay blind against that program's own rubric, with no access to the human score. Results were compared criterion by criterion rather than on the overall grade, and scored with quadratic weighted kappa on each rubric's native scale.
State-assessment alignment, six programs
Released state writing prompts with published anchor papers and official scores, across Texas STAAR, Florida B.E.S.T., MCAS, PSSA/Keystone, NY Regents, Smarter Balanced, graded on each state's own rubric to see whether our scores track the real ones.
Grading consistency: same essay, three times
If a teacher regraded the same essay, would the marks hold? We asked that of our own grader, using the same essays, the same rubrics and the same settings, graded three separate times, and measured how far the scores moved.
How we compare to human raters
The bar for essay scoring isn't perfection; it's a trained human rater, and two qualified teachers don't fully agree either. We set our agreement with official scores beside the human-to-human figures each state publishes for its own raters.
Feedback at the student's reading level
Feedback that's accurate but written above a student's reading level doesn't get read. We scored every piece of feedback blind against a grade-level-appropriateness evaluator, across elementary through college.
Actionable feedback, not just corrections
Education research separates feedback that names a fault from feedback that names a next step. We measured, criterion by criterion, which kind ours produces, against Hattie & Timperley's model of effective feedback.
Accuracy terms, defined
What our accuracy terms mean
Every figure on this page uses these definitions. Vague accuracy claims usually hide a loose one.
Agreement
The AI score matched the human score exactly, or within one point on the same rubric criterion. This is the figure we lead with, state programs use the same "within one point" bar to qualify their own raters, where it's called adjacent agreement.
Exact Match
The AI score and the human score were identical on that criterion, with no tolerance. A stricter measure, reported alongside agreement so the looser number is never read on its own.
Quadratic weighted kappa
A measure of how closely two raters agree, corrected for the agreement you'd get by chance. It penalises being badly wrong far more than being slightly wrong. Below 0.60 is weak, 0.60-0.70 is acceptable, 0.70-0.80 is solid working agreement, and above 0.80 is comparable to a well-trained human rater.
Criterion
A single line of a rubric, scored on its own scale. We compare criterion by criterion rather than on an overall grade, because a total can match while the underlying judgements disagree.
Consistency
Whether the same essay, graded again with identical inputs, gets the same score. Measured separately from accuracy, because a grader can be repeatable and repeatably wrong, or accurate on average and too noisy to defend a single grade.
Human score
The score a study measures against: for state assessments, the official score the program published with its released papers; for research corpora, the score that dataset's trained raters assigned. Each study names its source.
Exemplars
Scored example papers attached to a rubric as reference points. Each study states whether a run used them, and where used they were held out of the scored set, so no paper was graded against itself.
How we handle AI-writing and plagiarism detection
No AI detector on the market is accurate enough to justify an academic integrity decision on its own, and any vendor claiming otherwise is selling you risk. Paraphrased AI text defeats every detector we have tested, including ours, and non-native English writing is disproportionately misflagged industry-wide.
So our reports go to the teacher, never automatically to the student or a parent, and they show evidence rather than a verdict. Use them to decide whether a conversation is worth having that conversation, not the score, is what establishes what happened.
How we keep grading accurate
Accuracy is not a one-time benchmark. These are the checks that run continuously, and the design choices that make a wrong score visible rather than silent.
Every score carries its reasoning
The report cites the rubric language and the passage that earned each score, so a teacher can check the logic in seconds instead of trusting a number.
The teacher approves everything
No grade or comment reaches a student until a teacher releases it. Human review is not an optional setting it is how the product works.
We report the honest metric
We publish quadratic weighted kappa, the measure states use to qualify their own scorers, not just the flattering agreement scores.
Grading is tested blind
When we benchmark, the engine never sees the human score, so agreement is earned, not fed to it.
No model ships without a re-run
The full benchmark suite runs against every model change. If agreement or kappa drops on any rubric type or student group, the change does not ship.
Calibration support for districts
Send us the essays where scoring felt wrong. We investigate with your team and tune the rubric alignment for your district, not just our benchmark.
What our AI grader is not good at
Any vendor claiming their grader has no weaknesses has not measured carefully. These are the cases where we tell teachers to look closely.
Creative and highly personal writing
Rubrics reward structure. Deliberate rule-breaking that a human reads as voice can score lower than it deserves, so poetry and experimental narrative need your eyes.
Factual claims outside the source material
If you have not attached the reading, the engine cannot always tell a confident wrong claim from a correct one. Reference materials close most of this gap.
Very short responses
Under roughly 75 words there is not enough evidence for criterion-level scoring to be reliable. These essays are flagged rather than scored with false confidence.
AI-writing detection is a signal, not a verdict
No detector, ours included, is accurate enough to accuse a student. We show the evidence and leave the judgment where it belongs - you, the teacher.
Do not take our word for it — audit us

Frequently asked questions?
Our answers to the most common questions teachers ask us.
How accurate is EssayGrader?
Our score matched the official human score exactly or within one point on 94% of rubric criteria, across eight programs. Each essay was already scored by the program that released it: state test papers and public research corpora. That is the same bar states use to qualify their own human scorers. Benchmarks are re-run on every model release.
How is EssayGrader accuracy measured?
Each essay is graded blind on the benchmark's own rubric, with no access to the human score. We then compare criterion by criterion rather than on the overall grade, since a total can match while the individual criterion scores disagree. Consistency is measured separately, by grading the same essay three times with identical inputs.
How closely do AI essay scores match teacher scores?
Four state programs publish how well their own trained scorers agree with each other. On all four, our agreement with the official score was equal or higher, averaging 0.79 where the programs report an average of 0.70. Two qualified humans do not fully agree on an essay either, and this is the realistic bar for comparisons.
Will the same essay get the same score twice?
Yes, almost every time. We graded 940 essays three times each, holding every input constant. 99% stayed within 10% of themselves. Average movement was a fraction of a point on a rubric's scale. Consistency held across elementary through college work.
Does using my own rubric improve the accuracy?
Your rubric is what we grade against by default, not a rubric of ours. Every benchmark we publish was run on the assessment program's own rubric for the same reason. Attaching two or three scored examples improves agreement further.
Is EssayGrader accurate enough for state writing assessments like STAAR or Florida B.E.S.T.?
Yes, for practice and formative scoring. We benchmark against released papers from eight assessment programs, graded on each program's own published rubric. EssayGrader's agreement scores are equal or higher than among trained teachers. EssayGrader is not an official scoring engine, and its scores do not replace the state's own.
How accurate is the feedback, not just the score?
We measure the feedback separately from the score, against Hattie and Timperley's model of effective feedback. 83% of EssayGrader's suggestions name a method the student can reuse on the next assignment. We also check readability: 95% of comments are written at or below the reading level of the essay they respond to.
Does a new AI model change how my essays are graded?
Not without being checked first. A new model reaches classrooms only after the full benchmark has been re-run, so an upgrade cannot quietly shift how essays are scored. If agreement drops on any rubric type, the change does not ship.
Can an AI essay grader make mistakes?
Yes. Agreement with official scores runs about 94% on individual rubric lines, so some judgments differ from what a human rater would give. When our scores disagree with a teacher it is almost always by a single point on one criterion. Teachers can change any score before it reaches the student.
How often does an AI essay grader get the score wrong?
Across our benchmark, about 6 rubric criteria in every 100 landed more than a point from the human score. Misses cluster in predictable places: very short responses, highly personal writing, and anything the rubric does not cover. The rate varies by program, and it improves when the rubric includes released exemplars.
What are the limitations of an AI essay grader?
An essay grader can only be as good as the rubric it grades against. It scores what the criteria describe, so very short responses, highly personal writing, and anything outside the rubric need a human read. It also cannot say whether a student's claims are factually true.
Is AI essay grading accurate enough to use for real grades?
Yes, when a teacher issues the final grade. Agreement with official scorers on individual rubric lines runs 94%, and every score arrives attached to the rubric line and the reasoning behind it.
No grade is final until it is approved by the teacher.
Can an AI essay grader replace a teacher or professor?
No, and it is not built to. An AI grader applies a rubric and drafts feedback, which is one task inside teaching. Deciding what the class needs next, sitting with a student who has stopped trying, and setting the standards in the rubric all remain with the teacher, who also decides which scores to keep.
Can I see why it gave a score, and can I override it?
Yes to both. Every score shows the rubric criterion it maps to and the reasoning behind it, so what gets reviewed is an argument rather than a number. A single click changes any score, which usually settles a borderline line faster than re-reading the essay. The grade a student sees is the one the teacher released.
How can I check the accuracy on my own students' essays?
Run a small audit on a single class. Take 20 essays already graded by hand, upload the same rubric, and grade them again. Comparing the two sets shows where the scores diverge and by how much. Our team will review any that missed and explain the cause. The first 100 essays are free.