A new paper by UC Berkeley puts a number on AI and assessment in higher education that’s hard to dismiss. The paper by researcher Igor Chirikov shows that in courses where the probability of using AI is higher, like writing essays or coding, the share of A's rose by 13 percentage points after ChatGPT's release, which is roughly 30% above the 2022 baseline.
This supports the idea that using AI is not a cheating problem but a design problem, which is showing up in transcripts. None of this is really new, if you think about it.
Writing was always an indirect signal of comprehension and knowledge. But an indirect signal only works if it’s hard to fake, and this one got easy to fake almost overnight.
So now, whether students can still write a good essay is beside the point. The question is if that was ever the right yardstick to begin with.
Why is student AI use a design problem, not a cheating problem?
Until almost four years ago, a well-written essay used to require proper understanding, because there was no other way to produce one. But that is no longer applicable today. Now, a chatbot can write a better essay in the time it takes to read this paragraph.
So when a student turns in a strong essay now, it doesn’t tell you what it told you before. It never told you everything, honestly.
Jisc's recent report on assessment trends suggests that instead of trying to police AI usage, universities should focus on designing assessments using authentic indicators of learning. These can be process-based activities, oral tests, or reflective tasks that require judgment only a human can make.
That reframes the debate around AI in higher education assessment. So instead of asking:
- How do we stop students from using AI?
The focus should be:
- What skills are we looking to assess?
- What evidence would convince us that a student has developed that skill?
- Can AI make such a submission without the student demonstrating their ability?
- What do we need to know about students’ capacity for judgment?
By thinking through these kinds of questions first, and then creating assessments focused on developing those skills, AI can become just another part of our overall approach to assessment.
The skill worth measuring now is judgment
There are two recent research papers that reach close agreement on this point despite starting out from very different angles.
In Designing assessment in a digital world: an organizing framework (Bearman, Nieminen & Ajjawi, 2023, Assessment & Evaluation in Higher Education), the authors developed ‘an organizing framework to inform assessment design in a digital world.’ One particular statement was that:
‘Evaluative judgment enables the student to learn standards and qualities associated with the assessment task’ (Tai et al. 2018)
And in the second framework, Validity Architecture for AI Integrated Assessment (VAAI) (Woodworth, Kharbach & Doe, 2026, Intersection), the authors offer an alternative path that ends at the same conclusion. It uses the term “non-delegable work.” This refers to a specific thinking that a student cannot outsource to a chatbot without sacrificing whatever the assignment was supposed to prove.
The whole approach of this framework treats the AI problem as a question about evidence and inference.
‘...could a student satisfy the assessment standard through AI-mediated strategies’
So when you line these two ideas together, a clear criterion appears. That explains the skill worth building right now and assessing is the ability to look at a piece of work (someone else’s or your own) and say plainly what’s strong, what’s weak, and what would actually fix it.
This is how a good editor earns their keep, so does a good clinician or a good engineer. They’re valuable because they can tell good from bad and explain the difference.
And this is exactly the test every rubric should have to pass. If a student can hand your assignment to AI and still clear the bar, the bar was never measuring what you thought it was.
What does the research say about how students actually learn?
The clearest evidence of this comes from the PAIRR project. At UC Davis, Sperber, MacArthur, Minnilo, Stillman, and Whithaus created a formative feedback model they called “Peer and AI Review + Reflection” or “PAIRR.”
In this program, over 654 college writers took part in ten different writing courses. Every student got both peer and AI feedback on the same draft. And then they had to reflect on both before revising anything.
That meant students needed to make decisions about whether the AI's suggestion really applied to their writing and identify discrepancies between what the AI thought was important and their peer reviewer’s judgment.
In fact, a quarter of all reflections revealed that students were specifically calling out things that ChatGPT suggested. They realized that AI made an error or simply found its recommendations misguided. One such reflection was:
‘AI can be helpful, but it’s critical to evaluate the relevance of the feedback’ (HSW, P509)
So, this evaluation skill allowed students to get better at writing. But more importantly, “This evaluation process - assessing whether the feedback meets their goals rather than stockpiling it helped students develop an understanding of what AI could do,” according to the authors.
All in all, it seems the solution was to allow students to use the technology themselves. This meant the system did not provide additional lectures around AI in education assessment, but instead allowed them to build these skills into their natural writing process.
How to redesign an assignment around judgment
Once judgment is your criterion for assessment, the whole grading approach changes. Meaning you start grading based on how your students evaluated the outcome.
The same task, before and after AI

Here are two versions of assigning tasks to your students.
The first one is the old way that grades according to the quality of the essay. And the second one is a redesign of the first version that grades based on judgment.
Three questions to ask about any task

The core idea behind this comes from the VAAI framework, which encourages treating assignments less like a thing to grade and more like evidence.
This means that the assignment should be designed so that it gives you evidence that a student has done their part. So before giving any task, ask yourself:
- What am I actually claiming a student can do by asking them to complete this?
- What evidence would genuinely justify that claim?
- Could a student clear my bar using AI without ever demonstrating the skill I’m grading?
If the honest answer to that 3rd question is yes, then the assignment needs to change.
Try to break your own assignment
Use the AI tools that your students are already using and try to pass your own assignment with them. Run the prompt, read what comes back, and check it against your rubric.
If a chatbot clears your bar without showing the competence you’re actually grading for, then the assignment has failed on its own terms.
And no detection will patch that hole. The only way moving forward is to redesign your assignment, pass it through the adversarial testing model, and then assign it.
Where do the rubric and the teacher still matter?
Nonetheless, redesigning an assessment doesn’t take the teacher out of the equation. Instead, it just changes the focus of their judgment. The teacher still has to approve every single grade and is still tasked with determining what constitutes “good.”
What has changed is that simply defining “good” is not sufficient anymore. A rubric that values well-organized prose will continue to assess and reward that. And as we all know, AI has no problem producing polished content.
The solution is to make sure the rubric is weighted towards defensibility rather than fluency of the content.
An essay that is technically perfect but becomes irrelevant if one asks the student to justify their choice should score lower than a more flawed essay where the student’s reasoning is clear.
If done consistently, it makes AI in grading and assessment part and parcel of the process.
The goal is not to keep AI out
Trying to keep AI out of higher education is probably the wrong battle to fight. Students will use these tools because they'll eventually use them at work too.
The challenge is making sure the grade still reflects what only the student can do. Good assessment has always been about collecting the right evidence. AI hasn't changed that principle, but simply made weak assessment designs much easier to spot.
When assignments are built around judgment instead of polished writing alone, AI becomes another tool students and teachers learn to work with. One such tool for teachers is EssayGrader AI, which helps you grade with ease.
You can try EssayGrader AI for free and see for yourself the amount of grading time it saves you.
.png)

.png)
.png)


.avif)
.avif)