AI grading is often discussed as a single capability. It is not.
There is a meaningful difference between organizing rubric comments, suggesting formative feedback, scoring a multiple-choice item, recommending a writing score, and assigning a final course grade. The educational consequence and the need for professional judgment rise sharply across that sequence.
In brief: districts should classify AI-supported assessment tasks by consequence. Lower-risk clerical and formative support may be suitable for a controlled pilot. High-stakes scoring, final grades, placement, eligibility, and disciplinary consequences should not be delegated to an opaque automated output. A qualified educator must remain accountable, and students need a clear path to human review.
Why the category needs to be unpacked
“AI grading” can describe at least four different jobs:
- clerical support: sorting responses, formatting feedback, or mapping teacher-written comments to a rubric
- formative support: suggesting questions or feedback while learning is still in progress
- scoring assistance: recommending a score or proficiency level for educator review
- final evaluation: determining a grade, placement, credential, intervention, or other consequential outcome
A district that writes one rule for all four will either prohibit useful low-risk support or allow high-risk automation without adequate safeguards.
The right unit of governance is the task and consequence, not the marketing label.
A consequence-based decision framework
Level 1: clerical and organizational support
Examples include grouping similar misconceptions, converting teacher notes into a consistent format, or preparing a draft comment bank from district-approved rubric language.
These uses may be lower risk when:
- no protected information enters an unapproved system
- the output does not determine a score
- the educator can quickly verify accuracy
- the tool does not send feedback directly to students
The main question is whether the support actually reduces work after review time is counted.
Level 2: formative feedback support
Examples include suggesting a follow-up question, identifying a possible reasoning gap, or drafting feedback on a practice response.
The educator should check:
- whether the feedback matches what was taught
- whether it identifies the student's actual misconception
- whether tone and reading level are appropriate
- whether it gives away an answer instead of supporting thinking
- whether patterns differ across language backgrounds or student groups
Formative feedback can influence confidence and learning even when it does not affect a grade. “Low stakes” is not the same as “no stakes.”
Level 3: scoring recommendation
At this level, the system proposes a score, rubric level, or classification that a teacher may accept or change.
This requires a stronger pilot design:
- a clearly defined rubric
- representative samples across performance levels
- agreement testing between qualified human scorers and the system
- review of disagreements, not only average accuracy
- subgroup analysis
- documentation of when a teacher must ignore or override the recommendation
- a record of the human decision
