Get Started
Research questionHow can AI-assisted scoring reduce grading workload in national assessments without compromising human-judged scores or pass/fail decisions?National assessments must score thousands of short written responses consistently, yet automated judgments can diverge from expert ratings and affect pass/fail outcomes. The central difficulty is determining when AI assistance is reliable and when human review remains necessary.
AI
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing AssessmentThe evidence concerns a human-in-the-loop system for 150–200-word responses in a national assessment, using data from two test editions with about 5,000 responses each. It examines agreement with human raters across rubric dimensions and effects on pass/fail decisions; the input does not establish broader deployment conditions or longitudinal performance.research paper · Sep 4, 2026
Related questions
How do students distinguish AI-generated writing feedback’s usefulness from its authority to assign grades?How can rubric-based LLM grading resist prompt injection without compromising fair judgments?How can AI-assisted peer review reduce unsupported criticisms without overlooking consequential weaknesses?How can we distinguish AI assistance from AI authorship when the prompting dialogue is unavailable?
Home
Topics
Search
Library