Writing

Agent systems · October 3, 2026

Grading the Tutor: How We Evaluate AI Agents in Education

Evaluating AI agents in education requires moving beyond simple accuracy metrics to assess pedagogical effectiveness, ethical deployment, and long-term learning outcomes. This essay examines the technical and institutional principles behind evaluating educational AI for teachers and researchers.

Evaluating AI agents in education is a technical challenge that determines whether these systems actually help students learn or merely simulate helpfulness. This essay is for teachers and researchers who need to understand how we measure the effectiveness of AI tutors, from controlled studies to classroom deployments.

When an AI agent interacts with a student, it generates text, offers hints, and responds to questions. But generating fluent text is not the same as teaching. The distinction between a conversational chatbot and an effective educational tool lies entirely in evaluation. Without rigorous assessment frameworks, schools risk deploying systems that are engaging but pedagogically hollow, or worse, actively harmful to student development. Understanding how we evaluate these agents—both technically and institutionally—is essential for anyone tasked with integrating them into learning environments.

A student working at a desk with a laptop in a quiet classroom

The Legacy of Intelligent Tutoring Evaluation

The problem of evaluating educational technology is not new, even if the underlying models are. Decades before large language models became capable of holding open-ended conversations, researchers were building intelligent tutoring systems designed to guide students through specific domains like mathematics or physics. The foundational work in this area established strict benchmarks for what counts as evidence of learning.

A comprehensive meta-analysis by VanLehn (2011), published as Intelligent Tutoring Systems and Learning: A Review of the Literature, examined decades of research on how these systems are evaluated for learning effectiveness. VanLehn’s review highlights a critical principle: evaluating a tutor requires comparing student outcomes against specific baselines, such as traditional classroom instruction or human one-on-one tutoring. The research establishes that simply measuring whether a student completes a task while using the software is insufficient. Instead, evaluators must look at pre-test and post-test gains, the efficiency of learning (how much time it takes to reach mastery), and the transfer of knowledge to novel problems.

For modern AI agents, which are often built on general-purpose generative models rather than domain-specific architectures, these historical benchmarks remain vital. An AI agent might be able to explain a calculus concept clearly, but does interacting with that agent lead to durable learning? VanLehn’s review reminds us that effectiveness is not a single metric but a matrix of outcomes. When researchers evaluate a new AI agent today, they must still ask whether the system produces measurable learning gains compared to a control group, rather than relying solely on user satisfaction surveys or engagement metrics. Engagement is easy to engineer; deep learning is not.

Furthermore, the evaluation of step-based versus substep-based feedback, a major theme in earlier tutoring research, translates directly to how modern AI agents handle planning and scaffolding. If an agent evaluates a student's multi-step math problem, does it intervene at the right moment? Evaluating the timing and granularity of AI feedback requires the same careful experimental design that defined the earlier era of intelligent tutoring systems.

Stacked research papers and notebooks on a library table

Institutional Frameworks and Ethical Evaluation

Technical efficacy is only one dimension of evaluation. An AI agent might produce excellent test score gains in a controlled laboratory setting but fail entirely when deployed across diverse school districts. This is where broader institutional frameworks become necessary.

The UNESCO publication Artificial Intelligence in Education: Promises and Implications for Teaching and Learning outlines the ethical guidelines and institutional structures required for evaluating AI technologies in real-world educational settings. According to this framework, evaluating an AI agent cannot be separated from evaluating its safety, equity, and impact on the teaching profession. The document emphasizes that deployment must be assessed not just for what the AI teaches, but for how it handles student data, whether it introduces algorithmic bias, and how it alters the dynamic between teachers and students.

This institutional perspective shifts the evaluation burden from purely cognitive science to systemic oversight. For example, if an AI agent is evaluated only on its ability to answer questions correctly, developers might ignore how the system performs for students with different linguistic backgrounds or learning disabilities. The UNESCO guidelines argue that evaluation must include audits for fairness and inclusivity. A system that works brilliantly for native English speakers but consistently misunderstands the syntax of multilingual learners fails the institutional test of educational equity, regardless of its average performance metrics.

Moreover, the UNESCO framework stresses the importance of evaluating the role of the teacher. AI agents should be assessed on how well they support human instruction rather than replace it. If an evaluation shows that a tool saves teachers time but simultaneously reduces their awareness of individual student struggles, the tool has introduced a hidden cost. Effective evaluation frameworks therefore require input from educators, administrators, and policymakers, ensuring that the metrics used to judge the AI align with the broader goals of the educational community.

Designing Better Benchmarks for Generative Agents

Combining the technical rigor of historical tutoring research with the ethical breadth of institutional guidelines gives us a clearer picture of how to evaluate modern AI agents. Today’s systems are highly unpredictable compared to older, rule-based tutors. They can hallucinate facts, shift their pedagogical tone mid-conversation, and respond differently to the same prompt depending on subtle variations in phrasing. This non-determinism makes evaluation significantly harder.

Researchers and developers must build evaluation pipelines that account for this variability. This means running thousands of simulated student interactions to map out the boundaries of the agent's behavior before it ever reaches a classroom. It means designing rubrics that penalize confident misinformation heavily, because a student is more likely to internalize a wrong answer delivered fluently than one delivered hesitantly.

Ultimately, evaluating AI agents in education is not a problem that can be solved by computer scientists alone. It requires the empirical discipline outlined in VanLehn’s review of intelligent tutoring systems, ensuring that we measure actual learning rather than surface-level interaction. And it demands the holistic, ethically grounded approach championed by UNESCO, ensuring that our pursuit of technological innovation does not compromise equity, privacy, or the essential role of the human teacher. For educators and researchers, understanding these dual pillars of evaluation is the first step toward demanding AI tools that genuinely serve the classroom.

Sources