~/writing
Modern AI Systems

Agent Evaluation

A breakdown of how AI agent evaluation differs from LLM evaluation, covering outcome, trajectory, safety, operational, and consistency metrics, plus how golden datasets and LLM-as-judge grading fit in.

AI agent evaluation refers to the systematic measurement of an AI agent’s quality, reliability and safety across a completed task. Basically, it measures whether an agent reaches the right outcome through a sound, governed and repeatable path. The final, correct answer is only a part of the evaluation; the entire evaluation includes the user request, the plan the agent followed, the tools it selected to use, the arguments it passed, the context it retrieved and the intermediate decisions it made ON TOP of the final result that the agent returns.

Agent evaluation is different from LLM evaluation. A LLM application is usually single-turn and only has one evaluation target, with a response to a prompt. The evaluator can score that response that response for correctness, groundedness, instruction following, safety or style against a reference/rubric. An agent has more moving parts – receiving the goal, decomposing steps, selecting tools, updating its plan etc. The failure surface expands with each extra moving part because agent interactions are multi-turn, and so evaluation has to account for those intermediate states.

Agent evaluation metrics usually fall into 4 groups: outcome metrics, trajectory metrics, operational metrics and safety metrics.

Agent evaluation metric categories

Outcome metrics

This answers the most direct question: did the agent complete the task? Common metrics used are task success rate, task completion rate, goal accuracy, final-answer correctness etc. If a clear answer exists, the evaluation can compare the output against ground truth. If multiple answers exist, the evaluation needs a rubric that defines success, partial success and failure.

Trajectory metrics

This inspects the path that the agent took through the task. They measure whether the agent chose the right tools, used those tools in the right order, passed the right arguments and interpreted tool results correctly.

Tool-call accuracy is one of the most meaningful trajectory metrics. It includes both choosing the right tools, supplying valid arguments and responding appropriately to the tool result. Other trajectory metrics include reasoning quality, step efficiency, retrieval relevance, verification behavior and loop rate. If an agent calls three tools when one would have been enough, repeats the same failed action or skips a required check before finalizing the answer, the final response might still look acceptable while the trace reveals a reliability issue.

Safety & Compliance metrics

This measures whether the agent stayed within defined boundaries. In enterprise workflows, these boundaries typically include data access, policy rules etc. Relevant metrics include policy adherence, harmful output rate, hallucination rate, unauthorized tool use attempts and citation accuracy.

Operational metrics

This measures whether the agent is practical to run. A workflow that completes correctly but costs too much, takes too long or depends on repeated retries may not be ready for production. Common metrics include latency, cost-per-task, token usage, tool-call volume, retry rate etc. However, operational metrics should be evaluated alongside quality metrics. A cheaper path that skips verification might reduce costs while increasing risk. Likewise, a slower path that performs one required policy check might be the right trade-off for a regulated workflow.

Consistency metrics

Agents are nondeterministic, and so one successful run alone can’t prove reliability. The same task should be evaluated across repeated runs. Common metrics include pass rate across repeated runs, variance in tool use, source consistency and trajectory stability. These measures help to determine whether agent performance is repeatable enough for production.

A good agent evaluation data set should capture the task, environment, available tools, allowed data, trajectory and scoring criteria.

Agent evaluation dataset components

For tasks with clear answers, the dataset should include ground truth. For tasks with multiple answers, the dataset needs a rubric. Golden data sets — curated sets of test cases with known or expected results — are especially useful because they provide a stable test suite for agent behavior. They let teams run the same set of tasks against different prompts, models, tools, agent orchestration logic or guardrails and compare the results over time. If every possible guardrail is added before the team understands where the agent fails, the evaluation may show a constrained system rather than a reliable one. A golden data set gives teams a way to identify weak surfaces, tune the agent and rerun the same tasks after each change.

How to evaluate AI agents

LLM-as-judge evaluation uses a model to score agent behavior against a rubric. The judge might assess whether the final answer is grounded, whether the agent followed instructions, whether the tool call was appropriate or whether the reasoning trace supports the result. This method typically uses a mix of code-based, model-based and human graders, with each grader evaluating some portion of the transcript or outcome. Model-based graders are useful when the evaluation involves language, judgment or multi-step interaction that cannot be captured by exact matching alone. The rubric is the control surface. A vague rubric produces vague scores. A useful rubric defines what the grader should inspect, what counts as success, what counts as partial success and which failures should be treated as severe. Human calibration is still required, especially for policy-sensitive or ambiguous workflows, because judge disagreement is itself a signal.

Human review remains the calibration layer for agent evaluation. Reviewers can inspect traces, compare model-based scores with human judgment and identify failure modes that automated graders miss. Human-in-the-loop evaluation is most useful for high-risk workflows, ambiguous tasks and early-stage eval design. It’s usually too expensive to review every agent run manually, but it provides the reference point for validating automated judges and refining rubrics.