AI Agent Evaluation Framework: Metrics, Test Sets and CI Gate
Artificial Intelligence

AI Agent Evaluation Framework: Metrics, Test Sets and CI Gate

Author : Nishant Bijani
Make us preferred on Google
Read time:17 minsUpdated:October 5, 2026

TL;DR

  • Grade the outcome and the steps, not the reply. Agent evaluation checks what changed in your systems and which tools were called, not only what the agent said.
  • Track eight metrics. Task success, tool selection accuracy, tool parameter accuracy, consistency, groundedness, latency, cost per task and handoff rate.
  • Consistency is the metric most teams miss. An agent that succeeds 75% of the time on a single try succeeds on all three of three tries only about 42% of the time.
  • Build the test set from real work. Start with 20 to 50 tasks taken from real requests and real failures, then turn every production bug into a regression test.
  • Gate every release in CI. Prompt edits, model swaps and tool changes all run the suite, and the deploy is blocked if any metric drops below its threshold.
  • Someone edits one line in a support agent's prompt to make it sound friendlier. Every demo conversation still looks fine, so the change ships. Two days later the agent quietly stopped asking for an order number before it issues refunds.

    Nothing crashed, no error fired, and no single answer looked wrong. That is what makes agents hard to test. A chatbot gives you one reply to grade. An agent takes a dozen steps, calls tools and changes records, and it can fail in any of them while the final message reads perfectly.

    This guide lays out a practical AI agent evaluation framework in three parts: the metrics to measure, the test sets to build, and the CI gates that stop a bad change from reaching production.

    What is an AI agent evaluation framework?

    An AI agent evaluation framework is a repeatable system for testing whether an agent completes its tasks correctly, consistently and within budget. It defines what to measure, which test cases to run, how each result is graded, and which scores must pass before a change can be deployed.

    Every framework, whatever the tooling, is built from the same six parts. Anthropic's engineering guide to agent evals uses this vocabulary, and it is worth adopting because it makes conversations between engineers precise:

  • Task. One test scenario: the input, the starting state of your systems, and what success looks like.
  • Trial. One run of the agent on one task. Agents are not deterministic, so each task runs several times.
  • Grader. The check that decides whether a trial passed. A task can have several.
  • Transcript. The full record of what the agent did: every message, tool call and result.
  • Outcome. The final state of the environment after the agent finished, such as whether the refund record exists.
  • Eval runner. The test runner that loads tasks, resets the environment, runs trials, calls graders and adds up the scores.
  • How agent evaluation differs from LLM evaluation

    Most LLM evaluation grades a single response: did the model answer the question accurately and well? Agent evaluation has to grade a process, and that changes almost everything about how the test works.

    LLM evaluationAI agent evaluation
    What gets gradedOne responseThe final outcome plus the steps taken
    Typical failureA wrong or weak answerA wrong tool, a bad parameter, a skipped step
    Where you lookThe output textSystem state, tool calls and the transcript
    Runs per testOften oneSeveral, because results vary between runs
    Main riskInaccurate contentWrong actions in real systems

    The practical consequence: an agent can produce a polished, accurate-sounding final message while having called the wrong tool, updated the wrong record or skipped a required check. If your evaluation only reads the last message, it will pass that agent.

    AI agent evaluation metrics: what to measure

    Most LLM evaluation metrics score a single response for accuracy or relevance. Agentic AI evaluation metrics have to cover more: correctness, reliability, efficiency and safety across many steps. These eight cover most production agents:

    MetricWhat it measuresHow to grade it
    Task successDid the agent complete the user's goal?Check the end state of your systems
    Tool selection accuracyDid it call the right tool for each step?Compare tool calls against the expected ones
    Tool parameter accuracyDid it pass the right values to each tool?Check arguments against the task's ground truth
    Consistency (pass^k)Does it succeed on every repeated run?Run each task several times and count full passes
    GroundednessAre its claims supported by retrieved data?Check statements against sources, often with an LLM judge
    LatencyHow long does a task take end to end?Measure p50 and p95 per task
    Cost per taskWhat does one completed task cost?Sum model tokens and tool fees per run
    Handoff rateHow often does it escalate to a person?Count escalations, and check they were correct

    Task success and goal accuracy

    Task success is the headline number: did the agent finish what the user asked? Grade it on the outcome, not the wording. For a booking agent, the test is whether the reservation exists with the right date and guests, not whether the agent said “your booking is confirmed”. AWS's published framework for its own agents separates this into goal success, whether every user goal was met, and goal accuracy, whether the result matches the ground truth.

    Tool selection and tool parameter accuracy

    These two metrics catch most agent bugs before they reach a customer. Tool selection asks whether the agent chose the right function: fetching quarterly revenue instead of general market data. Parameter accuracy asks whether it filled that function correctly: the right region, the right quarter, in the right format.

    Both are easy to grade with code, because the expected call is known. That makes them ideal CI gates. AWS's production blueprint blocks a deploy when tool selection accuracy falls below 95%.

    Consistency: pass@k vs pass^k

    Agents give different results on different runs of the same task, so a single pass tells you little. Two metrics handle this, and they answer different questions:

  • pass@k is the chance that at least one of k attempts succeeds. It suits work where a retry is acceptable, such as a research or coding assistant.
  • pass^k is the chance that all k attempts succeed. It suits customer-facing agents, where every customer expects the same correct behavior.
  • The gap between them is large. An agent with 75% success per attempt scores about 98% on pass@3 but only about 42% on pass^3. Sierra introduced pass^k with τ-bench (tau bench), its tool-agent-user benchmark, in June 2024, after finding that agents with respectable single-run scores degraded sharply across repeated runs. If your agent talks to customers, pass^k is the reliability number to report.

    Hallucination and groundedness

    A grounded agent only states what its sources or tools returned. Check groundedness by comparing each claim in the final answer against the retrieved documents or tool outputs. This is one of the few places an LLM judge earns its cost, because matching a claim to a source is hard to do with exact string rules.

    Latency and cost per task

    Measure both per completed task, not per model call. An agent that makes eleven tool calls to finish a job can be slow and expensive even when each call is fast and cheap. Track p95 latency, the slowest 5% of tasks, because that is what a frustrated user experiences. A model swap often trades one against the other: AWS's own example shows a smaller model cutting median latency from 3.2 to 1.8 seconds while tool selection accuracy fell from 92% to 87%. Your eval suite is what makes that trade visible before customers find it.

    Handoff rate

    Handoff rate is how often the agent escalates to a person. Too high, and the agent saves little. Too low, and it is probably guessing when it should stop. Grade the quality of each handoff as well as the rate: did it escalate the cases your rules say it must, and did it pass enough context that the person did not start again?

    How to grade agent outputs: code, LLM as a judge and humans

    Use three kinds of graders, in this order of preference:

  • Code-based graders first. Exact checks on system state, tool calls and parameters. They are fast, cheap and reproducible, and they should cover everything that has a single right answer.
  • LLM as a judge for open-ended quality. A second model scores things code cannot, such as tone, helpfulness or whether a summary is faithful. Write a rubric with clear criteria, and prefer pass or fail verdicts over 1 to 10 scores, which drift.
  • Human graders to calibrate. People label a sample of transcripts so you can check whether the LLM judge agrees with them. Human review is too slow to run on every commit, but without it you cannot trust the judge.
  • LLM judges carry known biases. The MT-Bench research that established the method in 2023 found three: a preference for whichever answer appears first, a preference for longer answers, and a preference for output written by the judge's own model family. Shuffle answer order, cap response length in the rubric, and use a different model as judge than the one being tested.

    Two rules from Anthropic's guide apply across every grader. Grade the outcome rather than the exact path, because an agent may find a valid route you did not anticipate. And read transcripts regularly: you will not know whether your graders are fair until you look at what they passed and failed.

    How to build a test set for AI agent evaluation

    A good test set is small, real and growing. Build it in three layers.

    Golden tasks from real requests

    Start with 20 to 50 tasks, which is the starting range Anthropic recommends. Take them from real user requests and real failures, not from scenarios your team imagines. For each, record the input, the starting state, and the expected outcome and tool calls. This golden dataset is the core of everything that follows, so spend time making each success definition unambiguous.

    Edge cases

    Add the cases that break agents: missing information, contradictory instructions, a tool that returns an error, a user who changes their mind halfway, a request the agent should refuse. Each edge case should have a defined correct behavior, which is often a clean handoff rather than an answer.

    Regression cases from production failures

    Every time the agent fails in production, turn the failure into a test case. This is how the suite grows to match the real world, and it guarantees the same bug cannot ship twice. AWS describes the same practice: let real user behavior grow the evaluation suite.

    Capability suites vs regression suites

    Keep two separate suites. A capability suite holds hard tasks the agent cannot yet do reliably; its pass rate starts low and shows whether changes improve the agent. A regression suite holds tasks the agent already handles; its pass rate should sit near 100%, and any drop signals a break. When a capability task passes consistently, move it into the regression suite.

    One warning: a suite at 100% catches regressions but tells you nothing about improvement. If every score is perfect, add harder tasks.

    Most agent failures we see come from evaluation gaps, not model limits: no golden set, no consistency check, no gate before deploy. Codiste builds the test set, graders and CI gates alongside the agent, so the team that takes it over can change it safely.

    CI gates: what must pass before an agent deploys

    A CI gate turns your evaluation into a release rule: the suite runs automatically, and the deploy is blocked if a metric falls below its threshold. Set gates for each layer of the framework:

  • Regression suite pass rate. Near 100%. Any previously passing task that now fails blocks the release.
  • Tool selection and parameter accuracy. An absolute floor, such as the 95% AWS uses, because wrong tool calls cause wrong actions.
  • Task success against the baseline. No drop of more than a few points from the current production version.
  • Consistency on critical tasks. A pass^k threshold for the tasks customers depend on most.
  • Safety and refusal tests. 100%. An agent that starts answering what it should refuse never ships.
  • Latency and cost budgets. Soft gates that warn rather than block, unless the increase is large.
  • Set thresholds relative to your current production numbers rather than ideal targets. A gate that blocks every release gets switched off.

    What should trigger an evaluation run

    Run the suite on every change that can alter agent behavior, not only on code changes: a prompt edit, a model or model version swap, a new or changed tool, a change to retrieval data or the knowledge base, and an update to any connected API. Prompt and model changes are the ones teams most often ship without testing, and they are the most likely to cause quiet regressions like the refund example at the top of this guide.

    Agent observability and production monitoring

    CI gates protect releases. Agent observability and AI agent monitoring protect everything after release. Trace production runs, sample a share of live interactions, and score them with the same graders you use in CI. Some teams also run a new version in shadow mode first: it processes real traffic in parallel, without affecting users, so you can compare its scores against production before switching over. When monitoring finds a new failure, it goes into the regression suite, which closes the loop.

    AI agent evaluation tools, mapped to the framework

    Tools do not build the framework for you. They run it. Choose by which part of the framework you need covered:

    Framework stageToolWhat it covers
    Whole framework, built into a custom agentCodisteGolden test sets, graders, CI gates and monitoring designed with the agent and handed over to your team
    Test runner in CIDeepEvalOpen-source, pytest-style evaluation that fits existing CI pipelines
    Eval platform with release gatesBraintrustDatasets, scorers and a GitHub Action that blocks merges below set thresholds
    Tracing and observabilityArize Phoenix, LangfuseOpen-source tracing and evaluation, self-hostable
    LangChain and LangGraph agentsLangSmithTracing and evaluation native to the LangChain stack
    AWS-hosted agentsAmazon Bedrock AgentCore EvaluationsManaged evaluators for tool selection, correctness and goal success on live traffic
    Public reference pointτ-benchA benchmark for comparing models, not a replacement for your own test set

    Two notes on choosing. Pick tools that keep your test cases in your own repository or export them cleanly, because the test set is the asset you are building. And treat public benchmarks as orientation only: a model's τ-bench score says little about how it handles your tools, your data and your rules. For context on where agents are already in production, see our guide to AI agents for business.

    Common mistakes in AI agent evaluation

    Most problems in AI agent testing come from a handful of habits:

  • Testing on demo data. Clean, invented scenarios pass easily and miss the messy inputs real users send.
  • Grading only the final message. The answer reads well while the wrong record was updated.
  • Running each task once. A single pass hides inconsistency that customers will hit.
  • Trusting an uncalibrated LLM judge. If nobody has checked the judge against human labels, its scores are a guess.
  • Grading the exact path. Penalizing a valid alternative route makes the agent look worse than it is and trains the team to ignore failures.
  • Never reading transcripts. Aggregate scores move for reasons that only show up in the individual runs.
  • Conclusion

    An agent that works in a demo and an agent that works on Monday morning are different products, and the difference is evaluation. Agents fail quietly. They call the wrong tool, skip a step or drift after a prompt edit, and the final message still reads as if nothing went wrong.

    The framework that catches this is not complicated. Measure the outcome and the steps, not only the reply. Run each task more than once and report pass^k for anything customers touch, because single-run scores flatter every agent. Grade with code wherever a right answer exists, use an LLM judge only where it does not, and check that judge against people.

    The test set matters more than any tool. Twenty to fifty real tasks, a set of edge cases, and a regression case for every production failure will tell you more than a public benchmark ever will. That set grows with the agent, and after a year it becomes the most valuable thing your team owns about it.

    Then put it in the release path. Every prompt edit, model swap and tool change runs the suite, and a drop below the threshold blocks the deploy. Teams that do this ship agent changes with confidence. Teams that skip it find their regressions in customer complaints, usually days after the change that caused them, and spend the next week working out which change it was.

    If your agent is in production without a test set or a release gate, Codiste can build both around the agent you already have, starting from the failures your users have already found.

    FAQs

    What is the best AI agent evaluation framework? +
    The best framework is the one built around your own tasks: a golden test set from real requests, code-based graders for tool calls and outcomes, an LLM judge for open-ended quality, and CI gates on every change. Tools such as DeepEval, Braintrust and LangSmith run that framework but do not replace the test set.
    What metrics should I use to evaluate an AI agent? +
    Start with task success, tool selection accuracy, tool parameter accuracy and consistency measured as pass^k. Add groundedness for agents that cite sources, latency and cost per task for efficiency, and handoff rate to check the agent escalates when it should.
    How many test cases do I need to evaluate an AI agent? +
    Twenty to fifty well-defined tasks from real requests is enough to start, and is the range Anthropic recommends. Grow the set by turning every production failure into a regression test, and add harder tasks once your scores approach 100%.
    What is pass^k in AI agent evaluation? +
    Pass^k is the probability that an agent succeeds on all k repeated attempts at the same task. It was introduced by Sierra's τ-bench in 2024 to measure reliability. It matters most for customer-facing agents, because an agent at 75% success per attempt passes all three of three attempts only about 42% of the time.
    Is LLM as a judge reliable for evaluating agents? +
    It is useful for open-ended qualities that code cannot check, but it has known biases toward answers shown first, longer answers, and output from its own model family. Use clear pass or fail rubrics, shuffle answer order, use a different model as judge, and calibrate it against human labels.
    How do you evaluate an AI agent in production? +
    Trace live runs, sample a share of real interactions and score them with the same graders used in CI. Test new versions in shadow mode on real traffic before switching, and add every new failure to the regression suite so it cannot happen again.
    Nishant Bijani
    CTO & Co-Founder | Codiste
    Nishant is a dynamic individual, passionate about engineering and a keen observer of the latest technology trends. With an innovative mindset and a commitment to staying up-to-date with advancements, he tackles complex challenges and shares valuable insights, making a positive impact in the ever-evolving world of advanced technology.

    Relevant blog posts

    The Ultimate Development Guide for an AI Marketing Agent in 2026
    Artificial Intelligence
    February 20, 2025

    The Ultimate Development Guide for an AI Marketing Agent in 2026

    What is an AI Voice Agent for Customer Care? A Beginner’s Guide
    Artificial Intelligence
    April 09, 2025

    What is an AI Voice Agent for Customer Care? A Beginner’s Guide

    The Science Behind AI-Driven Customer Journey Mapping
    Artificial Intelligence
    March 01, 2025

    The Science Behind AI-Driven Customer Journey Mapping

    Talk to Experts About Your Product Idea

    Every great partnership begins with a conversation. Whether you're exploring possibilities or ready to scale, our team of specialists will help you navigate the journey.

    Contact Us

    Phone