Agent Evals: Benchmarking & CI/CD Testing
How to evaluate non-deterministic AI agents before shipping to production. Learn exact trajectory scoring, LLM-as-a-Judge grading rubrics, DeepEval integration, and automated regression testing in GitHub Actions.
01 The Three Evaluation Layers
Architecture
Standard unit tests check f(x) == y. But AI agents are non-deterministic, multi-step systems. Testing them requires three distinct evaluation layers:
Validates that the model emitted valid JSON matching the Pydantic schema, passed regex constraints, and invoked valid tool names with non-null required parameters.
Evaluates the agent's path: did it call tools in the correct logical order? Did it loop infinitely? Did it finish within 4 steps, or did it waste 15 unnecessary roundtrips?
A frontier model (e.g. Claude Opus 5.5 or GPT-6 Astra) grades the final synthesized answer against custom rubrics for hallucination, tone, completeness, and safety.
02 Writing Agent Test Suites with DeepEval
Pytest HarnessHere is a complete, runnable test file grading tool selection correctness and answer faithfulness:
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase, ToolCall
from deepeval.metrics import (
ToolCorrectnessMetric,
FaithfulnessMetric,
GEval
)
def test_financial_agent_tool_selection():
# 1. Simulate agent's executed trajectory
user_query = "What was Nvidia's gross margin in Q4 FY2026?"
actual_tool_calls = [
ToolCall(name="query_sec_database", input_parameters={"ticker": "NVDA", "period": "Q4-FY2026"}),
ToolCall(name="calculate_ratio", input_parameters={"metric": "gross_margin"})
]
expected_tool_calls = [
ToolCall(name="query_sec_database", input_parameters={"ticker": "NVDA", "period": "Q4-FY2026"})
]
actual_output = "Nvidia's gross margin in Q4 FY2026 was 75.8%, driven by record enterprise data center revenue."
retrieval_context = ["SEC 10-K: Nvidia Corporation recorded GAAP gross margin of 75.8% for the fourth quarter."]
test_case = LLMTestCase(
input=user_query,
actual_output=actual_output,
expected_tools=expected_tool_calls,
tools_called=actual_tool_calls,
retrieval_context=retrieval_context
)
# 2. Assert tool correctness (did it call required tools?)
tool_metric = ToolCorrectnessMetric(threshold=0.8)
# 3. Assert faithfulness (zero hallucinations against retrieved SEC context)
faithfulness_metric = FaithfulnessMetric(threshold=0.9, model="gpt-6-sol")
assert_test(test_case, [tool_metric, faithfulness_metric])
03 Automated CI/CD Regression Gating
GitHub ActionsWhenever prompt templates, system instructions, or tool signatures are modified, trigger an automated eval matrix in GitHub Actions to block breaking pull requests:
name: Agent Evals & Regression Suite
on:
pull_request:
paths:
- 'prompts/**'
- 'agents/**'
- 'tools/**'
jobs:
run-evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install Dependencies
run: |
pip install uv
uv pip install -r requirements-eval.txt
- name: Execute Eval Suite
env:
OPENAI_API_KEY: ${{ secrets.EVAL_OPENAI_API_KEY }}
DEEPEVAL_TELEMETRY: "0"
run: |
pytest tests/evals/ -v --junitxml=eval-results.xml
- name: Gate PR on Failure
if: failure()
run: |
echo "๐ฅ Agent evaluation scores regressed below 85% threshold. PR blocked."
exit 1
Frequently Asked Questions
How many evaluation test cases do I need for production confidence? โ
Start with a golden dataset of 50 to 100 high-variance real user queries containing edge cases, adversarial inputs, and multi-tool scenarios. Run a fast 20-sample smoke test on every pull request, and run the complete 100+ suite nightly or before production deployment.
How do I avoid bias when using LLM-as-a-Judge? โ
Never use the same model family as both the worker and the judge (e.g. if the agent runs GPT-6 Sol, grade with Claude Opus 5.5). Use structured Chain-of-Thought scoring rubrics where the judge must write step-by-step reasoning before providing a numerical score between 1 and 5.