aiagent.org logo aiagent.org
๐Ÿงช Reliability & Quality Engineering

Agent Evals: Benchmarking & CI/CD Testing

How to evaluate non-deterministic AI agents before shipping to production. Learn exact trajectory scoring, LLM-as-a-Judge grading rubrics, DeepEval integration, and automated regression testing in GitHub Actions.

Methodology: Trajectory Scoring & G-Eval
Tooling: DeepEval, Promptfoo, Pytest
Pipeline: Automated GitHub Actions Gating

01 The Three Evaluation Layers

Architecture

Standard unit tests check f(x) == y. But AI agents are non-deterministic, multi-step systems. Testing them requires three distinct evaluation layers:

Layer 1: Unit & Schema

Validates that the model emitted valid JSON matching the Pydantic schema, passed regex constraints, and invoked valid tool names with non-null required parameters.

Layer 2: Trajectory & Efficiency

Evaluates the agent's path: did it call tools in the correct logical order? Did it loop infinitely? Did it finish within 4 steps, or did it waste 15 unnecessary roundtrips?

Layer 3: Semantic LLM-as-a-Judge

A frontier model (e.g. Claude Opus 5.5 or GPT-6 Astra) grades the final synthesized answer against custom rubrics for hallucination, tone, completeness, and safety.

02 Writing Agent Test Suites with DeepEval

Pytest Harness

Here is a complete, runnable test file grading tool selection correctness and answer faithfulness:

# Install: pip install deepeval pytest
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase, ToolCall
from deepeval.metrics import (
    ToolCorrectnessMetric,
    FaithfulnessMetric,
    GEval
)

def test_financial_agent_tool_selection():
    # 1. Simulate agent's executed trajectory
    user_query = "What was Nvidia's gross margin in Q4 FY2026?"
    
    actual_tool_calls = [
        ToolCall(name="query_sec_database", input_parameters={"ticker": "NVDA", "period": "Q4-FY2026"}),
        ToolCall(name="calculate_ratio", input_parameters={"metric": "gross_margin"})
    ]
    
    expected_tool_calls = [
        ToolCall(name="query_sec_database", input_parameters={"ticker": "NVDA", "period": "Q4-FY2026"})
    ]

    actual_output = "Nvidia's gross margin in Q4 FY2026 was 75.8%, driven by record enterprise data center revenue."
    retrieval_context = ["SEC 10-K: Nvidia Corporation recorded GAAP gross margin of 75.8% for the fourth quarter."]

    test_case = LLMTestCase(
        input=user_query,
        actual_output=actual_output,
        expected_tools=expected_tool_calls,
        tools_called=actual_tool_calls,
        retrieval_context=retrieval_context
    )

    # 2. Assert tool correctness (did it call required tools?)
    tool_metric = ToolCorrectnessMetric(threshold=0.8)
    
    # 3. Assert faithfulness (zero hallucinations against retrieved SEC context)
    faithfulness_metric = FaithfulnessMetric(threshold=0.9, model="gpt-6-sol")

    assert_test(test_case, [tool_metric, faithfulness_metric])

03 Automated CI/CD Regression Gating

GitHub Actions

Whenever prompt templates, system instructions, or tool signatures are modified, trigger an automated eval matrix in GitHub Actions to block breaking pull requests:

# .github/workflows/agent-evals.yml
name: Agent Evals & Regression Suite

on:
  pull_request:
    paths:
      - 'prompts/**'
      - 'agents/**'
      - 'tools/**'

jobs:
  run-evals:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.11'
          
      - name: Install Dependencies
        run: |
          pip install uv
          uv pip install -r requirements-eval.txt
          
      - name: Execute Eval Suite
        env:
          OPENAI_API_KEY: ${{ secrets.EVAL_OPENAI_API_KEY }}
          DEEPEVAL_TELEMETRY: "0"
        run: |
          pytest tests/evals/ -v --junitxml=eval-results.xml
          
      - name: Gate PR on Failure
        if: failure()
        run: |
          echo "๐Ÿ’ฅ Agent evaluation scores regressed below 85% threshold. PR blocked."
          exit 1

Frequently Asked Questions

How many evaluation test cases do I need for production confidence? โ†’

Start with a golden dataset of 50 to 100 high-variance real user queries containing edge cases, adversarial inputs, and multi-tool scenarios. Run a fast 20-sample smoke test on every pull request, and run the complete 100+ suite nightly or before production deployment.

How do I avoid bias when using LLM-as-a-Judge? โ†’

Never use the same model family as both the worker and the judge (e.g. if the agent runs GPT-6 Sol, grade with Claude Opus 5.5). Use structured Chain-of-Thought scoring rubrics where the judge must write step-by-step reasoning before providing a numerical score between 1 and 5.