Technical reports
Agent evaluationEvaluation design№ 03

From answer quality to task completion

An evaluation protocol combining task outcomes, policy adherence, repeated runs and execution costs.

Central proposition

“Completed” is a conclusion that needs verification.

  1. 01Fixed task
  2. 02Controlled environment
  3. 03Repeated runs
  4. 04Outcome verification
Method structure

01Define success before execution

A tool-using model can produce a fluent explanation while updating the wrong data. Agent evaluation therefore starts with observable completion criteria. A data task may require a target schema, record scope and quality rules; a research task may require valid citations and coverage of specified questions.

A task set should fix the initial state, tool versions, source material and resource budget, and identify unacceptable actions. Scoring only the final answer can miss duplicate writes, unauthorized access and unfinished intermediate work.

02Observe outcomes, behavior and cost

τ-bench offers a useful perspective: agents operate with user interaction, tools and domain policies, while final database state is checked against a goal. It also examines reliability across repeated trials. We use the evaluation idea, not its third-party results as IDENIFE performance claims.

Enterprise evaluation can retain three layers: outcomes determine whether artifacts meet criteria; behavior checks tool permissions, constraints and evidence; cost records calls, latency and human intervention. Keeping them separate explains why an approach fits one task class but not another.

03Repeated execution reveals consistency

One successful run does not establish reliability. Repeat the same task in independently reset, reproducible environments and retain every outcome. Per-run success, success on every repeated attempt and success at least once answer different questions and should not be conflated.

Introduce disturbances such as tool timeouts, empty results, revision conflicts and insufficient permissions. Completing a happy path and responding appropriately to changed conditions through stopping, asking or recovering are different capabilities.

  • Outcome: are the task goals met and artifacts usable?
  • Consistency: does reliability persist across independent runs?
  • Recovery: is the response to a disturbance appropriate?

04Hold conditions constant when comparing topologies

Multiple agents are not inherently better than one. Comparisons of sequential work, parallel specialization and independent review should hold the model, tools and task set constant under comparable budgets. Collaboration can reduce some errors while adding communication loss and cost.

Ablations can remove the review node, external memory or structured artifact constraints to inspect changes in failure types. When reviewers disagree, preserve inspectable inputs and outputs rather than relying only on a model preference that cannot be audited.

05Make the evaluation record inspectable

An inspectable report describes task provenance, sampling, environment revisions, success criteria, run counts and failures. Human judgments need a rubric and a method for resolving disagreement. Cost and consistency tradeoffs should remain visible rather than disappearing into an aggregate score.

This is an evaluation design with no unverified performance figures. Future results should link to actual run records and scope before becoming an experimental report. That connection allows research conclusions and engineering capability to develop together.

References

  1. [1]
    τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

    Shunyu Yao · Noah Shinn · Pedram Razavi · Karthik Narasimhan

This note describes methods and architecture, not experimental performance or a released API specification. Referenced research and engineering materials are the work of their respective authors.

Continue exploringAgents & execution
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.