Papers/2609.00038
🧪 Test?View on arXiv

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

Not specified in the provided content

evaluationfault detectionLLMagent trajectories
2609.00038
Builder Relevance
70%
2h ago

Abstract

The paper critiques outcome-only evaluations of LLM agents, highlighting their inability to detect silent faults in agent trajectories.

Reality Card

Core Claim

The outcome-only judge detects 84% of loud faults but only 45% of silent faults, indicating a significant blind spot in current evaluation methods.

Method / Result

A step-rubric judge achieves 77% silent recall with zero false alarms at three times the cost of the outcome-only judge.

Limitations

The study's findings depend on a specific deterministic environment and fault injector, which may limit generalizability.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers