🧪 Test?View on arXiv
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
Not specified in the provided content
evaluationfault detectionLLMagent trajectories
2609.00038
Builder Relevance
2h ago70%
Abstract
The paper critiques outcome-only evaluations of LLM agents, highlighting their inability to detect silent faults in agent trajectories.
Reality Card
Core Claim
The outcome-only judge detects 84% of loud faults but only 45% of silent faults, indicating a significant blind spot in current evaluation methods.
Method / Result
A step-rubric judge achieves 77% silent recall with zero false alarms at three times the cost of the outcome-only judge.
Limitations
The study's findings depend on a specific deterministic environment and fault injector, which may limit generalizability.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.