Agent Psychometrics: Task-Level Performance Prediction in Agentic Coding Benchmarks
Abstract
This work extends item-response theory with agent and task features to predict task-level performance in agentic coding evaluations.
Reality Card
Benchmark and agent teams can estimate task difficulty and success probability using LLM, scaffold, and rich task features instead of relying only on one aggregate solve rate.
The method evaluates across SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0, and a software-optimization benchmark, while noting the cost of full agent-evaluation runs.
Predictions are conditioned on the available evaluation data, feature extraction, benchmarks, and agent/scaffold coverage; they are not a replacement for production evaluations.
Paper to code
Verified implementation resources so builders can test the paperβs claims instead of stopping at the abstract.