Papers/2604.00594
πŸ“– Read?View on arXiv

Agent Psychometrics: Task-Level Performance Prediction in Agentic Coding Benchmarks

coding-agentsbenchmarksevaluationswe-bench
2604.00594
Builder Relevance
90%
Jul 20

Abstract

This work extends item-response theory with agent and task features to predict task-level performance in agentic coding evaluations.

Reality Card

Core Claim

Benchmark and agent teams can estimate task difficulty and success probability using LLM, scaffold, and rich task features instead of relying only on one aggregate solve rate.

Method / Result

The method evaluates across SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0, and a software-optimization benchmark, while noting the cost of full agent-evaluation runs.

Limitations

Predictions are conditioned on the available evaluation data, feature extraction, benchmarks, and agent/scaffold coverage; they are not a replacement for production evaluations.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers