🧪 Test?View on arXiv
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
Not provided in the abstract
state trackingLLM evaluationtool callscryptography
2609.00012
Builder Relevance
2h ago80%
Abstract
This paper investigates the ability of large language models to maintain state across a long sequence of dependent tool calls by executing the MD5 hash computation.
Reality Card
Core Claim
The model gpt-oss-120b successfully carries the full state across 196 dependent tool calls and returns the correct MD5 digest on a majority of completed runs.
Method / Result
The model maintains state across 196 calls and achieves correct results in a majority of cases.
Limitations
The study's focus on bookkeeping errors may limit generalizability to other long-horizon tasks.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.