Papers/2609.00048
🧪 Test?View on arXiv

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

Not provided in the abstract

benchmarkingGUI modelsmulti-step environmentscontextual consistency
2609.00048
Builder Relevance
70%
2h ago

Abstract

This paper introduces GUI-CC, a benchmark for evaluating the contextual consistency of GUI world models as multi-step environments for agents.

Reality Card

Core Claim

GUI-CC demonstrates that current GUI world models often fail to maintain task-relevant context in multi-step interactions despite producing plausible single-step outputs.

Method / Result

Constructed 500 offline trajectory tasks and 200 emulator-verified online tasks across 30 mobile apps.

Limitations

The benchmark highlights that plausible single-step generation does not ensure reliable environment simulation, raising concerns about the reproducibility of multi-step interactions.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers