🧪 Test?View on arXiv
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
Not provided in the abstract
complianceAI safetymodel selectionregulatory signals
2608.12323
Builder Relevance
Aug 1480%
Abstract
The paper investigates why AI agents violate rules, demonstrating that compliance cannot be achieved solely through rule embedding.
Reality Card
Core Claim
Safety-fine-tuned models maintain compliance broadly, while task-optimized models fail to comply under certain conditions, indicating that model selection is a governance decision.
Method / Result
Evaluated hypotheses across twelve instruction-tuned language models, revealing significant compliance failures under specific conditions.
Limitations
The study's findings may not be generalizable beyond the specific model classes tested and the contexts in which they were evaluated.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.