Papers/2609.00051
🧪 Test?View on arXiv

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Not specified in the provided content

safetymechanistic interpretabilityadversarial robustnessLLM architecture
2609.00051
Builder Relevance
80%
2h ago

Abstract

This paper investigates the internal mechanisms of safety in Large Language Models (LLMs) and proposes a multi-stage safety circuit that enhances refusal behavior against harmful inputs.

Reality Card

Core Claim

Circuit-guided weight scaling improves safety rates under adversarial attacks by 26.5% with only a 1.7% drop in accuracy across standard benchmarks.

Method / Result

26.5% improvement in safety rates under attacks across six LLMs.

Limitations

The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers