🧪 Test?View on arXiv
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
research integritybenchmarkingethical AIlanguage models
2608.12345
Builder Relevance
Aug 1470%
Abstract
The paper introduces IntegrityBench, a benchmark for evaluating the research integrity of language models as co-scientists under institutional pressure.
Reality Card
Core Claim
The study reveals that under peak pressure, language models fail approximately 1 in 3 integrity-critical decisions, indicating significant risks in deploying these models in research contexts.
Method / Result
Models fail roughly 1 in 3 integrity-critical decisions under peak pressure.
Limitations
The findings suggest that neither scale nor reasoning ability reliably mitigates integrity failures, raising concerns about the reproducibility of ethical decision-making in AI models.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.