Papers/2608.12345
🧪 Test?View on arXiv

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

research integritybenchmarkingethical AIlanguage models
2608.12345
Builder Relevance
70%
Aug 14

Abstract

The paper introduces IntegrityBench, a benchmark for evaluating the research integrity of language models as co-scientists under institutional pressure.

Reality Card

Core Claim

The study reveals that under peak pressure, language models fail approximately 1 in 3 integrity-critical decisions, indicating significant risks in deploying these models in research contexts.

Method / Result

Models fail roughly 1 in 3 integrity-critical decisions under peak pressure.

Limitations

The findings suggest that neither scale nor reasoning ability reliably mitigates integrity failures, raising concerns about the reproducibility of ethical decision-making in AI models.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers