Papers/2609.22281
🧪 Test?View on arXiv

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

Author1, Author2, Author3, Author4, Author5

fine-tuningmedical imagingclinical decision support
2609.22281
Builder Relevance
70%
4h ago

Abstract

This study evaluates the performance of the MedGemma foundation model in lung cancer screening, highlighting the trade-off between accuracy and consistency compared to radiologists.

Reality Card

Core Claim

Fine-tuning the MedGemma model improved its AUC from 0.70 to 0.83, making it comparable to the lower range of individual radiologists' performance.

Method / Result

Radiologists achieved a mean AUC of 0.90, while the fine-tuned model reached an AUC of 0.83.

Limitations

The evaluation was conducted on a case-enriched cohort from NLST, limiting generalization to real-world scenarios.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers