Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Introduces novel interpretability techniques for clinical AI systems that researchers can apply to improve model transparency and reliability.
AI Summary
Researchers propose CAST, a concept-guided framework using sparse autoencoders to make clinical language models more auditable and robust against deployment shifts.
Excerpt
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-ass
