Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
Advanced ML research applying mechanistic interpretability techniques to particle physics models with rigorous validation protocols.
AI Summary
Researchers apply sparse autoencoders to interpret latent representations in a neutrino foundation model, discovering underutilized physical concepts that improve angular resolution from 20.2° to 3.2°.
Excerpt
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head bar
