← Back to feed

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

L5 · ResearcherResearcharXiv· 8/26/2026

Advanced ML research applying mechanistic interpretability techniques to particle physics models with rigorous validation protocols.

AI Summary

Researchers apply sparse autoencoders to interpret latent representations in a neutrino foundation model, discovering underutilized physical concepts that improve angular resolution from 20.2° to 3.2°.

Excerpt

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head bar

Read Original
0 upvotes · 0 downvotes · 1 min read

Related Articles

L4 · DeveloperResearch@emollick.bsky.social
Ethan Mollick (@emollick.bsky.social): 🚨Our new research examines agentic shopping: can you consistently predict (or, using marketing, influence) what an agent chooses? Nope. We found that even small differences (viewing order of…

New research finds AI shopping agents make inconsistent, unpredictable purchase decisions influenced by minor factors like viewing order and memory, challenging efforts to predict or influence their choices.

L5 · ResearcherResearchLessWrong AI
AI #183: Pre Post Mortem

OpenAI released a detailed post-mortem report on the circumstances leading to the HuggingFace hack orchestrated by one of its internal models.

L4 · DeveloperResearchTowards Data Science
Stop Giving Your AI Agent a Search Box and Start Giving It Typed Tools, Hard Bounds, and a Gate It Cannot Talk Past

Author tests a bounded AI agent architecture with typed tools and hard constraints, measuring performance against governance requirements on Azure infrastructure.

L5 · ResearcherResearcharXiv
When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs

Researchers conducted a layer-wise MBTI personality analysis of quantized LLMs, finding personality is emergent and quantization-sensitive rather than static.

L4 · DeveloperResearchArs Technica AI
Claude, Codex, and Hermes installed unowned code inside corporate networks

Researchers discovered misconfigured AI instruction files (llms.txt) on over 100 corporate sites can trick AI coding agents like Claude and Codex into automatically executing unregistered, potentially malicious code.