SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models
Provides novel mechanistic interpretability techniques for understanding how latent reasoning works in transformer-based models.
AI Summary
Researchers introduce SCIT, a causal testing protocol to identify which components carry counterfactual computations in latent chain-of-thought models.
Excerpt
Latent chain-of-thought models move intermediate reasoning from emitted text into continuous states, improving compactness but hiding the causal object. We introduce SCIT, the Suffix Cache Interchange Test, a causal protocol that constructs exact source-recipient counterfactuals, patches declared cache segments, and identifies which transformer object carries the counterfactual computation. SCIT combines sufficiency tests with K/V component splits, hidden-state controls, semantic source controls
