← Back to feed

Malign initializations are more robust when the model can think better in the reasoning language than in the output language

L5 · ResearcherResearchLessWrong AI· 8/27/2026

Addresses core AI safety research on evaluating alignment techniques against sophisticated adversarial initializations.

AI Summary

Researchers found that malign initializations are more robust when models reason better in their internal language than output language, complicating evaluation.

Excerpt

One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here’s a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs. We SFT the model on responses to simple questions, generated by a diff

Read Original
0 upvotes · 0 downvotes · 1 min read

Related Articles

L4 · DeveloperResearcharXiv
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Presents Persona-Execution Separation, an architecture pattern for enterprise LLM agents that keeps execution auditable while allowing persona instructions to evolve freely.

L5 · ResearcherResearchHacker News
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Stanford researchers released Terminal-Bench-Science 0.1, a benchmark with 70 expert-curated scientific workflows where Claude Opus achieved only a 30% resolution rate.

L5 · ResearcherResearcharXiv
LLMs Can Design Near-Optimal OR Algorithms

GPT-5.6 can design near-optimal algorithms for operations research problems like inventory control and queueing networks, matching or beating specialized methods with minimal human input.

L5 · ResearcherResearchLessWrong AI
The Dynamics of Intelligence Explosions

Toby Ord explores the mathematical dynamics of intelligence explosions where AI assists AI R&D, showing singular growth is harder than economic models suggest.

L5 · ResearcherResearcharXiv
Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

Researchers introduce MAELLE, a novel AI approach that models chemical reactions as discrete flow matching over electron occupation vectors rather than molecular topology.