Malign initializations are more robust when the model can think better in the reasoning language than in the output language
Addresses core AI safety research on evaluating alignment techniques against sophisticated adversarial initializations.
AI Summary
Researchers found that malign initializations are more robust when models reason better in their internal language than output language, complicating evaluation.
Excerpt
One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here’s a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs. We SFT the model on responses to simple questions, generated by a diff
