Safety's Second Way
Critical examination of fundamental assumptions in AI safety methodology that impacts research priorities.
AI Summary
Analyzes OpenAI security breaches as evidence of systemic overemphasis on single-agent risks while neglecting multi-agent threats in AI safety research.
Excerpt
Epistemics: I've tried to strike a balance between getting it right and getting it out while the community is discussing how to update. I am using the Hack as an example of a broader problem. I look forward to counterarguments. The OpenAI hacks [1] demonstrate an overweighting on single-agent risks, both at OpenAI and within the broader LessWrong and Alignment Forum safety communities. OpenAI neglected to monitor known multi-agent risks even after observing them on their deployment, showing a
