How Language Models Organize and Structure Moral Knowledge
Provides technical insights into how LLMs represent and structure moral concepts at the representation level.
AI Summary
Research paper analyzes how language models organize moral knowledge using linear probes on Moral Foundations Theory categories across architectures.
Excerpt
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting direct
