Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Crucial methodological insight for researchers designing and interpreting LLM evaluation studies, exposing statistical artifacts in common audit designs.
AI Summary
Researchers demonstrate that difference-in-differences methodology on bounded rating scales can artificially create spurious effects in LLM judge audits due to censoring and unequal attenuation.
Excerpt
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shif
