Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
Provides deep technical insights into LLM evaluation mechanisms for researchers developing or using automated NLG assessment tools.
AI Summary
Researchers mechanistically analyze how LLM evaluators judge summarization quality using causal tracing and attention-head knockout experiments on Llama-3-8B and Mistral-7B models.
Excerpt
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and explicit token-l
