TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
For researchers building or evaluating image generation models, this provides a novel diagnostic methodology beyond holistic scoring.
AI Summary
Researchers propose TRACE-Bench, a new evaluation framework that decomposes multi-reference image generation into four atomic operators to better diagnose model bottlenecks.
Excerpt
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor ($f$),
