CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
Provides crucial evaluation metrics for LLM performance on corporate-scale data with temporal consistency.
AI Summary
Researchers introduce CorporateBench, a large-scale Q&A with temporal knowledge bases evaluating LLMs on enterprise-scale document collections.
Excerpt
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (informati
