BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks
Provides novel methodology for creating reproducible agent evaluation benchmarks from real-world telemetry data.
AI Summary
Researchers introduce BTS-AgentBench, a deterministic pipeline for converting industrial telemetry logs into reproducible agent benchmarks with 532 test cases.
Excerpt
Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks. We present a telemetry-to-episode construction method instantiated as BTS-AgentBench. The pipeline normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes. The 532-row release adds clarific
