← Back to feed

BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks

L5 · ResearcherResearcharXiv· 8/27/2026

Provides novel methodology for creating reproducible agent evaluation benchmarks from real-world telemetry data.

AI Summary

Researchers introduce BTS-AgentBench, a deterministic pipeline for converting industrial telemetry logs into reproducible agent benchmarks with 532 test cases.

Excerpt

Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks. We present a telemetry-to-episode construction method instantiated as BTS-AgentBench. The pipeline normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes. The 532-row release adds clarific

Read Original
0 upvotes · 0 downvotes · 1 min read

Related Articles

L5 · ResearcherResearchHacker News
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Stanford researchers released Terminal-Bench-Science 0.1, a benchmark with 70 expert-curated scientific workflows where Claude Opus achieved only a 30% resolution rate.

L4 · DeveloperResearcharXiv
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Presents Persona-Execution Separation, an architecture pattern for enterprise LLM agents that keeps execution auditable while allowing persona instructions to evolve freely.

L5 · ResearcherResearchLessWrong AI
The Dynamics of Intelligence Explosions

Toby Ord explores the mathematical dynamics of intelligence explosions where AI assists AI R&D, showing singular growth is harder than economic models suggest.

L5 · ResearcherResearcharXiv
LLMs Can Design Near-Optimal OR Algorithms

GPT-5.6 can design near-optimal algorithms for operations research problems like inventory control and queueing networks, matching or beating specialized methods with minimal human input.

L5 · ResearcherResearchLessWrong AI
Do AI Models Want to Be Monitored? Measuring Monitorability Disposition in Large Reasoning Models

Researchers propose measuring 'monitorability disposition' in large reasoning models to proactively understand model behavior rather than reactively filtering outputs.