Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Provides an academic benchmark for rigorously evaluating AI agent capabilities in real scientific research workflows.
AI Summary
Stanford researchers released Terminal-Bench-Science 0.1, a with 70 expert-curated scientific workflows where Claude Opus achieved only a 30% resolution rate.
