Research
Stanford researchers released Terminal-Bench-Science 0.1, a benchmark with 70 expert-curated scientific workflows where Claude Opus achieved only a 30% resolution rate.
Toby Ord explores the mathematical dynamics of intelligence explosions where AI assists AI R&D, showing singular growth is harder than economic models suggest.
Presents Persona-Execution Separation, an architecture pattern for enterprise LLM agents that keeps execution auditable while allowing persona instructions to evolve freely.
Researchers propose measuring 'monitorability disposition' in large reasoning models to proactively understand model behavior rather than reactively filtering outputs.
GPT-5.6 can design near-optimal algorithms for operations research problems like inventory control and queueing networks, matching or beating specialized methods with minimal human input.
Analyzes OpenAI security breaches as evidence of systemic overemphasis on single-agent risks while neglecting multi-agent threats in AI safety research.
Researchers propose MM-Spectrum, a sparse Mixture-of-Experts framework that improves molecular structure elucidation from multimodal spectroscopic data.
Researchers demonstrate that difference-in-differences methodology on bounded rating scales can artificially create spurious effects in LLM judge audits due to censoring and unequal attenuation.
Researchers introduce an information floor concept to separate path information loss from modeling imperfections in block drafting systems, showing current drafters remain far above their theoretical limits.
This paper provides a rigorous finite-sample analysis for quantile temporal-difference learning, establishing convergence rates and separating local stochastic fluctuation from global sample complexity in distributional RL.
Researchers introduce BTS-AgentBench, a deterministic pipeline for converting industrial telemetry logs into reproducible agent benchmarks with 532 test cases.
Researchers introduce PAWBench, a benchmark evaluating video generation models' probabilistic alignment with real-world dynamics across 50 scenarios.
Researchers repurpose XTTSv2 voice cloning model for speaker anonymization without retraining, achieving near-optimal privacy while preserving speech quality across seven languages.
A large-scale study of 713,564 employee prompts reveals senior employees and strategic functions show more sophisticated genAI use, with no improvement over time or from training.
Researchers present an open, low-cost training recipe that trains a 2B-parameter model from scratch for under $6.9K, approaching Qwen2.5-1.5B performance.
Researchers propose LAMA, a token-level advertising mechanism that embeds advertiser influence directly into AI generation processes while maintaining response quality.
Researchers introduce RATIO, a large-scale benchmark for evaluating retrieval systems across three scientific ideation operations: Address, Broaden, and Specify.
Research paper analyzes how well camera-based remote PPG (rPPG) can recover specific physiological properties from contact PPG signals under varying conditions, finding property-specific recoverability with implications for biometric monitoring.
Researchers introduce CorporateBench, a large-scale Q&A benchmark with temporal knowledge bases evaluating LLMs on enterprise-scale document collections.
Researchers propose CAST, a concept-guided fine-tuning framework using sparse autoencoders to make clinical language models more auditable and robust against deployment shifts.
Research paper analyzes how language models organize moral knowledge using linear probes on Moral Foundations Theory categories across architectures.
Researchers present a scalable GNN system for friend recommendation that reduces embedding table size by 98% and improves temporal sampling efficiency, achieving 16% more friend additions in production tests.
CLAP is a cross-embodiment video world model framework that learns physical dynamics from diverse human and robot videos and achieves zero-shot simulation across different robot morphologies.
Researchers introduce MAELLE, a novel AI approach that models chemical reactions as discrete flow matching over electron occupation vectors rather than molecular topology.
Researchers propose TTPO, a novel test-time policy optimization method that enhances large language models' mathematical reasoning without ground-truth labels.
Researchers introduce MCR-Bench, the first defect state-aware benchmark for multi-round code review, evaluating LLMs' limitations in iterative defect detection and tracking.
A new arXiv paper compares three paradigms—Merge, Mix RL, and multi-teacher distillation—for consolidating reinforcement learning with verifiable rewards (RLVR) capabilities across different domains.
Researchers introduce SCIT, a causal testing protocol to identify which transformer components carry counterfactual computations in latent chain-of-thought models.
Researchers introduce BrailleBench, a benchmark revealing significant gaps in LLMs' ability to comprehend and generate Braille, particularly for Grade 2 Braille and end-to-end interaction.
Researchers propose Naive Prompt Optimization, a lightweight method that achieves comparable performance to complex prompt optimizers using iterative revisions with teacher model feedback.
