Today in AI
The most important AI developments from the past day. Through September 11, 2026
Target leakage is the great accuracy illusion of medical AI, and a new cardiovascular screening audit delivers a much-needed reality check. Researchers found that supposed breakthroughs in tabular foundation models were largely artifacts of data leakage, not architectural genius. The kicker: simpler glass-box models matched that inflated performance while running 104 times faster. For an industry pouring billions into ever-larger models for healthcare, this is a bracing reminder that dataset hygiene and rigorous validation matter more than parameter counts. The paper should force a reckoning across computational medicine, where splashy accuracy claims too often dissolve under proper scrutiny.
On the capabilities front, two strands hint at where reasoning research is heading. "Thinking with Looped Flows" proposes letting recurrent models spin longer at inference time through iterative denoising—a computed-time-over-parameters trade that feels increasingly central to the field's trajectory. Meanwhile, the provocatively titled "The Last AI Built by Humans" sketches genuine recursive self-improvement, moving past the current hype cycle toward systems that could actually bootstrap their own capabilities in scientific discovery and software engineering. It's speculative, but the framing is correct: today's fine-tuning loops are not RSI, and recognizing the gap is the first step toward closing it.
Less headline-grabbing but quietly important: a new framework for quantifying dataset shifts using entropic optimal transport gives practitioners tools to know when their models are operating outside their training distribution. And in speech recognition, standard word error rate metrics get exposed as fundamentally broken for English-Yoruba code-switching, masking failures that disproportionately hurt African language users. These aren't just technical corrections; they're about who AI serves and who it leaves behind.
Bottom line: The most important AI advances today aren't bigger models—they're sharper epistemology, from catching leaky benchmarks to measuring what actually matters.
Stories referenced
Research paper finds target leakage, not model class, explains high accuracy in cardiovascular screening models, with glass-box models performing comparably to foundation models while being 104x faster.
arXivA new evaluation of 11 ASR and audio language models on English-Yoruba code-switched speech shows that standard word error rate metrics hide serious problems in recognizing African languages and code-switch boundaries.
arXivResearchers propose 'looped flows'—training recurrent models with local denoising objectives to solve harder reasoning problems by spending more computation time during inference.
arXivA new reinforcement learning algorithm using multi-step transition lookahead achieves near-optimal performance despite NP-hard planning complexity.
arXivResearchers introduced a model-aware schedule for diffusion/flow-matching models using fiberwise optimal transport, achieving a 38.6% FID improvement on CIFAR-10 with fewer function evaluations.
arXivResearchers propose a novel framework using entropic optimal transport to quantify covariate and concept shifts with estimable error bounds and practical algorithms.
arXivThis arXiv paper proposes a roadmap for genuine recursive self-improvement (RSI) in AI systems, analyzing challenges across scientific discovery, embodied intelligence, and software engineering.
arXivResearchers propose 3D Point Splatting, a differentiable radar renderer that achieves 1.7-5.2x better performance than optical-NVS baselines for novel view synthesis with mmWave radar.
arXiv