Models & Releases
Google launches Gemini 3.5 Transcribe, a speech-to-text model with 4.0% WER for streaming and 2.6% for non-streaming use cases, now available via API.
Google released Gemini Omni 1.1 Flash, a production-ready video generation model with new controls for scene extension, frame interpolation, 4K upscaling, and faster prototyping.
Google released Gemini Omni 1.1 Flash with 10-second scene extension capability, 4K upscaling, and improved video generation controls.
Qwen3.8-Flash-Next is a new 125B token multimodal Mixture-of-Experts model that serves as an early preview of the Qwen4 architecture, with only 6B active parameters for performance efficiency.
Z.ai revealed it created the mysterious Ox Alpha model, an open-weight reasoning model for coding and agentic work that tops benchmarks.
Perceptron AI launched Isaac 0.5, an open-weight vision model that helps robots perceive, reason, and act in industrial settings like warehouses and factory floors.
Google's WeatherNext Cyclones model, used in real-time by the U.S. National Hurricane Center, provided 5-day advance warning for a Category 5 hurricane and offers accuracy improvements equivalent to a decade of meteorological progress.
IBM releases Granite 4.2 reasoning LLMs in 3B, 8B, and 30B sizes with 512K context, agentic RL training, and Apache 2.0 license.
DeepSeek added vision capabilities to its V4 model, Google introduced the EnvHarness training framework, and Etched shipped its first AI inference hardware to Jane Street.
Z.ai confirms Ox Alpha as a new GLM-series model and plans to open-source its weights.
Alibaba's Qwen 3.8 Flash-Next offers low-cost inference, but enterprises must evaluate other performance metrics before adoption.
Sam Altman suggests throwing another party for OpenAI's next model release, referencing the fun of their GPT-5.5 celebration.
Sam Altman announces OpenAI has developed a custom AI chip described as fast, signaling potential hardware acceleration for model inference.
Google launches Gemini 3.5 Transcribe, a multimodal transcription model that filters filler words and formats speech across 85+ languages.
Thomson Reuters developed a new AI model using an open-weight foundation model as its base architecture.
Google DeepMind released Gemini 3.5 Transcribe, a speech-to-text model with improved accuracy for phone numbers, custom vocabulary, and noise handling.
A mysterious 'stealth model' called Ox Alpha has been released on OpenRouter, sparking speculation about whether it's from Chinese developers or an unreleased Microsoft model.
Hugging Face announces Qwen3.8-Flash-Next, a preview of Qwen4 architecture, scheduled for release on August 26, 2026.
OpenAI released GPT-5.6 in Kiro with improved price-performance for software development tasks like planning, building, and testing.
An open-weight model, GLM-5.3, reportedly outperformed proprietary models like GPT-5.5 and Claude Sonnet/Opus in a leaderboard of real-world tasks while costing 80% less.
LiquidAI releases DSpark draft models for LFM2.5 family, achieving up to 3.2x faster inference through speculative decoding while maintaining output quality.
Inherent's Faraday AI agent outperformed Claude Opus 4.8 and GPT-5.5 at scientific paper replication using a smaller 27B parameter Qwen model.
Nvidia released SONIC, a model that enables humanoid robots to learn motion in real-time from human demonstrations.
OpenAI reduced GPT-5.6 Sol pricing by 20% and introduced new features like reasoning_effort controls and realtime prompt caching.
DeepSeek released v4-flash-vision-exp, a multimodal model that accepts images alongside text with three integration methods: base64 encoding, external URLs, and Files API references.
LiquidAI releases Q4_0 quantized checkpoints for LFM2.5 models using quantization-aware distillation, recovering 97% of BF16 accuracy while maintaining 4-bit efficiency.
OpenAI cuts GPT-5.6 Sol API pricing by 50% to $2.50/$15 per 1M tokens, making it more accessible for developers building complex reasoning and coding applications.
AlayaWorld v1.1 introduces major architectural revisions for interactive long-horizon world modeling, including a new 3D point-cache renderer and redesigned conditioning pipeline.
Qwen 3.8 27B is a new open-source vision-capable LLM that defaults to an 'xhigh' reasoning effort setting, causing extensive overthinking even for simple tasks.
Analyzes DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning, and NeMo Switchyard in technical depth, focusing on their architectural innovations and benchmark performance.
