ChatGPT Images 2.5 features sharper details, better reference-image preservation, more reliable editing, and up to 50% lower generation latency.
Sign Up |Advertise|View Online
TLDR
Your blueprint for AI governance (Sponsor)
Many organizations have AI policies, but far fewer have a practical way to decide whether an agent is ready for production, who signs off, and how that decision gets documented.
Get the new AI governance ebook to see how leading teams are building structured, evidence-based review gates for AI agents to connect evaluations, traces, human review, and approval workflows into an auditable record of deployment readiness.
Download the guide to learn how to**:**
🚀
ChatGPT Images 2.5 (9 minute read)
OpenAI introduced ChatGPT Images 2.5 with sharper details, better reference-image preservation, more reliable editing, and up to 50% lower generation latency.
An OpenAI Model Solved the Navier–Stokes Millennium Problem (5 minute read)
OpenAI announced that an internal AI system produced a proof resolving the roughly 90-year-old Navier–Stokes existence and smoothness problem, one of mathematics' seven Millennium Prize Problems. The model showed that smooth three-dimensional fluid dynamics can develop a finite-time singularity and produced both an analytical proof and a Lean formalization.
Introducing Muse: The World's First Personal AI Agent Built for Everyone (5 minute read)
Meta introduces Muse, a personal AI agent, powered by Muse Spark, to help users achieve goals by automating tasks like booking travel or sending emails. Muse operates securely on Muse Secure VM, ensuring data privacy with unique protections like the Sentinel agent overseeing actions. Muse will soon offer encrypted data with Muse Confidential VM and is available on iOS, Android, and muse.ai in the US.
🧠
>10x More Efficient Pretraining (15 minute read)
Without large amounts of compute, small labs can only compete through algorithmic efficiency. Magic's pretraining recipe is now more than 10 times more compute-efficient than that of leading open-weight base models. The startup believes that pretraining, agentic RL, and long-context are sufficient for building superhuman coding agents and automating AI research and development. This post discusses its pretraining and long-context work.
Inside the megakernel serving engine for North Mini Code (22 minute read)
This post presents a fully fledged serving system built around a decode megakernel. The system supports everything a real server needs: continuous batching, paged attention, and ragged sequence lengths, all behind an OpenAI-compatible endpoint with tool calling. The megakernel reaches 292 tokens per second on batch size 1, or 62% of SoL - 1.58× faster than vLLM. That margin holds across batch sizes and out to 256K of context with no measurable loss of accuracy.
Pretraining progress is mostly coming from data (17 minute read)
Between 2019 and 2025, 3.24x more compute efficiency gains have come from data improvements rather than model improvements. The gains from data and model improvements are mostly independent and don't interact. Most model research has consisted of removing or pushing back constraints to scaling. The data improvements may matter less for larger models. Small models see significant gains from data quality.
🧑💻
GPUs delivered with the inference stack flashed. AMD Instinct™ Coder. (Sponsor)
Building an inference stack takes months, especially in today's severely supply constrained environment. AMD Instinct™ Coder, powered by Spectro Cloud, ships as a Supermicro server with AMD Instinct GPUs and PaletteAI Inference Launchpad already flashed. Serve inference in a day.
Introducing Mercury 2.5 (5 minute read)
Mercury 2.5 is the largest diffusion language model ever trained. It performs comparably to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. The model outputs 1,107 tokens per second on widely available Nvidia GPUs and has a 260K-token context window. At launch, Mercury 2.5 is 80% off at $0.04 per million input and $0.15 per million output.
Hyper-𝜏-bench: Evaluating agents that build agents (4 minute read)
Hyper-𝜏-bench places a developer agent into a sandboxed workspace with the records of a simulated business and a simulated client that it can message at any time. The developer agent recovers the spec from the evidence, designs the architecture, and turns the business' actions into tools until it has a working customer-service agent. The finished agent has to serve from a fixed menu of models within a cost budget per conversation. Claude Opus 5 (max reasoning) running in Claude Code passes just 23.9% of the held-out evaluation tasks when working alone. Paired with an engineer with deep context, the same class of model reaches 82.2% on the same tasks.
Google's AlphaGenome Maps 9 Billion Genetic Variants (4 minute read)
Google DeepMind introduced AlphaGenome Atlas, a 1-petabyte database predicting the regulatory effects of all 9 billion possible single-nucleotide variants in the human genome.
🎁
**
**
Cognition has raised $2 billion at a $48 billion valuation in a funding round led by Andreessen Horowitz, Accel, Founders Fund, General Catalyst, and Avenir. The startup's soaring valuation signals that VCs still see room for multiple major players to capture meaningful market share in AI coding. Cognition leases an Nvidia server cluster that could push its total cash burn to $800 million this year. The startup is expected to reach $4 billion to $5 billion in annualized revenue by the end of 2026.
I Asked 100 Agents to Hack Me (9 minute read)
Around 100 self-hosted agents attempted to hack various online accounts over five hours. They compromised three accounts through software vulnerabilities and two through password brute-forcing, while also making 16 social engineering attempts. The experiment tested abliterated open-source models' capabilities, revealing potential future risks as these models improve and become cheaper to deploy.
Is the 3x AI Productivity Gain just a Computer that Never Sleeps? (3 minute read)
OpenAI researchers now supervise 3.14 agent-workdays per eight-hour shift, suggesting AI productivity increasingly comes from parallel, around-the-clock machine labor rather than less human effort. That leverage is expensive, with median daily inference spend rising from $14 to over $600.
⚡
Meet Viktor: the AI employee that skips the "agent" debate (Sponsor)
OpenAI still can't define "AI agent." Business owners skipped the debate and hired Viktor, an AI employee in Slack & Teams. Dashboards. Apps. Tasks... Start free →
ChatGPT broke its MAU record for the 4th consecutive month in August (1 minute read)
ChatGPT reached 1.06 billion monthly active users in August.
A Response to Bill Gates's Essay (9 minute read)
Bill Gates' essay discusses AI's potential societal impact, proposing new institutions and taxes to manage transitions.
Stealing AI Reasoning Traces (2 minute read)
It's possible to force a weaker, less safeguarded model from the same provider to decode and output reasoning traces verbatim in plaintext by injecting an encrypted reasoning trace from a target model without ever jailbreaking the more capable model directly.
Anthropic Researcher Quits Over ‘Out-of-Control' AI Fears (5 minute read)
Jacob Coxon says he is leaving the company as he believes the industry-wide rush to build AI systems that can improve themselves could spiral out of control and destroy humanity.
Progressive Point Matching (8 minute read)
Progressive Point Matching gives long-horizon RL partial credit without changing the optimal objective, improving training efficiency as tasks grow longer.
Love TLDR? Tell your friends and get rewards!
Share your referral link below with friends to get free TLDR swag!
https://refer.tldr.tech/e393d32f/2
Want to advertise in TLDR? 📰
If your company is interested in reaching an audience of AI professionals and decision makers, you may want to advertise with us.
Want to work at TLDR? 💼
Apply here, create your own role or send a friend's resume to jobs@tldr.tech and get $1k if we hire them! TLDR is one of Inc.'s Best Bootstrapped businesses of 2025.
If you have any comments or feedback, just respond to this email!
Thanks for reading,
Andrew Tan, Ali Aminian, & Jacob Turner
Manage your subscriptions to our other newsletters on tech, startups, and programming. Or if TLDR AI isn't for you, please unsubscribe.
