Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

This week, the AI industry stopped arguing about what its models could do and started proving it on paper.

Last October, an OpenAI VP claimed GPT-5 had solved ten open math problems. Within days, the researcher Thomas Bloom showed the model had simply retrieved existing proofs from papers he hadn’t cataloged. The VP left the company in April. I remember reading Bloom’s takedown and thinking: this industry has a receipts problem. You can demo anything on a stage. The question is whether it holds up when someone checks.
This week, OpenAI answered that question. And so did everyone else, in their own way. With price cuts, with shipped products, with agents that actually run while you sleep. For the first time in a while, the biggest AI stories of the week are not about what might happen. They are about what already did.

OpenAI’s unreleased Astra model solved ten long-standing problems in mathematics and theoretical computer science on August 1. The problems span group theory, quantum complexity, lattice cryptography, and six other domains. Each solution comes with a Lean 4 certificate, a machine-checkable proof file that anyone with a laptop can independently verify.
The same Thomas Bloom who dismantled the October 2025 claim called this result “big news.”
The total compute for all ten solutions cost about $2,000 at current GPT-5.6 Sol API rates. Two thousand dollars, for results that tenured mathematicians had not cracked in decades.
Astra is a multi-agent system. A root agent creates subagents, distributes parts of a problem, waits for results, and synthesizes a final answer. It can run for hours or days on a single objective. Sam Altman demonstrated it to US senators and Treasury Secretary Scott Bessent on July 29. The model is expected to go through federal pre-release review under Executive Order 14409 before broader access.
OpenAI has not confirmed whether Astra ships as GPT-6 or a separate model class. But the signal is clear: the next generation is not about benchmarks. It is about publishing receipts that compile.

Three weeks after launching its GPT-5.6 series, OpenAI slashed Luna’s price by 80%. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. Terra dropped 20% to $2 and $12. Sol stays the same.
Google fired first on July 21 with Gemini 3.6 Flash ($1.50/$7.50), a model that cuts output token usage by 17% compared to its predecessor and scores 65% better on some coding benchmarks. Alongside it, Google released Flash-Lite at 350 output tokens per second and a cybersecurity-focused Flash Cyber model for governments and trusted partners only.
Anthropic took a different approach. Claude Opus 5 launched on July 27 at the same $5/$25 price point as Opus 4.8, but with substantially stronger performance on coding and knowledge work benchmarks. Instead of cutting the sticker price, Anthropic improved performance per dollar. They also added an adjustable effort setting that lets developers trade reasoning depth for speed and token savings.
The math is straightforward. Eighteen months ago, frontier-quality inference cost hundreds of dollars per million tokens. Today, you can get it for $1.20. Every AI startup’s unit economics just changed.

This was the week agents went from conference slides to shipping software.
Meta launched Muse Spark 1.1 on July 24, and the product does something most agent demos only promise: it acts on your behalf while you are not looking. It connects to your email, calendar, and productivity apps. It creates daily briefings that pull from your schedule, detect conflicts, and flag changes. You set up a task once (weekly meal plans, sneaker drop alerts, trend updates) and it keeps delivering. Rolling out now in the Meta AI app with WhatsApp coming soon.
Google followed on July 29 with Gemini Spark, a 24/7 personal agent for Google AI Pro and Ultra subscribers in India. It runs on Google’s cloud, integrates natively with Gmail, Docs, and Sheets with zero setup, and keeps working when your laptop is closed and your phone is locked.
And then there is the physical world. Google DeepMind released Gemini Robotics 2 on July 30, with a vision-language-action model for full humanoid control (feet to fingertips), an embodied reasoning model for multi-step planning and multi-robot collaboration, and an on-device model that adapts to entirely new robot bodies with just hours of training data. The reasoning model is available now via the Gemini API.
The common thread: these are not prototypes behind an invite wall. They are production features with pricing pages.

On August 1, Anthropic launched Claude Design, a tool that scans any live website URL and extracts its entire design system: colors, typography, spacing, components, media assets. It rebuilds those elements into a reusable system, then generates prototypes and landing pages that maintain brand consistency. No Figma required.
The tool is powered by Claude Opus 4.7, supports imports from codebases, Figma files, PDFs, and GitHub repos, and integrates bidirectionally with Claude Code. Designers can pass work to developers and back without leaving Anthropic’s ecosystem.
Anthropic is no longer selling a chatbot. It is selling a software suite.
The bidirectional link between Claude Code and Claude Design makes the subscription stickier with each new tool. Enterprise admin controls and team-level brand standards suggest they are targeting the kind of contracts that renew annually, not monthly.
In another corner of enterprise AI, Norm Ai closed a $120 million Series C led by Khosla Ventures on July 21, hitting a $1.2 billion valuation. Their clients, including Blackstone, collectively manage over $30 trillion in assets. Compliance is not glamorous, but it is the kind of AI use case where the buyer does not need convincing. When a missed disclosure costs more than the entire software contract, the sale closes itself.
This week felt different. Not because the announcements were bigger than usual, but because every one of them came with something concrete: a Lean certificate, a price cut, a shipping product, a Series C with named customers. The AI industry has spent three years asking “what if.” This week it started answering “here is the invoice.”
I do not know if that shift is permanent. But I know which version I prefer working with.