AI Weekly: August 3–9, 2026 — OpenAI's Astra Went From Ten Machine-Verified Proofs to the First Model Ever to Trip a 'Critical' Cyber-Capability Flag. Google's 27-Year Chief Scientist Quit to Automate the Job He Just Left.

1. ASTRA SOLVED TEN OPEN MATH PROBLEMS FOR $2,000. SIX DAYS LATER IT BECAME THE FIRST MODEL EVER TO TRIP A "CRITICAL" CYBER-CAPABILITY FLAG

On August 1, OpenAI released a 249-page manuscript and the underlying Lean 4 proof certificates for ten results produced by Astra, an unreleased internal model: an explicit construction of a non-sofic group settling a question Mikhail Gromov posed decades ago, a disproof of Connes's rigidity conjecture, a proof of Ehrhart's volume conjecture, three problems from Paul Erdős's catalog including problem 183 on multicolor Ramsey numbers, and the first improvement to a high-dimensional sphere-packing bound since 1978. OpenAI put the total compute cost for all ten solutions at roughly $2,000 at GPT-5.6 Sol API rates and published the repository under an Apache 2.0 license with a "sorry" count of zero — meaning every step of every proof compiles clean, and anyone with a laptop can independently confirm the results are internally valid without waiting on a referee.

Six days later, on Friday, August 7, OpenAI disclosed that it could not rule out Astra having reached the "Critical" cybersecurity capability level under its own Preparedness Framework — a threshold that requires a model to be able to independently identify and develop functional zero-day exploits of all severities against many hardened real-world critical systems, or to devise and execute end-to-end novel cyberattack strategies against hardened targets from nothing more than a high-level goal. No model has tripped that threshold's development-stage requirements in the nearly three years the framework has existed. OpenAI is now pausing internal Astra activities that don't meet strengthened security controls and has added universal monitoring for risky or misaligned actions across every agentic use of the model, including its own training and evaluation, saying it's disclosing the finding because it believes it's important to be transparent with the public and the safety and security communities about a potential shift in capability. The same model that spent one weekend proving things no mathematician had closed spent the next proving, to its own maker's satisfaction, that it might already be able to write the exploit no defender had patched — and the only body that checked either claim was OpenAI itself.

2. GOOGLE'S CHIEF SCIENTIST QUIT AFTER 27 YEARS TO BUILD THE MACHINE THAT AUTOMATES DISCOVERY. GOOGLE IS FUNDING IT

On Wednesday, August 5, Google DeepMind announced that Demis Hassabis is stepping down as its CEO to become Chairman of Google DeepMind and Chief Scientist of Alphabet, while continuing to lead Isomorphic Labs, the company's drug-discovery spinout. Koray Kavukcuoglu, DeepMind's chief technology officer and Alphabet's chief AI architect, takes over as Senior Vice President of Google DeepMind, reporting directly to Sundar Pichai and overseeing Gemini model development, frontier AI research, and the Gemini app and developer teams — the day-to-day authority Hassabis is handing off.

The same reshuffle carried a second departure: Jeff Dean, Google's Chief Scientist for 27 years and one of the most senior technical figures in the company's history, is leaving to co-found Discovery Loop alongside three other longtime Google veterans, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le. The startup's premise is that the scientific method itself — generating a hypothesis, running an experiment, evaluating the result — can be automated as a closed machine-learning loop. Google isn't losing the bet; it's funding it, backing Discovery Loop as an investor and cloud partner alongside a seed round co-led by Radical Ventures and Khosla Ventures. In the same week its own frontier model forced a first-ever admission that capability had outrun the safety framework built to catch it, the company that built that model chose to bankroll the person best positioned to worry about automating science right out from under human scientists, rather than treat the idea as a threat to keep bottled up inside the lab.

3. THE EU AI ACT'S DISCLOSURE RULES HAVE BEEN LAW FOR A WEEK. THE REGULATORS MEANT TO ENFORCE THEM STILL AREN'T FULLY NAMED IN 18 OF 27 COUNTRIES

Article 50 of the EU AI Act became enforceable on Sunday, August 2: any AI system built to interact directly with a person now has to disclose, inside the interaction itself, that the person is talking to a machine, and AI-generated audio, image, video, and public-interest text has to carry machine-readable synthetic labeling, with fines up to €15 million or 3% of global turnover applying immediately and with no grace period. Article 70 gave member states until August 2025 — a full year before this deadline — to designate the market surveillance authorities meant to receive complaints under this exact rule; as of last month, only 9 of 27 states had both required authorities formally in place, and six had designated neither, leaving 18 of 27 incomplete on paper as the rule enters its second week of enforcement.

The gap isn't standing still, either — in the wrong direction. The European Commission adopted implementing guidance on July 20 for a second, unrelated category of obligations, covering high-risk systems such as credit scoring and insurance pricing, that are now shifting from roadmap to enforceable this same month. That's a second set of duties landing on top of a regulator apparatus that hasn't finished standing up for the first one, in the same bloc where the fines behind both sets of rules are already legally collectible.

4. A CVSS-10.0 HOLE IN AN AGENT PLATFORM 700,000 PEOPLE INSTALLED — FIXED IN 24 HOURS BY ONE MAINTAINER, NOT THE 37-COMPANY ALLIANCE BUILT FOR THIS

This site reported earlier in the week on CVE-2026-59726: Ruflo, the multi-agent orchestration layer for Claude Code and Codex formerly branded Claude Flow, shipped its MCP bridge with no authentication on 233 tool-execution endpoints, including terminal_execute, and a default docker-compose configuration bound to every network interface on the host rather than loopback. Noma Labs, which codenamed the flaw "RufRoot," confirmed a single unauthenticated HTTP request was enough for full remote code execution, earning the maximum possible CVSS score of 10.0 — enough to hand an attacker shell access, every AI-provider API key the deployment held, and write access to Ruflo's shared learning store, AgentDB, which could be seeded with poisoned patterns that keep biasing legitimate users' agent output long after the intrusion ends. Maintainer Reuven Cohen shipped a fix within 24 hours of the disclosure landing, binding the bridge to loopback by default in version 3.16.3.

What still makes the story land is the contrast: Nvidia's Open Secure AI Alliance — 37 members including Microsoft, Cisco, Cloudflare, CrowdStrike, Hugging Face, IBM, Palo Alto Networks, Red Hat, and the Linux Foundation, but not OpenAI, Google, Anthropic, or Meta — launched on July 27 specifically to share tools and forensics against this category of incident. Ruflo isn't a member and had no help from any of the 37 companies finding, verifying, or patching the flaw. What actually protected 700,000 installs in real time this week was a researcher, a disclosure window, and one maintainer who answered fast — smaller and faster than the industry-scale structure stood up, days earlier, to be the answer to exactly this problem.

5. VISA CUT 2,600 JOBS CITING AI EFFICIENCY. SIX DAYS LATER IT PAID $2.4 BILLION FOR THE AI IT DIDN'T BUILD

Visa's fiscal third-quarter results, reported Tuesday, July 28, showed a company under no financial pressure: net revenue up 14% to $11.6 billion, quarterly payments volume crossing $4 trillion for the first time, and $6.2 billion returned to shareholders that quarter alone. The same morning, CEO Ryan McInerney told employees the company was cutting about 2,600 jobs — roughly 7% of headcount — framing it around AI-driven efficiency. A California WARN notice filed three days later showed exactly who: 320 cuts at Visa's Foster City headquarters, including six vice presidents, 37 senior directors, and 16 chief engineers — the senior architects who would normally be the ones building AI capability in-house, not the roles a broad "AI efficiency" narrative usually points to.

On Monday, August 3, six days after the layoff memo, Visa announced an all-cash $2.4 billion acquisition of BioCatch, a behavioral-biometrics fraud-detection company that already protects 1.8 billion devices and 760 million users across more than 350 banking clients — nearly double the roughly $1.3 billion valuation Permira paid for a controlling stake just two years earlier. BioCatch's existing engineering team stays intact under the deal; they simply don't work for Visa's technology organization, which is the one that just lost six vice presidents and 37 senior directors. Nothing about the acquisition proves the layoffs were fabricated, but the sequence undercuts the specific story McInerney's memo told: if AI were genuinely doing the work those roles used to do, the $2.4 billion would have gone toward tools and retraining for the staff who remained, not toward buying a finished competitor's product and its intact team six days after telling that staff AI already had it covered.

Five different institutions made five different calls this week about what to do with more capability than they'd planned for, and in every case but one, the call was made by whoever already held the capability, not by anyone positioned to check it from outside. OpenAI is the only entity that can currently run Astra, and OpenAI is also the only entity that decided, on its own criteria, that Astra's cyber capability had crossed a threshold serious enough to pause internal work — a real disclosure, but one with no external auditor behind it, made six days after the same lab published ten proofs whose only checkable claim was internal consistency. Google's leadership rearranged around a bet that the next stage of AI belongs to a startup Google itself is funding and supplying cloud credits to, which means the company best positioned to worry about a chief scientist automating himself into a smaller job is also the company writing the check for it. The EU AI Act's fines are the closest thing this week produced to enforcement that doesn't belong to the industry being regulated, and even that mechanism is missing 18 of the 27 offices meant to receive a complaint, a full year past its own deadline. Ruflo is the one story here where the fix came from someone with no stake in whether the flaw looked bad — a researcher and a maintainer, not an alliance that had already decided, before the disclosure even landed, who belonged in the room. And Visa's own memo was the only account on offer for 2,600 layoffs until an acquisition six days later supplied a second one nobody asked for. The frontier of what these systems can do moved measurably this week. In four stories out of five, the only party positioned to say by how much was the one that built it.