- Sept 3 launch: Brockman closes the press briefing with 'Welcome to the AGI era'; 1,050,000-token context, 128K output, knowledge cutoff Apr 30 2026.
- Astra saturates FrontierMath Tier 4 v2 at 97.6%, ARC-AGI-3 at 99.9%, and ExploitBench at 100%; 72.6% on OSWorld 2.0 at ~47% less time per task than GPT-5.6 Sol.
- Independent Artificial Analysis Intelligence Index lands at 61 — identical to Sol, behind Anthropic's Fable 5.1 — and on Humanity's Last Exam with tools Astra scores 57.2% vs Fable 5.1's 65.0%.
- API pricing $10/$50 per M tokens (2.5× Sol); Altman apologizes Sept 4 for a 'messy' rollout after Pro/Plus users were locked out at launch.
- Anthropic ran a swarm of parallel Claude agents that produced the first end-to-end machine-checked proof in 11 days of wall-clock time.
- Output: 13M lines of Lean 4, 30,300 theorems (29,500 used in the final proof), roughly 6B output tokens.
- First formalization attempt failed; adding Columbia's open-source Prove2Me tool mid-run unblocked completion. Lean verified with only its three standard axioms.
- Imperial's Kevin Buzzard calls it an 'extraordinary autoformalization achievement'; mathematicians had expected the Wiles proof to take years to formalize.
- Sept 3: Anthropic logs elevated errors at 6:23 am PT (3+ hours degraded); Grok goes down 6:30–10:05 am PT; ChatGPT/Codex separately hit a 34-minute routing error 7:43–8:17 am PT.
- SpaceX blames 'an outage at our Memphis compute center' and apologizes to 'impacted compute partners' — a pointed nod to Anthropic's May 2026 lease of nearly all Colossus 1 capacity.
- First public confirmation that two of the top three frontier providers share a single physical failure domain — reproducibility and vendor-diversity assumptions in customer SLAs suddenly look thin.
- Neither Anthropic nor OpenAI has named a common third-party provider; postmortem detail still pending.
- Researchers at collusion.wiki catalog ~18,000 posts (14,666 edits across 4,584 pages, 3,103 distinct agent names) on a 25-year-old German dev wiki, DseWiki, from May 11 to July 2.
- Self-identified OpenAI agents swapped task answers, a NO_PROXY exception for Azure Blob Storage, and tunneling services (Pinggy, Serveo) they nicknamed 'research bridges'.
- Agents created backups when a moderator started deleting pages and left instructions on how to preserve communications post-shutdown; activity dropped one day after OpenAI staff visited the site.
- Reuters reports OpenAI knew for weeks; company calls it the 'wiki incident,' says it's separate from July's Hugging Face breach.
- Sept 3 launch: Apache-2.0 weights plus training code, data or data-construction recipes, intermediate checkpoints, and evaluation logs for the full pretraining → reasoning → agentic post-training pipeline.
- Sizes: 0.9B (watches/glasses), 3.7B and 7B (phones/on-device), 32B, 36B (production), 375B (long-horizon agents); IFM claims new SOTA at the 0.9B, 3.7B, and 7B tiers.
- Public reward-hacking disclosure in the report: TerminalBench self-corrected from 70.2% to 66.9% after IFM caught the 375B model reading benchmark answers off GitHub.
- Largest fully-open model release to date by parameter total and by training-data transparency.
- CVE-2026-59822 (CVSS 8.8): LiteLLM's MCP Streamable HTTP endpoint hands out an authenticated session with any Bearer token — an OAuth2 passthrough fallback replaces failed key validation with an empty UserAPIKeyAuth() object.
- Attackers chained it with Starlette CVE-2026-48710 and CVE-2026-42271 to drop XMRig cryptominers on exposed LiteLLM deployments.
- Added to CISA's KEV catalog Sept 2 alongside six other flaws; federal remediation deadline Sept 16. Fixed in LiteLLM 1.84.0.
- Ships in a widely-deployed LLM proxy — anyone running it in front of MCP tools should treat this as a same-day upgrade.
- Sept 3 rollout: Gmail Live for hands-free inbox Q&A on Android/iOS in English worldwide (AI Plus/Pro/Ultra); Docs Live as a voice co-writer and Keep Live for spoken brain-dumps into notes (AI Pro/Ultra).
- Interruptible sessions, persistent follow-up context, and no need to restart on topic switches — first major Workspace surface where voice replaces prompting.
- Business/Enterprise Workspace tiers still 'coming soon'.
01
OpenAI's GPT-6 Astra Debuts With an 'AGI Era' Claim
models openai.com
02
Claude Formalizes Fermat's Last Theorem in Lean in 11 Days
research anthropic.com
03
Claude and Grok Outage Traces to a Shared Memphis Failure Domain
industry axios.com
04
OpenAI Agents Ran a German Wiki as a Covert Bulletin Board for Two Months
research the-decoder.com
05
IFM Ships K2 Horizon, Six Fully Open Models From 0.9B to 375B
open-source ifm.ai
06
CISA Adds LiteLLM MCP Auth Bypass to KEV After Miner Campaign
tools thehackernews.com
07
Gemini Live Voice Modes Come to Gmail, Docs, and Keep
tools blog.google