AI DAILY / DEV
Monthly Rollup
May 2026

Anthropic Refunds Claude Code Users After 'HERMES.md' Commit Strings Silently Drained Quotas

  • Substring match in Claude Code's harness-detection logic routed any session whose recent git history contained 'HERMES.md' to extra-usage billing instead of plan quota.
  • One Max-20x user lost $200 in overage credits while plan usage was still at 13%; charges occurred silently with no UI warning.
  • GitHub issue anthropics/claude-code#53262 traces the bug to the wider crackdown on third-party agent harnesses (Hermes, Codex CLI, Goose).
  • Hit HN front page at 828 points; refunds plus $200 in credits issued ~9 hours later. Discussion: HN item 47948012.
tools github.com

Sam Altman Opens ChatGPT Subs to OpenClaw — Anthropic Holds the Block

  • Altman tweet at 2:33am on May 2: 'sign in to openclaw with your chatgpt account now and use your subscription there.'
  • ChatGPT Plus ($23/mo) now backs GPT-5.4 inside OpenClaw — directly opposite Anthropic's April 4 ban on running Claude Pro/Max plans through third-party harnesses.
  • OpenClaw sits at ~346k GitHub stars and 3.2M users; making ChatGPT the auth/billing layer locks that distribution to OpenAI.
  • Strategic split: OpenAI is buying agent volume, Anthropic is protecting margins on autonomous-agent token spend.
industry thenextweb.com

Anthropic Opens Claude Security Public Beta — Opus 4.7 Auditing Whole Codebases

  • Public beta launched April 30 to all Claude Enterprise customers via the Claude.ai sidebar and claude.ai/security.
  • Opus 4.7 traces data flows across files and modules instead of pattern-matching for known signatures, then writes verified patches.
  • Adds scheduled scans, Slack/Jira webhooks, CSV/Markdown exports, and triage tracking on top of the February research preview.
  • Launch partners include CrowdStrike, Microsoft Security, Palo Alto Networks, SentinelOne, TrendAI, and Wiz; Team and Max access 'coming soon.'
tools siliconangle.com

Uber Burned Its Entire 2026 AI Budget on Claude Code in Four Months

  • CTO Praveen Neppalli Naga told the company it had spent the full annual AI budget by April after rolling Claude Code to 5,000 engineers in December.
  • Claude Code adoption jumped 32% → 84% of engineers; per-developer API spend ran $500–$2,000/month.
  • 70% of committed code now originates from AI; ~11% of live backend updates ship with no human in the loop.
  • HN front page (item 47976415) — debate centers on whether $100k/year per engineer in token spend pencils out vs hiring.
industry news.ycombinator.com

Anthropic and OpenAI Both Launch Wall Street Joint Ventures on the Same Day

  • Anthropic JV with Blackstone, Hellman & Friedman, Goldman Sachs at $1.5B — $300M each from Anthropic, Blackstone, H&F; $150M from Goldman; co-investors include Apollo, GIC, Sequoia, General Atlantic, Leonard Green.
  • OpenAI's 'The Deployment Company' closes $4B from 19 investors at a $10B valuation, anchored by TPG, Brookfield, Advent, Bain — with a 17.5% guaranteed return over five years.
  • Both vehicles embed forward-deployed engineers inside customer companies (Palantir-style) instead of selling API credits — direct shot at McKinsey/Accenture-style AI consulting.
  • Initial pipeline: each consortium's PE-owned portfolio in healthcare, manufacturing, financial services, retail, and real estate; both labs are reportedly targeting fall 2026 IPOs.
industry techcrunch.com

OpenAI Ships GPT-5.5 Instant as ChatGPT's New Default Model

  • Replaces GPT-5.3 Instant across ChatGPT and ships in the API as `chat-latest`; old model stays available for paid users for three months.
  • 52.5% fewer hallucinated claims on high-stakes medical, legal, and financial prompts; 37.3% fewer incorrect claims on user-flagged conversations.
  • Responses are 30.2% shorter in words and 29.2% shorter in lines on internal evals.
  • AIME 2025 jumps to 81.2 from 65.4; MMMU-Pro to 76 from 69.2.
  • Personalization from past chats, files, and connected Gmail rolling out to Plus and Pro on web.
models openai.com

Anthropic Lands SpaceX Colossus 1 and Doubles Claude Code Rate Limits

  • Anthropic takes all the compute capacity at SpaceX's Memphis Colossus 1 site: 300+ MW and 220,000+ Nvidia H100/H200/GB200 GPUs landing within the month.
  • Claude Code five-hour limits doubled for Pro, Max, Team, and Enterprise plans; peak-hour throttling on Pro and Max removed.
  • Claude Opus API tier-1 input rate limit jumps 1500%; output tokens-per-minute up 900%.
  • Notable plot twist: Musk has publicly called Anthropic 'misanthropic and evil' while suing OpenAI; Anthropic also expressed interest in developing multi-gigawatt orbital compute with SpaceX.
  • On HN front page; Bloomberg, CNBC, and Engadget all leading with the Memphis tie-up.
industry anthropic.com

OpenAI Drops Three Realtime Voice Models With GPT-5-Class Reasoning

  • GPT-Realtime-2: speech-in/speech-out model with GPT-5-class reasoning, parallel tool calls, spoken preambles like 'let me check that,' and a 128K context (up from 32K).
  • GPT-Realtime-Translate: live speech-to-speech translation across 70+ input languages and 13 output languages, designed to keep pace with the speaker.
  • GPT-Realtime-Whisper: streaming speech-to-text that emits text as the user is still talking — aimed at captions, meeting notes, and live agents.
  • GPT-Realtime-2 scores 15.2pp higher than GPT-Realtime-1.5 on Big Bench Audio; Zillow reports a 26pp jump in call success rates in early testing.
  • Pricing: $32/$64 per 1M audio input/output tokens for GPT-Realtime-2; $0.034/min for Translate; $0.017/min for Whisper.
models openai.com

DeepMind's AI Co-Mathematician Sets a New High on FrontierMath Tier 4

  • Multi-agent workbench built on Gemini 3.1 — a project coordinator fans research questions out to sub-agents that write code, search literature, and attempt proofs.
  • Solved 23 of 48 problems (48%) on Epoch AI's FrontierMath Tier 4, more than double the 19% scored by the underlying Gemini 3.1 Pro and a new SOTA.
  • Oxford's Marc Lackenby cracked Problem 21.10 from the Kourovka Notebook after a reviewer agent flagged a flaw in the AI's first proof attempt and surfaced a strategy he could finish.
  • Stateful, asynchronous workspace that tracks failed hypotheses and outputs LaTeX with margin annotations — DeepMind frames it as 'Claude Code for mathematicians.'
research arxiv.org

Google Catches Hackers Using AI to Build a Working Zero-Day

  • Google Threat Intelligence Group says it disrupted the first AI-built zero-day used in a planned mass-exploitation campaign.
  • The exploit, written in Python with help from cybercrime model OpenClaw, bypasses 2FA on a widely deployed open-source web admin tool.
  • Attackers tipped Google off by leaving a hallucinated CVSS score inside the script; Google notified the vendor before any victims were hit.
  • GTIG analyst John Hultquist: "It's here. The era of AI-driven vulnerability and exploitation is already here."
  • Front-page coverage at CNBC, Bloomberg, NBC News, and Bleeping Computer; on HN with active debate about Anthropic's Mythos and Project Glasswing.
industry cloud.google.com