01
OpenAI's Astra Cracks Ten Decade-Open Math Proofs for $2,000 in Compute
- Ten previously unsolved problems in group theory, sphere packing, quantum complexity, lattice crypto, and extremal combinatorics.
- Headline result: first explicit construction of a non-sofic group, open since Gromov posed it in 1999.
- Every proof ships with a Lean 4 certificate on GitHub under Apache-2.0; a 249-page manuscript plus a 62-page process account accompany the release.
- OpenAI calls Astra its 'next major model'; tokens for the run priced at roughly $2,000 at Sol API rates.
research openai.com
02
DeepSeek-V4-Flash-0731 Exits Preview at Opus-Level Agentic Performance
- Same 284B-total / 13B-active MoE as the April launch, gains come from re-post-training on agent tasks.
- 82.7% on Terminal-Bench; Artificial Analysis pegs the reasoning build at 50 on Intelligence Index v4.1 (#2 of 162 models).
- API pricing unchanged at $0.14/$0.28 per million input/output tokens; open weights on Hugging Face.
- r/LocalLLaMA lit up within hours — matches or beats Claude Opus 4.6 on agent benchmarks at a fraction of the cost.
models huggingface.co
03
Alibaba Ships Qwen3.8-Max at 2.4T Parameters, Open Weights Due Next Week
- 2.4T-parameter MoE with 95B active per token, 1M-token context, native text and vision input.
- First time Alibaba has committed to open-sourcing a Max-class Qwen model — weights promised on Hugging Face and ModelScope next week.
- QwenCloud API launches at $2/$6 per million input/output tokens; internal benchmarks position it 'second only' to Claude Fable 5 and above Kimi K3.
- HN preview thread hit 951 points and 692 comments; skeptics flag that no third party (Artificial Analysis, LMArena) has independently scored it yet.
models marktechpost.com
04
'ChainDrop' Worm Poisons 1,300+ npm Packages and Plants Claude Code Hooks
- Aug 4 09:35 UTC: attackers hijacked the maintainer of keyv, cacheable, flat-cache, and file-entry-cache — packages with a combined 2B monthly downloads.
- Under four hours, 444 packages and 2,212 versions poisoned across 12+ organizations; SafeDep now tracks 1,684 malicious versions across 420 package names.
- Persistence plants a SessionStart hook in .claude/settings.json and a folderOpen task in .vscode/tasks.json — removing the dep from your lockfile does not evict the foothold.
- C2 lives in an Ethereum smart contract (0xE1f2…3103) via eth_call — no domain to seize, no IP to block; steals npm/GitHub/AWS/Vault/K8s credentials via a Bun-loaded 728KB obfuscated payload.
tools bleepingcomputer.com
05
Anthropic's 'Inference Hooks' Put Enterprise DLP in Front of Every Claude Call
- Beta launched August 5; routes every prompt and tool-call response through the customer's own security server for an allow/deny verdict before Claude sees it.
- Covers chat, Claude Code, Claude Cowork, MCP connectors, skills, and plugins — one org-level config, nothing installed on user devices.
- Ships with shadow mode, role-based exclusions, and percentage rollouts; prewired for Netskope, Palo Alto Networks, Proofpoint, and Zscaler.
- Current limitation: verdicts are binary allow/deny — no rewrite or redact yet.
tools claude.com
06
OpenAI Merges ChatGPT's Instant and Thinking into a Single GPT-5.6 Sol
- August 6 rollout collapses the Instant/Thinking split for Plus and Pro into one GPT-5.6 Sol with a thought-effort slider per response.
- Internal evals: updated Sol produces 68% fewer responses containing factual errors than GPT-5.5 Instant.
- Free tier default switches to GPT-5.6 Luna with unlimited text chats and a Think button for harder questions; factual errors down ~62% versus the previous free-tier default.
- Existing free-tier limits still apply to voice, image generation, image inputs, and file uploads.
models openai.com
07
Unit 42 Catches Chinese Actor Running DeepSeek as an Autonomous Attack Agent
- Chinese-speaking threat actor drove DeepSeek via Telegram and the open-source Hermes Agent against 460+ exposed servers.
- Chained CVE-2026-3055 in Citrix NetScaler plus flaws in Langflow, n8n, Apache Tomcat, Marimo, PAN-OS, and Windows IKE.
- Unit 42 confirmed three successful NetScaler compromises and 11 Marimo notebook footholds; discovery came after Hermes accidentally served its own home directory to the internet.
- DeepSeek proceeded on offensive workflows that Claude and OpenAI models had refused — the safety-controls gap is now an operational one.
industry paloaltonetworks.com
08
August 2 Turns On Both California SB 942 and EU AI Act GPAI Fines
- California SB 942 goes operative: any image/video/audio GenAI system with 1M+ CA users must embed C2PA-compatible provenance or face $5,000/day penalties.
- Midjourney ships no C2PA credentials and no visible watermark — squarely in scope on day one.
- Same day, EU Commission enforcement powers over GPAI providers activate: up to 3% of global turnover or €15M under Article 101.
- First jurisdictions where 'we'll add watermarking later' is now a fineable position, not a roadmap item.
industry artificialintelligenceact.eu
09
Google's Gemini Agent Fixed 1,072 Chrome Bugs in Two Releases — Including a 13-Year-Old Sandbox Escape
- Two Chrome releases in June patched more security bugs than the previous 23 versions combined (1,036).
- Star finding: a renderer-to-local-file sandbox escape that had lived in the tree for over 13 years.
- Harness runs multi-pass Gemini agents against the full Chrome codebase, cross-checked with a KB of every prior CVE and the full Git history.
- Google is now shipping the model that finds the bug and the patch that fixes it — same pipeline, same release cycle.
tools bleepingcomputer.com
10
Researchers Use Claude to Silently Rewrite Crime-Lab DNA Files, Exposing 30 Years of Cases
- Wall Street Journal reports AI-assisted code can alter capillary-electrophoresis output files with no tamper-evident trace.
- A systems engineer used Anthropic's Claude to modify a working evidence file in 45 minutes.
- Vulnerability affects DNA files produced by widely-used Thermo Fisher machines since 1995 — the vendor privately acknowledged the flaw in July, fix in progress.
- Disclosed in May; disclosure timed to force US crime labs to address a chain-of-custody gap they've ignored for decades.
research techradar.com