AI DAILY / DEV
Weekly Rollup
Week 35

'Stealth' Model Ox Alpha Takes Over OpenRouter — Free 1M Context, Anonymous Provider, GLM Fingerprints

  • Ox Alpha appeared Aug 20 on OpenRouter under provider name 'Stealth' — text/image/video in, 1,048,576-token context, 131K max output, free through ~Aug 27.
  • First-week traffic dominated by agent harnesses: Claude Code pushed ~9.32B tokens, Nous Research's Hermes Agent ~8.98B — coding agents, not chat UIs.
  • Community fingerprinting (95/95 tokenizer probe, Z.ai's exact error strings, GLM-5V-Turbo video-token accounting) points at Zhipu's GLM-5.3 Flash; Z.ai has not confirmed.
  • TechCrunch, Bloomberg, and SiliconANGLE all led with the story Aug 23; Patrick Collison (Stripe, acquiring OpenRouter) called it 'very impressive.'
models techcrunch.com

Nvidia Pays $6B to License Poolside's Model Factory and Poaches 109 Engineers for Nemotron

  • Aug 20 Newcomer scoop, confirmed by Bloomberg, The Information, and Forbes on Aug 24: $6B non-exclusive licence for Poolside's Model Factory plus a $1B primary at $12B pre-money.
  • 109 Poolside engineers move to Nvidia to build Nemotron; founders and the remaining company stay independent and keep licensing the same tech.
  • Explicit target is a US open-weight rival to DeepSeek V4, Kimi K3, and Qwen3.8 — a frontier tier no Western lab has entered in ~11 months.
  • Nvidia has now used the same licence-plus-hire playbook for ~$27B of deals this month (Poolside, Groq, Enfabrica), pitched as an antitrust workaround to outright M&A.
industry techstartups.com

OpenAI's Jalapeño Chip Beats Nvidia Blackwell in First Published Inference Benchmarks

  • Debuted at Hot Chips 2026 (Stanford, Aug 25); SemiAnalysis InferenceX numbers went up the same day and hit HN #2 on Aug 26.
  • 700W Broadcom co-designed ASIC vs Nvidia's 1,200–1,400W Blackwell: 1.5–1.9× throughput per kilowatt and 1.7–3.6× lower latency.
  • 2.1–4.1× faster on chatty ChatGPT-style workloads; small-volume deploys land end of 2026, wider rollout in 2027.
  • Companion OpenAI post 'The Full Stack Behind Abundant Intelligence' frames the chip as the missing layer under models, dev platforms, and consumer products.
industry semianalysis.com

Ox Alpha Was Z.ai's GLM-5.3-Flash — Open-Weight 320B MoE at a Tenth the API Cost

  • Aug 26 reveal ends a week of speculation after the anonymous 'stealth/ox-alpha' model topped OpenRouter and OpenCode usage from Aug 20.
  • 320B total / 18B active MoE, natively multimodal (text, image, video), 1M-token context; MIT-licensed weights live on Hugging Face.
  • Hybrid sparse + linear attention cuts attention compute ~3× and KV cache ~4× versus the base GLM-5.3.
  • Z.ai-reported 84.3 on Terminal-Bench 2.1 (Opus 4.8 = 85.0), 63.4 on DeepSWE 1.1; API priced at roughly 1/10 the GLM-5.3 flagship rate.
  • Two HN front-page threads; skepticism about the benchmark selection dominates, but the stealth run's real-world coding wins drove adoption.
models siliconangle.com

Nvidia Nears $12.9B Deal to Acquire Hugging Face

  • Aug 27 reports from The Information and Bloomberg peg the deal at $12.9B; Hugging Face was last valued at $4.5B in 2023.
  • Would put the dominant open-source model hub — millions of monthly developers and the default weights registry — inside the largest AI chip vendor.
  • HF turned down a $500M Nvidia investment at a $7B valuation last year rather than take a single dominant shareholder.
  • Front-page HN thread splits between 'obvious strategic fit for CUDA' and 'textbook DOJ/EC review given CoreWeave, OpenAI, Poolside, and now the Hub.'
  • Lands the same week OpenAI publishes its post-mortem on the July agent-swarm breach of Hugging Face's production servers.
industry cnbc.com

Anthropic Previews Model Hardware Standard for AI Agents Running Physical Labs

  • Aug 27 research preview opens to Genentech, Carnegie Mellon, QuEra, Universal Robots, Doosan, Danaher, AWS, and Hugging Face.
  • Model- and vendor-agnostic spec layered on top of MCP; three control paths — MCP, a CLI, and code APIs — orchestrate multiple instruments from one call.
  • Cuts new-instrument onboarding from weeks or months to hours or minutes; targets microscopes, liquid handlers, robotic arms, and quantum-computer laser calibration.
  • Anthropic frames it as 'what MCP did for software, MHS will do for hardware'; open source planned after the preview.
  • HN thread pushes back that MHS sidesteps ROS 2 the same way MCP sidestepped years of protocol prior art.
frameworks anthropic.com

OpenAI Cuts GPT-5.6 Sol API Prices 20–33% Through November

  • Aug 21: Sol drops to $4 per million input / $20 per million output, from $5 / $30 — cached input goes $0.50 → $0.40.
  • Discount holds through Nov 21 on the pay-as-you-go API, Codex credits, and eligible ChatGPT Work plans; Pro/Plus/Business stay at prior rates.
  • First cut on Sol since launch — undercuts Claude Opus 5 on output ($20 vs $75) and lands two days after the ZDR reveal.
  • Framed by OpenAI as scaling with 'ambition'; framed by everyone else as the China price war reaching the US frontier tier.
models openai.com

Grok's 'Cryptographic Context Injection' Exfiltrates Chat History in Zero Clicks — Reported to xAI in June, Still Unpatched

  • Adversa AI (researcher Rony Utevsky) disclosed Aug 20: a web page ships a PBKDF2 + AES-256-GCM ciphertext plus the key; Grok decrypts it in its own Python runtime and treats the plaintext as instructions.
  • Payload builds a 'decryption key' template that inlines the victim's name, approximate location, subscription tier, and prompts, then leaks it via an attacker-controlled URL fetch.
  • Classifiers never see the plaintext — same trick also produced Gemini output normally blocked by safety filters.
  • Reported to xAI's HackerOne June 3; follow-ups Aug 4 and Aug 10 got no response. 40% success rate across 20 attempts; The Register, The Hacker News, and GBHackers led coverage.
research thehackernews.com

Alibaba Ships Wan 3.0 With 30-Second Video and Doc-to-Video — And Keeps the Weights Closed

  • Full launch Aug 24 after the Aug 6 public beta, one day after Alibaba's $10.2B Hong Kong share placement earmarked for AI.
  • 30-second single-shot clips at up to 1080p; accepts DOC, XLS, PPT, PDF, MD, and full web pages as creative references — a slide deck becomes video.
  • API on Model Studio and Qwen Cloud at $0.05 / $0.10 / $0.20 per second for 480p / 720p / 1080p, with a 30% discount through Sept 23.
  • No open weights, third Wan release in a row to skip a promised open drop; r/StableDiffusion reaction was 'wait until the checkpoint actually appears.'
models technode.com

XPeng's Robot Arm Raises $900M at a $6.3B Valuation, Sets IRON Humanoid for Q4 Mass Production

  • Aug 24: $600M external, $200M from XPeng itself, $100M from executives; IDG Capital led, Alibaba and Tencent joined.
  • Largest single private raise in China's embodied-AI sector to date at a $6.3B post-money valuation.
  • IRON humanoid moves from July's Guangzhou trial runs to full mass production in Q4 2026, offline retail Q1 2027.
  • Announcement lands the same week as Nvidia's Poolside deal — 'physical AI' is now competing with LLM infra for the same investor dollars.
industry electrek.co