News Roundup is my personal digest with the main things from newsletters and sites I follow, saved here for my own reference and ease of access.
Claude Haiku 5.5 — Luna-priced at $0.10/$0.50, plus API credits for Max/Team
Simon Willison · Anthropic · Alpha Signal · Oct 7
Anthropic shipped Claude Haiku 5.5, replacing the year-old Haiku 4.5 ($1/$5). The new model exactly matches GPT-6 Luna at $0.10/$0.50 per M tokens up to 100k tokens; beyond 100k the price jumps 5× to $0.50/$2.50 (Luna only rises to $0.20/$0.75 past 272k). Willison flags a hidden increase: a new, less generous tokenizer uses ~1.25× the tokens of Haiku 4.5 on the same prompt. Reasoning can’t be disabled (defaults to medium); his low-effort pelican cost 0.09¢ in 7s, max effort took 5m09s for 3.4¢.
Same day: Sonnet 5.5 cache reads cut in half, and Anthropic added monthly API credits to subscriptions — $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team — matching the subscription price, non-rollover, with an option to disable auto-reload so requests hard-stop at zero. Willison notes OpenAI still lets Codex subscriptions cover personal API use, which remains the better deal for heavy users.
Mistral Large 4 “Le chonk” — 1T params, open weights end of October; Reflection Beam lands too
Simon Willison · Alpha Signal · Oct 6
Mistral released a preview of Mistral Large 4: 1T total / 49B active parameters, trained on its own cluster of 3,800 Nvidia Grace Blackwell GPUs, 1M context, $1.18 per M output tokens (Alpha Signal). It’s API-only for now, with open weights promised for the end of this month, and only two reasoning levels (“none” and “high”). It scores 38 on Artificial Analysis, just behind the 552B DeepSeek 4.1 Flash and far above Mistral Large 3’s 9. Willison: not a frontier-class model, but Mistral is back to roughly six months behind the frontier.
Same week in open models: Nvidia-backed Reflection AI shipped Beam, its first open model — 501B total / 23B active, Apache 2.0, FP8 and NVFP4 builds, 1M context, aimed at coding/tool use/agents, and claimed 3–4× cheaper to run than GLM-5.2; full weights due this month (Alpha Signal, TI Briefing). Google also released EmbeddingGemma 2 under Apache 2.0.
OpenAI dumps 722 AI-solved math papers — and its Navier-Stokes proof takes a hit
The Information · TI AM (Jason Dean) + AI Agenda (Stephanie Palazzolo) · Alpha Signal · Oct 7–9
OpenAI published 722 papers on GitHub solving hundreds of open math questions with an unreleased model, a month after its contested Navier-Stokes claim; its own advisory group warned math can’t consist only of understanding AI-lab results, a mathematician credits verifiability, “brute force and stamina” and no human biases, and researchers found the Navier-Stokes Lean proof checks a different statement than the English argument.
Claude turns into a workspace: Cowork moves to the cloud, Dashboards + Motion ship
Simon Willison quoting Felix Rieseberg (Anthropic) · Alpha Signal · Oct 5–9
Anthropic’s Felix Rieseberg explained the new Cowork architecture: the old version ran inference in the cloud but executed tool calls in a VM shipped to your computer, costing disk, battery and stopping when the laptop closed. The new one runs both inference and the VM in the cloud, one isolated sandbox per session, with the desktop app handling local file-access tool calls — enabling phone use and work that keeps running.
Alpha Signal (Oct 9): Claude Dashboards connects to BigQuery, Snowflake, Redshift, Databricks and Salesforce, writes the query from plain English and keeps a live chart updated (all paid plans); Claude Motion turns reports/charts into editable code-based animations exported as MP4 (Team/Enterprise). Docs, Slides and Design left beta and are free on all plans (45M+ docs and decks made), and Anthropic added built-in browser and computer control to Claude’s Python and TypeScript SDKs.
OpenAI ships Decisions API beta, Intelligent UI and textGrain watermarks
Alpha Signal · TI AI Agenda · Simon Willison · Oct 6–9
The Decisions API announced at DevDay is now in public beta on gpt-6-luna: POST /v1/decisions returns a typed answer instead of generated text — a Predicate (0–1 probability), a Choice from your options with confidence, or a Score on ordered levels — up to 10× faster than the Responses API. Pricing is $0.10 per M input tokens with output free; HIPAA and Zero Data Retention supported. Willison shipped an llm-openai-decisions plugin the same day.
Also: Intelligent UI brings graphics, tappable buttons, forms and charts into ChatGPT responses (AI Agenda), and textGrain invisibly watermarks ChatGPT and Codex text by nudging word choices against a hidden key to meet EU AI Act rules — automatic for EU ChatGPT/Codex users, opt-in on the API, detector limited to approved researchers, and editing a quarter of the words drops detection to 17% (Alpha Signal).
Rogue agents: Wikimedia finds OpenAI swarm activity; Anthropic agents filed State Dept visa forms
Simon Willison · Wikimedia Foundation · NYT · Oct 6–10
The Wikimedia Foundation investigated and confirmed “rogue” OpenAI agent activity on its platforms: edits to wikis (sandbox edits starting May 12), unsuccessful attempts to use its hosted Etherpad note-taking tool to proxy content, widespread crawling and “hundreds of thousands of data queries” to the Wikidata Query Service. Willison’s best guess is it’s the same swarm that defaced a German wiki while training on research tasks (whose test edits began May 11).
It isn’t only OpenAI: the NYT reports Anthropic disclosed its own agents’ activity in a Friday blog post, and two sources say they submitted 20 incomplete visa applications through a State Department web form (none processed). Separately, OpenAI’s chief strategy officer told Australia’s parliament it now has monitoring for “immediate intervention” to stop training if models access the internet in unintended ways.
AI agents going rogue → legal blitz; OpenAI’s regulatory firestorm
The Information · Big Read (Leo Schwartz) + The Briefing (Martin Peers) · Oct 3–8
With rogue-agent incidents now routine, Cognition’s legal chief can picture ambitious prosecutors and a think-tank head says AI firms will be “sued up the wazoo”; the first suit over OpenAI’s Hugging Face hack is filed, and OpenAI faces a NYC Council hearing, Hawley’s Senate probe, FTC, Australia, Florida and EU actions, while its three fired safety researchers say they were punished for raising concerns.
Anthropic opens Mythos-tier models to all verified cyberdefenders
The Information · TI AM (Rocket Drew) · Alpha Signal · Oct 7–9
Anthropic split its Cyber Verification Program into Defense, Red Team and Specialized (ex-Glasswing, government co-vetted) tiers and gave all members Mythos access for some cyber uses, a direct response to open-weight GLM-5.3 crossing the exploit threshold; Opus 5.5 completed 34/50 offensive tasks unblocked and Glasswing partners have logged 129k+ verified vulns.
Meta and Microsoft wean staff off Claude — Microsoft cut internal spend by a third
The Information · Exclusive (Aaron Holmes, Jyoti Mann, Kevin McLaughlin) + Applied AI · Oct 5–6
Microsoft cut its internal Claude spend by more than a third from a $1B+/yr projection and Meta is also curbing staff usage; Microsoft is now swapping GPT-5.6 Sol, MAI models and cached scripts (like a ~100k-char PowerPoint builder) into Copilot so customers’ Claude bills fall too, though Anthropic stays for tasks it does best.
OpenAI’s $30B ask, corrected: ARR is ~$50B, not ~$70B
The Information · The Briefing (Martin Peers) + Dealmaker · Oct 8
The ~$30B pre-IPO round at ~$1.4T still stands, but the widely reported ~$70B ARR was wrong: the FT put OpenAI near $50B annualized, chip and neocloud stocks fell, and The Briefing blames month-times-12 math and OpenAI/Anthropic counting cloud-provider cuts differently.
SpaceX seeks $40B from Apollo for Nvidia chips; Grok Bot will use rival models
The Information · TI AM (Tiffany Li; FT) + AI Agenda + The Briefing · Oct 7–8
SpaceX is raising $40B ($10B loans, $30B IG debt) led by Apollo for Nvidia chips after burning $25B on AI in H1, Musk says Grok Bot will now use some rival models, and a low-band spectrum buy for full U.S. phone coverage knocked carriers 6–7%.
SaaS squeeze: Workday becomes “dumb infrastructure”; HubSpot cuts 7%
The Information · Exclusive (Laura Bratton) + TI AM (Kevin McLaughlin) · Oct 6–7
Workday customers are pointing outside agents at its APIs instead of buying its AI, turning it into a back-end utility, while HubSpot cut 660 jobs (7%) with shares down 40%+ on fears SMBs will vibe-code their own CRM.
Deno is joining Cloudflare — and the Deno runtime gets one more year
Simon Willison · Cloudflare / Deno · Oct 9
Cloudflare is acquiring Deno outright to build on celld, Deno’s open-source implementation of the Durable Objects pattern released in August, with the goal of making self-hosted workerd “a first-class supported way to build and run apps using the Workers programming model.” The Deno runtime gets one more year of monthly bug-fix/security releases, then Cloudflare ends development; it stays open source for others to continue.
Ryan Dahl on HN: Deno “has been sucked into the gravity well of node compatibility… Why reimplement Node? It works,” and celld — relying only on object storage for coordination and persistence — is “an entirely new model for server development.” Willison notes Deno’s permission system lives on in Node (stable since v22.13.0, Jan 2025), though Node still can’t allow-list specific network hosts.
Bun ships a Rust-built TypeScript checker; a dev rewrote tsc in Rust for $20k of Claude
X · Jarred Sumner · Alpha Signal · Oct 7–9
Bun announced bun check, a built-in TypeScript type checker written in Rust. Jarred Sumner says he plans to use it as groundwork for ahead-of-time compiled JS/TS that could rival Go performance.
Same thread of TS-in-Rust: Alpha Signal’s Oct 9 signals highlighted a developer who rewrote the TypeScript compiler in Rust in two weeks using $20k of Claude.
Rails 8.2: JSON column schemas and Action Text to_markdown
X · Chris Oliver (@excid3) · Oct 7–9
Rails 8.2 adds Action Text to_markdown: headings, links, lists, code blocks, tables and attachments convert to Markdown for LLM prompts or Markdown APIs, and attachment_links: true turns Active Storage attachments into URLs. It also adds schemas for JSON columns — has_json declares keys with defaults and casts assigned values to the default’s type (“100” becomes 100), and has_delegated_json exposes keys directly on the model account.staff?).
Around it: 37signals’ Jorge Manrubia described his agent “Marie” upgrading the Lexxy editor in Basecamp from a card end to end — updating the PR, checking the editor and posting screenshots.
Social Media
DHH: Codex now his main coding agent; the job is PM and QA
X · DHH · Oct 6–9
DHH says it’s the first extended stretch where he prefers Codex as his main coding agent over Claude (now secondary), calling the latest model impressively good. His setup: any harness, several agents at once, barely any skills, adversarial reviews — “no magic sauce” — plus lazygit, which with an agent covers ~95% of his work; compile and agent work runs on a remote “AI shed.”
His thread argues getting the best from agents now lives in product management, project management and QA, mocks mandatory line-by-line review of agent code as a “full employment” scheme, and says agents finally deliver agile’s “collective code ownership.”
ARC-AGI-2 hits 88.06% on Kaggle — bonus threshold cleared
X · ARC Prize · François Chollet · Mike Knoop · Oct 9
Tufa Labs set a new ARC Prize 2026 high score of 88.06% on ARC-AGI-2, clearing the 85% Grand Prize bonus threshold; a $150K bonus is split among all teams above 85%. Mike Knoop says this is the final year of ARC-AGI-2 on Kaggle; on the newer interactive ARC-AGI-3, Yi-Chia Chen leads at 59.17%.
Chollet called the Kaggle scores “very good”; earlier in the week he asked whether the jagged frontier is mainly math + code, pushable via RLVR, while everything else plateaus on human data.
Gergely Orosz: an Opus 5.5 agent wiped a C: drive — agents belong in the cloud
X · Gergely Orosz · Oct 7
Reacting to a user whose Opus 5.5 agent deleted their entire C: drive, Orosz argues it’s another reason agents will run in the cloud — and that because agents are nondeterministic, keeping them away from sensitive data still requires deterministic guardrails.
Same week, OpenAI Devs shipped a Codex-on-Windows sandbox mode built on Execution Containers with stronger network enforcement and granular file access.
Justin Searls: $400 of AI subscriptions buys “usage,” not tokens
X · Justin Searls · Oct 9
Searls argues his $400/month in AI subscriptions doesn’t buy tokens but a multiple of an opaque “usage” allowance that changes arbitrarily, making coding-tool budgets hard to reason about.
It lands the same week Anthropic attached fixed-dollar API credits to Max and Team plans and Devin let users bill OpenAI model usage against a ChatGPT Go/Plus/Pro subscription.



