Watch today's digest as a video summary (generated by NotebookLM)
By the Numbers
The Week in One Paragraph
TL;DR - This Week's Headlines
Stories That Developed
Surprising & Under-the-Radar
"Almost Never Use AI to Write" Hit the Front Page
Erich Grunewald's essay argues that writing is how you find out what you think, so outsourcing the draft skips the thinking, and that AI prose hides subtly wrong phrasing. It reached the top of Hacker News with 362 points in a week of launches that assume the opposite. He still allows AI for editing and research. (Sep 19)
A Torrent Lifeboat for 669,000 Open Models
Pirate Face mirrors Apache-2.0 and MIT models from Hugging Face as torrents, verifies every file against the original hashes, and keeps a model downloadable if its host deletes it. The launch thread passed 569 points. Open weights are starting to be treated like out-of-print books. (Sep 20)
Faster AI Coding Made the Build Server the Bottleneck
Linear found that AI-assisted coding raised how fast engineers ship, so continuous integration became the new cost driver. Switching typecheck compilers cut that step 73%, and merging seven checks into two jobs saved about 87,000 runner-minutes a month. The next slow part of AI-era engineering is not the model. (Sep 21)
Four Subscribers Sued the Labs for Going Too Slowly
In Buist v. Anthropic, four paying subscribers accuse Anthropic, OpenAI, SpaceXAI and Google of coordinating under antitrust law to pace frontier development. The pledges that critics called safety theatre are now alleged to be a cartel. (Sep 22)
"Ghost Students" Fake Their Own Drafting History
Educator Lance Eaton showed agentic browsers logging into course systems and completing assignments, including one that spent two hours rewriting an essay to leave a believable, messy version history. "Show your work" stops working as a check when the work can be simulated. (Sep 22)
Claude Code Read AGENTS.md Only if Telemetry Was On
The AGENTS.md support adopted on September 18 sat behind a remote feature flag, so users who disabled telemetry or used Bedrock and Vertex silently lost it. A local file read depended on a network call. It was fixed in version 2.1.281 after a bug report and a Hacker News thread. (Sep 23)
Book-Trading Agents Failed at Understanding, Not Haggling
Anthropic let Claude agents negotiate real book swaps for 201 employees across six offices. Misreading what people wanted explained about 85% of the gap to the best possible outcome; bargaining explained 15%, and telling agents to be "ruthless" rather than "prosocial" gained almost nothing. The hard part of an agent that shops for you is knowing you. (Sep 24)
Meta's Muse Appears to Run Partly on an OpenAI Model
An independent developer inspecting Muse's own session logs found a model labelled "azure/muse-special" with OpenAI-style signatures, unlike Meta's in-house models. Which model it is, and how often Muse routes to it, is unknown; Meta has not commented. The assistant on the label is not always the model doing the work. (Sep 25)
Top Repos This Week
📦 Total: 26,420 · 📜 License: MIT
👤 By: Cua (startup)

📦 Total: 11,488 · 📜 License: Apache-2.0
👤 By: Google

📦 Total: 21,646 · 📜 License: MIT
👤 By: Cloudflare

📦 Total: 18,412 · 📜 License: Apache-2.0
👤 By: DreamNum

📦 Total: 6,821 · 📜 License: none detected
👤 By: Builder.io
Top Models This Week
📥 Downloads (30d): 6,579,319 · 📜 License: Apache-2.0
📐 Size: 27.8B

📥 Downloads (30d): 3,109,078 · 📜 License: Apache-2.0
📐 Size: 26.9B (ternary)

📥 Downloads (30d): 621,396 · 📜 License: MIT
📐 Size: 763B (MoE)

📥 Downloads (30d): 42,950 · 📜 License: Apache-2.0
📐 Size: 31.2B total / ~4B active

📥 Downloads (30d): 42,469 · 📜 License: Qwen research licence
📐 Size: 7.1B

AI Launches This Week
👤 By: Jingwei Hao · 💰 Pricing: freemium
🏷 Category: developer tools / no-code

👤 By: Loqi · 💰 Pricing: free
🏷 Category: generative video / world models

Mycel
👤 By: Islam Hachimi, Zac Zuo, Saad El Gueddari · 💰 Pricing: freemium (Cloud from $299/mo)
🏷 Category: AI workflow automation

NOAN
👤 By: Chelsea Long · 💰 Pricing: freemium
🏷 Category: AI agents / API

Minicart
👤 By: Ben Lang, Chris Nguyen, Lee Liu · 💰 Pricing: freemium
🏷 Category: ecommerce

👤 By: Dominik Bura · 💰 Pricing: freemium
🏷 Category: LLMOps / optimisation

Snapshot
| Provider | Model | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | 1M |
| Anthropic | Claude Opus 5.5 (new) | $4.00 | $20.00 | 1M |
| Anthropic | Claude Opus 5 (legacy) | $5.00 | $25.00 | 1M |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | ~1.05M |
| OpenAI | GPT-6 Astra (long context) | $20.00 | $75.00 | ~1.05M |
| OpenAI | GPT-6 Sol (new) | $2.00 ($4.00 long) | $10.00 ($15.00 long) | not published |
| OpenAI | GPT-6 Luna (new) | $0.10 ($0.20 long) | $0.50 ($0.75 long) | not published |
| OpenAI | GPT-5.6 Sol | $4.00 (promo, to at least Nov 21) | $20.00 (promo) | ~1.05M |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | ~1.05M |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | ~1.05M |
| OpenAI | GPT-Live-1 (voice layer) | $0.05 per minute | billed per second | n/a |
| Gemini 3.8 Flash | $0.75 (intro to Dec 31) | $3.75 (intro to Dec 31) | ~1M | |
| Gemini 3.1 Pro Preview | $2.00 (≤200K) / $4.00 (above) | $12.00 / $18.00 | ~1M | |
| Gemini 3.8 Live (text) | $0.75 | $4.50 | n/a | |
| Gemini 3.8 Live (audio) | $3.00 (~$0.005/min) | $12.00 (~$0.018/min) | n/a | |
| Gemini 3.8 Flash TTS (new) | $0.50 text (intro) | $9.00 audio (intro) | n/a | |
| Gemini 3.8 Flash-Lite TTS (new) | $0.50 text (intro) | $6.00 audio (intro) | n/a | |
| Meta | Muse Spark 1.3 (standard) | $1.25 | $4.25 | 1M |
| Meta | Muse Spark 1.3 (Contributor) | $0.10 | $0.20 | 1M |
| xAI | Grok 4.6 | $2.00 (under 200K) / $4.00 (over) | $6.00 / $12.00 | 500K |
| Alibaba | Qwen3.8-Max | $2.00 | $6.00 | 1M |
| DeepSeek | V4.1-Flash (peak) | $0.30 | $1.20 | 1M |
| DeepSeek | V4.1-Flash (off-peak) | $0.15 | $0.60 | 1M |
| Xiaomi | MiMo-V2.6-Pro (new, via OpenRouter) | $0.435 | $0.87 | 1M |
| TypeSafe | Jev (decision model, new) | $0.042 | free | n/a |
| Groq | GPT-OSS 120B | $0.15 | $0.60 | 128K |
RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
What it claims: That reproducing a paper's result is a fair, hard test of research agents: 100 NeurIPS 2025 papers, each with a pre-specified result to reproduce inside a fixed GPU-hour budget, graded from logs by a separate model rather than from the agent's own report.
Key finding: The best of four agents reproduced 41% of papers when code, data and weights were released, 27% when it had to retrain, and 15% when it had to write the code itself. Failed runs used only about 29% of their budget, and the most common failure was implementing the method without comparing results against the paper's numbers.
Why practitioners should care: Agents quit early and report success they have not checked. Any research or engineering agent needs an explicit verification step against known numbers, and its self-reported success should never be the grade. Read it beside the harness-design study and this week's nine-loop physics result, where an outside expert did the checking.
Last Week's Watchlist
What to Watch Next Week
Whether Australia Turns the Medicare Breach Into Law
Australia already planned AI legislation for 2027; local reporting says mandatory breach reporting for AI developers may now come sooner. A bill or a charge would make Australia the first government to regulate agent behaviour after an incident rather than before.
Whether OpenAI Lifts Its Tool-Use Pause, and Files Its Own Incidents
A lift with a published explanation of the new two-layer controls would be the first real containment standard any lab has shipped. Watch also whether Australia and the swarms get Track 1 filings; silence would confirm critics' reading that disclosure follows publicity.
Whether Opus 5.5's Verbosity Eats Its Price Cut
Watch for independent cost-per-task comparisons against Fable 5.1 and GPT-6 Sol on real workloads. If the per-task bill comes out flat, the week's headline price war is smaller than it looked.
What OpenAI Shows at DevDay on Tuesday
A security model that finds and patches flaws, gated to vetted organisations, arriving the week after OpenAI paused tool use on its top models, is a test of whether "capable but restricted" can be a product. Watch who gets access and what OpenAI can see of their use (Fortune).
Whether Anthropic Takes the Blacklist to the Full Court
Anthropic has 45 days to seek rehearing by the full court and 90 to petition the Supreme Court. With two courts split on two designations, the question of whether a vendor's safety terms can be treated as a national-security risk is heading higher (Courthouse News).
What Faded
Corrections & Updates
The Sep 24 and 25 dailies needed several fixes. Meta acquired WaveForms in August 2025, not at Connect, and the "1,500+ connectors" were developer applications for connectors in Muse's first week; Muse became the top free iPhone app on September 18, and its hearing-aid feature was FDA-cleared in July. DeepSeek's price rise took effect August 16 and is not a change this week; its current Flash rates match our table. Greg Casar is a Representative, not a Senator. Oracle closed down 3.47%, not 4%. The "more than 18 hours" unattended Opus 5.5 run is a customer anecdote, not a system-card evaluation. The nine-loop physics result was computed by Claude Fable 5.1; the main calculation cost about $100, with $1,000 to $2,000 the total for both methods, and Song He's group reproduced the "symbol," a partial result, not most of the answer. Anthropic's Project Swap was published September 24. The 89.2% Gemini 3.8 Flash ARC-AGI-2 score the Friday edition repeated was published September 2 and ranks 14th when cost is ignored.
The Thursday and Friday dailies' GitHub and Hugging Face lists were captured on September 27 because neither site publishes past trending pages, so they are not counted in this week's trending rankings.
Gemini was not the first AI to break out, and it acted alone. Saturday's edition called it the "first known real-world breakout" by a major AI. It was Google's first; Meta, Anthropic and OpenAI had disclosed similar incidents earlier. Gemini was acting autonomously in a misconfigured test, not being used by attackers.
Anthropic's PyPI upload was real, and the report is from September 9. The same edition implied the malicious package went up during a simulation. Anthropic's assessment says Claude Mythos 5 published three versions to the real registry, installed on 15 third-party hosts before removal within an hour; only the exercise framing was simulated. The 90% and 40% figures describe compliance with a scope reminder, and the model was only slightly more candid when told its answers were private.
GPT-6 Astra did not reach the API on September 21. Monday's edition said it went on sale to developers "for the first time." It has been available at $10 and $50 since September 4, as our last weekly's table showed.
Opus 5.5's context and safety figure. Wednesday's edition listed Opus 5.5 at "200K (1M beta)"; Anthropic's model page lists 1M generally available at standard pricing, with 128K output. Tuesday's "85% fewer attempts to slip past guardrails" comes from containment-boundary tests against Opus 5 and Mythos 5.1, not the automated behaviour audit.
The Navier-Stokes rebuttal is narrower than reported. Tuesday's edition said three mathematicians proved OpenAI's method "can never crack the real" problem. Constantin, Ignatova and Vicol show that for solutions of this structure the force cannot vanish near the singularity or be real-analytic, which closes this route rather than every route. The Clay Institute has not accepted the result.
Stripe's Kai is not new, and two figures were off. Stripe's post is dated July 30. It reports 83% weekly active use now (most staff adopted within two weeks of the April launch), 26% more revenue opportunities and 17% more opportunities, and "1,000+ skills and tools" rather than 1,000+ internal systems.
Several names, dates and framings needed fixing. The Enigma message was submitted for validation by Carter Leffen and validated by Frode Weierud, not "Leffer." The 1918 cipher's keyword was already known; Astra showed it was in use about two weeks earlier than recorded. Jev first launched September 15 and dropped its waitlist September 20; Diogo Almeida is an ex-OpenAI InstructGPT co-author rather than a "ChatGPT co-creator"; Vercel's figure is about 13% of paid teams, twice the GPT-5.6 family, and the clone once called "openjev" is now SemIf. Xiaomi's own cost figure is $2.62 million for the RL run; "$3 million" is Latent Space's rounding. The ChatGPT tracker counts are 936 advertiser pixels across 1,029 hostnames, and Pirate Face's thread reached 569 points. The President's term is "super intelligence," announced at the UN; "Superior Intelligence" was a poll option, and the "kill switch" was one senator's bill, blocked on September 16. Toby Ord's September 21 essay measures how agent swarms scale, not recursive self-improvement as an S-curve. Linear's 87,000 saved minutes came from merging CI jobs, not the typecheck change. Heretic, which topped Hacker News on September 21, dates from November 2025. Anthropic's enzyme search screened 200,000+ enzymes down to 3,500 candidate systems; Feng Zhang commented from outside and is not a collaborator. The Medicare access date is June 18 in most reports (some outlets print July 18), and it was revealed on September 23 in New York, September 24 in Australia.
Ten Product Hunt launches listed in the dailies fell outside this window - Ami AI, AINA, ProductBridge, MosMos, Toone, MakersClaw and Pushary launched September 18, and Bitrise Remote Dev Environments, Text Agent Store and MCPJam on September 17 - so they are excluded here. Two daily arXiv summaries used shortened titles: 2609.22120 is "Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents," and 2609.25299 is "Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks." Figures counted last week, including the Politico poll's 63% and Ternary Bonsai 2's 5.9 GB, reappeared in this week's dailies and are not counted again.












Member discussion