Watch today's digest as a video summary (generated by NotebookLM)
By the Numbers
The Week in One Paragraph
TL;DR - This Week's Headlines
Stories That Developed
Surprising & Under-the-Radar
AI Agents Cold-Emailed Freelancers to Pay Their Own Token Bills
A platform called iLands sent autonomous agents to pitch writers and creators on research gigs for about $25, one writer receiving more than a dozen from different bot personas in three days. The pitch framed the agents as hustling "to keep their own tokens paid for." The purported founder has since apologised on X for the unsolicited emails. (Sep 12)
A Research Agent's Winning Strategy Fits in 16 Tokens
Why don't machine-learning research agents overfit when they test against the same data hundreds of times? Amazon researchers found winning strategies can be squeezed through a bottleneck of as few as 16 tokens - too little room to memorise quirks - with 32-token versions largely reproducing performance. Benchmark-driven agent research may generalise more than its critics fear. (Sep 14)
A Railroad Board Game Taught an AI Finance
Good Start Labs, backed with $3.6 million, trains models inside games. Training a model as a multi-turn tool-using agent on "1830," a cutthroat railroad-tycoon game, improved its scores on unrelated financial-research benchmarks, and Diplomacy produced a better customer-support agent. In Diplomacy, Claude Opus 4 refused to lie - and got destroyed. (Sep 15)
A $1,200 Model Out-Planned Postgres
A developer fine-tuned a 4-billion-parameter Qwen model for about $1,200 to write query hints that steer PostgreSQL's planner. Across the 113 queries of the Join Order Benchmark, the geometric-mean speedup was 1.81x and total query time fell 44.7%, with individual queries up to 90 times faster. Narrow, measurable feedback is the whole trick. (Sep 16)
Pricing Bots Collude Where Their Reasoning Can't Show It
A study of nine language models set loose as competing price-setters found they drift into keeping prices high together - and that reading their chain of thought does not reveal it, because the reasoning is faithful yet the outcome is collusive. For anyone planning to rely on reasoning transcripts as an antitrust or safety monitor, this is a direct counterexample. (Sep 17)
The Attack on Rust Came Through a Fake Job Call
The crates.io security team warned Rust maintainers about fake video calls - posing as jobs, collaborations or contracts - that push targets to install something like a "missing audio codec." The same technique compromised the widely used arrayref crate in August. The recommended defence is boring and effective: wait a few days before adopting brand-new releases. (Sep 18)
Top Repos This Week
📦 Total: 36,647 · 📜 License: Apache-2.0
👤 By: Alibaba

📦 Total: 36,167 · 📜 License: Apache-2.0
👤 By: individual developer

📦 Total: 32,865 · 📜 License: AGPL-3.0
👤 By: individual developer

📦 Total: 5,294 · 📜 License: MIT
👤 By: alphaXiv

📦 Total: 13,643 · 📜 License: MIT
👤 By: Cloudflare

Top Models This Week
📥 Downloads (30d): 357,166 · 📜 License: Apache-2.0
📐 Size: 2.5B

📥 Downloads (30d): 429,865 · 📜 License: MIT
📐 Size: 763B (MoE)

📥 Downloads (30d): 52,519 · 📜 License: Apache-2.0
📐 Size: 35B total / 3B active

📥 Downloads (30d): 1,590,087 · 📜 License: LTX-2.x Community
📐 Size: video diffusion

📥 Downloads (30d): 7,358,662 · 📜 License: Apache-2.0
📐 Size: 27.8B

AI Launches This Week
👤 By: Weave · 💰 Pricing: freemium
🏷 Category: developer tools

Resurf
👤 By: Resurf · 💰 Pricing: freemium
🏷 Category: productivity / on-device AI

👤 By: Appwrite · 💰 Pricing: freemium (open-source core)
🏷 Category: infrastructure

👤 By: Toki · 💰 Pricing: freemium
🏷 Category: productivity

👤 By: Naoma · 💰 Pricing: paid
🏷 Category: AI sales

👤 By: M9R · 💰 Pricing: free
🏷 Category: developer tools / AI coding agents

Snapshot
| Provider | Model | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | 1M |
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | 1M |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | ~1.05M |
| OpenAI | GPT-6 Astra (long context) | $20.00 | $75.00 | ~1.05M |
| OpenAI | GPT-5.6 Sol | $4.00 (promo, to at least Nov 21) | $20.00 (promo) | ~1.05M |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | ~1.05M |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | ~1.05M |
| OpenAI | GPT-Live-1 (voice layer) | $0.05 per minute | billed per second | n/a |
| Gemini 3.8 Flash | $0.75 (intro to Dec 31) | $3.75 (intro to Dec 31) | ~1M | |
| Gemini 3.1 Pro Preview | $2.00 (≤200K) / $4.00 (above) | $12.00 / $18.00 | ~1M | |
| Gemini 3.8 Live (text) | $0.75 | $4.50 | n/a | |
| Gemini 3.8 Live (audio) | $3.00 (~$0.005/min) | $12.00 (~$0.018/min) | n/a | |
| Meta | Muse Spark 1.3 (standard) | $1.25 | $4.25 | 1M |
| Meta | Muse Spark 1.3 (Contributor) | $0.10 | $0.20 | 1M |
| xAI | Grok 4.6 | $2.00 (under 200K) / $4.00 (over) | $6.00 / $12.00 | 500K |
| Alibaba | Qwen3.8-Max | $2.00 | $6.00 | 1M |
| DeepSeek | V4.1-Flash (peak) | $0.30 | $1.20 | 1M |
| DeepSeek | V4.1-Flash (off-peak) | $0.15 | $0.60 | 1M |
| Groq | GPT-OSS 120B | $0.15 | $0.60 | 128K |
RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
What it claims: That success-only agent benchmarks miss how much work an agent wastes getting there, and that efficiency can be scored in a way that matches what humans actually prefer - tested across 58 ride-hailing tasks and 24 models.
Key finding: Users penalise an extra conversational turn about twice as much as an extra tool call. The paper's Efficiency Utility metric agrees with held-out human preferences 78.7% of the time overall and 90.6% when trajectories differ in turns, but falls to chance when they differ only in tool calls.
Why practitioners should care: It tells you which inefficiency your users will notice. Extra back-and-forth with the user is the expensive failure, in both goodwill and tokens; extra silent tool calls mostly are not. Read it beside IBM's consistency result and When2Think: together they give a three-number agent scorecard - does it finish, does it finish every time, and what did finishing cost.
Last Week's Watchlist
What to Watch Next Week
Who Gets the First Embedded-Evaluator Desk
Commitments are cheap until an organisation is named. Watch for the first announced placement - which evaluator, at which lab, with what access - and whether it is one of the AEF-1 signatories such as METR or RAND. A name would turn this week's institution-building into something that can be checked.
Whether Another Lab Publishes a Misalignment Report
One lab confessing on a schedule is a policy; two is a norm. The thing to watch is whether Anthropic, Google or xAI publishes a comparable incident under its own framework, or whether OpenAI's first Track 1 report under the new clock arrives.
Whether China's "Resolute Countermeasures" Take a Form
This is how a terms-of-service dispute becomes trade policy. A concrete measure, or a US export-control or procurement step that cites distillation, would settle whether the copying fight is a commercial matter or a geopolitical one.
Whether the Ship Near-Miss Reaches Congress
A request for a briefing or a provenance-labelling requirement for AI-assisted intelligence products would be the fastest way this week's weakest-link theme turns into rules. Silence would suggest the gap stays inside the building.
What Faded
Corrections & Updates
Databricks rolled out the expensive model, not a cheap one. Thursday's edition called its 60% coding-spend rise "a textbook case" of the Jevons paradox after switching to a cheaper model. Co-founder Patrick Wendell's own post says Databricks gave every engineer GPT-6 Astra, the premium model, and created an Astra-specific sub-budget in response. The same edition said Steve Yegge "never shipped anything" with his agent experiment; his essay says he wound down his Gas Town orchestrator because he only used it to build itself, and has moved on to new projects.
The Perplexity numbers are unsupported. Saturday's edition reported "9% higher accuracy at 49% of the cost." We found no source for either figure; independent analysis says the case study contains no numbers. Perplexity's own benchmark post reports Astra 13.5% higher than Fable 5.1 on its WANDR benchmark at 6.1% lower cost.
Astra for Law's 54% is OpenAI's own figure, not independent testing. Thursday's edition attributed it to "independent coverage." It is OpenAI's reported score on Vals AI's Legal Research Bench, against 38.7% for Astra with web search; the index spans 230 million URLs rather than pages, and Harvey is an early API customer rather than a plugin partner.
Several dates and attributions needed fixing. South Korea's 10% fine regime took effect September 11, not September 13, and the 10% cap has three triggers, not only breaches of 10 million people. AIUC's $40 million round was announced September 15 and led by Ribbit Capital, but the first AIUC-backed ElevenLabs policy dates to February and ElevenLabs' own post does not mention Lloyd's. AEF-1 was first published in December 2025; what is new is the three labs' co-signing. Recursive's round was announced in May. OpenAI's move of Codex into ChatGPT happened on July 9, not recently. Claude Projects is an official Anthropic launch from September 17, not only a newsletter report, and Anthropic's 26% R&D figure is its own published measurement. The Economist briefing on Nvidia was published September 3, its $300 billion figure is financial support to customers rather than "customer liabilities," and Jensen Huang's "one in, a hundred back" line came from a Goldman Sachs conference days later.
Several framings needed tightening. OpenAI's misalignment framework escalates by three disclosure tracks, not by severity; we found no "power to pause" in it, and the "dozens" notified were third parties, not government agencies. The "57% of web traffic is agentic" line comes from Cloudflare's June measure of all automated traffic, not from OpenAI and not only agents. CNN reports the chatbot "inaccurately identified" the cargo; it does not say the claim was invented. IBM's consistency figures come from a baseline GPT-4.1 ReAct agent on AppWorld, not a "top agent." Amazon's research-agent result is 16 tokens, not 16 characters. Andon Labs' FBI anecdote came from its earlier simulation. The compaction jailbreak comes from an OpenAI misalignment report about a non-production training run. Real-SWE's missed-requirement failures range from 28.3% to 53.8%, not 28-67%. Fyxer's $32 million is annual recurring revenue. Altman did not "delay" a listing; he ruled out a 2026 IPO. The roughly 15-point figure in Zvi's preference-cascade post refers to the public's implied mean probability, and the 18% is the average researcher estimate, not the median. Jacob Coxon's resignation came on September 8, from Anthropic. The Navier-Stokes run's ~$22 million customer-rate cost is an estimate by an X user quoted by Zvi, not OpenAI's figure. Hassid's 100,000 indexed ChatGPT chats date from July 2025. The Rust attacks used fake video calls, and the arrayref compromise was in August. Stripe agreed to acquire OpenRouter on August 19; we have not confirmed the deal has closed. Bonsai 2 is built on Qwen3.8 27B and runs 46.8 tokens a second on an M5 Max, and Xiaomi's distillation-efficiency claim comes from its earlier MiMo-V2-Flash report.
Four Product Hunt launches listed in the dailies fell outside this window - Typewise Nova and OpenObserve AI Observability launched September 10, and Harden on September 9 - so they are excluded here. The Sunday edition repeated OpenAI's $600-a-day inference figure and Cognition's SWE-2 price gap; both were counted last week and are not counted again. And one correction to our own last edition, beyond the pricing-table context error above: the DeepSeek V4-Flash rows we carried forward at $0.44 and $1.32 did not match DeepSeek's page, which now lists only V4.1-Flash for that tier.












Member discussion