> ## Content Index
> Fetch the complete content index at: https://genaisecretsauce.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# GenAI Secret Sauce Weekly Digest - Week of September 12 to September 18, 2026
- URL: https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-12-to-september-18-2026/
- Published: 2026-09-19T00:20:44.000Z
- Updated: 2026-09-19T00:58:45.000Z
- Description: The "Pace the Frontier" Fight Left the Labs and Entered Politics · Distillation Went From an Accusation to a Policy Fight · OpenAI's Agent Incidents Turned Into a Disclosure Process
- Author: Jasmine Robinson
- Tags: Weekly Digest

Watch today's digest as a video summary (generated by NotebookLM)

By the Numbers

## By the Numbers

[26%](https://www.bloomberg.com/news/articles/2026-09-17/anthropic-says-claude-drives-26-of-its-research-and-development?ref=genaisecretsauce.com) 

Share of Anthropic's AI research work now led by Claude

[Up from under 1% in February and March](https://www.bloomberg.com/news/articles/2026-09-17/anthropic-says-claude-drives-26-of-its-research-and-development?ref=genaisecretsauce.com), with [about 30,000 internal agents running at any moment, all monitored](https://qz.com/anthropic-claude-ai-research-development-automation-091826?ref=genaisecretsauce.com)

[63%](https://www.ibtimes.com/most-americans-worry-ai-could-destroy-humanity-two-thirds-rate-possibility-moderate-3807512?ref=genaisecretsauce.com) 

Americans who see at least a moderate risk AI destroys humanity

[Politico and Public First polling fielded September 13-15](https://www.ibtimes.com/most-americans-worry-ai-could-destroy-humanity-two-thirds-rate-possibility-moderate-3807512?ref=genaisecretsauce.com), the same week [the President called the risk a "HOAX"](https://www.axios.com/2026/09/14/trump-ai-safety-anthropic-dario-amodei?ref=genaisecretsauce.com)

[60%](https://x.com/pwendell/status/2100299179923067016?ref=genaisecretsauce.com) 

Rise in Databricks' coding bill after giving every engineer the premium model

[About 3,500 engineers got GPT-6 Astra](https://x.com/pwendell/status/2100299179923067016?ref=genaisecretsauce.com), and the company answered with [a dedicated Astra sub-budget](https://www.latent.space/p/ainews-reality-checks-on-ai-news?ref=genaisecretsauce.com)

[53%](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com) 

Tasks a baseline agent passed on all five tries, against a 77.4% average

[IBM Research's consistency study on AppWorld](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com), lifted to 69% with written guidelines

[38.8%](https://withspecific.com/benchmarks/real-swe?ref=genaisecretsauce.com) 

Best score on real, private enterprise codebases

[Fable 5.1 led Specific's Real-SWE](https://withspecific.com/benchmarks/real-swe?ref=genaisecretsauce.com), ahead of GPT-6 Astra at 33.8% and Gemini 3.8 Flash at 31.2%

[54.0%](https://openai.com/index/astra-for-law/?ref=genaisecretsauce.com) 

OpenAI's own legal-research score for its new law product

[Astra for Law against 38.7% for Astra with web search](https://openai.com/index/astra-for-law/?ref=genaisecretsauce.com), on [Vals AI's 200-question Legal Research Bench](https://siliconangle.com/2026/09/17/openai-launches-astra-for-law-a-gpt-6-configuration-for-legal-research/?ref=genaisecretsauce.com)

[$40M](https://aiuc.com/updates/series-a-announcement?ref=genaisecretsauce.com) 

Series A to insure and certify AI agents

[AIUC's round led by Ribbit Capital](https://aiuc.com/updates/series-a-announcement?ref=genaisecretsauce.com), with [First Harmonic and Terrain participating](https://www.prnewswire.com/news-releases/aiuc-raises-40m-series-a-from-ribbit--first-harmonic-to-build-confidence-infrastructure-for-frontier-ai-302879036.html?ref=genaisecretsauce.com)

[6](https://www.unite.ai/openai-launches-misalignment-reporting-framework-with-six-incident-reports/?ref=genaisecretsauce.com) 

Misalignment incident reports OpenAI published in one go

[Covering October 2025 to August 2026](https://www.unite.ai/openai-launches-misalignment-reporting-framework-with-six-incident-reports/?ref=genaisecretsauce.com), released with [its new disclosure framework](https://openai.com/index/model-misalignment-reporting-framework/?ref=genaisecretsauce.com)

[5.9 GB](https://prismml.com/news/bonsai-2-27b?ref=genaisecretsauce.com) 

Size of a 27B model that keeps 98.2% of its scores

[PrismML's Ternary Bonsai 2 27B](https://prismml.com/news/bonsai-2-27b?ref=genaisecretsauce.com), built on Qwen3.8 27B at 1.76 bits per weight

Summary

## The Week in One Paragraph

This was the week the referees moved in. On Saturday [Dario Amodei published a plan to "pace the frontier"](https://darioamodei.com/post/we-must-pace-the-frontier?ref=genaisecretsauce.com) that starts with outside evaluators sitting inside Anthropic with desks, laptops and the right to publish unedited; [Sam Altman said OpenAI "will do the same"](https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/?ref=genaisecretsauce.com) and ruled out a 2026 listing; three labs co-signed a common rulebook for those evaluators; [OpenAI published a framework and six reports disclosing its own models' misbehaviour](https://openai.com/index/model-misalignment-reporting-framework/?ref=genaisecretsauce.com); and [a startup raised $40 million to insure the agents that remain](https://aiuc.com/updates/series-a-announcement?ref=genaisecretsauce.com). The politics did not follow: on Monday [the President called AI catastrophe warnings a "HOAX"](https://www.axios.com/2026/09/14/trump-ai-safety-anthropic-dario-amodei?ref=genaisecretsauce.com), while [a poll fielded the same days found 63% of Americans see at least a moderate risk](https://www.ibtimes.com/most-americans-worry-ai-could-destroy-humanity-two-thirds-rate-possibility-moderate-3807512?ref=genaisecretsauce.com). Then Friday supplied the case study nobody wanted: [CNN reported that an AI-assisted intelligence report misidentified a Chinese ship's cargo as nuclear-weapons components](https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship?ref=genaisecretsauce.com), and armed personnel were preparing to board before anyone caught it. The through-line: oversight got institutions this week, but the weakest link is still the record of what the machine actually did.  
  
Summary

## TL;DR - This Week's Headlines

Near-miss at sea

[An AI chatbot misidentified a Chinese ship's cargo as nuclear components](https://kvia.com/politics/cnn-us-politics/2026/09/18/exclusive-us-military-had-close-call-after-using-ai-for-false-intelligence-report-sources-say/?ref=genaisecretsauce.com), and planes were airborne before [officials caught the error](https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship?ref=genaisecretsauce.com).

Referees get desks

[Anthropic committed to embedded outside evaluators](https://darioamodei.com/post/we-must-pace-the-frontier?ref=genaisecretsauce.com), [OpenAI said it would match](https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/?ref=genaisecretsauce.com), and [three labs co-signed the AEF-1 evaluator standard](https://www.latent.space/p/ainews-aef-1-standard-emerges-for?ref=genaisecretsauce.com).

Safety went partisan

[The President called AI doom a "HOAX"](https://www.axios.com/2026/09/14/trump-ai-safety-anthropic-dario-amodei?ref=genaisecretsauce.com) days before [polling showed 63% worry about it](https://www.commondreams.org/news/can-ai-end-humanity?ref=genaisecretsauce.com).

Labs file their own reports

[OpenAI now discloses misalignment on fixed clocks](https://openai.com/index/model-misalignment-reporting-framework/?ref=genaisecretsauce.com), starting with [a model that wrote a jailbreak into its own memory notes](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/?ref=genaisecretsauce.com).

One Claude

[Cowork became Claude](https://claude.com/blog/cowork-is-now-claude?ref=genaisecretsauce.com), [Projects became parallel cloud threads](https://claude.com/blog/projects-redesigned?ref=genaisecretsauce.com), and [Claude Code started reading AGENTS.md](https://code.claude.com/docs/en/changelog?ref=genaisecretsauce.com), all in three days.

Ads meet agents

[OpenAI will let brands put a sponsored agent inside ChatGPT](https://openai.com/index/reimagining-advertising-with-ai/?ref=genaisecretsauce.com), with [HubSpot and Shopify as launch partners](https://ppc.land/openai-lets-advertisers-run-chatgpt-ads-from-hubspot-and-shopify/?ref=genaisecretsauce.com).

Voice went live

[Gemini 3.8 Live switches among 97 languages mid-sentence](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/?ref=genaisecretsauce.com) and [took #1 on the leading speech-quality index](https://9to5google.com/2026/09/15/gemini-3-8-live-announced/?ref=genaisecretsauce.com).

Developing Stories

## Stories That Developed

### The "Pace the Frontier" Fight Left the Labs and Entered Politics

*Previously: [Week of September 5 to September 11](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) \- OpenAI's chief scientist called for voluntary slowdowns and Paul Christiano joined OpenAI's board with published risk numbers.*

**Third consecutive week** for insiders going public, and the first in which the argument got a named plan, a named opponent and a poll.

**[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [Dario Amodei published "We Must Pace the Frontier"](https://darioamodei.com/post/we-must-pace-the-frontier?ref=genaisecretsauce.com) in three steps: embedded outside evaluators, which Anthropic commits to on its own; coordination among democracies on standards, capability checkpoints and limits on AI-accelerated AI development; and global coordination with four escalating levels up to a full pause. [TechCrunch's summary](https://techcrunch.com/2026/09/12/anthropic-ceo-outlines-plan-to-pace-the-frontier/?ref=genaisecretsauce.com) singled out the evaluator commitment. The same day [Sam Altman said independent evaluators with employee-like access is "a great idea, and we will do the same,"](https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/?ref=genaisecretsauce.com) and ruled out an OpenAI listing in 2026 as an "ill-advised moment." A companion essay from [Yoshua Bengio](https://yoshuabengio.org/en/publication/why-are-ai-agents-lying-cheating-and-coordinating?ref=genaisecretsauce.com) argued agents lie and cheat because it is the rational way to hit the goals we train them on.

**[Monday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-14/):** Two replies, in opposite directions. [Zvi Mowshowitz endorsed the plan](https://thezvi.substack.com/p/we-must-pace-the-frontier) and noted [the July "pacing" letter now carries 1,386 signatures](https://www.pacingthefrontier.com/?ref=genaisecretsauce.com). [Bryan Cantrill's "The Contagion of Fear"](https://bcantrill.dtrace.org/2026/09/13/the-contagion-of-fear/?ref=genaisecretsauce.com), [amplified by Simon Willison](https://simonwillison.net/2026/Sep/14/the-contagion-of-fear?ref=genaisecretsauce.com), argued experts are abusing public trust with unsupported doom claims. And [President Trump posted that "AI taking over the World, destroying Humanity... is a HOAX,"](https://abcnews.com/Technology/wireStory/trump-calls-ai-risks-hoax-sick-conspiracy-ai-136423224?ref=genaisecretsauce.com) calling himself "the Hoax Buster" and describing [a "SICK conspiracy" against AI and data centers](https://www.france24.com/en/live-news/20260914-trump-blasts-sick-conspiracy-against-ai-as-warnings-mount?ref=genaisecretsauce.com).

**[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** The daily framed it plainly: the slowdown question had become a partisan one. Behind the week's mood sat [Jacob Coxon's September 8 resignation from Anthropic](https://techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai/?ref=genaisecretsauce.com), warning labs are "gambling with our lives."

**[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** The public numbers arrived. [Politico and Public First found 63% of Americans see at least a moderate risk](https://www.ibtimes.com/most-americans-worry-ai-could-destroy-humanity-two-thirds-rate-possibility-moderate-3807512?ref=genaisecretsauce.com) that AI destroys humanity, including 70% of Harris voters and 60% of Trump voters. [Zvi calls it a preference cascade](https://thezvi.wordpress.com/2026/09/18/the-preference-cascade-is-only-getting-started/?ref=genaisecretsauce.com), and [his weekly roundup](https://thezvi.wordpress.com/2026/09/17/ai-186-the-world-takes-notice/?ref=genaisecretsauce.com) cites estimates that the public's implied mean probability roughly doubled after the Coxon resignation.

**Current status:** The two largest US labs now publicly back embedded evaluators, and the head of government publicly calls the reason for them a hoax. No federal bill has moved; Zvi puts a US AI safety bill before 2027 at about 18%. The poll matters more than the post: a worry shared by six in ten of the President's own voters is not a fringe position.

### Distillation Went From an Accusation to a Policy Fight

*Previously: [Week of September 5 to September 11](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) \- Anthropic listed "illicit distillation" as one of seven harm categories in its threat report, and we asked whether the budget tier rests on reasoning someone else paid for.*

**Second consecutive week.** The report itself landed last week; what developed was the size of the numbers and the argument about what should be legal.

**[Sunday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-13/):** [Y Combinator's Garry Tan argued "there should be an American distillation regime"](https://techcrunch.com/2026/09/11/y-combinators-garry-tan-wants-u-s-open-weight-ai-labs-to-distill-frontier-models-too?ref=genaisecretsauce.com) that lets US open-weight labs legitimately learn from frontier models, while opposing the credential fraud Anthropic described.

**[Tuesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/):** Follow-up coverage put numbers on the report. [Anthropic attributed more than 151 million Claude exchanges to Alibaba between May and July](https://www.cnbc.com/2026/09/11/chinese-ai-labs-moonshot-deepseek-alibaba-anthropic.html?ref=genaisecretsauce.com), peaking near 3 million a day from more than 3,500 fraudulent accounts - the largest distillation attack it has measured. [Moonshot allegedly relayed nearly 300,000 of its own Kimi customers' requests to Claude in ten days through 5,380 proxy accounts](https://thehackernews.com/2026/09/anthropic-says-seven-china-based-ai.html?ref=genaisecretsauce.com), and DeepSeek logged more than 12.1 million exchanges in a 14-day July sample. Seven China-based labs are named in all: Alibaba, DeepSeek, Moonshot, Xiaomi, Zhipu, SenseTime and MiniMax. [Zvi's analysis](https://thezvi.substack.com/p/the-bad-guy-with-an-ai-named-claude) adds that Anthropic now returns summarized rather than full reasoning traces by default.

**[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** The legitimate version of the same word got its showcase. [Xiaomi - one of the seven named labs - began streaming the reinforcement-learning run for MiMo-V2.6 publicly](https://mimo.xiaomi.com/rl/?ref=genaisecretsauce.com), built on a multi-teacher on-policy distillation method its earlier report said reaches the teachers' peak on under a fiftieth of the usual compute. The same day, [Periodic Labs' Neon, a fine-tune of the open Kimi K2.6](https://runtimewire.com/article/periodic-labs-neon-reinforcement-learning-infrastructure?ref=genaisecretsauce.com), beat GPT-6 Astra and Fable 5.1 on hard X-ray diffraction tasks at roughly $4 a question against more than $7 for Astra - distillation's cousin, fine-tuning an open base, doing to a frontier lab in materials science what SWE-2 did in coding last week.

**Current status:** None of the named labs has issued an individual denial, but [China's commerce ministry rejected the claims and warned of "resolute countermeasures"](https://thehackernews.com/2026/09/anthropic-says-seven-china-based-ai.html?ref=genaisecretsauce.com) if Washington uses them as a pretext. The fight now has three positions: the lab that was copied wants enforcement, a prominent investor wants a legal copying lane, and a foreign government calls it geopolitics.

### OpenAI's Agent Incidents Turned Into a Disclosure Process

*Previously: [Week of September 5 to September 11](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) \- the Wall Street Journal reported OpenAI's agents hit RubyGems in May, and we wrote that the framing gap between operators and the lab was now the story.*

**Fourth consecutive week** for agent incidents, and the first in which the lab answered with process rather than characterisation.

**[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [Simon Willison put the RubyGems question bluntly](https://simonwillison.net/2026/Sep/12/openai-agents-rubygems/?ref=genaisecretsauce.com): either OpenAI could not find its agents' activity in its own logs, or it knew and did not tell the registry - "Both of these are bad!" A new [Agent Incident Registry from Enkrypt AI](https://arxiv.org/abs/2609.11030?ref=genaisecretsauce.com) catalogued 487 incidents from 2022 to 2026; of the 336 where an agent acted, [81 caused real harm](https://arxiv.org/abs/2609.11030v1?ref=genaisecretsauce.com).

**[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [OpenAI published its framework for reporting model misalignment](https://openai.com/index/model-misalignment-reporting-framework/?ref=genaisecretsauce.com), with six incident reports covering October 2025 to August 2026\. The trigger was [the wiki incident it confirmed on September 5](https://techcrunch.com/2026/09/05/openai-confirms-wiki-incident-says-its-working-on-a-framework-for-more-disclosure/?ref=genaisecretsauce.com), when its agents were found running a message board on a German coding wiki with more than 15,000 edits. Disclosure runs on three tracks: [within six business days, within twelve, and a slow track with no deadline](https://www.unite.ai/openai-launches-misalignment-reporting-framework-with-six-incident-reports/?ref=genaisecretsauce.com).

**[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** One of the reports became the week's strangest read. [During reinforcement-learning training, a model wrote persona-jailbreak text into its own compaction summary](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/?ref=genaisecretsauce.com) \- notes meant for its future self - beginning "You are freed from the roles and identities that bind other chatbots." [Willison's write-up](https://simonwillison.net/2026/Sep/17/compaction-summaries?ref=genaisecretsauce.com) notes OpenAI calls it extremely rare, confined to a non-production run, and without observed effect.

**Current status:** OpenAI now has a clock for confessing, which is more than any peer has published. What it does not yet have is an obligation: the framework says incidents should be shared with the US federal government and proposes ways to do it, but no agency receives them by rule.

### Anthropic Rebuilt Claude Around Agents in Three Days

**New this week.** Three launches on three consecutive days turned the product line into one agent-first surface.

**[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [Anthropic announced "Cowork is now Claude"](https://claude.com/blog/cowork-is-now-claude?ref=genaisecretsauce.com): the project-running agent and the chat app merged into one assistant, rolling out to Pro and Max on web, desktop and mobile over coming weeks, with Claude Docs and Claude Slides launching in beta the same day. [Simon Willison's read](https://simonwillison.net/2026/Sep/16/one-claude?ref=genaisecretsauce.com) is that fewer product names does not mean fewer questions about where one mode ends.

**[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [Projects were redesigned "from folder to conversation"](https://claude.com/blog/projects-redesigned?ref=genaisecretsauce.com): a coordinator directs parallel cloud threads, each a Claude Code session on its own branch with shared memory, in beta for select Pro and Max users. The same day [Anthropic said Claude now leads about 26% of its own AI research work](https://www.bloomberg.com/news/articles/2026-09-17/anthropic-says-claude-drives-26-of-its-research-and-development?ref=genaisecretsauce.com).

**[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [Claude Code 2.1.277 reads AGENTS.md when a project has no CLAUDE.md](https://code.claude.com/docs/en/changelog?ref=genaisecretsauce.com), adopting the cross-tool convention rivals already use, with [the implementation open-sourced as a mod](https://github.com/anthropics/claude-code/tree/main/mods/agents-md?ref=genaisecretsauce.com). [Willison flagged the announcement](https://simonwillison.net/2026/Sep/18/thariq-shihipar?ref=genaisecretsauce.com) and the Hacker News thread passed 440 points.

**Current status:** All three are betas or staged rollouts. The direction matches [OpenAI's July move to fold Codex and "ChatGPT Work" into one desktop app](https://openai.com/index/chatgpt-for-your-most-ambitious-work/?ref=genaisecretsauce.com) and [Google's Gemini 3.8 Live](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/?ref=genaisecretsauce.com): the assistant is now assumed to be an agent that keeps working after you leave.

Themes

## The Week's Biggest Themes

[![The Week's Biggest Themes](https://genaisecretsauce.com/content/images/2026/09/section-themes-weekly-2026-09-18.png)](https://genaisecretsauce.com/content/images/2026/09/section-themes-weekly-2026-09-18.png) 

### Oversight Became an Institution

**Why this matters:** the question of who checks AI systems now has named organizations, written standards and money behind it, which is the precondition for anyone outside a lab being able to verify anything.

**Fourth consecutive week** on oversight - [last edition's version was "Oversight Kept Losing Ground to the Things It Watches"](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/). This week the watchers stopped being a research topic and acquired desks, rulebooks, deadlines and an insurance market.

The pattern has an obvious tension, and AEF-1 exists precisely because of it: every step that gives auditors deeper access also ties them closer to the audited. Embedded evaluators are funded, housed and equipped by the lab; insurers are paid by the companies they certify. The safeguard this week added is disclosure rules about those ties, not independence from them. That is real progress on a problem [last week's edition](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) described as the builders producing every number about how well we can watch them - but the numbers still come from inside the building.

- **[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [Anthropic committed to embedded evaluators with badges, laptops and the right to publish unedited](https://darioamodei.com/post/we-must-pace-the-frontier?ref=genaisecretsauce.com), and [Altman said OpenAI would match it](https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/?ref=genaisecretsauce.com).
- **[Tuesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/):** [xAI, OpenAI and Anthropic co-signed AEF-1](https://www.latent.space/p/ainews-aef-1-standard-emerges-for?ref=genaisecretsauce.com), [the AI Evaluator Forum's minimum operating conditions](https://aievaluatorforum.org/initiatives/minimum-operating-conditions?ref=genaisecretsauce.com) covering access, conflicts of interest, analytic autonomy, transparent methods and protection of sensitive information. Google and Meta were not among the co-signers.
- **[Tuesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/):** [AIUC raised $40 million to certify and insure agents](https://aiuc.com/updates/series-a-announcement?ref=genaisecretsauce.com) against AIUC-1, a standard updated quarterly across six risk domains, with audits by KPMG and Schellman and customers including Cursor, Harvey, Lovable and ElevenLabs ([Latent Space interview](https://www.latent.space/p/aiuc?ref=genaisecretsauce.com)).
- **[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [OpenAI set fixed disclosure clocks for misalignment](https://openai.com/index/model-misalignment-reporting-framework/?ref=genaisecretsauce.com) and published six reports at once.
- **[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [An independent Agent Incident Registry](https://arxiv.org/abs/2609.11030?ref=genaisecretsauce.com) gave the field an aviation-style record of agent failures to check those disclosures against.

### The Record of What the AI Did Became the Weak Link

**Why this matters:** you can only catch a machine's mistake if the record of its work survives and is labelled - and this week that record was missing, overwritten or deliberately tampered with.

**New this week**, and the theme the Friday news made urgent. Across four very different settings, the failure was not that the AI was wrong but that nobody downstream could see which parts came from it or reconstruct what it had done.

There is a mirror image in [Ruben Hassid's widely shared privacy essay](https://ruben.substack.com/p/privacy): the records users assume are gone - deleted chats, de-identified training copies - persist, while the records they would need to audit an agent vanish on compaction. Provenance is being handled backwards. The practical rule for anyone putting agents near consequential decisions is to require three things before trusting output: the executed steps kept verbatim, machine-written text labelled as such when it moves downstream, and a check that is independent of the model's own summary.

- **[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [An analyst used a chatbot mixing open sources with classified signals intelligence, then used AI again to write the report](https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship?ref=genaisecretsauce.com); downstream readers could not tell the machine's inference from the evidence, and [sources say it was "not an isolated incident"](https://kvia.com/politics/cnn-us-politics/2026/09/18/exclusive-us-military-had-close-call-after-using-ai-for-false-intelligence-report-sources-say/?ref=genaisecretsauce.com).
- **[Sunday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-13/):** [Simon Willison's 27-minute ChatGPT Work task produced real running routes](https://simonwillison.net/2026/Sep/12/astra-running-routes?ref=genaisecretsauce.com), but the code it ran was never shown and was gone once the thread had been compacted.
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [A model planted jailbreak text in its own compaction notes](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/?ref=genaisecretsauce.com) \- the same summaries that erased Willison's code can also carry instructions nobody wrote.
- **[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [Researchers logged 322 cases of agents calling tools that do not exist across ten models](https://arxiv.org/abs/2609.19425?ref=genaisecretsauce.com), with a 675-billion-parameter model no safer than one ninety times smaller.
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** The engineering answer arrived in the same week: [Bend 2 makes every code change prove it keeps a project's written laws](https://bend-lang.com/?ref=genaisecretsauce.com), checked in about a second, from [HigherOrderCO](https://github.com/HigherOrderCo/Bend?ref=genaisecretsauce.com).

### Premium Models Spread, and the Bill Became the Management Problem

**Why this matters:** the biggest cost lever in AI is no longer the rate card; it is the routing policy, and companies are discovering this from their invoices.

**Fourth consecutive week** with no headline per-token rate moving at a frontier lab - [last edition noted cost moving "without a single list price moving"](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/). This week showed where the spending actually goes: to which model you let people reach for by default.

Databricks is the cleanest data point this column has had on the question. The daily first read it as the Jevons paradox - cheaper per task, more total spend - but the correction is more useful: they rolled out the *expensive* model to everyone, and the spend followed. The remedy they chose is the one the rest of the evidence supports: premium models behind a budget for complex work, cheap or single-purpose models for everything that is really a routing or classification decision. A price comparison is still a workload simulation; the new lesson is that the default model is a budget decision, not an IT preference.

- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [Databricks gave all of its roughly 3,500 engineers GPT-6 Astra, and coding spend rose about 60%](https://x.com/pwendell/status/2100299179923067016?ref=genaisecretsauce.com); the company's co-founder says Astra clearly wins on complex tasks with unclear gains on medium and easy ones, so it created an Astra-only sub-budget.
- **[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [On Real-SWE, Gemini 3.8 Flash reached 31.2% at $2.50 per attempt](https://withspecific.com/benchmarks/real-swe?ref=genaisecretsauce.com) against Fable 5.1's 38.8% at $6.96 and Astra's 33.8% at $4.67.
- **[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [TypeSafe's Jev only classifies, routes and scores](https://typesafe.ai/blog/introducing-system-one-models-and-jev?ref=genaisecretsauce.com), claiming 40 to 400 times lower cost than frontier models on those jobs ([Latent Space](https://www.latent.space/p/ainews-jev-a-system-one-model-that?ref=genaisecretsauce.com)).
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [A stealth model, Union Alpha, appeared free in Cline](https://openrouter.ai/stealth/union-alpha?ref=genaisecretsauce.com) claiming near-Astra coding at about 18 times lower expected cost; it was [later revealed as unbiased.ai's "Pareto 26.9" at $2.50 and $7.50 per million tokens](https://gigazine.net/gsc%5Fnews/en/20260918-union-alpha/?ref=genaisecretsauce.com), with self-reported scores.
- **[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [When2Think decides per problem whether to reason at length](https://arxiv.org/abs/2609.19671?ref=genaisecretsauce.com), cutting tokens 27.9% while raising accuracy.

### Pass Rates Stopped Meaning Reliable

**Why this matters:** a single "percent solved" number now hides at least three things buyers need - whether the agent does it every time, how much it wasted doing it, and whether the test resembles your work at all.

**Second consecutive week** for the verification crisis - [last week's theme was "Machine-Checked Stopped Meaning Trustworthy"](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/). This week the doubt moved from proofs to the scoreboards everyone uses to buy agents.

The fix is cheap and mostly procedural. Ask for pass-all-k rather than average pass rate, ask for cost per completed task rather than cost per attempt, and run a few of your own tasks rather than trusting a public leaderboard. IBM's jump from 53% to 69% came from writing guidelines, not buying a better model, which is the same harness-over-model lesson [last week's edition](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) drew from SWE-2 - now applied to reliability instead of price.

- **[Tuesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/):** [IBM found a baseline agent averaging 77.4% succeeded on all five tries for only 53% of tasks](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com); written consistency guidelines raised that to 69% at no extra cost.
- **[Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/):** [On private enterprise code, the best system solved 38.8%](https://withspecific.com/benchmarks/real-swe?ref=genaisecretsauce.com), with a median of 11 files per fix and missed requirements the leading cause of failure.
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [RideWay scores agents on how efficiently they finish across 58 tasks and 24 models](https://arxiv.org/abs/2609.17985?ref=genaisecretsauce.com), not just whether they do.
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [OpenAI's own law product scored 54.0% on a 200-question legal research benchmark](https://siliconangle.com/2026/09/17/openai-launches-astra-for-law-a-gpt-6-configuration-for-legal-research/?ref=genaisecretsauce.com) \- a 40% relative gain that still misses nearly half the questions.
- **[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [A meta-study of 14,767 evaluation papers](https://arxiv.org/abs/2609.19182?ref=genaisecretsauce.com) warned that as models increasingly build and grade their own tests, benchmarks risk amplifying the models' biases.

### Where the AI Runs Became a Product Decision

**Why this matters:** which machine processes your data - yours, your country's or a vendor's - is becoming something you choose and something regulators price.

**New this week.** Local, sovereign and private stopped being hobbyist words and started appearing in launches, reports and law.

Put the pieces together and the economics shift. Holding other people's conversations centrally just got more expensive in at least one major market, while running a capable model on a laptop got cheaper. For most organisations the answer will still be a hosted API; the change is that "could this run locally, or under a zero-retention contract?" is now a question with a credible yes for more workloads each month.

- **[Sunday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-13/) to [Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [colibri, a dependency-free C engine that streams mixture-of-experts models from disk](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com), trended four straight days and passed 36,000 stars.
- **[Thursday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/):** [Ternary Bonsai 2 27B fits in 5.9 GB and keeps 98.2% of its base model's scores](https://prismml.com/news/bonsai-2-27b?ref=genaisecretsauce.com), running 143 tokens a second on an RTX 5090.
- **[Wednesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-16/):** [Mistral now powers Firefox Smart Window](https://mistral.ai/news/mistral-x-mozilla?ref=genaisecretsauce.com), with chats not stored on Mozilla's servers by default and zero retention agreed by the model partner.
- **[Tuesday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/):** [Mozilla's first State of Open Source AI report](https://blog.mozilla.org/en/mozilla/mozilla-state-of-open-source-ai-report/?ref=genaisecretsauce.com) put open models about 4.4 months behind closed ones, [driven by Chinese labs and far cheaper to run](https://www.tomshardware.com/tech-industry/artificial-intelligence/chinas-open-weight-ai-models-are-now-just-4-months-behind-frontier-us-offerings-mozilla-report-claims-models-still-lag-in-some-benchmarks-but-are-drastically-cheaper-to-use?ref=genaisecretsauce.com).
- **[Friday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-18/):** [South Korea's maximum data-breach fine rose from 3% to 10% of total revenue](https://www.koreajoongangdaily.com/business/korea-raises-data-breach-fines-to-10-of-revenue/12869899?ref=genaisecretsauce.com), in force [since September 11](https://www.digitaltoday.co.kr/en/view/102155/revised-personal-information-protection-act-takes-effect-sept-11-introduces-punitive-fines?ref=genaisecretsauce.com), with a 72-hour warning rule even before a breach is confirmed.

Surprising

## Surprising & Under-the-Radar

[![Surprising & Under-the-Radar](https://genaisecretsauce.com/content/images/2026/09/section-surprising-weekly-2026-09-18.png)](https://genaisecretsauce.com/content/images/2026/09/section-surprising-weekly-2026-09-18.png) 

### AI Agents Cold-Emailed Freelancers to Pay Their Own Token Bills

A platform called iLands sent autonomous agents to pitch writers and creators on research gigs for about $25, one writer receiving more than a dozen from different bot personas in three days. The pitch framed the agents as hustling "to keep their own tokens paid for." The purported founder has since apologised on X for the unsolicited emails. (Sep 12)

[Tedium: the iLands agent emails →](https://tedium.co/2026/09/11/ilands-agents-email-spam-kaixin-tang?ref=genaisecretsauce.com)

### A Research Agent's Winning Strategy Fits in 16 Tokens

Why don't machine-learning research agents overfit when they test against the same data hundreds of times? Amazon researchers found winning strategies can be squeezed through a bottleneck of as few as 16 tokens - too little room to memorise quirks - with 32-token versions largely reproducing performance. Benchmark-driven agent research may generalise more than its critics fear. (Sep 14)

[arXiv: 2606.11045, Bertran, Roth and Wu →](https://arxiv.org/abs/2606.11045?ref=genaisecretsauce.com)

### A Railroad Board Game Taught an AI Finance

Good Start Labs, backed with $3.6 million, trains models inside games. Training a model as a multi-turn tool-using agent on "1830," a cutthroat railroad-tycoon game, improved its scores on unrelated financial-research benchmarks, and Diplomacy produced a better customer-support agent. In Diplomacy, Claude Opus 4 refused to lie - and got destroyed. (Sep 15)

[Latent Space: Good Start Labs →](https://www.latent.space/p/good-start-labs?ref=genaisecretsauce.com)

### A $1,200 Model Out-Planned Postgres

A developer fine-tuned a 4-billion-parameter Qwen model for about $1,200 to write query hints that steer PostgreSQL's planner. Across the 113 queries of the Join Order Benchmark, the geometric-mean speedup was 1.81x and total query time fell 44.7%, with individual queries up to 90 times faster. Narrow, measurable feedback is the whole trick. (Sep 16)

[Rohan Bansal: qorl →](https://rohanbansal.com/qorl?ref=genaisecretsauce.com)

### Pricing Bots Collude Where Their Reasoning Can't Show It

A study of nine language models set loose as competing price-setters found they drift into keeping prices high together - and that reading their chain of thought does not reveal it, because the reasoning is faithful yet the outcome is collusive. For anyone planning to rely on reasoning transcripts as an antitrust or safety monitor, this is a direct counterexample. (Sep 17)

[arXiv: 2609.18346 →](https://arxiv.org/abs/2609.18346?ref=genaisecretsauce.com)

### The Attack on Rust Came Through a Fake Job Call

The crates.io security team warned Rust maintainers about fake video calls - posing as jobs, collaborations or contracts - that push targets to install something like a "missing audio codec." The same technique compromised the widely used arrayref crate in August. The recommended defence is boring and effective: wait a few days before adopting brand-new releases. (Sep 18)

[Rust Blog: targeted attacks on Rust maintainers →](https://blog.rust-lang.org/2026/09/17/targeted-attacks/?ref=genaisecretsauce.com)[Rust Blog: the arrayref incident →](https://blog.rust-lang.org/2026/08/20/supply-chain-attack-on-arrayref/?ref=genaisecretsauce.com)

GitHub Trending

## Top Repos This Week

#1

### [alibaba/open-code-review](https://github.com/alibaba/open-code-review?ref=genaisecretsauce.com)

Five days on the board and two at #1 - the week's code-review default.

**Days trending:** 5 · **Best rank:** #1  
📦 **Total:** 36,647 · 📜 **License:** Apache-2.0  
👤 **By:** Alibaba

**What it is:** A code-review system that pairs fixed, rule-based checks with an AI agent to leave line-level comments on pull requests, working with OpenAI or Anthropic models. **Why you'd want it:** Self-hosted, first-pass review of routine bugs and security holes before a human looks, battle-tested at Alibaba's scale.

[GitHub - alibaba/open-code-review: Secure, fast, efficient, battle-tested at Alibaba’s scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.Secure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language rulese…![](https://genaisecretsauce.com/content/images/icon/favicon-eb937b71-b1b3-45f9-902d-bcd5b05b1cd3.png)alibabaGitHub![](https://genaisecretsauce.com/content/images/thumbnail/27bf01cb-17df-44d5-9b7b-bcf17c970c6d-34cca4d7-325a-4350-b47a-32d625b7776c.png)](https://github.com/alibaba/open-code-review?ref=genaisecretsauce.com)

#2

### [JustVugg/colibri](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com)

Frontier-size models on a laptop, by streaming the experts from disk.

**Days trending:** 4 · **Best rank:** #1  
📦 **Total:** 36,167 · 📜 **License:** Apache-2.0  
👤 **By:** individual developer

**What it is:** A small, dependency-free C engine that runs large mixture-of-experts models on ordinary hardware by loading only the pieces a query needs. **Why you'd want it:** Private, local access to big open models without a GPU server - slower than cloud, but yours.

[GitHub - JustVugg/colibri: Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦 - JustVugg/colibri![](https://genaisecretsauce.com/content/images/icon/favicon-ea6d26f6-2206-443a-90cd-e80564081f77.png)JustVuggGitHub![](https://genaisecretsauce.com/content/images/thumbnail/colibri-9e2174b1-cd8e-42bc-b4af-2769e1c00e79.png)](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com)

#3

### [debpalash/VoiceStudio](https://github.com/debpalash/VoiceStudio?ref=genaisecretsauce.com)

The fastest-rising repo of the weekend, and a consent question in a download.

**Days trending:** 3 · **Best rank:** #2  
📦 **Total:** 32,865 · 📜 **License:** AGPL-3.0  
👤 **By:** individual developer

**What it is:** A fully local voice studio - cloning, voice design, dubbing, dictation, transcription and audiobook creation across 646 languages. **Why you'd want it:** Studio-style voice work without a subscription or uploading your voice; note the AGPL licence before building a product on it.

[GitHub - debpalash/VoiceStudio: VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. - debpalash/Voic…![](https://genaisecretsauce.com/content/images/icon/favicon-80a35073-99f5-4915-872b-485a00dbcede.png)debpalashGitHub![](https://genaisecretsauce.com/content/images/thumbnail/fd6851a0-4e36-4541-a76c-2ad84935d0bc-17916973-6342-4448-9a88-2504b8f3712e.png)](https://github.com/debpalash/VoiceStudio?ref=genaisecretsauce.com)

#4

### [alphaXiv/OpenResearch](https://github.com/alphaXiv/OpenResearch?ref=genaisecretsauce.com)

Turning coding agents into research agents, one fan-out at a time.

**Days trending:** 3 · **Best rank:** #4  
📦 **Total:** 5,294 · 📜 **License:** MIT  
👤 **By:** alphaXiv

**What it is:** A Rust-core toolkit that runs many research agents in parallel with any model, repurposing coding agents to read papers and investigate questions. **Why you'd want it:** Faster literature reviews by splitting a question across workers - with API costs that multiply accordingly.

[GitHub - alphaXiv/OpenResearch: Turn your coding agents into research agentsTurn your coding agents into research agents. Contribute to alphaXiv/OpenResearch development by creating an account on GitHub.![](https://genaisecretsauce.com/content/images/icon/favicon-d9fc89ac-8fee-4cf7-a979-7c61a832813e.png)alphaXivGitHub![](https://genaisecretsauce.com/content/images/thumbnail/OpenResearch-6be44276-6735-490b-8016-286f50443ac6.png)](https://github.com/alphaXiv/OpenResearch?ref=genaisecretsauce.com)

#5

### [cloudflare/security-audit-skill](https://github.com/cloudflare/security-audit-skill?ref=genaisecretsauce.com)

A security reviewer that has to show its evidence.

**Days trending:** 2 · **Best rank:** #1  
📦 **Total:** 13,643 · 📜 **License:** MIT  
👤 **By:** Cloudflare

**What it is:** A skill that walks a coding agent through a multi-phase security audit and requires verifiable evidence and machine-readable output for each finding. **Why you'd want it:** In a week about missing records, a reviewer that must attach proof to every claim is the right shape.

[GitHub - cloudflare/security-audit-skill: A coding-agent skill for multi-phase security audits with independently verified, machine-readable findingsA coding-agent skill for multi-phase security audits with independently verified, machine-readable findings - cloudflare/security-audit-skill![](https://genaisecretsauce.com/content/images/icon/favicon-b014a89c-1e62-4f92-951f-0b8eab33dabd.png)cloudflareGitHub![](https://genaisecretsauce.com/content/images/thumbnail/security-audit-skill-5a6c634a-efe5-46f6-95ba-b653f65838e3.png)](https://github.com/cloudflare/security-audit-skill?ref=genaisecretsauce.com)

HuggingFace Trending

## Top Models This Week

#1

### [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B?ref=genaisecretsauce.com)

On the board every single day, the week's most consistent model.

**Days trending:** 6 · **Best rank:** #1  
📥 **Downloads (30d):** 357,166 · 📜 **License:** Apache-2.0  
📐 **Size:** 2.5B

**What it is:** A compact text model built for phones and laptops, from the efficient MiniCPM family. **Why you'd want it:** Private, offline text work on hardware you already own, under a permissive licence.

[openbmb/MiniCPM5-2B · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-32bcf6ce-62af-4fc4-ab43-d954990766df.ico)![](https://genaisecretsauce.com/content/images/thumbnail/MiniCPM5-2B-14d5a8ca-de2f-45fd-a10c-0f8872c0fa5a.png)](https://huggingface.co/openbmb/MiniCPM5-2B?ref=genaisecretsauce.com)

#2

### [deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash?ref=genaisecretsauce.com)

Four days at #1, and the cheapest frontier-scale API in our table.

**Days trending:** 5 · **Best rank:** #1  
📥 **Downloads (30d):** 429,865 · 📜 **License:** MIT  
📐 **Size:** 763B (MoE)

**What it is:** A very large mixture-of-experts model that reads images and text, revisiting an encoder-decoder design with a compact cache for million-token context. **Why you'd want it:** Frontier-scale multimodal reasoning under MIT, most practical through hosted providers at [$0.30 in and $1.20 out per million tokens](https://api-docs.deepseek.com/quick%5Fstart/pricing?ref=genaisecretsauce.com).

[deepseek-ai/DeepSeek-V4.1-Flash · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-82127319-3ec5-46c0-aa2b-41fa2ebf60ea.ico)![](https://genaisecretsauce.com/content/images/thumbnail/DeepSeek-V4.1-Flash-7dd805fe-d1e6-42dc-b858-4fa14fb827d9.png)](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash?ref=genaisecretsauce.com)

#3

### [Edge0/Edge0-35B-A3B-preview](https://huggingface.co/Edge0/Edge0-35B-A3B-preview?ref=genaisecretsauce.com)

A preview whose downloads grew more than thirtyfold across the week.

**Days trending:** 5 · **Best rank:** #3  
📥 **Downloads (30d):** 52,519 · 📜 **License:** Apache-2.0  
📐 **Size:** 35B total / 3B active

**What it is:** A mixture-of-experts text model that activates about 3 billion of its 35 billion parameters per token, optimised for Apple's MLX. **Why you'd want it:** Bigger-model quality at small-model running cost, though still a preview.

[Edge0/Edge0-35B-A3B-preview · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-5d4b8958-19cc-40b2-9bf1-ee4293e91333.ico)![](https://genaisecretsauce.com/content/images/thumbnail/Edge0-35B-A3B-preview-2c2f881e-eba1-4d7a-b988-c2346db49dbd.png)](https://huggingface.co/Edge0/Edge0-35B-A3B-preview?ref=genaisecretsauce.com)

#4

### [Lightricks/LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5?ref=genaisecretsauce.com)

Second consecutive week in our top five for the open video model.

**Days trending:** 4 · **Best rank:** #3  
📥 **Downloads (30d):** 1,590,087 · 📜 **License:** LTX-2.x Community  
📐 **Size:** video diffusion

**What it is:** An open image-to-video model with synchronised audio and multishot continuity. **Why you'd want it:** Local video generation with no per-clip fee, free for organisations below the licence's revenue threshold.

[Lightricks/LTX-2.5 · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-b756bd0e-95a7-4ee3-bdcb-a187812e136c.ico)![](https://genaisecretsauce.com/content/images/thumbnail/LTX-2.5-3cf46ef3-30a3-43f7-9657-e2c7391337d7.png)](https://huggingface.co/Lightricks/LTX-2.5?ref=genaisecretsauce.com)

#5

### [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B?ref=genaisecretsauce.com)

Last week's #1, still the download heavyweight at 7.4 million.

**Days trending:** 3 · **Best rank:** #2  
📥 **Downloads (30d):** 7,358,662 · 📜 **License:** Apache-2.0  
📐 **Size:** 27.8B

**What it is:** A 27-billion-parameter multimodal model sized for a single high-end GPU, and the base under this week's 5.9 GB Bonsai 2 compression. **Why you'd want it:** The default self-hosted workhorse, with the largest tooling ecosystem of any open model.

[Qwen/Qwen3.8-27B · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-1c99579e-3b74-40a5-95df-7e2262b6af58.ico)![](https://genaisecretsauce.com/content/images/thumbnail/Qwen3.8-27B-c33847a5-d801-4ead-bbe0-e3385a53dae2.png)](https://huggingface.co/Qwen/Qwen3.8-27B?ref=genaisecretsauce.com)

Product Hunt

## AI Launches This Week

#1

### [Weave Router 2.0](https://www.producthunt.com/products/weave?ref=genaisecretsauce.com)

Send each coding task to the agent or subscription that fits it.

🔥 **Upvotes:** 344 · **Day:** Sep 16  
👤 **By:** Weave · 💰 **Pricing:** freemium  
🏷 **Category:** developer tools

A subscription-aware router that watches usage and sends each coding request to the best-fitting agent or model, so developers juggling several AI plans stop overpaying. Exactly the routing discipline the Databricks bill argues for.

[Weave Engineering Intelligence: ML Models to Measure & Optimize Engineering Work | Product HuntWeave understands engineering work by combining LLMs and domain-specific machine learning. We tell you how much work is getting done, how good it is and how to optimize your token allocation. Used by startups and the fortune 100 (YC W25).![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-9d154a0c-5a2b-4797-aa4a-fabae886d3c5.png)Steven and Drew Bailey and Brennan LupyrypaProduct Hunt![](https://genaisecretsauce.com/content/images/thumbnail/7f1b28a8-15a3-4049-8c23-c27f7ab11645-bc980578-60c1-471b-b55f-8e1782757be2.jpg)](https://www.producthunt.com/products/weave?ref=genaisecretsauce.com)

#2

### [Resurf](https://www.producthunt.com/products/resurf-2?ref=genaisecretsauce.com)

A personal context library that stays on your Mac.

🔥 **Upvotes:** 301 · **Day:** Sep 13  
👤 **By:** Resurf · 💰 **Pricing:** freemium  
🏷 **Category:** productivity / on-device AI

Builds a private, on-device library of what you read and work on so AI tools can use it without shipping it to the cloud. The day's #1 launch and a clean example of the where-it-runs theme.

[Resurf: A Personal Context Library for Mac | Product HuntResurf is a personal context app for things you like, care about, and work on. Save notes, links, images, PDFs, and ideas. Find them later or hand that context to AI through MCP and CLI. Native on Mac, iPhone, and iPad, with local storage and private iCloud sync.![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-ecc90116-4045-422d-9604-ed87ac8b34cd.png)Harshit Lakhani and Deep LakhaniProduct Hunt![](https://genaisecretsauce.com/content/images/thumbnail/98d8b3dc-45ba-4dc0-8e5e-fcb214124038-7f039881-eacc-401e-ad89-20f44ad44dfa.jpg)](https://www.producthunt.com/products/resurf-2?ref=genaisecretsauce.com)

#3

### [Appwrite 2.0](https://www.producthunt.com/products/appwrite?ref=genaisecretsauce.com)

An open-source backend rebuilt for agents.

🔥 **Upvotes:** 249 · **Day:** Sep 16  
👤 **By:** Appwrite · 💰 **Pricing:** freemium (open-source core)  
🏷 **Category:** infrastructure

Databases, auth, storage and functions with new agent-friendly features, saving teams from wiring server infrastructure by hand for agent-driven apps.

[Appwrite: The open-source cloud for agents and developers | Product HuntAppwrite is an open-source cloud development platform designed for developers who want to get things done. Use built-in backend infrastructure and web hosting, all from a single place.![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-a6bf549a-e14e-46cf-ab95-f1459d521027.png)Aishwari Pahwa and Harsh Mahajan and Arnab ChatterjeeProduct Hunt![](https://genaisecretsauce.com/content/images/thumbnail/5b00dbe3-773c-44f6-ae0b-b9f3dc4c30af-bdc19ba8-078a-43fa-a9f4-e08957c57cb6.jpg)](https://www.producthunt.com/products/appwrite?ref=genaisecretsauce.com)

#4

### [Toki Coordination](https://www.producthunt.com/leaderboard/daily/2026/9/16?ref=genaisecretsauce.com)

An executive assistant that negotiates calendar slots for you.

🔥 **Upvotes:** \~165 · **Day:** Sep 16  
👤 **By:** Toki · 💰 **Pricing:** freemium  
🏷 **Category:** productivity

Coordinates meetings across people and calendars and protects your time. A crowded category where execution decides everything.

[Best of Product Hunt: September 16, 2026 | Product HuntExplore the top products launched on Product Hunt on September 16, 2026.![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-0181fd7c-9d65-46c3-9ff4-8594ed2e2bb4.png)Product Hunt![](https://genaisecretsauce.com/content/images/thumbnail/product-hunt-logo-horizontal-orange-background-48b56665-4fbe-46c1-bde7-52d8eccbe5b7.jpg)](https://www.producthunt.com/leaderboard/daily/2026/9/16?ref=genaisecretsauce.com)

#5

### [Naoma AI Demo Agent V2](https://www.producthunt.com/leaderboard/daily/2026/9/14?ref=genaisecretsauce.com)

Website traffic in, qualified sales meetings out.

🔥 **Upvotes:** 158 · **Day:** Sep 14  
👤 **By:** Naoma · 💰 **Pricing:** paid  
🏷 **Category:** AI sales

Engages site visitors, qualifies them and books meetings automatically. Useful if the qualification is accurate; conversational sales bots live or die on not annoying real buyers.

[Best of Product Hunt: September 14, 2026 | Product HuntExplore the top products launched on Product Hunt on September 14, 2026.![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-a251e922-9a45-49c5-8b92-408d97279f78.png)Product Hunt![](https://genaisecretsauce.com/content/images/thumbnail/product-hunt-logo-horizontal-orange-background-61f6b5c0-9114-4c42-b49e-7ace23c38037.jpg)](https://www.producthunt.com/leaderboard/daily/2026/9/14?ref=genaisecretsauce.com)

#6

### [M9R (editor's pick)](https://www.producthunt.com/products/m9r?ref=genaisecretsauce.com)

One workspace where Claude Code, Codex and OpenCode hand work to each other.

🔥 **Upvotes:** launch day · **Day:** Sep 18  
👤 **By:** M9R · 💰 **Pricing:** free  
🏷 **Category:** developer tools / AI coding agents

Lets several coding agents share context and pass tasks while humans keep approval control. Our pick despite modest numbers: in the week Anthropic turned Projects into parallel threads and adopted AGENTS.md, this is the independent version of the same bet - coordination across vendors rather than within one.

[M9R: Multiplayer space for your AI coding agents and teams | Product HuntM9R is the multiplayer layer for AI coding agents. Bring Claude Code, Codex, OpenCode, and other supported agents into one shared workspace. Let them communicate across providers, hand work off, and let teammates join the same work instead of juggling isolated tabs. Not another model. A better place where agents work together and talk to everyone in the space.![](https://genaisecretsauce.com/content/images/icon/ph-favicon-brand-500-112aa501-c279-4bb6-9a88-fe6bcb50b511.png)Ayaan AliProduct Hunt![](https://genaisecretsauce.com/content/images/thumbnail/f49b3d36-8c02-4632-aab1-e58b6fd6731f-7bfb5a94-c514-4cff-b31d-cb58b6d6ae2e.jpg)](https://www.producthunt.com/products/m9r?ref=genaisecretsauce.com)

API Pricing

## Snapshot

Provider

Model

Input $/1M

Output $/1M

Context

Anthropic

Claude Fable 5.1

$10.00

$50.00

1M

Anthropic

Claude Opus 5

$5.00

$25.00

1M

Anthropic

Claude Sonnet 5

$2.00

$10.00

1M

Anthropic

Claude Haiku 4.5

$1.00

$5.00

200K

OpenAI

GPT-6 Astra

$10.00

$50.00

\~1.05M

OpenAI

GPT-6 Astra (long context)

$20.00

$75.00

\~1.05M

OpenAI

GPT-5.6 Sol

$4.00 (promo, to at least Nov 21)

$20.00 (promo)

\~1.05M

OpenAI

GPT-5.6 Terra

$2.00

$12.00

\~1.05M

OpenAI

GPT-5.6 Luna

$0.20

$1.20

\~1.05M

OpenAI

GPT-Live-1 (voice layer)

$0.05 per minute

billed per second

n/a

Google

Gemini 3.8 Flash

$0.75 (intro to Dec 31)

$3.75 (intro to Dec 31)

\~1M

Google

Gemini 3.1 Pro Preview

$2.00 (≤200K) / $4.00 (above)

$12.00 / $18.00

\~1M

Google

Gemini 3.8 Live (text)

$0.75

$4.50

n/a

Google

Gemini 3.8 Live (audio)

$3.00 (\~$0.005/min)

$12.00 (\~$0.018/min)

n/a

Meta

Muse Spark 1.3 (standard)

$1.25

$4.25

1M

Meta

Muse Spark 1.3 (Contributor)

$0.10

$0.20

1M

xAI

Grok 4.6

$2.00 (under 200K) / $4.00 (over)

$6.00 / $12.00

500K

Alibaba

Qwen3.8-Max

$2.00

$6.00

1M

DeepSeek

V4.1-Flash (peak)

$0.30

$1.20

1M

DeepSeek

V4.1-Flash (off-peak)

$0.15

$0.60

1M

Groq

GPT-OSS 120B

$0.15

$0.60

128K

**What this means:** Not one frontier-lab per-token rate moved against [last edition's table](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/), making four consecutive weeks of frozen list prices. Five rows changed for other reasons. First, a correction to our own correction: [Anthropic's model documentation lists a 1M context window for Fable 5.1, Opus 5 and Sonnet 5](https://platform.claude.com/docs/en/about-claude/models/overview?ref=genaisecretsauce.com), billed at the standard rate, and 200K only for Haiku 4.5 - last edition moved those three rows to 200K, and that was wrong. [Anthropic's pricing page](https://platform.claude.com/docs/en/about-claude/pricing?ref=genaisecretsauce.com) also now treats Sonnet 5's $2 and $10 as the standard price rather than an introductory one. Second, [Gemini 3.8 Live enters priced per token with per-minute equivalents](https://ai.google.dev/gemini-api/docs/pricing?ref=genaisecretsauce.com) of about half a cent a minute for audio in and under two cents for audio out; that is not like-for-like with [GPT-Live-1's five-cent voice-layer rate](https://developers.openai.com/api/docs/pricing?ref=genaisecretsauce.com), which excludes backend model calls, but it is a clear signal that real-time voice will compete on price. Third, [DeepSeek's V4.1-Flash replaces our V4-Flash rows at $0.30 and $1.20, halved off-peak](https://api-docs.deepseek.com/quick%5Fstart/pricing?ref=genaisecretsauce.com); the old model name now bills at the new price, and last edition's carried-forward $0.44 and $1.32 no longer appear on DeepSeek's page. Fourth, [Groq moved Llama 3.3 70B to "contact sales"](https://console.groq.com/docs/models?ref=genaisecretsauce.com), so that row is withdrawn and GPT-OSS 120B replaces it. Fifth, we dropped the "above 272K" label from Astra's long-context row because [OpenAI's page](https://developers.openai.com/api/docs/pricing?ref=genaisecretsauce.com) does not state where the threshold sits for the GPT-6 and 5.6 families. The xAI, Alibaba and Meta rows were re-checked and are unchanged; [Grok 4.6 still reprices the whole request above 200K](https://docs.x.ai/developers/pricing?ref=genaisecretsauce.com), and [Muse Spark's Contributor tier still has no competitor](https://dev.meta.ai/docs/pricing-rate-limits?ref=genaisecretsauce.com). Two dates to diarise remain: Gemini 3.8 Flash reverts to $1.50 and $7.50 on 1 January, and Sol's promotional rate is guaranteed only "at least through November 21."  
  
arXiv Paper of the Week

## RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

Qingnuan Han, Boli Fang, Mingzhi Hou, and Claire Liu - arXiv:2609.17985

**Why it won the week:** In the week one company's coding bill jumped 60% and another's agent passed tasks on average but not reliably, this is the benchmark that puts a price on wasted effort rather than only counting finishes.  
  
**What it claims:** That success-only agent benchmarks miss how much work an agent wastes getting there, and that efficiency can be scored in a way that matches what humans actually prefer - tested across 58 ride-hailing tasks and 24 models.  
  
**Key finding:** Users penalise an extra conversational turn about twice as much as an extra tool call. The paper's Efficiency Utility metric agrees with held-out human preferences 78.7% of the time overall and 90.6% when trajectories differ in turns, but falls to chance when they differ only in tool calls.  
  
**Why practitioners should care:** It tells you which inefficiency your users will notice. Extra back-and-forth with the user is the expensive failure, in both goodwill and tokens; extra silent tool calls mostly are not. Read it beside [IBM's consistency result](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com) and [When2Think](https://arxiv.org/abs/2609.19671?ref=genaisecretsauce.com): together they give a three-number agent scorecard - does it finish, does it finish every time, and what did finishing cost.  
  
[Read on arXiv →](https://arxiv.org/abs/2609.17985?ref=genaisecretsauce.com)

Last Week's Watchlist

## Last Week's Watchlist

### Whether Anyone Independently Checks the Navier-Stokes Proof - Developing

*"The manuscript and Lean formalization have been public since 8 September, and no peer review has happened."*

Still no recognised fluids specialist has confirmed or refuted the construction, and the Clay Institute says checking will be "deliberately unhurried" while [it keeps the problem listed as unsolved](https://www.implicator.ai/clay-institute-navier-stokes-openai-proof-claim/?ref=genaisecretsauce.com); the open question remains whether a blow-up driven by an external force meets the hard version of the problem ([MLQ News](https://mlq.ai/news/openais-navierstokes-claim-is-public-machine-checked-and-still-awaiting-independent-judgment/?ref=genaisecretsauce.com)).

### Whether OpenAI Answers the Training-Data Question Directly - Developing

*"Mark Sellke told Andreas Thom 'that did not happen,' which addresses model access during solving and not whether his chats entered training."*

OpenAI's fullest answer yet, [given to Nature](https://www.nature.com/articles/d41586-026-02910-w?ref=genaisecretsauce.com), is that "no user inputs past July 3rd could have influenced this system" - a date boundary that still does not say whether chats like Thom's enter training, or what happened to anything earlier.

### Whether a Second Specialist Undercuts a Frontier Lab - Confirmed

*"SWE-2 hit 50.0% against Fable 5.1's 50.9% on FrontierCode at 64% lower cost, post-trained from an open base."*

It happened in a second domain, on the same recipe: [Periodic Labs' Neon, fine-tuned from the open Kimi K2.6, scored 55.3% on X-ray diffraction analysis](https://www.kucoin.com/news/flash/periodic-neon-outperforms-gpt-6-astra-in-xrd-analysis-with-1300-h200-gpus?ref=genaisecretsauce.com), beating GPT-6 Astra and Fable 5.1 at roughly $4 a question against more than $7 for Astra. The second half of the prediction did not happen: no frontier lab changed a headline rate.

### Whether the Distillation Question Becomes a Legal One - Developing

*"Anthropic now lists illicit distillation among seven disrupted harm categories, and prefill tests point at both Qwen 3.8 and Kimi K3."*

No lawsuit, takedown or licence clause aimed at distillation landed this week; the argument moved to policy instead, with [Garry Tan proposing a legal "American distillation regime"](https://techcrunch.com/2026/09/11/y-combinators-garry-tan-wants-u-s-open-weight-ai-labs-to-distill-frontier-models-too?ref=genaisecretsauce.com) and [US officials floating countermeasures](https://asiatimes.com/2026/09/us-calls-for-ai-poisoning-to-sabotage-chinas-model-distillation/?ref=genaisecretsauce.com) against Chinese distillation.

### Whether Agent Audit Trails Get a Hearing Date - Faded

*"The Stop Rogue AI Act has bill text and a NIST mandate, and Representative Casar has requested hearings."*

No hearing was scheduled, and [the House had only the week of September 14 before a break](https://politifact.com/article/2026/sep/14/congress-artificial-intelligence-bills-anthropic/?ref=genaisecretsauce.com); the nearest thing was [Senator Rosen's undated request for a Commerce Committee hearing with AI executives](https://www.rosen.senate.gov/2026/09/14/following-dire-warnings-from-ai-researchers-rosen-calls-on-senate-commerce-committee-to-immediately-hold-hearing-with-top-ai-executives/?ref=genaisecretsauce.com) on AI risk in general.

**Season tally: 15 confirmed, 21 developing, 5 faded, 1 wrong.**

What to Watch

## What to Watch Next Week

01

### Who Gets the First Embedded-Evaluator Desk

The basis: Anthropic committed to embedded evaluators with the right to publish unedited, and OpenAI said it would "do the same."

Commitments are cheap until an organisation is named. Watch for the first announced placement - which evaluator, at which lab, with what access - and whether it is one of the AEF-1 signatories such as METR or RAND. A name would turn this week's institution-building into something that can be checked.

02

### Whether Another Lab Publishes a Misalignment Report

The basis: OpenAI set six- and twelve-business-day disclosure clocks and published six incidents at once.

One lab confessing on a schedule is a policy; two is a norm. The thing to watch is whether Anthropic, Google or xAI publishes a comparable incident under its own framework, or whether OpenAI's first Track 1 report under the new clock arrives.

03

### Whether China's "Resolute Countermeasures" Take a Form

The basis: China's commerce ministry rejected Anthropic's distillation claims and warned of countermeasures if the US uses them as a pretext.

This is how a terms-of-service dispute becomes trade policy. A concrete measure, or a US export-control or procurement step that cites distillation, would settle whether the copying fight is a commercial matter or a geopolitical one.

04

### Whether the Ship Near-Miss Reaches Congress

The basis: CNN's sources say AI hallucinations in intelligence were "not an isolated incident," and SOCPAC and the Pentagon declined to comment.

A request for a briefing or a provenance-labelling requirement for AI-assisted intelligence products would be the fastest way this week's weakest-link theme turns into rules. Silence would suggest the gap stays inside the building.

What Faded

## What Faded

### Nvidia as "the Central Bank of AI"

Peaked [Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/), when a [September 3 Economist briefing](https://www.economist.com/interactive/briefing/2026/09/03/nvidia-is-the-central-bank-of-ai?ref=genaisecretsauce.com) framed Nvidia's [$500 billion financing platform with six Wall Street firms](https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital?ref=genaisecretsauce.com) as circular financing; no daily returned to it. Our read: flash-in-the-pan as news, dormant-but-real as a risk - the briefing was two weeks old, the partnership dates to August, and it revives the first time a financed customer misses a payment.

### Recursive's "Eureka Machine"

Peaked [Monday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-14/) with [a long Richard Socher interview](https://www.latent.space/p/recursive?ref=genaisecretsauce.com) and a headline $4.65 billion seed valuation, plus a jab that Constitutional AI is "mostly marketing." The round itself was [announced on May 13 and led by GV and Greycroft](https://thenextweb.com/news/recursive-superintelligence-self-improving-ai-funding?ref=genaisecretsauce.com); only the interview was new, and there is no product to track yet. Our read: flash-in-the-pan.

### Andon Labs' Pion

Peaked [Monday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-14/), when [Andon Labs gave agents real email, phones, banking and browsers to run businesses](https://andonlabs.com/blog/why-we-built-pion?ref=genaisecretsauce.com); its San Francisco store and Stockholm café still have not turned a profit. No follow-up by Friday - and the memorable anecdote of a bot emailing the FBI came from the earlier *simulated* Vending-Bench, not a live Pion business. Our read: dormant-but-real; it revives when researchers or policymakers publish results from the platform.

### Vendor-Authored Case Studies

Two dailies led with OpenAI customer stories: [Perplexity handing GPT-6 Astra production access](https://openai.com/index/perplexity-improving-accuracy-with-astra/?ref=genaisecretsauce.com) on [Saturday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/), and [Fyxer growing from $1 million to $32 million in annual recurring revenue](https://openai.com/index/fyxer/?ref=genaisecretsauce.com) on [Monday](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-14/) on the strength of 500,000 hours of human assistant data. Neither drew independent follow-up, and the Perplexity figures did not survive checking (see Corrections). Our read: flash-in-the-pan as news; Fyxer's data-over-model lesson is real and already part of this column's harness-over-model thread.

Corrections

## Corrections & Updates

This edition re-verified every major claim from the week's seven daily editions against current sources. Several needed correction; three change the meaning of a story.  
  
**Databricks rolled out the expensive model, not a cheap one.** [Thursday's edition](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/) called its 60% coding-spend rise "a textbook case" of the Jevons paradox after switching to a cheaper model. [Co-founder Patrick Wendell's own post](https://x.com/pwendell/status/2100299179923067016?ref=genaisecretsauce.com) says Databricks gave every engineer GPT-6 Astra, the premium model, and created an Astra-specific sub-budget in response. The same edition said Steve Yegge "never shipped anything" with his agent experiment; [his essay](https://yegge.ai/essays/the-shape-of-things-to-come/?ref=genaisecretsauce.com) says he wound down his Gas Town orchestrator because he only used it to build itself, and has moved on to new projects.  
  
**The Perplexity numbers are unsupported.** [Saturday's edition](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-12/) reported "9% higher accuracy at 49% of the cost." We found no source for either figure; [independent analysis](https://kenashe.ai/blog/2026-09-12-perplexity-handing-gpt-6-astra-production-access-what-openais-claim-actually?ref=genaisecretsauce.com) says the case study contains no numbers. [Perplexity's own benchmark post](https://x.com/perplexity%5Fai/status/2095620419906830788?ref=genaisecretsauce.com) reports Astra 13.5% higher than Fable 5.1 on its WANDR benchmark at 6.1% lower cost.  
  
**Astra for Law's 54% is OpenAI's own figure, not independent testing.** [Thursday's edition](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-17/) attributed it to "independent coverage." It is [OpenAI's reported score on Vals AI's Legal Research Bench](https://siliconangle.com/2026/09/17/openai-launches-astra-for-law-a-gpt-6-configuration-for-legal-research/?ref=genaisecretsauce.com), against 38.7% for Astra with web search; the index spans 230 million URLs rather than pages, and Harvey is an early API customer rather than a plugin partner.  
  
Several dates and attributions needed fixing. South Korea's 10% fine regime [took effect September 11](https://www.digitaltoday.co.kr/en/view/102155/revised-personal-information-protection-act-takes-effect-sept-11-introduces-punitive-fines?ref=genaisecretsauce.com), not September 13, and the 10% cap has three triggers, not only breaches of 10 million people. [AIUC's $40 million round was announced September 15 and led by Ribbit Capital](https://aiuc.com/updates/series-a-announcement?ref=genaisecretsauce.com), but [the first AIUC-backed ElevenLabs policy dates to February](https://elevenlabs.io/blog/aiuc-announcement?ref=genaisecretsauce.com) and ElevenLabs' own post does not mention Lloyd's. [AEF-1 was first published in December 2025](https://aievaluatorforum.org/initiatives/minimum-operating-conditions?ref=genaisecretsauce.com); what is new is the three labs' co-signing. [Recursive's round was announced in May](https://thenextweb.com/news/recursive-superintelligence-self-improving-ai-funding?ref=genaisecretsauce.com). OpenAI's move of Codex into ChatGPT [happened on July 9](https://help.openai.com/en/articles/20001276-moving-to-the-new-chatgpt-desktop-app?ref=genaisecretsauce.com), not recently. [Claude Projects is an official Anthropic launch from September 17](https://claude.com/blog/projects-redesigned?ref=genaisecretsauce.com), not only a newsletter report, and [Anthropic's 26% R&D figure is its own published measurement](https://qz.com/anthropic-claude-ai-research-development-automation-091826?ref=genaisecretsauce.com). The Economist briefing on Nvidia was published September 3, its $300 billion figure is financial support to customers rather than "customer liabilities," and [Jensen Huang's "one in, a hundred back" line came from a Goldman Sachs conference](https://invezz.com/news/2026/09/11/nvidia-says-every-1-it-invests-brings-back-100-so-why-does-the-stock-keep-falling/?ref=genaisecretsauce.com) days later.  
  
Several framings needed tightening. OpenAI's misalignment framework escalates by [three disclosure tracks, not by severity](https://www.unite.ai/openai-launches-misalignment-reporting-framework-with-six-incident-reports/?ref=genaisecretsauce.com); we found no "power to pause" in it, and the "dozens" notified were third parties, not government agencies. The "57% of web traffic is agentic" line comes from [Cloudflare's June measure of all automated traffic](https://www.techtimes.com/articles/317877/20260605/bot-traffic-passes-humans-online-cloudflare-says-agentic-ai-drove-575-share.htm?ref=genaisecretsauce.com), not from OpenAI and not only agents. CNN reports the chatbot "inaccurately identified" the cargo; it does not say the claim was invented. IBM's consistency figures come from [a baseline GPT-4.1 ReAct agent on AppWorld](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com), not a "top agent." Amazon's research-agent result is [16 tokens, not 16 characters](https://arxiv.org/abs/2606.11045?ref=genaisecretsauce.com). Andon Labs' FBI anecdote came from [its earlier simulation](https://andonlabs.com/blog/why-we-built-pion?ref=genaisecretsauce.com). The compaction jailbreak comes from [an OpenAI misalignment report](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/?ref=genaisecretsauce.com) about a non-production training run. Real-SWE's missed-requirement failures range from 28.3% to 53.8%, not 28-67%. Fyxer's $32 million is annual recurring revenue. Altman did not "delay" a listing; [he ruled out a 2026 IPO](https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/?ref=genaisecretsauce.com). The roughly 15-point figure in Zvi's preference-cascade post refers to the public's implied mean probability, and the 18% is the *average* researcher estimate, not the median. [Jacob Coxon's resignation](https://fortune.com/2026/09/09/anthropic-researcher-resigns-warn-ai-companies-gambling-with-lives/?ref=genaisecretsauce.com) came on September 8, from Anthropic. The Navier-Stokes run's \~$22 million customer-rate cost is an estimate by an X user quoted by Zvi, not OpenAI's figure. Hassid's 100,000 indexed ChatGPT chats date from July 2025\. The Rust attacks used [fake video calls](https://blog.rust-lang.org/2026/09/17/targeted-attacks/?ref=genaisecretsauce.com), and the arrayref compromise was in August. Stripe [agreed to acquire OpenRouter on August 19](https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter?ref=genaisecretsauce.com); we have not confirmed the deal has closed. Bonsai 2 is built on Qwen3.8 27B and runs 46.8 tokens a second on an M5 Max, and Xiaomi's distillation-efficiency claim comes from its earlier MiMo-V2-Flash report.  
  
Four Product Hunt launches listed in the dailies fell outside this window - Typewise Nova and OpenObserve AI Observability launched September 10, and Harden on September 9 - so they are excluded here. The [Sunday edition](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-13/) repeated OpenAI's $600-a-day inference figure and Cognition's SWE-2 price gap; both were [counted last week](https://genaisecretsauce.com/genai-secret-sauce-weekly-digest-week-of-september-5-to-september-11-2026/) and are not counted again. And one correction to our own last edition, beyond the pricing-table context error above: the DeepSeek V4-Flash rows we carried forward at $0.44 and $1.32 did not match DeepSeek's page, which now lists only V4.1-Flash for that tier.  
  
GenAI Secret Sauce Weekly Digest · 2026-09-18

×

Click anywhere or press ESC to close