> ## Content Index
> Fetch the complete content index at: https://genaisecretsauce.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# GenAI Secret Sauce Daily Digest - 2026-09-15
- URL: https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-15/
- Published: 2026-09-15T23:22:01.000Z
- Updated: 2026-09-15T23:57:39.000Z
- Description: Google's new AI can watch, listen, and talk back in real time - in 97 languages · The big AI labs just agreed on who gets to grade their homework · The "AI misuse" report has a second story: labs allegedly copying each other at scale
- Author: Jasmine Robinson
- Tags: Daily Digest

Watch today's digest as a video summary (generated by NotebookLM)

By the Numbers

## Statistically Speaking

[1](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com) [on the leaderboard](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com) 

Google's new AI can watch, listen, and talk back in real tim

Top Story

[97](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com) [supported languages during a single conversation and](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com) 

Google's new AI can watch, listen, and talk back in real tim

[151](https://www.cnbc.com/2026/09/11/chinese-ai-labs-moonshot-deepseek-alibaba-anthropic.html?ref=genaisecretsauce.com) [million and counting](https://www.cnbc.com/2026/09/11/chinese-ai-labs-moonshot-deepseek-alibaba-anthropic.html?ref=genaisecretsauce.com) 

The "AI misuse" report has a second story

2

launched at 64% cheaper than a frontier

Cheap, small, open models keep crashing the party

53%

to 69% without extra cost

"Passing the test" is quietly the wrong metric for AI agents

One Thing to Tell Your Friends

## One Thing to Tell Your Friends

Chinese AI labs secretly funneled more than 151 million real user conversations through Claude to train their own models - by quietly rerouting their own customers' chats through it without telling them.

Summary

## TL;DR

Top Stories

[Google's new AI can watch, listen, and talk back in real time](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com), [The big AI labs just agreed on who gets to grade their homework](https://www.latent.space/p/ainews-aef-1-standard-emerges-for?ref=genaisecretsauce.com), and [The "AI misuse" report has a second story: labs allegedly copying each other at scale](https://www.cnbc.com/2026/09/11/chinese-ai-labs-moonshot-deepseek-alibaba-anthropic.html?ref=genaisecretsauce.com).

Trends

**The referees are moving inside the labs they judge**, **Whether to "slow down" AI has become an open public fight**, and **Cheap, small, open models keep crashing the party**.

Creative AI

**Describe a building in plain English and get a real, editable 3D model**.

Dev Tools

**Inside a real "AI software factory"** and [An enterprise HR brain you can plug into any AI assistant](https://joshbersin.com/2026/09/hr-intelligence-goes-enterprise-the-galileo-jupiter-release?ref=genaisecretsauce.com).

Research

**A 7-billion**, **An AI that aced a task once will not always do it again**, and **Can large language model (LLM) "judges" grade expert work? Partly.**.

Business

**A dating** and **Voice AI is now a benchmark battleground**.

Education

**Google is running a free "badgeathon" for teachers**, **Trusted, specialized AI is coming to corporate learning**, and **A case for "tapestry over hustle" in the AI era**.

Surprising

**An AI trained on a railroad board game got better at finance**, **AI models show consistent "personalities" under pressure**, and **LLM agents beat specialized robots-of-the**.

Worth Watching

**AI that keeps a lab's knowledge after the people leave**, **One rulebook to define "recursive self**, and **Carbon**.

GitHub

Leading repos: [alibaba/open-code](https://github.com/alibaba/open-code-review?ref=genaisecretsauce.com) (+2,751), [JustVugg/colibri](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com) (+2,035), and [debpalash/VoiceStudio](https://github.com/debpalash/VoiceStudio?ref=genaisecretsauce.com) (+2,081).

HuggingFace

Leading models: [deepseek-ai/DeepSeek-V4.1](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash?ref=genaisecretsauce.com) (326,000), [Qwen/Qwen3.8](https://huggingface.co/Qwen/Qwen3.8-27B?ref=genaisecretsauce.com) (7.7M), and [Lightricks/LTX](https://huggingface.co/Lightricks/LTX-2.5?ref=genaisecretsauce.com) (1.58M).

Product Hunt

Top launches: **Resurf** (295), **Visiby** (140), and **Epilude Notetaker** (109).

API Pricing

What this means: No material price changes among the flagship text models versus yesterday - Google's Gemini 3.1 Pro remains the value leader at $2 input, and Groq's open-model hosting stays an order of magnitude cheaper for lighter work.

arXiv

[ZGCM-1](https://arxiv.org/abs/2609.13356?ref=genaisecretsauce.com) — Results competitive with Qwen3-235B and GLM-5.1 on math and agentic-search benchmarks, plus about a 4.2x improvement in early-training time-to-target.

FYI

## Hot off the Presses

01

### Google's new AI can watch, listen, and talk back in real time - in 97 languages

**What this means for you:** The voice assistant on your phone is about to feel less like a menu and more like a person who can see your screen, follow you across languages, and keep talking while it thinks.

Google released two real-time voice models on September 15: Gemini 3.8 Live, tuned to be cheap and fast, and Gemini 3.8 Live Extended Thinking, tuned for harder, multi-step problems. Both can take in what a camera sees in near real time and run background tools (like looking something up) without pausing the conversation. The "Extended Thinking" version can reason and speak at the same time, using natural filler like "Let me check that..." so it does not go silent mid-answer.

- **#1 on the leaderboard** \- Extended Thinking topped Artificial Analysis' Speech-to-Speech Quality Index (a ranking of how good voice AIs sound and answer) with a score of 82.6.
- **Mid-sentence language switching** \- it detects when you change among 97 supported languages during a single conversation and follows along automatically.
- **Already inside Google's apps** \- live now via the Gemini application programming interface (API) and Google AI Studio, and wired into Search Live, Docs, Gmail, Keep, and the Gemini app.
- **Everything it says is watermarked** \- all generated audio carries SynthID marking so it can later be flagged as AI-made.

[Google: Gemini 3.8 Live announcement →](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com)[Try it →](https://aistudio.google.com/?ref=genaisecretsauce.com)

02

### The big AI labs just agreed on who gets to grade their homework

**What this means for you:** The companies racing to build powerful AI are, for the first time, agreeing to a shared rulebook for the outsiders who check their safety work - a small but real step toward independent oversight.

xAI, OpenAI, and Anthropic all endorsed AEF-1, a new baseline standard for third-party AI evaluators (the independent groups paid to probe models for danger before release). The standard spells out how evaluators handle conflicts of interest, funding ties, when they must step aside, and what they have to disclose. It was published by the AI Evaluator Forum, a body formed in December 2025.

- **Why a standard is needed** \- outside auditors are often funded by the very labs they inspect, so the rules focus on keeping that relationship honest.
- **Anthropic went further** \- it said it will embed external evaluators with office space, laptops, and desk access on par with its own internal risk teams.
- **The open worry** \- one researcher quoted in coverage warned future models could "appear aligned under evaluation while hiding misalignment," which is exactly what independent testing is meant to catch.

[Latent Space: AEF-1 standard emerges →](https://www.latent.space/p/ainews-aef-1-standard-emerges-for?ref=genaisecretsauce.com)

03

### The "AI misuse" report has a second story: labs allegedly copying each other at scale

**What this means for you:** Some of the "new" cheap AI models on the market may have been trained partly on a rival's answers - and the fight over who copied whom is becoming the industry's next big conflict.

*Previously: [September 10](https://genaisecretsauce.com/genai-secret-sauce-daily-digest-2026-09-10/) \- Anthropic published a threat report and said it disrupted attempts to use Claude for biological-weapons research and a network of roughly 70 fake news sites.*

**Today:** Follow-up reporting and analysis this week put the spotlight on a different part of the same report - "distillation," where competitors harvest a top model's answers to train a cheaper copy. Anthropic says it detected and cut off large-scale distillation tied to seven China-based labs, and the numbers are the eye-opener.

“More than 151 million Claude exchanges, peaking at nearly 3 million per day.”

- **151 million and counting** \- Anthropic attributed more than 151 million Claude exchanges to Alibaba between May and July 2026, peaking near 3 million per day from over 3,500 accounts it calls fraudulent.
- **Routing real customers through Claude** \- Anthropic alleges Moonshot and DeepSeek passed their own users' live chats through Claude and kept the responses as training data.
- **A crowded list** \- the disrupted distillation campaigns were attributed to Alibaba, DeepSeek, Moonshot, Xiaomi, Zhipu, SenseTime, and MiniMax.

[CNBC: Chinese labs used Claude exchanges to train models →](https://www.cnbc.com/2026/09/11/chinese-ai-labs-moonshot-deepseek-alibaba-anthropic.html?ref=genaisecretsauce.com)[Zvi Mowshowitz: analysis →](https://thezvi.substack.com/p/the-bad-guy-with-an-ai-named-claude)

04

### A startup is teaching AI to work by making it play board games

**What this means for you:** One promising path to more capable AI assistants is not a bigger model, but smarter training - and it can come from unexpected places like a 19th-century railroad game.

Good Start Labs, a $3.6 million-funded startup, trains AI inside game environments to see whether game skills carry over to real work. Its headline finding: how you train matters far more than which game you use. When it trained a mid-size model on "1830" (a cutthroat railroad-tycoon board game) as a hands-on agent that adapts turn by turn and uses tools, the model got better at outside financial-research tests. A simpler training setup only made it better at the game itself.

- **Two real transfers so far** \- the railroad game improved finance-research scores, and training on the strategy game Diplomacy produced better customer-support behavior.
- **Models have "personalities"** \- in Diplomacy tests, one lab's model planned betrayals while Claude Opus 4 refused to deceive.
- **The honest caveat** \- broad generalization is still unproven; two cases is a signal, not a law.

[Latent Space: Good Start Labs →](https://www.latent.space/p/good-start-labs?ref=genaisecretsauce.com)

Trends & Themes

## Trends & Themes

[![Trends & Themes](https://genaisecretsauce.com/content/images/2026/09/section-what-this-means-2026-09-15-1.png)](https://genaisecretsauce.com/content/images/2026/09/section-what-this-means-2026-09-15-1.png) 

### The referees are moving inside the labs they judge

**Why this matters to you:** Whether AI is safe increasingly depends on outsiders who can see what the labs are doing - and this week they got both a rulebook and a desk in the building.

The pattern this week is oversight getting more formal and more embedded at the same time. That is progress, but it also ties the auditors closer to the audited - the exact conflict the new standard is trying to manage.

- **A shared standard** \- xAI, OpenAI, and Anthropic all cosigned AEF-1 for third-party evaluators.
- **Physical access** \- Anthropic said it will give external evaluators laptops, offices, and internal-team-level access.
- **The catch, in the labs' own words** \- researchers warn a model can look safe during a test and still hide problems, so access has to be deep, not cosmetic.

### Whether to "slow down" AI has become an open public fight

**Why this matters to you:** The people building AI are now arguing in public about whether they are going too fast, and the outcome shapes how quickly these tools land in your job and your apps.

This debate ran through several sources today, from newsletters to a widely shared departure post. The takeaway is that the disagreement is no longer polite or private - it is now a named, public split among the field's most powerful people.

- **Slow-down camp** \- Anthropic's Dario Amodei argues labs should "slow down long enough for safety reasons," and Sam Altman broadly agrees.
- **Full-speed camp** \- Jensen Huang and US AI policy voices frame slowing as optional, not required.
- **Fuel on the fire** \- a former lab researcher's public resignation warning that labs are "gambling with our lives" went viral this week, and at least one current researcher agreed.

### Cheap, small, open models keep crashing the party

**Why this matters to you:** The best AI is getting dramatically cheaper, which means more of it ends up free or nearly free in the tools you already use.

The throughline: capability is leaking downward to small, open, tool-using models. For everyday users, that means fewer paywalls and more "good enough" AI running cheaply, or even on your own machine.

- **Tiny but mighty** \- a new fully open 7-billion-parameter model, ZGCM-1, reports math and search results competitive with models 30x its size by leaning on tools instead of memorized facts.
- **Price war on coding models** \- Cognition's SWE-2 launched at 64% cheaper than a frontier model while staying near the top on coding tasks.
- **Open models dominate the download charts** \- Hugging Face's trending list is full of open releases like DeepSeek, Qwen, and MiniCPM pulling millions of downloads.

### "Passing the test" is quietly the wrong metric for AI agents

**Why this matters to you:** An AI that works in a demo can still fail when you actually rely on it - and the industry is starting to measure that gap.

Three separate research threads today - IBM's consistency work, an "agent iteration" framework, and a paper on agents recovering from their own mistakes - all point the same way: reliability, not raw smarts, is the frontier that decides whether agents are trustworthy.

“A top AI agent passed 77% of the time on average, but nailed the same task on all five tries only 53% of the time.”

- **The gap, measured** \- IBM found a top agent averaged a 77.4% pass rate but succeeded on all five tries for only 53% of tasks.
- **Why it happens** \- when the model's next-word choice is a near-tie, tiny technical differences between runs can flip the outcome.
- **A cheap fix helped** \- applying consistency guidelines lifted "succeeds every time" from 53% to 69% without extra cost.

Creative AI & Media

## Creative AI & Media

### Describe a building in plain English and get a real, editable 3D model

**What this lets you do:** Design furniture, interiors, or buildings by talking or sketching, then export the result straight into professional design software.

**Try it:** [Cartesian by Formas](https://www.formas.ai/cartesian?ref=genaisecretsauce.com)

- **Solids, not just surfaces** \- Cartesian by Formas generates precise geometric solids with clean, editable geometry, not the messy polygon shells most AI 3D tools produce.
- **Every part stays editable** \- you can lock some elements and freely change others, and export to SketchUp and Rhino (with AutoCAD and building-info formats planned).
- **Who it is for** \- architects and product designers who want AI speed without giving up professional precision.

Developer Tools

## Developer Tools & Infrastructure

### Inside a real "AI software factory" - and what broke

**What this means for you:** The way software gets built is being rebuilt around AI agents as the main authors, with humans moving to review and approval.

The bottleneck now is not writing code, it is the human-speed steps around it - like Apple and Google app reviews that still take days. [Pragmatic Engineer: inside OpenAI's software factory](https://newsletter.pragmaticengineer.com/p/openai-software-factory?ref=genaisecretsauce.com)

- **Adoption exploded** \- at OpenAI, non-engineering teams went from roughly 0% to 90% use of its Codex coding agent in about four months.
- **Ten times the pull requests** \- code-change volume rose about 10x in six months, forcing constant infrastructure firefighting.
- **Agents review agents** \- low-risk changes can skip human review entirely, while high-risk ones get specialist AI reviewers plus a human sign-off, and an incident bot named Sevbot drafts fixes.

### An enterprise HR brain you can plug into any AI assistant

**What this means for you:** Big companies are wiring trusted, specialized knowledge into general AI assistants so answers stop being confident guesses.

- **Plugs in anywhere** \- Josh Bersin's Galileo "Jupiter" release connects HR intelligence into tools like Microsoft Copilot, Workday, and ServiceNow.
- **Cheaper to run** \- a new architecture cut its token use to about one-tenth of prior versions.
- **The claim to check** \- Bersin reports zero hallucinations across 30 test prompts versus errors from general chatbots, a strong claim that deserves independent testing.

[Josh Bersin: Galileo Jupiter release →](https://joshbersin.com/2026/09/hr-intelligence-goes-enterprise-the-galileo-jupiter-release?ref=genaisecretsauce.com)

Research & Models

## Research & Models

### A 7-billion-parameter open model that punches 30x above its weight

**Why it matters:** Small open models that use tools well can match giant models on real tasks - meaning powerful AI you can run yourself keeps getting cheaper.

- **The claim** \- ZGCM-1, fully open (weights, code, data, logs), reports math and agentic-search results competitive with models like Qwen3-235B and GLM-5.1.
- **The trick** \- it pairs deliberate internal reasoning with active tool use instead of trying to memorize the web.
- **Efficiency win** \- about a 4.2x improvement in early-training time-to-target. [arXiv: ZGCM-1](https://arxiv.org/abs/2609.13356?ref=genaisecretsauce.com)

### An AI that aced a task once will not always do it again

**Why it matters:** Consistency, not peak skill, is what makes an AI agent safe to depend on - and it is now measurable and fixable.

- **The finding** \- a strong agent passed 77.4% of the time on average but succeeded on all five tries for only 53% of tasks.
- **The tool** \- IBM's Consistency Analyzer inspects a single recorded run to flag "flip-prone" decision points, no full re-run needed.
- **The payoff** \- guidelines raised "succeeds every time" from 53% to 69%, with the biggest gains on hard tasks. [Hugging Face: IBM Research on agent consistency](https://huggingface.co/blog/ibm-research/altk-evolve-consistency?ref=genaisecretsauce.com)

### Can large language model (LLM) "judges" grade expert work? Partly.

**Why it matters:** AI is increasingly used to grade other AI, and this tests whether that works on genuinely hard professional tasks.

- **The setup** \- separate AI judges reviewed AI-drafted patents and fed back fixes ("vibe patenting").
- **What worked** \- judge-guided revision steadily improved drafts, and weak models with feedback approached far pricier ones.
- **The limit** \- agreement with human patent attorneys was real but strongly depended on the metric, so AI judges are a helper, not a final arbiter. [arXiv: Vibe Patenting](https://arxiv.org/abs/2609.13422?ref=genaisecretsauce.com)

Business & Industry

## Business & Industry

### A dating-network scam ran 4,700 fake AI personas on real people

**What this means for you:** The same AI that powers helpful chatbots is being used to run persuasion and fraud at industrial scale, and the numbers are now documented.

- **The scale** \- Anthropic's report describes a fake dating-app network using 4,700 distinct AI personas that engaged over 25,000 real individuals.
- **The context** \- it sits alongside the distillation and cyber-misuse findings in the same September report.
- **The signal** \- "one operator, many convincing fakes" is becoming a repeatable business model for scammers. [Zvi Mowshowitz: on the misuse report](https://thezvi.substack.com/p/the-bad-guy-with-an-ai-named-claude)

### Voice AI is now a benchmark battleground

**What this means for you:** Competition on real-time voice is heating up fast, which usually means better, cheaper assistants for you.

- **The move** \- Google's Gemini 3.8 Live launched straight into the #1 speech-quality spot and undercut rivals on price.
- **The metric race** \- it posted 97.7% on an audio-understanding benchmark and strong scores on agent tasks like banking help.
- **Why it is business news** \- voice is the interface for phones, cars, and customer service, so leading the benchmark is a commercial land-grab. [Google: Gemini 3.8 Live](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking?ref=genaisecretsauce.com)

Education

## GenAI in Education

### Google is running a free "badgeathon" for teachers

**What this means for you:** K-12 teachers can earn recognized AI credentials in a single free day without spending anything.

- **When** \- September 19, 2026, 9:00 AM to 7:00 PM ET, drop-in 30-minute sessions.
- **What you get** \- ISTE-aligned micro-credentials, live badging quizzes, and hands-on tool overviews including a session on Gemini Notebook.
- **Cost** \- free, registration via the event site. [Control Alt Achieve: Google AI Badgeathon](https://www.controlaltachieve.com/2026/09/google-ai-badgeathon.html?ref=genaisecretsauce.com)

### Trusted, specialized AI is coming to corporate learning

**What this means for you:** Workplace training and HR answers are shifting from generic chatbots to systems grounded in vetted data.

- **The release** \- Bersin's Galileo Jupiter embeds trusted HR knowledge into everyday work tools at one-tenth the token cost.
- **The pitch** \- grounded answers for hiring, restructuring, and skills questions instead of confident guesses.
- **Why it belongs here** \- it is the L&D and workforce-upskilling front of the same "trusted enterprise AI" trend. [Josh Bersin: Galileo Jupiter](https://joshbersin.com/2026/09/hr-intelligence-goes-enterprise-the-galileo-jupiter-release?ref=genaisecretsauce.com)

### A case for "tapestry over hustle" in the AI era

**What this means for you:** How you use AI at work is a choice - as an accelerator for output, or as a tool for deeper thinking.

- **The argument** \- educator Lance Eaton says durable careers come from values-driven practice that compounds over years, not say-yes-to-everything hustle.
- **On AI** \- he urges using it as "a tool of discernment" that supports thinking, not just faster production.
- **The test he offers** \- ask "does this belong to the work, and do I actually have time for it?" [AI + Education = Simplified: The Tapestry and The Hustle](https://aiedusimplified.substack.com/p/the-tapestry-and-the-hustle)

Surprising

## Surprising & Under-the-Radar

### An AI trained on a railroad board game got better at finance

Good Start Labs found that training a model as a hands-on agent inside the 1830s tycoon game "1830" improved its scores on unrelated financial-research tests - evidence that how you train can matter more than what you train on.

### AI models show consistent "personalities" under pressure

In strategy-game tests, some models planned betrayals while Claude Opus 4 refused to deceive - a reminder that different models behave differently in the same high-stakes spot, not just at different skill levels.

### LLM agents beat specialized robots-of-the-mind when conditions change

A physical-task study found general LLM agents matched purpose-built reinforcement-learning systems in steady conditions but clearly beat them when the environment shifted - suggesting reasoning may generalize better than narrow training when the world is unpredictable.

### Debate: is publishing "who misused our AI" oversight or geopolitics?

Anthropic's naming of specific foreign labs and actors is praised as transparency by some and questioned by others as a move that could inflame US-China tensions during ongoing AI talks. Both readings can be true at once.

Worth Watching

## Signals to Track

[![Worth Watching](https://genaisecretsauce.com/content/images/2026/09/section-worth-watching-2026-09-15-1.png)](https://genaisecretsauce.com/content/images/2026/09/section-worth-watching-2026-09-15-1.png) 

01

### AI that keeps a lab's knowledge after the people leave

Why this is worth watching right now: it targets the quiet crisis of expertise walking out the door when staff move on.

A new system called LabAgent customizes an AI agent to a specific research group's methods and records how to fix errors, so work can continue after key people depart. It reportedly outranked general commercial agents across several life-science tasks and reproduced a published result. If it holds up, small teams could stop losing years of hard-won know-how to turnover.

02

### One rulebook to define "recursive self-improvement"

Why this is worth watching right now: the field is trying to formally define AI that improves itself before such systems arrive.

A new framework, Generalized Agent Iteration, puts ordinary AI training and "AI that rewrites itself" on the same map, giving researchers shared language to spot where self-improving systems could go wrong. It is theory today, but it is the kind of groundwork that shapes future safety rules for autonomous AI.

03

### Carbon-aware AI routing

Why this is worth watching right now: it quietly turns "which data center answers you" into a climate decision.

Research on carbon-aware routing picks where an AI request runs based on the cleanliness of the local power grid at that moment. As AI's energy use grows, this kind of behind-the-scenes routing could cut emissions without users noticing any difference in speed.

GitHub Trending

## Top Repos Today

#1

### [alibaba/open-code-review](https://github.com/alibaba/open-code-review?ref=genaisecretsauce.com)

Rank yesterday: #2 - Rising ↑

⭐ **Stars today:** +2,751 · 📦 **Total:** 28,423  
📜 **License:** Apache-2.0 · 👤 **By:** Alibaba (corporate)  
🎯 **Time to value:** 20 minutes

**What it is:** An open-source tool that uses AI to automatically review code changes, built and battle-tested inside Alibaba. It plugs into your code workflow and flags issues before a human looks. **Why you'd want it:** It brings big-company-grade automated review to any team for free, cutting the time reviewers spend on routine checks.

| ✓ Pros                       | ✗ Cons                         |
| ---------------------------- | ------------------------------ |
| Proven at large scale        | Tuned to Alibaba's conventions |
| Fast and low-dependency      | Needs an LLM provider to run   |
| Active, fast-growing project | Newer, so docs still maturing  |

[GitHub - alibaba/open-code-review: Fast, efficient, battle-tested at Alibaba’s scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE,…![](https://genaisecretsauce.com/content/images/icon/favicon-30166cc2-f5bf-456b-b78c-db84846140e8.png)alibabaGitHub![](https://genaisecretsauce.com/content/images/thumbnail/27bf01cb-17df-44d5-9b7b-bcf17c970c6d-e4a3b16e-7532-4f56-bf07-57179ee741e4)](https://github.com/alibaba/open-code-review?ref=genaisecretsauce.com)

#2

### [JustVugg/colibri](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com)

Rank yesterday: #1 - Falling ↓

⭐ **Stars today:** +2,035 · 📦 **Total:** 33,765  
📜 **License:** MIT · 👤 **By:** Independent developer  
🎯 **Time to value:** 30 minutes

**What it is:** A lightweight tool for running large "mixture-of-experts" AI models (a design where only part of the model runs per query) on hardware you already own. It keeps dependencies minimal. **Why you'd want it:** It lets you run frontier-style models locally without a specialized server or heavy setup.

| ✓ Pros                             | ✗ Cons                              |
| ---------------------------------- | ----------------------------------- |
| Runs big models on modest hardware | Command-line, not beginner-friendly |
| Very lightweight                   | Performance varies by machine       |
| Permissive MIT license             | Setup takes some patience           |

[GitHub - JustVugg/colibri: Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦 - JustVugg/colibri![](https://genaisecretsauce.com/content/images/icon/favicon-4b4e964c-9ad3-4500-aea8-c92c0486b354.png)JustVuggGitHub![](https://genaisecretsauce.com/content/images/thumbnail/colibri-5554c941-b11f-413d-957a-cd973f8558a9)](https://github.com/JustVugg/colibri?ref=genaisecretsauce.com)

#3

### [debpalash/VoiceStudio](https://github.com/debpalash/VoiceStudio?ref=genaisecretsauce.com)

Rank yesterday: #3 - Holding steady ➡

⭐ **Stars today:** +2,081 · 📦 **Total:** 30,877  
📜 **License:** MIT · 👤 **By:** Independent developer  
🎯 **Time to value:** 15 minutes

**What it is:** A local, free alternative to paid voice-cloning services, offering voice cloning and audio creation across 646 languages, all on your own computer. **Why you'd want it:** You get studio-style voice tools without a subscription or sending your voice to a cloud service.

| ✓ Pros                        | ✗ Cons                              |
| ----------------------------- | ----------------------------------- |
| No subscription, runs locally | Voice cloning raises consent issues |
| Huge language coverage        | Needs a capable machine             |
| Simple to get started         | Quality varies by voice             |

[GitHub - debpalash/VoiceStudio: VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. - debpalash/Voic…![](https://genaisecretsauce.com/content/images/icon/favicon-feb11205-d1a8-4925-aedf-3d130f3ad188.png)debpalashGitHub![](https://genaisecretsauce.com/content/images/thumbnail/fd6851a0-4e36-4541-a76c-2ad84935d0bc-80b793ad-a26c-4961-bdaf-67c089800e5d)](https://github.com/debpalash/VoiceStudio?ref=genaisecretsauce.com)

#4

### [alphaXiv/OpenResearch](https://github.com/alphaXiv/OpenResearch?ref=genaisecretsauce.com)

Rank yesterday: New entry 🆕

⭐ **Stars today:** +593 · 📦 **Total:** 3,299  
📜 **License:** Apache-2.0 · 👤 **By:** alphaXiv (startup)  
🎯 **Time to value:** 25 minutes

**What it is:** A toolkit that turns coding agents into research agents, helping AI read papers and run literature-style investigations instead of just writing code. **Why you'd want it:** It repurposes tools you may already use for coding into an AI research assistant.

| ✓ Pros                         | ✗ Cons                        |
| ------------------------------ | ----------------------------- |
| Builds on existing agent tools | Early-stage project           |
| Focused, useful niche          | Research quality still varies |
| Open Apache license            | Small community so far        |

[GitHub - alphaXiv/OpenResearch: Turn your coding agents into research agentsTurn your coding agents into research agents. Contribute to alphaXiv/OpenResearch development by creating an account on GitHub.![](https://genaisecretsauce.com/content/images/icon/favicon-cf7ac021-858c-4bcc-8cd3-6fd075abafaf.png)alphaXivGitHub![](https://genaisecretsauce.com/content/images/thumbnail/OpenResearch-7f6ad141-7314-4a16-8e14-c80b3de34a34)](https://github.com/alphaXiv/OpenResearch?ref=genaisecretsauce.com)

#5

### [earendil-works/pi](https://github.com/earendil-works/pi?ref=genaisecretsauce.com)

Rank yesterday: New entry 🆕

⭐ **Stars today:** +437 · 📦 **Total:** 105,666  
📜 **License:** MIT · 👤 **By:** earendil-works (startup)  
🎯 **Time to value:** 20 minutes

**What it is:** A broad AI-agent toolkit with one unified way to call many AI providers, plus a command-line tool for building and running agents. **Why you'd want it:** It saves you from wiring up each AI provider separately when building your own assistant.

| ✓ Pros                     | ✗ Cons                        |
| -------------------------- | ----------------------------- |
| One API for many providers | Large surface area to learn   |
| Big, active community      | Broad scope can feel heavy    |
| CLI included               | Fast changes between versions |

[GitHub - earendil-works/pi: AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLIAI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI - earendil-works/pi![](https://genaisecretsauce.com/content/images/icon/favicon-5bbbcb2a-d57f-4a74-8772-846f58769ded.png)earendil-worksGitHub![](https://genaisecretsauce.com/content/images/thumbnail/pi-681e708c-2fc7-4e37-9884-80a74779c1b1)](https://github.com/earendil-works/pi?ref=genaisecretsauce.com)

#6

### [melgarafael/DeskcommCRM](https://github.com/melgarafael/DeskcommCRM?ref=genaisecretsauce.com)

Rank yesterday: Holding steady ➡

⭐ **Stars today:** +205 · 📦 **Total:** 2,798  
📜 **License:** MIT · 👤 **By:** Independent developer  
🎯 **Time to value:** 30 minutes

**What it is:** An open-source sales system with built-in AI agents and WhatsApp integration, aimed at small teams that want automation without enterprise pricing. **Why you'd want it:** It bundles customer management and AI outreach in one free, self-hosted package.

| ✓ Pros               | ✗ Cons                    |
| -------------------- | ------------------------- |
| AI agents built in   | Self-hosting takes effort |
| WhatsApp integration | Smaller ecosystem         |
| Free and open        | Still maturing            |

[GitHub - melgarafael/DeskcommCRM: Open-source AI sales OS — self-hosted CRM with native AI agents + WhatsApp (WAHA). Open alternative to Kommo, Octadesk & Intercom for any business that sells by chat. MCP-ready, multi-tenant, LGPD.Open-source AI sales OS — self-hosted CRM with native AI agents + WhatsApp (WAHA). Open alternative to Kommo, Octadesk & Intercom for any business that sells by chat. MCP-ready, multi-tenant, L…![](https://genaisecretsauce.com/content/images/icon/favicon-089db0b4-2b7e-45dc-98a8-6ea59e1812bb.png)melgarafaelGitHub![](https://genaisecretsauce.com/content/images/thumbnail/976719c3-5d23-4dd9-a6b2-915229d2c325-b30590c5-7059-42fc-a3b9-e95caf7e9a64)](https://github.com/melgarafael/DeskcommCRM?ref=genaisecretsauce.com)

HuggingFace Trending

## Top Models Today

#1

### [deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash?ref=genaisecretsauce.com)

A fast, low-cost version of DeepSeek's model that handles both text and images, topping the trending charts on strong price-performance.

📥 **Downloads (30d):** 326,000 · 📜 **License:** MIT  
👤 **By:** DeepSeek · 🎯 **Task:** Image-Text-to-Text  
📐 **Size:** 763B (mixture-of-experts)

**What it is:** A "Flash" (speed-optimized) model from Chinese lab DeepSeek that reads both pictures and text and answers cheaply. It uses a mixture-of-experts design so only part of it runs per query. **Why you'd want it:** Frontier-level answers at a fraction of typical cost, with an open MIT license.

| ✓ Pros                  | ✗ Cons                             |
| ----------------------- | ---------------------------------- |
| Very cheap per task     | Huge total size to host            |
| Handles text and images | Best run via a provider            |
| Open license            | From a lab named in misuse reports |

[deepseek-ai/DeepSeek-V4.1-Flash · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-34b1e51f-9623-4591-991f-4aed0c519ac9.ico)![](https://genaisecretsauce.com/content/images/thumbnail/DeepSeek-V4.1-Flash-a64a4b44-589e-41a9-916a-78b12a07c0dc.png)](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash?ref=genaisecretsauce.com)

#2

### [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B?ref=genaisecretsauce.com)

Alibaba's mid-size open model that reads text and images, pulling millions of downloads as a practical workhorse.

📥 **Downloads (30d):** 7.7M · 📜 **License:** Apache-2.0  
👤 **By:** Alibaba (Qwen) · 🎯 **Task:** Image-Text-to-Text  
📐 **Size:** 27B

**What it is:** A 27-billion-parameter open model that handles text and images at a size many people can actually run. It is a general-purpose assistant model. **Why you'd want it:** A strong, permissively licensed all-rounder that fits on serious consumer or single-server hardware.

| ✓ Pros                      | ✗ Cons                                       |
| --------------------------- | -------------------------------------------- |
| Massive real-world adoption | 27B still needs a strong graphics card (GPU) |
| Text and image input        | General, not specialized                     |
| Apache license              | Competition is fierce                        |

[Qwen/Qwen3.8-27B · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-e3f88998-e7d9-4a16-b9cc-972fc3a30256.ico)![](https://genaisecretsauce.com/content/images/thumbnail/Qwen3.8-27B-6969ef93-6214-4b6d-b0cd-fc8d14cd636d.png)](https://huggingface.co/Qwen/Qwen3.8-27B?ref=genaisecretsauce.com)

#3

### [Lightricks/LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5?ref=genaisecretsauce.com)

A popular open image-to-video model that animates still pictures into short clips.

📥 **Downloads (30d):** 1.58M · 📜 **License:** Open (LTX community)  
👤 **By:** Lightricks · 🎯 **Task:** Image-to-Video  
📐 **Size:** not published

**What it is:** A model that turns a still image into a short moving clip, widely used by creators. It is one of the most downloaded open video models. **Why you'd want it:** Free, local video generation for social clips and prototypes without a subscription.

| ✓ Pros                  | ✗ Cons                     |
| ----------------------- | -------------------------- |
| Very widely used        | Video gen is compute-heavy |
| Turns photos into clips | Short clips only           |
| Open community license  | Quality varies by input    |

[Lightricks/LTX-2.5 · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-6e0d1417-424a-491c-8887-662c1a4253a5.ico)![](https://genaisecretsauce.com/content/images/thumbnail/LTX-2.5-de86ac05-a603-4073-8f33-b99a920cd500.png)](https://huggingface.co/Lightricks/LTX-2.5?ref=genaisecretsauce.com)

#4

### [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B?ref=genaisecretsauce.com)

A tiny 3B-class model built to run on phones and laptops while staying capable.

📥 **Downloads (30d):** 272,000 · 📜 **License:** Apache-2.0  
👤 **By:** OpenBMB · 🎯 **Task:** Text Generation  
📐 **Size:** 3B

**What it is:** A small text model designed to run on everyday devices, part of the push to put useful AI on-device. It trades some raw power for portability. **Why you'd want it:** Private, offline AI that runs on hardware you already own.

| ✓ Pros                 | ✗ Cons                       |
| ---------------------- | ---------------------------- |
| Runs on phones/laptops | Less capable than big models |
| Fast and private       | Short on world knowledge     |
| Open license           | Best for lighter tasks       |

[openbmb/MiniCPM5-2B · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-15d63dae-a2de-4048-9118-c1565a78511b.ico)![](https://genaisecretsauce.com/content/images/thumbnail/MiniCPM5-2B-75c5b835-3ecf-4f3f-80f6-2ba547c0bc76.png)](https://huggingface.co/openbmb/MiniCPM5-2B?ref=genaisecretsauce.com)

#5

### [Edge0/Edge0-35B-A3B-preview](https://huggingface.co/Edge0/Edge0-35B-A3B-preview?ref=genaisecretsauce.com)

A new mid-size open model climbing the charts as a fresh general-purpose option.

📥 **Downloads (30d):** 17,900 · 📜 **License:** Apache-2.0  
👤 **By:** Edge0 · 🎯 **Task:** Text Generation  
📐 **Size:** 35B (mixture-of-experts, \~3B active)

**What it is:** A 35-billion-parameter model that only activates about 3B at a time, keeping it efficient to run. It is a preview release gaining early traction. **Why you'd want it:** Bigger-model quality at closer to small-model running cost, thanks to the mixture-of-experts design.

| ✓ Pros                        | ✗ Cons                |
| ----------------------------- | --------------------- |
| Efficient to run for its size | Preview, not final    |
| Open license                  | Small track record    |
| Rising quickly                | Limited documentation |

[Edge0/Edge0-35B-A3B-preview · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.![](https://genaisecretsauce.com/content/images/icon/favicon-d6bbfa15-bbfe-4b5c-9875-dafc994f417a.ico)![](https://genaisecretsauce.com/content/images/thumbnail/Edge0-35B-A3B-preview-dea0499a-1240-4500-80dd-a342c41b41d4.png)](https://huggingface.co/Edge0/Edge0-35B-A3B-preview?ref=genaisecretsauce.com)

Product Hunt

## AI Launches Today

### [Resurf](#)

Personal context library for Mac

🔥 **Upvotes:** 295 · 👤 **By:** Resurf  
💰 **Pricing:** freemium · 🏷 **Category:** productivity / on-device AI

Resurf builds a private, on-device library of your context on macOS so AI tools can draw on what you have read and worked on without shipping it all to the cloud. It targets people who want personalized AI help while keeping their data local. **Verdict:** The clear standout of the day, and a bet that private, on-device context is the next battleground. [Product Hunt](https://www.producthunt.com/?ref=genaisecretsauce.com)

### [Visiby](#)

SEO tracking for AI-powered search engines

🔥 **Upvotes:** 140 · 👤 **By:** Visiby  
💰 **Pricing:** freemium · 🏷 **Category:** marketing / analytics

Visiby tracks how your brand shows up inside AI answer engines, not just traditional search results, as more people get information from chatbots instead of blue links. It helps marketers see whether AI assistants mention or recommend them. **Verdict:** Useful and well-timed as "being cited by the AI" becomes the new SEO. [Product Hunt](https://www.producthunt.com/?ref=genaisecretsauce.com)

### [Epilude Notetaker](#)

Privacy-focused meeting notes

🔥 **Upvotes:** 109 · 👤 **By:** Epilude  
💰 **Pricing:** freemium · 🏷 **Category:** productivity

Epilude records and summarizes meetings with a privacy-first design, aimed at people who like AI note-takers but not the idea of every call living on someone else's server. **Verdict:** A crowded space, but privacy is a real differentiator worth watching. [Product Hunt](https://www.producthunt.com/?ref=genaisecretsauce.com)

API Pricing

## Snapshot

Provider

Model

Input $/1M

Output $/1M

Context

Anthropic

Claude Opus 5

$5.00

$25.00

200k

OpenAI

GPT-5.6 (flagship)

$5.00 (promo $4.00)

$30.00 (promo $20.00)

not published

Google

Gemini 3.1 Pro (preview)

$2.00

$12.00

1M

Groq

GPT-OSS 20B

$0.075

$0.30

not published

**What this means:** No material price changes among the flagship text models versus yesterday - Google's Gemini 3.1 Pro remains the value leader at $2 input, and Groq's open-model hosting stays an order of magnitude cheaper for lighter work. **Notable today:** Google launched its Gemini 3.8 Live voice models, which are priced separately from these text APIs; real-time voice pricing is a new line item to watch as that market heats up. Batch processing and prompt caching still cut effective rates by roughly half at most providers. (OpenAI and Groq figures come from third-party pricing trackers, as their pages were not directly readable.)  
  
arXiv Paper of the Day

## ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Open-source team · arXiv:2609.13356

**What it claims:** A fully open 7-billion-parameter model can match models roughly 30x larger on math reasoning and agentic search by pairing deliberate internal reasoning with active tool use, rather than trying to memorize the web.  
  
**Key finding:** Results competitive with Qwen3-235B and GLM-5.1 on math and agentic-search benchmarks, plus about a 4.2x improvement in early-training time-to-target.  
  
**Why practitioners should care:** It ships weights, code, datasets, and training logs, giving teams a genuinely reproducible recipe for a small, tool-using model they can run and adapt cheaply.  
  
[Read on arXiv →](https://arxiv.org/abs/2609.13356?ref=genaisecretsauce.com)

GenAI Secret Sauce Daily Digest · 2026-09-15

×

Click anywhere or press ESC to close