Watch today's digest as a video summary (generated by NotebookLM)
Statistically Speaking
One Thing to Tell Your Friends
TL;DR
Hot off the Presses
ChatGPT has been quietly tracking you across the wider web
Previously: September 16 - OpenAI began building advertising into ChatGPT.
Today: An independent researcher took apart the network traffic around ChatGPT and found OpenAI is running a full advertising and conversion-tracking system, internally named "Bazaar," at a hidden web address (bzr.openai.com). It works like the ad-tracking machinery already run by Meta (Facebook's owner) and Google: a small piece of tracking code, called a pixel, sits on advertiser websites and reports your activity back. The new and unusual part is that this one is tied to an artificial intelligence chat product, where people share far more intimate details than they ever would on social media.
OpenAI's support team acknowledged the researcher's questions but did not answer them. The test only covered Chrome on Android, and the link back to your account was strongly inferred rather than directly proven - but the tracking cookie itself is plainly visible.
- Your account gets a hidden ID - ChatGPT plants a cookie (a small tracking file) called
__obithat lasts a full year and is deliberately set to travel with you to other sites. - 936 advertiser trackers, 1,029 websites - the researcher counted OpenAI's pixel across a huge span of the web, including named retailers like Chewy, Wayfair, HelloFresh, and Coursera.
- Logged out does not mean private - signed-out users still get an "anonymous" device ID that persists at least 27 days.
- Some data travels in the clear - email, phone and names are scrambled first, but your country, region, city and postal code are sent as plain readable text.
A new site "rescues" free AI models by turning them into torrents
Pirate Face is a preservation service that takes open-source AI models from Hugging Face (the main public library for free AI models) and converts them into peer-to-peer torrents - the same file-sharing technology used to share large files without a central server. While a model is still hosted normally, downloads come straight from Hugging Face. The moment a model is removed, a swarm of volunteer computers keeps it alive, and models marked "Rescued" stay downloadable indefinitely.
The team frames it as "permanent infrastructure for sovereign AI," protecting open models from takedowns, license changes, or corporate policy shifts. It treats open model weights the way libraries treat out-of-print books: public heritage worth keeping alive.
- Over 669,000 models are eligible - the service mirrors freely-licensed models (Apache-2.0 and MIT) synced live from Hugging Face.
- Downloads are guaranteed identical - every file carries an official checksum (a digital fingerprint), so rescued weights exactly match the originals.
- No code changes needed - a drop-in setting lets existing projects pull from Pirate Face instead, so nothing breaks.
A firsthand account of a company where AI writes nearly everything
Developer and writer Simon Willison highlighted a firsthand account from a software engineer describing a large company where a coding assistant (Claude Code) now produces almost all technical work: the specifications, the code, the tests, the documentation, even the task tickets. Willison, who built widely-used open-source data tools, flagged the post as a cautionary tale, not a success story.
This is a single account, not a study. But it is a concrete, September-2026 snapshot of the difference between AI that supplements human judgment and AI that replaces it.
- 12 to 13 hour days, little real work - staff mostly execute the AI's suggestions rather than solving problems themselves.
- Everyone leans on it, junior to senior - engineers at every level default to asking the assistant instead of using their own judgment.
- The core risk is nobody understands the code - when AI output ships without real review, quality, knowledge, and long-term maintainability quietly erode.
The real AI cost story is not your token bill
Analyst Nate argues that fixating on the price of AI usage (measured in "tokens," the small chunks of text an AI processes) misses the bigger prize: AI lowers the cost of serving each customer. He calls the trap a "visibility problem" - the AI bill is a clear, single line item, while the human and operational costs it replaces are scattered and hidden, so a naive comparison always makes AI look expensive.
He backs it with a customer complaint relayed by a Salesforce executive: despite promises that AI would "change everything," many products still feel unchanged - the gap between what AI can do and what businesses have actually rebuilt.
- Old pricing rules are up for grabs - service tiers and account minimums were built around expensive human attention, and all of it becomes renegotiable as automation gets cheap.
- Savings are not automatic profit - a rival can pour the same efficiency into lower prices or better service, so who keeps the gains is a fight.
- Real change means new outcomes - genuine transformation is offering customers things that were impossible before, not bolting an AI feature onto an old product.
Trends & Themes
AI companies are quietly becoming data and advertising businesses
The pattern: the value in consumer AI is drifting from the cleverness of the model toward the data it can collect, sort, and act on. That is exactly the shift that turned search and social into advertising empires.
- The tracking is now visible - the ChatGPT ad network (a Top Story today) mirrors the cross-site tracking Meta and Google pioneered.
- Classification-for-hire is booming - a new tool called Jev (covered September 19) sorts and tags bulk personal data like contacts and inboxes for pennies.
- Analysts are reframing AI economics around the customer - Nate's cost-to-serve argument (a Top Story today) treats customer data and outcomes, not model size, as the real battleground.
Sorting beats writing for bulk work
When the job is bucketing, tagging, or routing, a "multiple-choice" AI call is dramatically cheaper than a "write me a paragraph" call. Expect more products to quietly swap expensive text generation for cheap classification behind the scenes.
- 200x faster, 400x cheaper - the Jev classification tool answers with a label and a confidence score instead of an essay, and reportedly tagged 348 research papers for 18 cents.
- The research agrees - today's arXiv paper of the day finds that smart handling of context, not brute force, drives efficiency in AI coding agents.
- Cost guides are going mainstream - Nate's newsletter shipped a 15-technique "token saver" guide aimed at cutting waste.
Open AI models are being treated as public heritage
Centralized hosting means a takedown, license change, or policy shift can make widely-used weights disappear. The response is content-addressed, censorship-resistant distribution - the same idea that keeps other digital archives alive.
- A torrent lifeboat launched - Pirate Face (a Top Story today) mirrors 669,000+ open models so takedowns cannot erase them.
- Open models keep topping the charts - freely downloadable models from Qwen, DeepSeek, and others dominate this week's trending list (see Sections 7 and 13).
- The framing is political - Pirate Face calls its work "sovereign AI," a hedge against any single company controlling access.
The hard question is shifting from "can AI do it?" to "should it, and who checks?"
The frontier of the conversation has moved. The interesting fights now are about guardrails, disclosure, and who is accountable when an AI acts.
- Should your AI ever refuse you? - writer Zvi Mowshowitz argues AI in professional roles should follow ethical codes (like a lawyer or doctor), not blind loyalty.
- Over-reliance has a human cost - the "AI writes everything" account (a Top Story today) shows what happens when nobody reviews the output.
- Verification is becoming a product - a security-audit tool that independently double-checks AI findings is among the fastest-rising projects on GitHub today (see Section 12).
Creative AI & Media
Open video generation keeps climbing the charts
- LTX-2.5 turns a still image into short video and is one of the most-downloaded media models this week (1.61 million downloads).
- What it lets you do - animate a single photo or piece of art into a short clip without a studio or subscription.
- Try it: Lightricks/LTX-2.5 on Hugging Face
Open-source music generation you can run yourself
- YuE2-3B is a text-to-audio model for generating music, trending with strong community interest.
- What it lets you do - create original backing tracks or song sketches from a text description, on your own computer, no per-song fee.
- Try it: m-a-p/YuE2-3B on Hugging Face
Developer Tools & Infrastructure
Research & Models
What actually makes an AI coding assistant good
- The finding - a large study varied how coding agents plan, act, and manage information across four different AI models.
- Key result - managing what the AI "remembers" matters more as its memory budget shrinks, and mixing simple filtering with AI summarizing works best.
- The twist - for weaker models, planning improves accuracy; for stronger models, planning mainly cuts cost.
- See Section 16 for the full paper.
New open models keep raising the bar
- DeepSeek-V4.1-Flash and Qwen3.8-27B are among the most-downloaded models this week, both freely available (see Section 13).
- Why it matters - these are the kinds of open models that services like Pirate Face (Top Story) are racing to preserve, and they power a large share of independent AI projects.
Business & Industry
Surprising & Under-the-Radar
Signals to Track
Secure "key vaults" for AI agents
Today's llm-keys-ui is a tiny example of a pattern that will matter a lot: giving autonomous AI agents access to sensitive credentials without exposing those secrets in a chat log. Expect this to become standard plumbing. For ordinary people, it is the difference between AI helpers that are convenient and ones that quietly leak your passwords.
"Sovereign AI" as a movement, not a slogan
The idea that open models are public heritage worth preserving in censorship-resistant networks is moving from talk to code. If it catches on, it changes who ultimately controls whether a useful AI stays available. For everyday users, it means the free tools you depend on are less likely to disappear at a company's whim.
Classification endpoints replacing chatbots behind the scenes
Tools like Jev hint at a broad shift: for sorting, tagging, and routing, companies are switching from expensive text generation to cheap "multiple-choice" AI. You will rarely see it, but it will make many apps faster and cheaper - and it means more of your data is being auto-sorted than you might realize.
Top Repos Today
📜 License: MIT · 👤 By: company (Cloudflare)
🎯 Time to value: 15 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Independent verification reduces false positives | Requires an AI coding agent to run |
| Backed by a major infrastructure company | New project, still maturing |
| Free and open (MIT license) | Security output still needs human sign-off |

📜 License: MIT · 👤 By: startup
🎯 Time to value: 30 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Safe, isolated environments for risky automation | Computer-use AI is still error-prone |
| Cross-platform (Mac VMs plus cloud) | Setup takes some technical comfort |
| Includes benchmarks to measure quality | Broad scope can feel complex |

📜 License: MIT · 👤 By: individual (Google engineer)
🎯 Time to value: 10 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Huge, well-organized set of proven workflows | You still supply the AI agent to run them |
| Plain-text, easy to read and adapt | Value depends on your agent's quality |
| Very popular and actively used | Not a standalone app |

📜 License: Apache-2.0 · 👤 By: startup
🎯 Time to value: 60 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Fault tolerance for long, expensive runs | Only relevant if you train large models |
| Scales to very large model sizes | Requires serious graphics processing unit (GPU) hardware |
| Open license (Apache-2.0) | Steep learning curve |

📜 License: AGPL-3.0 · 👤 By: startup
🎯 Time to value: 45 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Keeps AI agents on your own infrastructure | Self-hosting takes real setup effort |
| Standardized, reproducible workspaces | AGPL license can deter some businesses |
| Mature and widely deployed | Overkill for solo developers |

📜 License: MIT · 👤 By: startup
🎯 Time to value: 20 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Purpose-built for agent applications | Young project, small community |
| Clean, modern developer experience | Fewer examples than mature frameworks |
| Free and open (MIT) | Best for new builds, not retrofits |

Top Models Today
👤 By: Alibaba (Qwen team) · 🎯 Task: text + image understanding
📐 Size: 27B
| ✓ Pros | ✗ Cons |
|---|---|
| Handles text and images together | 27B size needs a strong computer |
| Free to download and run | Setup requires technical skill |
| Huge, active user base | Not as simple as a hosted chatbot |

👤 By: DeepSeek · 🎯 Task: text + image understanding
📐 Size: very large (mixture-of-experts)
| ✓ Pros | ✗ Cons |
|---|---|
| Fast responses at large scale | Very large; demanding to self-host |
| Free and open | Best on serious hardware |
| Strong momentum and support | Newer, less battle-tested |

👤 By: Lightricks · 🎯 Task: image-to-video
📐 Size: mid-size
| ✓ Pros | ✗ Cons |
|---|---|
| Free image-to-video generation | Clips are short |
| Very popular and well-supported | Needs a capable GPU |
| Runs on your own machine | Quality varies by input |

👤 By: M-A-P community · 🎯 Task: text-to-audio (music)
📐 Size: 3B
| ✓ Pros | ✗ Cons |
|---|---|
| Smaller size, easier to run | Not studio-quality yet |
| Free and open | Music generation is hit-or-miss |
| Good for prototyping audio | Smaller community than big models |

👤 By: prism-ml · 🎯 Task: text generation
📐 Size: 27B (compressed)
| ✓ Pros | ✗ Cons |
|---|---|
| Runs a 27B model on modest hardware | Compression can slightly reduce quality |
| Very high download count | Still needs decent memory |
| Free and open | Format requires compatible software |

AI Launches Today
💰 Pricing: freemium · 🏷 Category: sales / AI assistant
💰 Pricing: freemium · 🏷 Category: careers / coaching
💰 Pricing: freemium · 🏷 Category: productivity / transcription
💰 Pricing: freemium · 🏷 Category: education / language
💰 Pricing: freemium · 🏷 Category: productivity / macOS
Snapshot
| Provider | Model | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Anthropic | Opus 5 | $5.00 | $25.00 | up to 1M |
| Anthropic | Sonnet 5 | $2.00 | $10.00 | up to 1M |
| Anthropic | Haiku 4.5 | $1.00 | $5.00 | up to 1M |
| OpenAI | GPT-5.6 Sol | $5.00 | $30.00 | large |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | large |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | large |
| Gemini 3.8 Flash | $0.75 | $3.75 | large | |
| Gemini 3.1 Pro | $2.00 | $12.00 | large |
What this means: The cheapest capable models keep getting cheaper. Google's Gemini 3.8 Flash and OpenAI's budget "Luna" tier are priced for high-volume, everyday tasks, while the premium tiers (Opus 5, GPT-5.6 Sol) cost 5 to 25 times more for the hardest work. Establishing baseline; changes will be flagged in future editions. Batch processing typically halves these prices, and caching can cut input costs further.
An Empirical Study of Harness Design for Coding Agents
Key finding: Managing the model's working context becomes more valuable as its memory budget tightens, and combining simple rule-based filtering with AI summarization delivers the best efficiency.
Why practitioners should care: It suggests you can get better, cheaper results from an AI coding assistant by improving its scaffolding rather than paying for a bigger model - and that planning helps weaker models get answers right but mainly helps stronger models save money.






Member discussion