Statistically Speaking
One Thing to Tell Your Friends
TL;DR
Hot off the Presses
An AI hallucination nearly started a naval confrontation
CNN reports that during the spring war with Iran, a US intelligence analyst used an AI chatbot that combined public information with classified signals data and concluded a Chinese vessel was carrying nuclear-weapon components. The analyst then used AI again to write it up as a standard intelligence report, which moved up the chain. The military drew up plans to intercept the ship, and armed personnel were preparing to board before officials discovered the cargo claim was an AI fabrication.
- Not an isolated case - sources told CNN that AI hallucinations have recurred across the intelligence community as these tools spread through government.
- Targeting is the danger zone - officials specifically flagged the military's rapid adoption of AI for targeting, where a hallucination can be fatal.
- The missing piece was provenance - no one downstream could easily tell which parts of the report came from a machine that makes things up.
South Korea sets data-breach fines at 10% of revenue
South Korea's privacy regulator (the Personal Information Protection Commission) enacted a new enforcement decree, effective September 13, 2026, that raises the maximum fine for serious data-breach violations from 3% to 10% of a company's annual revenue. The harshest penalties target firms that leak data on 10 million or more people through intent or gross negligence. A new rule also forces companies to warn affected users within 72 hours whenever there is a high risk of exposure, even before a breach is confirmed.
- The math is staggering - e-commerce giant Coupang paid 624.6 billion won for exposing 37.55 million records in 2025; under the new cap a comparable fine could reach into the trillions of won.
- Carrots as well as sticks - fines can be cut up to 40% for prior security investment and up to another 40% for fast detection and user notification.
- Accountability climbs the org chart - privacy officers at large firms now need board approval to be appointed or dismissed.
Claude Code now reads AGENTS.md, nudging the industry toward one standard
Anthropic's Claude Code (an AI coding assistant that runs in the terminal) added support for AGENTS.md in version 2.1.277, released September 18, 2026. When a project folder has no CLAUDE.md file, Claude Code now reads an AGENTS.md file instead - the same cross-tool convention that rival coding agents already use. Anthropic open-sourced its implementation as a reference for the community.
- A de facto standard is forming - AGENTS.md gives AI coding agents repository-specific instructions, and one shared file now works across competing tools.
- The HN crowd noticed - the changelog drew 309 points and 131 comments on Hacker News.
- Same release, quieter wins - background tasks now show a "waiting" status when they finish, and several crash and hang bugs were fixed.
Anthropic reportedly launched Claude Projects for parallel work
According to the AINews roundup, Anthropic shipped Claude Projects, described as multi-threaded cloud sessions coordinated from a single conversation, so multiple tasks run in parallel while sharing context. The same roundup notes Google released updated Gemini Managed Agents with credential-management tools that keep secrets out of the model's view, claiming a 30% cost reduction and 22% better cache efficiency.
- The shape of the shift - AI tools keep moving from one-off chats toward coordinated, always-on work sessions.
- Reported figures, single source - the parallel-sessions and cost claims come from a newsletter roundup and are presented as reported, not independently confirmed.
Creative AI & Media
Describe a building and get a 3D form that fits its neighborhood
- What it does - "CoMa" uses a vision-language model to generate a building's early 3D bulk so it matches the surrounding streets, scale, and density.
- Why it is notable - it was trained on 12,845 real Melbourne buildings paired with parcel outlines and aerial views, and combining map, geometry, and 3D inputs beat any single input.
- Who it helps - architects and planners doing early massing studies, where fitting the context is half the job.
Open video models keep climbing the charts
- What is moving - image-to-video models "LTX-2.5" and "MiniMax-H3" are among the most-downloaded models this week (see the Hugging Face section below).
- Why it matters - free, downloadable video generation keeps improving, narrowing the gap with paid tools.
Developer Tools & Infrastructure
Research & Models
Business & Industry
Surprising & Under-the-Radar
Public worry about AI risk is cascading into the open
Zvi Mowshowitz argues a "preference cascade" is underway, where people who privately feared AI danger now feel free to say so. He cites polling that nearly two-thirds of Americans see at least a moderate AI extinction risk, up roughly 15 points, and a median AI-researcher estimate near 18%. Why it is surprising: the shift is happening fast, yet he argues it is still far too weak for the stakes. The Zvi: The Preference Cascade
More reasoning did not make AI agents harder to manipulate
In a study of 3,600 shopping agents across six models, "nudges" like default options and social-proof messaging swayed them - and extra reasoning did not reliably help. It reduced vulnerability to default-option nudges but increased it for social-influence ones. Why it is surprising: smarter agents were not safer, just manipulable in different ways. arXiv: Nudge Susceptibility in GUI Agents
Attackers are targeting the humans behind open-source code
A security warning to the Rust community describes fake job offers used to trick prominent package maintainers into compromising their systems, linked to a supply-chain compromise of the widely-used arrayref package. The defensive tip: wait a few days before adopting brand-new releases. Why it matters: the weak point is people, not code. Simon Willison: attacks on Rustaceans
Debate: should you use any words an AI suggests?
Writer Thomas Ptacek argues for a hard rule - "you may not use a single word an LLM suggests to you" - treating AI only as a fact-checker and grammar tool, never a ghostwriter. The counter-view, from Ethan Mollick's essay on AI's untapped "capability overhang," is that today's models can already do weeks of skilled work and the real waste is under-using them. The tension: preserving an authentic voice versus leaving value on the table. Simon Willison: How to Write With an LLM · One Useful Thing: The Overhang
Signals to Track
Synthetic study material for tiny AI models
A new open dataset called QVAC Genesis III packs 191 billion tokens of synthetic STEM content, built by turning a small model's mistakes and successes into study material. A 1.7B model trained on it improved over 28% on one reasoning test. If this holds, capable AI that runs on cheap hardware gets easier to build - useful for anyone who wants private, on-device AI.
Web agents that remember how to use a website
"EconSkills" turns successful browsing sessions into reusable, step-by-step procedures with placeholders, so an agent can apply a known recipe to a similar task. It used fewer steps when a stored skill matched. For everyday users, this points to assistants that get faster and more reliable at repetitive online chores.
A quantum classifier doing a real industrial job
Researchers used a small two-qubit quantum classifier to diagnose electrical-transformer faults from gas analysis, folding in engineering domain knowledge and reporting strong accuracy on shallow, near-term hardware. It is early, but it is a concrete industrial use rather than a demo. If quantum-assisted diagnostics pan out, utilities could catch equipment failures sooner.
Mathematically proving AI-written code is safe
"MAGS" runs AI-generated code through a formal verifier (using the Dafny proof language) and only ships programs that pass, hitting a 100% success rate at producing guaranteed-safe programs across 220 examples in CUDA, shell, and robotics. As agents write more code than people can review, machine-checkable proof may be the only scalable safety net.
Top Repos Today
📜 License: MIT · 👤 By: company (Cloudflare)
🎯 Time to value: 15 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Structured, repeatable audit output | Only as good as the underlying agent |
| Free and open (MIT) | Findings still need human review |
| Backed by a major infra company | Narrow, security-only focus |

📜 License: Apache-2.0 · 👤 By: company (Alibaba)
🎯 Time to value: 20 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Hybrid checks reduce false positives | Setup takes some configuration |
| Permissive Apache-2.0 license | Best suited to teams already on PR workflows |
| Very popular and active | Quality varies by language |

📜 License: MIT · 👤 By: company (Tencent)
🎯 Time to value: 15 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Reuses existing logins safely | Browser automation can be brittle |
| Non-disruptive to your own browsing | Requires trust in the agent's actions |
| Free and open (MIT) | Still early and evolving |

📜 License: MIT · 👤 By: individual
🎯 Time to value: 10 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Large, well-maintained collection | You must match skills to your stack |
| Free and open (MIT) | Not a standalone app |
| Maintained by a respected engineer | Assumes familiarity with agent tools |

📜 License: MIT · 👤 By: company (Tencent Cloud)
🎯 Time to value: 30 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Self-hosted for privacy | Requires server setup and upkeep |
| Multi-user and multi-agent | Younger project, smaller community |
| Free and open (MIT) | More moving parts than a hosted tool |

📜 License: Proprietary (Anthropic commercial terms) · 👤 By: company (Anthropic)
🎯 Time to value: 10 minutes
| ✓ Pros | ✗ Cons |
|---|---|
| Deep terminal and repo integration | Not open-source (commercial terms) |
| Large, fast-moving user base | Requires a paid Anthropic plan |
| Now reads the shared AGENTS.md | Terminal-first workflow is not for everyone |

Top Models Today
👤 By: DeepSeek · 🎯 Task: image-text-to-text
📐 Size: 763B
| ✓ Pros | ✗ Cons |
|---|---|
| Permissive MIT license | 763B size is impractical to self-host for most |
| Frontier-scale quality | Best accessed via hosted providers |
| Very active community | Heavy compute to run locally |

👤 By: Alibaba (Qwen) · 🎯 Task: image-text-to-text
📐 Size: 27B
| ✓ Pros | ✗ Cons |
|---|---|
| Runs on a single strong GPU | Not frontier-level on the hardest tasks |
| Apache-2.0 permissive license | Multimodal setup adds complexity |
| Huge download base and tooling | Competition in this size is fierce |

👤 By: Alibaba (Qwen) · 🎯 Task: image-text-to-text
📐 Size: 180B
| ✓ Pros | ✗ Cons |
|---|---|
| Large, capable multimodal model | Custom (non-standard) license |
| Open weights, self-hostable | 180B size needs serious hardware |
| From a very active model family | Preview-stage, expect changes |

👤 By: MiniMax · 🎯 Task: image-text-to-video
📐 Size: 33B
| ✓ Pros | ✗ Cons |
|---|---|
| Strong open video generation | Custom community license |
| Large, active user base | Video generation is compute-heavy |
| Self-hostable | Clip length and control still limited |

👤 By: Lightricks · 🎯 Task: image-to-video
📐 Size: n/a
| ✓ Pros | ✗ Cons |
|---|---|
| Fast generation | Custom community license |
| Backed by a consumer-app maker | Quality trails the largest video models |
| Widely used and documented | GPU still required |

👤 By: OpenBMB · 🎯 Task: text-generation
📐 Size: ~3B
| ✓ Pros | ✗ Cons |
|---|---|
| Runs on phones and laptops | Limited vs large models on hard tasks |
| Apache-2.0 permissive license | Small size caps capability |
| Low cost to run | Best for narrow, on-device tasks |

AI Launches Today
💰 Pricing: freemium · 🏷 Category: AI workflow / productivity
💰 Pricing: freemium (10 free counts) · 🏷 Category: computer vision
💰 Pricing: free · 🏷 Category: developer tools / AI coding agents
💰 Pricing: open-source core (cloud in beta) · 🏷 Category: backend / LLM dev tools
💰 Pricing: free · 🏷 Category: developer analytics
Snapshot
| Provider | Model | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | Up to 1M |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | Up to 1M |
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | - |
| Gemini 3.1 Pro (Preview) | $2.00 | $12.00 | ≤200k tier | |
| Groq | GPT-OSS 120B | $0.15 | $0.60 | - |
What this means: No price changes for Anthropic or OpenAI's flagships versus September 17. The spread remains huge: Groq's open-model hosting is over 60x cheaper on input than OpenAI's flagship, so matching the model to the task still matters far more than any single price cut. (OpenAI and Groq figures are from third-party pricing pages and may lag official updates.)
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Key finding: On the AIME24 math benchmark, accuracy (Pass@3) rose 10.0% while token usage fell 27.9%; on AIME25 it reached 40.0% Pass@3, beating both compression-only and routing-only methods.
Why practitioners should care: It improves accuracy and cost at the same time rather than trading one for the other, which is directly useful for anyone paying per token for a reasoning model.







Member discussion