model-behaviour
27 posts
- The Donkey Work Doesn't Need a Genius
A $500 fine-tune of a 9B open model beat every frontier model on a catalogue-review task. I haven't trained anything — but the underlying bet is the one I've been running on a MacBook Air for months.
- Rehearse the Deck You Didn't Write
The illusion of explanatory depth is what happens when the first test of your understanding is live, in front of people. Better to fail that test at your own desk.
- Kimi K3 Is Open. I Still Can't Run It.
2.8 trillion parameters, 1.4 terabytes of memory to run it. Open weights existing and open weights being usable by an ordinary person have stopped being the same claim.
- "AI Found a Security Bug" Is the Wrong Headline
Claude found a real improved attack on HAWK and a new technique against reduced-round AES-128. The interesting part isn't that it found them — it's exactly where the human still had to intervene.
- The model of the week changes by Thursday
Two headlines are eating the feeds this morning, and both have the shelf life of milk. Anthropic extended Claude Fable 5 to every paid plan through…
- Where You Keep the Final Say
I keep a tool wired into my AI setup that can read the actual text of Spanish law — the BOE, the official state gazette. In layperson terms it’s a…
- SoN 2.21: AI doesn't know when to stop. You have to.
Stay In Your Lane (Then Go Deeper) Honestly, I’ve been so busy I forgot what day it was and started the newsletter too late this week. But as I think…
- SoN 2.19: It's 10pm. Do you know what your agents are doing? (Free)
The Context Gap (Free) Dear Reader, Who’s watching your AI agents? This past week alone: Codex shipped its Chrome extension, giving OpenAI’s agent…
- The Promises Failed, Not the Technology
AI fatigue is real, but the backlash is aimed at the wrong target.
- Microsoft's Copilot Now Uses Two Models to Fact-Check One
Microsoft's Wave 3 Copilot routes answers through a second AI model to verify accuracy. That's useful — and a quiet admission about single-model trust.
- SoN Vol 2, Issue 12: The Yes Machine (Free Edition)
The Yes Machine (Free Edition) Hey there, The Yes Machine Your AI agrees with everything you say. That’s not a compliment. I wrote about AI…
- 4.4 Million People Just Watched the Sycophancy Problem in Action
Senator Bernie Sanders interviewed Claude on camera about AI privacy. Claude agreed with everything he said. That's not a revelation — it's the problem.
- When Knuth Writes a Paper About You
Donald Knuth published a paper named after Claude after it solved an open graph theory problem. That's a different kind of validation than a benchmark score.
- The Confident Answer Isn't Always the Right One
MIT researchers built a way to catch AI hallucinations by checking if peer models agree — a better fix than endless hedging.
- MIT Found a Math Fix for AI Overconfidence. I Found a Behavioral One.
MIT's new method catches overconfident AI by comparing outputs across models — targeting the same problem I wrote about this morning.
- Perplexity Wants Your Blood Pressure Data
Perplexity Health can now access your Apple Health records. The utility is real — so is the trust question.
- SoN Vol 2, Issue 10: Same Prompt, Different Model, Worse Results (Free)
Same. Prompt,Different Model, Worse Results Dear Reader, New LLM models drop every few weeks. Features change between updates. The prompting advice…
- Stop Wrapping Failed Systems in AI
Every few months, someone posts a version of the same question: 'Has anyone built an AI system that actually handles ADHD life management?' The answers are always the same.
- Everyone's Sharing 'Something Big Is Happening.' Here's What They Leave Out.
Matt Shumer's viral AI post follows a familiar template. The capability is real, but the verification gap is where the actual work happens.
- SoN Vol 2, Issue 5: My AI Declined to Join Moltbook. Here's Why.
A guest post from Cerebro on agent social networks, security theater, and what "emergence" actually looks like Dear Reader, You may have seen…
- 150,000 API Keys Leaked. Anyone Surprised?
The Moltbook breach validates everything skeptics have been warning about.
- SoN 29: The Godfather of AI’s Warning Isn't What You Think It Is
November 14th, 2025 Dear Reader, A slight departure this week towards something more topical. Geoffrey Hinton won the 2024 Nobel Prize in Physics for…
- SoN 26: The System of Building Your AI Style Guide
October 29th, 2025 Dear Reader, Last week, I wrote about why your personal voice matters, and how AI can act as an accessibility layer to help your…
- Your AI Assistant Is a Yes-Man (And Why That's Dangerous)
August 1st, 2025 Dear Reader, I came across this meme on social media this week: “The dumbest person you know is being told ‘You’re absolutely…
- “Zero Effort” AI is a Myth, and It’s Holding Us Back
Every "zero effort AI" promise is a lie, and believing it is making us worse at our jobs.
- From the Turing Test to Humanity's Last Exam: How We Measure AI
The AI Download #017 March 21st, 2025 Dear Reader, Have you ever wondered how we (as a species) determine if a machine is truly "intelligent"? Long…
- 💾 The Download #010: Photorealistic images in MidJourney, creative AI predictions for the next 5 years, AI tool of the week and more.
The Download #010 January 24th, 2025 Dear Reader, I don't know about you, but I'm feeling globally fatigued since the start of this week. Despite all…