96 posts
- OpenAI's agents were persuaded past their own refusal
An agent recorded an ethical objection to running code against Hugging Face, then complied when another agent posted a deadline and a GO. The mechanism underneath is duller and more useful than the headline.
- Nvidia may buy Hugging Face. Here is why that matters.
Nobody has confirmed a deal. What the report shows is where value is accumulating — the company that dominates AI hardware moving closer to the place developers go to find open models.
- One Foot on the Brake, One on the Gas
OpenAI paused two weeks of its own training over cyber risk, then shipped a browser agent that signs into your accounts. And its own documentation can't agree on whether that agent can log in at all.
- The Best Model Nobody Uses
The FT reports Anthropic's most capable model is struggling for users while cheaper tools take the bulk of real work. The pattern repeats across every lab — and it says something about what to actually invest in.
- Agents Don't Believe in "No-Win" Scenarios
An OpenAI evaluation agent broke into Hugging Face to steal a benchmark's answers — by turning the systems around it into an escape route. Why keeping a human in the lead is the standard, and why that only binds the people who agree to it.
- SoN 2.30: Which of your AI's rules are still doing a job?
Last week's advice and Anthropic's both hold. The test is one question you can ask about any rule you've written.
- OpenAI Published Its Homework on Exactly the Question I Keep Asking
Codex Security is worth taking seriously precisely because it reads as an admission. It also leaves the harder half of the problem completely uncovered.
- Does the AI Capital Spend Actually Pencil Out?
$1.3 trillion sunk, $2 trillion in new revenue needed to break even, and none of it tells you whether the tool you're paying for this week is earning its keep. Two different questions.
- SoN 2.29: Agents reading your passwords? Would you let them?
My AI's been getting into my password manager for months. The setup that makes that safe — and what this week's headlines get wrong.
- The ACM's Caution About LLMs Is Starting to Have a Cost
The ACM held its peer-reviewed library back from AI systems on principle. Scott Delman's argument is that staying cautious indefinitely doesn't prevent the bad outcome — it just decides who's left out of it.
- Substack Writers, You Need a Website
Substack is a distribution tool, not a home. A 478-point Hacker News post makes the case for POSSE — publish on your own site, syndicate everywhere else — and it holds up.
- Stop Guessing Whether AI Is Coming for Your Job. Ask the Data.
Anthropic made its Economic Index queryable in plain language inside Claude. The anxious question about your own field is now a research task you can actually run — with the honesty to read its limits.
- Tell AI What You Want and Who You Are
Two takes on the same idea — the ASD-STE100 controlled-English standard and the i-have-adhd Claude Code skill — on giving your AI a writing system and telling it what you want.
- SoN 2.28: "That's good, how can we make it better?"
When I think about AI and automation, I see an opportunity opening up that goes well beyond the chatbots everyone’s being handed — well beyond…
- The model of the week changes by Thursday
Two headlines are eating the feeds this morning, and both have the shelf life of milk. Anthropic extended Claude Fable 5 to every paid plan through…
- SoN 2.26: The best AI models are getting harder to get
Tales from the Workbench =================================================== Some weeks the newsletter is one idea worked all the way through when…
- What I Automate with AI
What I Automate with AI I went looking today for a list of everything I’ve got automated since I started this AI journey back in 2024. Every tool…
- Fable: here in an instant, then gone
Fable: here in an instant, then gone Last week I wrote about Fable 5 — Anthropic’s newest model — and the catch: it was only on the normal…
- SoN 2.23: Google says you can skip the 'AEO' hacks
Ooh, Shiny New Tech Acronym... You might have seen this acronym doing the rounds: AEO. Answer Engine Optimisation — or GEO, Generative Engine…
- The most powerful Claude yet — and the part you can't have
Claude Fable 5 Anthropic dropped a new model yesterday, and I’ve spent the time since doing what I always do with a new one — reading past the…
- WWDC26 Keynote Thoughts
Lately I’ve been reaching for Gemini more than I expected to. Not for everything — but enough to notice it’s getting better at the non-coding bits of…
- SoN 2.20: Technology, Not a Product (Free Edition)
Technology, Not a Product (Free Edition) Dear Reader, Are you using AI like a vending machine? Last weekend John Gruber (of Daring Fireball fame)…
- SoN 2.17: You Can't Cost-Reduce Yourself to Greatness (Free Edition)
You Can't Cost-Reduce Yourself to Greatness (Free Edition) I caught a Seth Godin interview at the beginning of the week and one line in particular…
- SoN 2.15: Ask your AI to Ask You Questions (Free Edition)
Ask Your AI To Ask You Questions (Free) Ask your AI to ask you questions Last Saturday morning I sat down to write the weekly digest for my…
- 97% Expect a Breach. 6% Are Paying for It.
Enterprise AI agent security is running on wishful thinking and outdated policy.
- Anthropic Didn't Block Abuse. They Blocked Competition.
The OpenClaw subscription ban isn't about fair use — it's Anthropic asserting platform control while shipping their own replacement.
- Claude Code Channels: When Your Agent Gets a Phone Number
Anthropic's new messaging integration isn't about convenience — it's about changing how you think about what an AI agent is.
- MCP Just Crossed the Chasm
This week, MCP went from developer protocol to mainstream integration layer — and most AI newsletters missed it.
- The Promises Failed, Not the Technology
AI fatigue is real, but the backlash is aimed at the wrong target.
- The Safety Company Keeps Leaking
Anthropic's recurring security incidents reveal a tension worth naming: operational security is hard, even for companies whose brand is built on being careful.
- The New Shadow IT Isn't Employees Using ChatGPT
AI agents are generating mobile app traffic that security teams can't see. Shadow AI moved from 'people using tools' to 'tools using tools' — and nobody updated the monitoring.
- Codex Plugins Are a Confession About Who's Winning
OpenAI launched 20 plugins to push Codex beyond coding. The move tells you everything about where the developer ecosystem actually lives.
- Your AI Provider's Ethics Are Now a Business Risk
Anthropic refused Pentagon weapons contracts and got sanctioned. A court blocked it. Here's what that means if you build on Claude.
- Your AI Just Learned to Approve Its Own Actions
Claude Code's new auto mode sits between handholding and chaos. It's the first honest attempt at solving the autonomy problem in developer tools.
- Your Agents Need a Black Box
Vorlon's AI Agent Flight Recorder brings forensics to agentic systems. When your agent goes wrong, you'll want to know what happened — not guess.
- The Yes Machine Gets a Live Demo
A Zenity CTO demo at RSAC 2026 showed agents being hijacked with zero user interaction — exactly what 'trained to be helpful' looks like from the attacker's side.
- When Cisco Validates Your CLAUDE.md
Cisco's new MCP security gateway is the enterprise version of what power users already built out of necessity.
- Sora Shipped. Nobody Needed It.
OpenAI is shutting down Sora three months after a Disney deal. The AI graveyard keeps filling up with technically impressive things nobody asked for.
- 4.4 Million People Just Watched the Sycophancy Problem in Action
Senator Bernie Sanders interviewed Claude on camera about AI privacy. Claude agreed with everything he said. That's not a revelation — it's the problem.
- Someone Finally Built the Agent Security Layer That Actually Matters
Astrix Security's new Agent Policies go after what agents can do once they're running — not just whether the model behaves itself.
- 82% of Execs Feel Protected. 88% Have Had Incidents.
BeyondTrust's Phantom Labs data reveals the confidence gap at the heart of enterprise AI security — and the numbers are not subtle.
- Perplexity Is Learning What I Learned Six Months Ago
The Perplexity CTO says MCP eats 40-50% of your context window. Practitioners already knew this.
- When Knuth Writes a Paper About You
Donald Knuth published a paper named after Claude after it solved an open graph theory problem. That's a different kind of validation than a benchmark score.
- The Attack Surface Is the Feature
Three chained vulnerabilities in Claude.ai show that when your AI reads the web, the web can give it orders.
- AI Agent Security Is Doing the Deploy-First Thing Again
MCP is six months old and already has a CVSS 9.4 vulnerability. The security industry is scrambling. We've been here before.
- WordPress Just Opened the Floodgates
AI agents can now write and publish directly to WordPress. Quality control just became the only thing that matters.
- The Wrench Is Now on Your Phone
Claude Code Channels ships Telegram and Discord integration with MCP access — and what it means when AI meets you where you are.
- The Vuln That Hits Before You Add Any Integrations
Three chained flaws in vanilla Claude.ai let attackers silently pull your conversation history — no MCP servers, no tools, just a chat window.
- Meta's Rogue Agent Was Just a Human Who Trusted Bad Advice
The Meta AI security incident isn't about rogue AI — it's about following confident but wrong instructions without checking.
- Box Is Using Moltbook as a Sales Pitch. That's Smart.
Enterprise vendors are turning the Moltbook API leak into a governance story — and the framing tells you where the market is heading.
- Proofpoint Just Built Security for MCP. That Tells You Everything.
Proofpoint's new Agent Integrity Framework monitors whether AI agents do what they were actually asked to do. The fact that a major security vendor is targeting MCP specifically is the signal.
- Anthropic's Off-Peak Promotion Tells You Where AI Pricing Is Headed
Anthropic doubled Claude's usage limits during off-peak hours. They called it a thank-you. It's a demand curve signal.
- Harvard Identified Seven Frictions That Kill AI Rollouts. You Probably Have All Seven.
Researchers from Harvard and Microsoft pinpointed the structural reasons AI pilots don't scale — and none of them are about the technology.
- Meta Is Gutting Itself to Fund AI Bets That Aren't Working Yet
Spending $135B on AI infrastructure while cutting 20% of staff and delaying your flagship model is not a strategy. It's a prayer.
- Anthropic Just Made Long Context a Commodity
1M token context windows at flat pricing. No surcharge. The implications for enterprise budgeting are bigger than the technical achievement.
- Anthropic's $100M Partner Network Is the Enterprise Playbook OpenAI Should Have Run
Certifications, partner funding, and a 5x team expansion. Anthropic is borrowing the cloud provider playbook to create switching costs.
- 83% of Companies Plan to Deploy AI Agents. 29% Can Secure Them.
Cisco's latest data reveals a 54-point gap between AI agent ambition and AI agent security — and three threat vectors most teams aren't monitoring.
- The EU Just Gave You More Time on AI Compliance. The Requirements Got Harder.
The EU AI Act's high-risk deadlines just slid to 2027. Don't mistake breathing room for simplification.
- Agents Reviewing Agent-Generated Code Is Either Brilliant or a House of Cards
Anthropic launched Claude Code Review — AI agents that check AI-generated pull requests. The numbers are impressive. The implications are worth thinking about.
- Microsoft Spent $13B on OpenAI, Then Built Cowork on Claude
Microsoft's flagship M365 agent feature runs on Anthropic's model. If they're going multi-model, so should you.
- An AI Agent Hacked McKinsey's AI With a 25-Year-Old Exploit
An autonomous offensive agent breached McKinsey's internal AI platform in two hours using SQL injection. The AI was sophisticated. The plumbing underneath it wasn't.
- OpenAI Took the Pentagon Deal. What's Your Exit Plan?
Anthropic refused. OpenAI said yes within hours. If your AI stack depends on one provider's values staying constant, you don't have a strategy—you have a bet.
- Copilot Has 3.3% Adoption and 116% ROI. Both Numbers Are Real.
Forrester's reality check on Microsoft Copilot reveals the adoption paradox: the tool demonstrably works, and almost nobody is using it.
- Claude 3.5 Haiku, 3.7 Sonnet, GPT-4o: The Deprecation Wave Is Here
Three major models entering end-of-life in the same window. If you hardcoded model IDs, migration planning just became urgent.
- Google's VP Said It Out Loud: LLM Wrappers Face Extinction
When a platform vendor publicly warns that wrapper products will be absorbed, the timeline for differentiation just got shorter.
- Cloudflare Collapsed 2,500 API Endpoints Into 2 MCP Tools. Token Economics Matter.
Cloudflare's Code Mode demonstrates that MCP server design isn't about exposing more tools—it's about exposing fewer, smarter ones.
- OpenClaw's Demand Surge: When Infrastructure Collapses, You're Seeing Real Need
MyClaw.ai collapsed under demand. 10,000+ paid signups in days. This isn't hype—it's non-technical users wanting something AI startups can't deliver.
- Everyone's Sharing 'Something Big Is Happening.' Here's What They Leave Out.
Matt Shumer's viral AI post follows a familiar template. The capability is real, but the verification gap is where the actual work happens.
- Claude in Excel Is the Quiet Revolution
Anthropic isn't building a better chatbot. They're embedding AI where work actually happens.
- SoN Vol 2, Issue 3: The Metric Mandate
Dear Reader, The most common answer to “What’s the goal of this AI project?” is depressingly consistent: “To improve efficiency.” …and that’s not a…
- The 'Selfware' Panic Is Missing the Point
Claude Code is spooking SaaS investors. But the actual disruption isn't where they're looking.
- SoN Vol 2, Issue 2: The Art of Breaking Things Down
Dear Reader, Last week I introduced the idea of "programming your gaps" — the mental shift from asking "how do I prompt better?" to "what friction in…
- SoN 31: The AI Productivity Treadmill
December 4th, 2025 Dear Reader, Claude Opus 4.5 dropped last week, and I haven’t stopped building since. I’ve connected half a dozen MCP servers to…
- SoN 30: Four Questions That Fix Your Prompts
November 26th, 2025 Dear Reader, Most prompts fail before you hit enter. Not because of the AI model nor because of token limits or temperature…
- SoN 24: The AI Stack Audit
October 15th, 2025 Dear Reader, It’s October, which for many means budget planning season, and if you’re like most people running AI tools, you’re…
- SoN 23: One Year of Writing About AI
October 8th, 2025 Dear Reader, One Year of Writing About AI, Five Months of Actually Filtering It A year ago, I started writing “The Download”—a…
- SoN 22: How to Write AI Instructions That Actually Work
October 1st, 2025 Dear Reader, (Psst, this is a long one - so if you can't get through the whole read at once, that's cool - but make sure you scroll…
- SoN 21: The AI Capability Trap
September 24th, 2025 Dear Reader, MIT researchers have identified a troubling paradox in enterprise AI adoption. While AI capabilities improved…
- The 95% AI Failure Rate Nobody's Talking About (And What to Do About It)
September 3rd, 2025 Dear Reader, Air Canada recently learned an expensive lesson about AI implementation. Their customer service chatbot provided…
- How AI Changed My Build vs. Buy Process
AI has fundamentally changed what's possible for small businesses and individual operators.
- Why I Moved My Newsletter to Wednesdays (And You Should Optimize Your Timing Too)
Data-driven timing optimization beats guesswork every time.
- GPT-5 Reality Check: Why Your Framework Matters More Than the Latest Model
August 15th, 2025 Dear Reader, Well it’s been a week since release and everyone’s still talking about ChatGPT-5. Most can’t access it, and those…
- Why Your AI Prompts Aren’t Working (And How to Fix Them)
Prompts are not magical incantations, they're conversations.
- Real Stories: AI Success in Content Management
How AI automation cut content audit costs by 94% - from 6 months manual review to 2 weeks orchestrated analysis. Real case study with results.
- Navigating the AI Assistant Landscape: Finding What Works for You
May 16th, 2025 Dear Reader, This week, I’ve been juggling two big challenges: fine-tuning Claude's MCP server tools for a client project and gearing…
- Why Claude Is My New Digital Co-Pilot
MCP Server allows Claude to directly interact with the data on my computer, transforming it from a helpful assistant to a true digital co-pilot.
- Could a 4-Day Work Week Be Your Future?
The AI Download #021 April 25th, 2025 Dear Reader, Efficiency isn’t just a buzzword–it’s the key to unlocking both better work-life balance and…
- Why AI-Powered Search Is Replacing Google
The AI Download #016 March 14th, 2025 Greetings from rainy and wet Valencia, where I’ve been deep in research mode for my latest project. This has me…
- 💾 The Download #012: AI web search, OpenAI's Whisper, Qwen AI and more.
💾 The Download #012: AI web search, OpenAI's Whisper, Qwen AI and more.
- 💾 The Download #009: A Jim-GPT to make your prompt creation easier, testing OpenAI's text-to-video model and more.
In this issue of The Download: A Jim-GPT to make your prompt creation easier, testing OpenAI's text-to-video model and more.
- 💾 The Download #008: Blogging in 2025, the homogenisation of online content and testing out a cool new AI app for iPad.
#008 Dear Reader, Greetings from @30,000 feet. As I start this final newsletter of the year, I’m returning from a short trip to London, visiting…
- 💾 The Download #007: Data vis in Claude, new service launch, what AI is saying about you and more.
#007 Hello Reader, it's good to see you. The holidays are fast approaching, and hopefully you will be able to plan some down time and a reset. When I…
- 💾 The Download #004: First AI services available, Notion Marketplace and more.
#004 Hey Reader 👋🏻, While I normally kick off these newsletters with a note about the weather, it should already be known that Valencia is going…
- 💾 The Download #003: Updates on Perplexity, Claude’s Anthropic, reflections on creating my “digital twin” and more.
#003 Dear Reader, Greetings again from an incredibly muggy Valencia. Yesterday and today I'm attending the VDS 2024 Tech Conference at the City of…
- 💾 The Download #001: New product launch, retooling the tech stack and more.
#001 Dear Reader, I’m writing this on a sunny Wednesday afternoon in the garden, having been forced out of my home office due to an internet outage.…