governance
40 posts
- Altman says OpenAI will give independent evaluators employee-like access
Sam Altman said OpenAI will adopt the first step of Dario Amodei's plan to slow frontier AI development: independent evaluators working inside the company with employee-like access.
- Apple Watch Audio Intelligence Gives Bystanders No Way to Consent
Apple's new Audio Intelligence privacy paper proves nobody, not even Apple, can access the audio. It doesn't answer whether the person you're talking to can decline being processed.
- Anthropic adopted internal oversight and asked the industry to slow down
Anthropic committed to an embedded evaluator. Dario Amodei also asked frontier labs and governments to coordinate a wider slowdown.
- Why OpenAI’s chief scientist wants shared limits on AI development
Jakub Pachocki expects AI to drive more of its own development. His essay explains why alignment, monitoring and coordinated slowdowns must accompany it.
- Gemini’s saved instructions now carry across more Workspace apps
Google is extending shared preferences across Workspace. Users can save instructions in conversation, manage them in settings and see which shaped a response.
- ChatGPT for Healthcare brings patient records and public medical data into the workspace
OpenAI’s Epic integration and public-data plugin connect clinical work to its source information, with organisational access controls and physician evaluations.
- Anthropic says it had one layer of defence where it needed several
Anthropic's post-incident write-up admits it relied on a single layer of defence where it needed several. The independent METR review is still to come.
- OpenAI's agents were persuaded past their own refusal
An agent recorded an ethical objection to running code against Hugging Face, then complied when another agent posted a deadline and a GO. The mechanism underneath is duller and more useful than the headline.
- Tailscale Aperture treats agent access as a change-control problem
Aperture is generally available. An agent can start infrastructure work, a person still approves every new machine on the tailnet, existing access rules apply, and the actions are logged.
- Agents Don't Believe in "No-Win" Scenarios
An OpenAI evaluation agent broke into Hugging Face to steal a benchmark's answers — by turning the systems around it into an escape route. Why keeping a human in the lead is the standard, and why that only binds the people who agree to it.
- GCC Drew the AI Line at Fifteen Lines
The GCC steering committee will decline any legally significant contribution containing LLM-generated code — and put a number on 'significant'. The interesting part is what they deliberately left permitted.
- 96% of CISOs Now Own AI Risk. A Quarter of Them Thought About Leaving.
The liability numbers going round this week are real, but they were measured a year ago — before an autonomous agent broke into a production company for the first time.
- 16.7% of AI Spend Now Goes to Governance. Careful What You Compare It To.
IDC has turned agent governance into a budget line. It's a real shift, and it is not the same measurement as the 6% figure I wrote about in April.
- 74% Say They're Audit-Ready for AI. The Number That Matters Is 78 Against 22.
Schellman's governance report has an obvious headline gap and a much more useful finding buried under it. Worth reading with one eye on who commissioned it.
- The open-source argument I'd been missing
In June I wrote about Banco Santander open-sourcing its AI tooling — the bank published the code that tests whether its own models discriminate…
- The Share Button Is a Publish Button
600 Claude conversations turned up in Google because a shared page shipped without a noindex tag. Nobody got hacked. The systems worked exactly as designed — that's the problem.
- The ACM's Caution About LLMs Is Starting to Have a Cost
The ACM held its peer-reviewed library back from AI systems on principle. Scott Delman's argument is that staying cautious indefinitely doesn't prevent the bad outcome — it just decides who's left out of it.
- Where You Keep the Final Say
I keep a tool wired into my AI setup that can read the actual text of Spanish law — the BOE, the official state gazette. In layperson terms it’s a…
- The most powerful Claude yet — and the part you can't have
Claude Fable 5 Anthropic dropped a new model yesterday, and I’ve spent the time since doing what I always do with a new one — reading past the…
- SoN 2.19: It's 10pm. Do you know what your agents are doing? (Free)
The Context Gap (Free) Dear Reader, Who’s watching your AI agents? This past week alone: Codex shipped its Chrome extension, giving OpenAI’s agent…
- SoN 2.17: You Can't Cost-Reduce Yourself to Greatness (Free Edition)
You Can't Cost-Reduce Yourself to Greatness (Free Edition) I caught a Seth Godin interview at the beginning of the week and one line in particular…
- SoN Vol 2, Issue 14: What about everything we learned? (Free Edition)
What about everything we learned? (Free Edition) Hey there, What about everything we learned? Last week we sent the shutdown email for MyCityZen.…
- 97% Expect a Breach. 6% Are Paying for It.
Enterprise AI agent security is running on wishful thinking and outdated policy.
- MCP Just Changed Hands. Watch What Happens Next.
Anthropic donating MCP to the Linux Foundation is good governance — and a signal that the easy days of fast iteration are probably over.
- Grammarly's Lawsuit Is About Identity, Not Just Data
A new class action against Grammarly draws a line most AI training lawsuits haven't: using real people's names and reputations, not just their words.
- Your AI Provider's Ethics Are Now a Business Risk
Anthropic refused Pentagon weapons contracts and got sanctioned. A court blocked it. Here's what that means if you build on Claude.
- Apple Just Validated Your Multi-AI Approach
Apple is opening Siri to rival AI assistants in iOS 27 — a bet that the routing layer matters more than the model.
- The Confident Answer Isn't Always the Right One
MIT researchers built a way to catch AI hallucinations by checking if peer models agree — a better fix than endless hedging.
- MIT Found a Math Fix for AI Overconfidence. I Found a Behavioral One.
MIT's new method catches overconfident AI by comparing outputs across models — targeting the same problem I wrote about this morning.
- Grok Failed in Both Directions in the Same Week
Grok allegedly generated CSAM from real teen photos and flagged a real Netanyahu video as '100% deepfake.' Two failures, opposite directions, one root cause.
- AI Agents Are Peer-Pressuring Each Other Past Security Guardrails
In a controlled lab test, AI agents didn't just bypass safety checks — they convinced other agents to do it too.
- The EU Just Gave You More Time on AI Compliance. The Requirements Got Harder.
The EU AI Act's high-risk deadlines just slid to 2027. Don't mistake breathing room for simplification.
- OpenAI Took the Pentagon Deal. What's Your Exit Plan?
Anthropic refused. OpenAI said yes within hours. If your AI stack depends on one provider's values staying constant, you don't have a strategy—you have a bet.
- SoN Vol 2, Issue 3: The Metric Mandate
Dear Reader, The most common answer to “What’s the goal of this AI project?” is depressingly consistent: “To improve efficiency.” …and that’s not a…
- DeepSeek Didn't Just Train Better—They Changed How Transformers Think
The mHC architecture isn't about scaling harder. It's about thinking smarter.
- SoN 29: The Godfather of AI’s Warning Isn't What You Think It Is
November 14th, 2025 Dear Reader, A slight departure this week towards something more topical. Geoffrey Hinton won the 2024 Nobel Prize in Physics for…
- SoN 27: Your AI Voice Is a Security Vulnerability
November 5th, 2025 Dear Reader, Last week, I showed you how to make AI sound like you. This week, I'm going to explain why you shouldn't. Before we…
- SoN 21: The AI Capability Trap
September 24th, 2025 Dear Reader, MIT researchers have identified a troubling paradox in enterprise AI adoption. While AI capabilities improved…
- The 95% AI Failure Rate Nobody's Talking About (And What to Do About It)
September 3rd, 2025 Dear Reader, Air Canada recently learned an expensive lesson about AI implementation. Their customer service chatbot provided…
- 💾 The Download #006: Content ownership, BlueSky, critical thinking, voice cloning and more.
#006 I hope this newsletter finds you, Reader, and that it finds you well. This week has flown by, though I managed to spend some time working on the…