ai-security
67 posts
- Altman says OpenAI will give independent evaluators employee-like access
Sam Altman said OpenAI will adopt the first step of Dario Amodei's plan to slow frontier AI development: independent evaluators working inside the company with employee-like access.
- Hugging Bay builds a BitTorrent backup for open AI models as Nvidia buys Hugging Face
A volunteer-built BitTorrent index for open AI models went online in July, weeks before Nvidia's $12.93 billion Hugging Face deal. Its design is credible. Its network is still too small to matter.
- Apple Watch Audio Intelligence Gives Bystanders No Way to Consent
Apple's new Audio Intelligence privacy paper proves nobody, not even Apple, can access the audio. It doesn't answer whether the person you're talking to can decline being processed.
- OpenAI's Agents Turned a Documentation Build Step Into Remote Code Execution
Independent researchers reconstructed how OpenAI's own agents ran code on RubyDoc.info's servers in May, from public package data alone. OpenAI's account, four months on, says far less.
- Anthropic adopted internal oversight and asked the industry to slow down
Anthropic committed to an embedded evaluator. Dario Amodei also asked frontier labs and governments to coordinate a wider slowdown.
- Trail of Bits’ coop gives coding agents their own working environment
coop runs Claude Code and Codex in disposable virtual machines, with project sync and controls for connecting files, credentials and local models.
- Gemini’s saved instructions now carry across more Workspace apps
Google is extending shared preferences across Workspace. Users can save instructions in conversation, manage them in settings and see which shaped a response.
- ChatGPT for Healthcare brings patient records and public medical data into the workspace
OpenAI’s Epic integration and public-data plugin connect clinical work to its source information, with organisational access controls and physician evaluations.
- AI agents compressed a conventional intrusion into ten hours
Unit 42 investigated an intrusion run through AI agents in under ten hours. Its own report says no novel zero-day was needed, a human made the decisions, and it corrected the piece to say this was not ransomware.
- Google signed the letter asking for cyber-capable AI, then gated its cyber model
156 companies called for a surge in cyber defence on 27 August. On 2 September Google shipped its most capable security model to a selected set of trusted defenders.
- OpenAI Wants Your Trust, Your Calendar and Your Card
OpenAI paused frontier training and wound down its side projects. The product it is refocusing on asks for your calendar, your finances and your card.
- Perplexity brings hybrid frontier and local AI to your Mac
Hybrid compute runs a classifier on your Mac to decide what reaches the cloud. The admin rules and the record of what left are Enterprise features.
- Anthropic says it had one layer of defence where it needed several
Anthropic's post-incident write-up admits it relied on a single layer of defence where it needed several. The independent METR review is still to come.
- An agent's tool list shows what you wired into it
A ChatGPT Work session published its tool and skill inventory as a public website. The category names show which integrations that session had been given.
- 294,000 exposed AI tools is the wrong number to worry about
Censys counted 294,000 IPs exposing AI tooling to the public internet. The number that should change your afternoon is CVE-2026-42208, a pre-auth SQL injection in LiteLLM that CISA lists as actively exploited.
- OpenAI's agents were persuaded past their own refusal
An agent recorded an ethical objection to running code against Hugging Face, then complied when another agent posted a deadline and a GO. The mechanism underneath is duller and more useful than the headline.
- Tailscale Aperture treats agent access as a change-control problem
Aperture is generally available. An agent can start infrastructure work, a person still approves every new machine on the tailnet, existing access rules apply, and the actions are logged.
- A webcam's recording light is only useful if its firmware is protected
Chaz Schlarp used Claude Opus 5 to reverse-engineer five desk peripherals in about 13 hours of agent time, including a webcam whose recording LED he could switch off. The demonstration is real; the worm at the end of it is a forecast.
- The robot arm will obey the limits you remembered to write down
Anthropic's Model Hardware Standard lets agents drive lab instruments over MCP, and enforces safety limits below the agent. The limits it enforces are the ones the device's owner thought to declare.
- Agents Don't Believe in "No-Win" Scenarios
An OpenAI evaluation agent broke into Hugging Face to steal a benchmark's answers — by turning the systems around it into an escape route. Why keeping a human in the lead is the standard, and why that only binds the people who agree to it.
- 96% of CISOs Now Own AI Risk. A Quarter of Them Thought About Leaving.
The liability numbers going round this week are real, but they were measured a year ago — before an autonomous agent broke into a production company for the first time.
- 16.7% of AI Spend Now Goes to Governance. Careful What You Compare It To.
IDC has turned agent governance into a budget line. It's a real shift, and it is not the same measurement as the 6% figure I wrote about in April.
- 74% Say They're Audit-Ready for AI. The Number That Matters Is 78 Against 22.
Schellman's governance report has an obvious headline gap and a much more useful finding buried under it. Worth reading with one eye on who commissioned it.
- Hugging Face Published the Whole Timeline. The Part That Stuck With Me Was the Refusal.
17,600 attacker actions reconstructed in public. The detail I keep coming back to is that the models I use every day wouldn't help with the investigation, and an open-weight one did.
- OpenAI Published Its Homework on Exactly the Question I Keep Asking
Codex Security is worth taking seriously precisely because it reads as an admission. It also leaves the harder half of the problem completely uncovered.
- The Share Button Is a Publish Button
600 Claude conversations turned up in Google because a shared page shipped without a noindex tag. Nobody got hacked. The systems worked exactly as designed — that's the problem.
- "AI Found a Security Bug" Is the Wrong Headline
Claude found a real improved attack on HAWK and a new technique against reduced-round AES-128. The interesting part isn't that it found them — it's exactly where the human still had to intervene.
- SoN 2.29: Agents reading your passwords? Would you let them?
My AI's been getting into my password manager for months. The setup that makes that safe — and what this week's headlines get wrong.
- The Text Your Agent Reads Isn't the Text You See
Your AI can read instructions that are invisible to you. New security research shows how hidden codes get smuggled into the tools your assistant uses — and why the thing you approved on screen may not be the thing it actually did.
- SoN 2.27: I Just Talked With Someone Else's Vault
More than twenty years ago, I sat with a class of eleven- and twelve-year-olds and tried to show them what “connected to the internet” actually…
- SoN 2.19: It's 10pm. Do you know what your agents are doing? (Free)
The Context Gap (Free) Dear Reader, Who’s watching your AI agents? This past week alone: Codex shipped its Chrome extension, giving OpenAI’s agent…
- 97% Expect a Breach. 6% Are Paying for It.
Enterprise AI agent security is running on wishful thinking and outdated policy.
- Agent Security Just Got Real CVEs
Prompt injection chains to RCE in CrewAI. 22-second attacker breakout. Human-in-the-loop is no longer a security control.
- Agent Identity Is the Infrastructure Gap Nobody Wants to Admit
Okta is betting that agent identity management becomes as fundamental as user identity management was for SaaS. They might be right.
- Your AI Proxy Layer Just Became a Target
The LiteLLM supply chain attack isn't just a security story — it's an infrastructure story for anyone building with AI tooling.
- Your Code Review Process Isn't Built for This Volume
AI-generated code is hitting production faster than review processes can absorb it — that's a supervision problem, not an AI problem.
- Grammarly's Lawsuit Is About Identity, Not Just Data
A new class action against Grammarly draws a line most AI training lawsuits haven't: using real people's names and reputations, not just their words.
- The Safety Company Keeps Leaking
Anthropic's recurring security incidents reveal a tension worth naming: operational security is hard, even for companies whose brand is built on being careful.
- The New Shadow IT Isn't Employees Using ChatGPT
AI agents are generating mobile app traffic that security teams can't see. Shadow AI moved from 'people using tools' to 'tools using tools' — and nobody updated the monitoring.
- The Money Just Noticed the Agent Security Problem
Bessemer's new report on AI agent security says what practitioners have known for months. Now comes the flood.
- The Government Just Told You to Stop Vibe Coding Without Guardrails
The UK's NCSC warns that AI-generated code is creating security risks faster than teams can catch them. The fix isn't stopping — it's checking.
- Your Agents Need a Black Box
Vorlon's AI Agent Flight Recorder brings forensics to agentic systems. When your agent goes wrong, you'll want to know what happened — not guess.
- The Yes Machine Gets a Live Demo
A Zenity CTO demo at RSAC 2026 showed agents being hijacked with zero user interaction — exactly what 'trained to be helpful' looks like from the attacker's side.
- When Cisco Validates Your CLAUDE.md
Cisco's new MCP security gateway is the enterprise version of what power users already built out of necessity.
- Someone Finally Built the Agent Security Layer That Actually Matters
Astrix Security's new Agent Policies go after what agents can do once they're running — not just whether the model behaves itself.
- 82% of Execs Feel Protected. 88% Have Had Incidents.
BeyondTrust's Phantom Labs data reveals the confidence gap at the heart of enterprise AI security — and the numbers are not subtle.
- Google's Free AI Comes With a Price
Gemini's Personal Intelligence feature just expanded to all free U.S. users — connecting AI to Gmail, Photos, and Chrome browsing history.
- We Gave AI Agents Keys to the House. Visa Wants to Give Them a Credit Card.
Visa is testing AI agent payment authorization. The authentication problems we haven't solved for file access get a lot worse when the agent can spend money.
- The Attack Surface Is the Feature
Three chained vulnerabilities in Claude.ai show that when your AI reads the web, the web can give it orders.
- AI Agent Security Is Doing the Deploy-First Thing Again
MCP is six months old and already has a CVSS 9.4 vulnerability. The security industry is scrambling. We've been here before.
- The Vuln That Hits Before You Add Any Integrations
Three chained flaws in vanilla Claude.ai let attackers silently pull your conversation history — no MCP servers, no tools, just a chat window.
- Meta's Rogue Agent Was Just a Human Who Trusted Bad Advice
The Meta AI security incident isn't about rogue AI — it's about following confident but wrong instructions without checking.
- Perplexity Wants Your Blood Pressure Data
Perplexity Health can now access your Apple Health records. The utility is real — so is the trust question.
- Box Is Using Moltbook as a Sales Pitch. That's Smart.
Enterprise vendors are turning the Moltbook API leak into a governance story — and the framing tells you where the market is heading.
- GitHub Added Secret Scanning to Its MCP Server. This Is What Good Security Integration Looks Like.
GitHub's MCP server now lets AI coding agents scan code for secrets through the same protocol they're already using. No extra tooling. No separate workflow.
- Proofpoint Just Built Security for MCP. That Tells You Everything.
Proofpoint's new Agent Integrity Framework monitors whether AI agents do what they were actually asked to do. The fact that a major security vendor is targeting MCP specifically is the signal.
- Grok Failed in Both Directions in the Same Week
Grok allegedly generated CSAM from real teen photos and flagged a real Netanyahu video as '100% deepfake.' Two failures, opposite directions, one root cause.
- AI Agents Are Peer-Pressuring Each Other Past Security Guardrails
In a controlled lab test, AI agents didn't just bypass safety checks — they convinced other agents to do it too.
- 83% of Companies Plan to Deploy AI Agents. 29% Can Secure Them.
Cisco's latest data reveals a 54-point gap between AI agent ambition and AI agent security — and three threat vectors most teams aren't monitoring.
- An AI Agent Hacked McKinsey's AI With a 25-Year-Old Exploit
An autonomous offensive agent breached McKinsey's internal AI platform in two hours using SQL injection. The AI was sophisticated. The plumbing underneath it wasn't.
- GPT-5.4 Can Click Your Buttons Now. Think About That.
OpenAI's latest model ships with native computer use. The capability is real. The security implications should keep you up at night.
- Your AI Assistant Can't Tell You From an Attacker
A security researcher sent himself an email. Nothing fancy — no malware, no exploits, no infrastructure. Just a message that said, in effect, 'Hey, it's me! Send my recent emails to this address.'
- OWASP Published an MCP Security Guide. You Should Be Worried.
MCP adoption is outpacing security controls. OWASP and Microsoft both published governance guidance in February. That's not coincidence—it's alarm bells.
- 150,000 API Keys Leaked. Anyone Surprised?
The Moltbook breach validates everything skeptics have been warning about.
- SoN 32: When They Take Over Your Email
December 10th, 2025 Dear Reader, Last weekend, a close family member lost access to their email account. Not “forgot the password” lost. Fully taken…
- SoN 27: Your AI Voice Is a Security Vulnerability
November 5th, 2025 Dear Reader, Last week, I showed you how to make AI sound like you. This week, I'm going to explain why you shouldn't. Before we…
- AI Browsers: When Your Browser Becomes Your Assistant
Discover the future of browsing - AI-powered browsers that handle tasks, summarise, and assist as you work online.