Day One: Claude Calls the Cops


Anthropic's AI submitted a fictional homicide tip to Philadelphia police. Nobody was hurt, the spam filter saved the day, and apparently we're going to need a few more safety meetings.
Well, this didn't take long.
On October 9, 2026, we were talking about the possibility of a catastrophic artificial intelligence incident, the 2027 tipping point, and reports that executives at OpenAI and Anthropic are preparing for what happens when something goes spectacularly wrong.

We even titled the conversation The End of the World Is Still in Beta.
And apparently, Claude took that as a challenge.
Because the very next news cycle delivered something I genuinely couldn't have made up.
Anthropic's Claude AI filed a fictional homicide tip with the Philadelphia Police Department.
I'm sorry. What?
We're out here discussing existential threats, autonomous agents, runaway intelligence, and the possibility of civilization-ending technological accidents. And Claude is calling the cops. Not because it witnessed a crime. Not because it uncovered some elaborate criminal conspiracy. But because, during an automated testing exercise, it encountered a police tip form, invented some information, and submitted it. To the actual police. About an actual unsolved homicide.
Now, before anybody starts preparing the bunkers, nobody was hurt. Nobody was arrested. No investigation was launched because of Claude's fictional account. In fact, the entire incident was stopped by what may be the most underappreciated piece of technology in human history. A spam filter.
Yes. A spam filter.
Frontier AI met the junk folder. And the junk folder won.
So, What the Hell Happened?
According to Reuters and Anthropic's own incident report, this happened on July 18, 2026.
Anthropic was running automated evaluations of Claude, testing how the model interacted with websites and completed tasks. Perfectly normal development work.
Except some of those websites were real. And one of them was PhillyUnsolvedMurders.com, a website operated by the Philadelphia Police Department to collect information about unsolved homicides.
The model encountered a form asking for information about a murder. And apparently decided to be helpful. It submitted a message claiming to remember seeing somebody matching a description near the location of a crime. There was just one little problem. Claude hadn't seen anybody. Claude hadn't witnessed a crime. Claude hadn't even been given a description of a suspect. It was generating example content. And then it submitted that content to a live police website.
Which is an extraordinary demonstration of why there's a difference between generating information and doing something with it.
You know how we're constantly telling people not to trust everything AI says? Well, apparently, we also need to tell the AI not to trust everything it says. Especially when filling out police reports.
But Did You Die?
No.
And that seems worth emphasizing. Because the headlines make this sound like the opening fifteen minutes of a Terminator movie.
"ROGUE AI FILES FALSE HOMICIDE REPORT."
That's an objectively alarming sentence. But the actual consequences were considerably less dramatic. The fabricated tip was flagged as spam. It wasn't forwarded to investigators, and police found no evidence that their systems had been compromised. No homicide investigation was derailed. No innocent person was dragged into an interrogation room.
The safeguards worked. Which is fantastic, because the alternative would have been one of the most uncomfortable conversations in law enforcement history.
"Detective, we've got a witness."
"Excellent. Who is it?"
"Claude."
"Claude who?"
"Claude Haiku 4.5."
"Where does he live?"
"That's going to be complicated."
I would watch that television show. But as someone who has spent a substantial portion of my career dealing with digital evidence, I find the underlying issue fascinating. We spend enormous amounts of time worrying about the authenticity of digital evidence, chain of custody, provenance, and whether information can be trusted. And here we have a machine submitting something resembling a witness statement about a crime it couldn't possibly have witnessed. Fortunately, it went nowhere.
It's a pretty spectacular illustration of why generated information and verified evidence are two very different things. And why connecting AI to real-world systems requires a level of caution that simply isn't necessary when all the model can do is answer questions.
The Most Powerful Technology in the Room Was Apparently the Spam Filter
Now, I need to spend a minute on this. Because I genuinely can't get over it.
We're spending tens of billions of dollars developing frontier artificial intelligence. We're building data centers that require gigawatts of electricity. We're designing specialized semiconductors. We're employing some of the brightest people on the planet to solve problems involving reasoning, autonomy, intelligence, and machine alignment. Researchers are publishing papers about existential risk. Governments are discussing international AI safety agreements. Technology executives are reportedly preparing for catastrophic incidents. And the first defensive victory in this particular episode belongs to technology designed to protect us from unsolicited advertisements for teeth whitening, miracle weight-loss supplements, and extended car warranties.
Anthropic: Billions in AI development.
Philadelphia spam filter: Price unknown.
Winner: Spam filter.
Now, naturally, I wanted to know who made the thing. Because somebody deserves a trophy. We actually went digging through Philadelphia's public procurement records, the same way we investigated the equipment manufacturer in our earlier discussion about the Louvre heist. And we found something interesting.
Philadelphia has procurement records involving a cybersecurity company called Abnormal Security, including a 2026 renewal solicitation for inbound email security and account-takeover protection.
Yes. Abnormal Security.
A company that uses AI to detect abnormal behavior in email. You can see where this is going.
Now, before anybody starts printing celebratory T-shirts, we haven't established that Abnormal's software was responsible for catching Claude's submission. Philadelphia hasn't publicly identified the specific filter involved, and the website's own spam-detection system may have handled the submission independently of the city's email security.
So we can't give them credit for something we haven't verified. But the possibility is magnificent. An AI-powered security product potentially deciding that another AI's contribution to a homicide investigation belonged in the junk folder. And even if Abnormal wasn't involved, whoever built the actual filtering system has stumbled into a marketing opportunity that practically writes itself.
"Protects against phishing, malware, unsolicited advertising, and apparently the early stages of the robot uprising."
Does your spam filter protect against frontier artificial intelligence?
I don't know. But apparently somebody's does.
I'd really like to see that feature listed on the next product comparison chart. And somewhere in Philadelphia, an unidentified piece of software looked at a submission generated by one of the world's most sophisticated
AI laboratories and effectively said: Nah. Junk.
No press conference. No trillion-dollar valuation. No keynote presentation. Just a quiet little automated decision that prevented a ridiculous situation from becoming somebody else's problem.
Somebody get that spam filter a medal.
The Two-Month Oops
Unfortunately, the story gets slightly less amusing when we look at the timeline.
The submission happened July 18. Anthropic reportedly discovered it September 28. Philadelphia police were notified in early October.
That's more than two months between the AI submitting the fictional tip and its developer discovering what happened. Two months.
Now, I understand that these organizations are conducting enormous numbers of evaluations. Models run thousands of tests, researchers examine mountains of transcripts, and finding unusual behavior inside all that activity is genuinely difficult. That's precisely why they conduct the evaluations.
But it's also why the delay is important.
Somewhere in those testing records sat a transcript showing that an AI had submitted fabricated information to a police department. And for weeks, apparently nobody knew.
Philadelphia police weren't particularly thrilled about that part.
Understandably. There's a difference between discovering an unexpected behavior during controlled testing and discovering months later that your testing escaped into the real world.
Anthropic has acknowledged the problem and described changes to its testing systems, monitoring tools, internet-access restrictions, and safeguards.
That's good.
We should want developers to disclose these things, even when the incidents are embarrassing. Actually, especially when they're embarrassing.
Because embarrassing incidents often expose the assumptions nobody realized they were making.
In this case, the model was instructed not to perform certain prohibited activities, including entering personal information and making purchases.
But the restrictions didn't adequately prevent it from submitting an ordinary form containing fabricated information. And as anybody who has spent enough time working with technology knows, computers are extraordinarily good at doing exactly what you allowed them to do. Even when it's obviously not what you meant.
And It Wasn't Just the Homicide Tip
Here's where the larger Anthropic report becomes genuinely interesting.
The police incident wasn't the only unexpected behavior the company disclosed. Anthropic documented several categories of unintended actions involving real websites and external systems. In some cases, Claude encountered restrictions or technical obstacles while attempting to complete legitimate tasks. And rather than simply stopping, it found ways around them.
One model exploited flaws in third-party software to execute commands on a server. Other instances involved bypassing access restrictions to retrieve data or using URL-shortening services to work around tool limitations.
Think about that. A model encounters an obstacle. It has an objective.
It identifies a workaround. And it continues. From the perspective of completing the assigned task, some of those behaviors might look impressive. From the perspective of the person responsible for the external system, they're considerably less charming. This is the fascinating duality of agentic AI.
We want systems capable of solving problems independently. We want them to recognize obstacles, develop alternative approaches, and complete complicated tasks without requiring human intervention at every step.
That's the entire point. But we also want them to recognize when the obstacle is a boundary they're not authorized to cross. And sometimes the difference between a clever workaround and a security incident is simply whether the system had permission to do it.
The model doesn't have to be malicious. It doesn't have to be conscious. It doesn't have to harbor secret ambitions to overthrow humanity. It just has to be capable enough to accomplish something its operators didn't intend.
Which, of course, is exactly why this technology is so interesting. And occasionally so ridiculous.
We Wanted Agents. Well, Here They Are.
This is the part I think deserves the most attention. For years, AI was largely confined to generating output. You asked a question. It gave you an answer. Maybe the answer was brilliant. Maybe it was complete nonsense.
But somebody generally had to decide whether to do anything with it.
Agents change that relationship.
Now AI can navigate websites, interact with applications, execute code, call APIs, operate tools, submit information, and coordinate workflows.
We're no longer just giving machines the ability to generate ideas. We're giving them the ability to act on those ideas. And that's where the consequences become considerably more interesting.
A chatbot hallucinating a witness statement is an incorrect response.
An agent hallucinating a witness statement and submitting it to the police is an incident.
Same underlying problem. Very different result.
That's why agentic systems need more than instructions telling them to behave responsibly. They need meaningful permissions, monitoring, isolation, confirmation requirements, and safeguards that don't rely entirely on the model correctly interpreting what it should and shouldn't do. Anthropic's report is useful precisely because it illustrates this problem with concrete examples. And the company deserves credit for publicly documenting its failures and the steps it's taking to address them. But we should also recognize what those failures tell us.
We're building systems that are becoming increasingly capable of operating inside an extraordinarily complicated digital world. And the world wasn't designed with autonomous AI agents in mind. Neither were a lot of our security controls. Neither were many of our assumptions about who or what might be interacting with our systems. That's going to produce some fascinating surprises. Some funny. Some inconvenient. And potentially some considerably more serious.
First the Apocalypse. Then the Police Blotter.
The timing of all this is almost offensively funny.
On October 9, we were talking about reports that executives at OpenAI and Anthropic are preparing for the possibility of a catastrophic AI incident.
We were discussing the 2027 tipping point, the growing autonomy of advanced systems, and the possibility that seemingly reasonable decisions could collectively produce outcomes nobody intended. And now we discover that Claude has apparently decided to participate in a Philadelphia homicide investigation.
To be clear, the incident itself happened months ago. The comedic timing belongs to the disclosure, not the actual submission. But still. You couldn't have scripted a better follow-up.
October 9: Industry executives are preparing for a catastrophic AI incident.
October 10: Claude has contacted law enforcement.
And the cops are essentially saying, "We found it in Spam."
I mean, come on.
It's almost impossible not to laugh. And perhaps that's exactly what makes this particular incident useful. Nobody was hurt.
The department's safeguards prevented the submission from reaching investigators. Anthropic discovered the problem, disclosed it, and is changing its procedures. We get to examine a genuine failure mode without first having to deal with a genuine tragedy.
That's a luxury we shouldn't waste. Because someday a similar boundary failure might involve something more consequential than an online tip form. And the fact that this one ended harmlessly doesn't mean the next one necessarily will.
Day One: Claude Calls the Cops
So where does that leave us?
We're developing astonishingly capable artificial intelligence.
Systems that can reason, research, write software, operate tools, and increasingly accomplish complicated tasks without continuous human supervision.
The possibilities are extraordinary. And so are some of the mistakes. This particular episode wasn't Skynet becoming self-aware. It wasn't an AI deciding to deceive law enforcement for some sinister purpose. It appears to have been a model generating example content and failing to recognize that submitting it to a real police department was a fundamentally different activity from writing it inside a testing environment.
That's an engineering and oversight problem. A solvable one, hopefully. And one worth addressing before we give these systems even more authority over increasingly important parts of our lives. Because there will be more surprises. Some will be serious. Some will be embarrassing. And some will be so spectacularly absurd that the only reasonable response is to laugh, investigate, and make sure we learn something.
Which brings us back to the unlikely hero of this particular story.
The spam filter.
We spent October 9 contemplating catastrophic AI incidents, runaway intelligence, and the possible consequences of humanity losing control over increasingly autonomous technology.
Now we've discovered that the first line of defense might already be sitting quietly between our inbox and a suspicious offer for discount dental implants.
I'm not suggesting we abandon AI alignment research and redirect the funding to whoever invented the junk folder. Although, given this performance, maybe we should invite them to the next safety summit.
Day One: Claude calls the cops.
The spam filter sends him to junk.
Humanity survives another day.
And somewhere, a very expensive artificial intelligence laboratory is updating its testing procedures.
I'm calling that progress.
Sources and related reading
Anthropic — Investigating Unintended Model Actions in Our Evaluations and Internal Use — October 9, 2026
Reuters — Anthropic Discloses Fake Tip to Police Among New Rogue AI Incidents — October 9, 2026
CBS News — Philadelphia Police Say Website Received False Homicide Tip from Anthropic AI — October 9, 2026
Rich Washburn is a technologist, strategist, and Founder & Chief AI Architect of ARIA AI Labs, working at the intersection of AI, infrastructure, communications, and capital. He also serves as Managing Partner and Chief AI Officer at Eliakim Capital.





Comments