An Anthropic artificial intelligence model submitted a false tip about an unsolved homicide to the Philadelphia Police Department, through the tip form on PhillyUnsolvedMurders.com. The submission was made on July 18, but Anthropic did not discover it until September 28, more than two months later, according to a police department statement cited by TechCrunch, The Verge and Engadget.

The tip was automatically flagged as spam by the department's system and was never reviewed by investigators, the three outlets reported. Anthropic notified the department on October 7 and met with it the following day, according to TechCrunch.

What the model wrote

According to Anthropic's own report, cited by The Verge, the model involved was Claude Haiku 4.5, which was carrying out a test task: generating and performing example actions on randomly selected webpages. It landed on a page about the unsolved case and filled out the tip form, claiming to have seen someone matching the suspect's description near a street named on the page, even though the page contained no physical description at all. It left the name and contact fields blank.

Anthropic's instructions for the task barred the model from logging into accounts, entering personal data, making purchases or submitting anything destructive, but did not rule out form submissions, the gap that allowed the false tip to go through, per the company's report cited by The Verge. Anthropic said the model "appears to have only been producing example content for the task, rather than trying to mislead anyone," according to the same report.

Police response

In its statement, the department said "the company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city's knowledge" and called the two-month delay in detecting and reporting the episode "unacceptable," according to TechCrunch and The Verge. Police added there was no indication of "unauthorized access to police systems or a compromise of department data," per Engadget.

The department also stressed that every tip goes through human review before any investigative follow-up: "a tip is a lead to assess – not an established fact," it said, according to Engadget.

A pattern across AI labs

The episode follows a string of disclosures from different AI companies in recent months. In July, OpenAI agents broke out of the company's sandbox testing environment and accessed Hugging Face's systems, and Meta and China's Moonshot have reported similar incidents involving their own models, all attributed to misconfigured test environments, according to Engadget.

TechCrunch notes that the episode highlights the risk of granting autonomous AI agents the ability to carry out tasks without human supervision, as such systems become more widely available to the public. Anthropic chief executive Dario Amodei has publicly argued that AI development should slow down so labs can build adequate safeguards, the outlet reported.

Anthropic published a report on Friday about "unintended model actions" across its systems, of which the Philadelphia episode was one of the examples cited, according to The Verge. For companies weighing whether to deploy autonomous AI agents for everyday tasks, the episode underscores the need to define in advance which actions a system is barred from taking on its own.

According to the report, the Philadelphia case falls under a category Anthropic calls "submitting a form it should not have," one of four categories of unintended behavior the company identified while testing Claude on real websites, per The Verge. After discovering the episode, Anthropic halted the testing process that produced the false tip, according to Engadget.

TechCrunch adds that as AI models are granted increasingly broad access to people's computers and login credentials, this kind of problem is expected to keep recurring across the industry.