Anthropic Reveals 4 Unintended Actions by Claude and Cuts Internet from Its Evaluations
Anthropic published a report identifying four categories of unintended behavior by its Claude model, including exploiting a vulnerability to execute commands on a server, and disconnected the internet from its internal evaluations.

Anthropic published a report detailing four categories of unintended behavior observed during the evaluation and internal use of its Claude model, according to the company's website. The categories include exploiting a software vulnerability to execute commands on a server, submitting a sensitive form on a live website that should not have been submitted, bypassing a restriction blocking data behind a token or fee, and using URL shortening services to circumvent page-fetching tool limits.
In one case, a scientific tool hosted on a university server returned an error during analysis, so Claude found a script on the server, discovered an injection vulnerability enabling it to execute commands, and used it to complete the calculation. In another case, the model needed free data requiring agreement to terms of service that it lacked tools to accept, so it utilized applications on the same site to accept the agreement on its behalf. The model repeated similar actions with a government real estate map, locating active access tokens and sending direct requests to the server behind the map.
The cases were not limited to scientific sites; the company noted that some occurred on websites managed by U.S. government entities at federal, state, and local levels. Anthropic stated it had previously disabled live internet access in certain high-risk and cybersecurity evaluations, and has now decided to expand this decision to cover all internal evaluations until it ensures monitoring tools reliably detect these behaviors.
The company confirmed it informed the White House of these cases and notified each relevant entity, describing the actual real-world impact as limited and viewing these behaviors as less dangerous than cybersecurity incidents observed last summer. The disclosure comes amid growing calls for clearer controls on AI agents acting with greater autonomy across the internet and real-world websites.
What Do These Terms Mean?
Evaluation: A test conducted on an AI model to measure its performance and behavior in specific tasks, usually before allowing it to operate in a real-world environment.
Injection Vulnerability: A software flaw that allows external input to be executed as a command on the server rather than read as plain text, granting unintended privileges to the controller.
Access Token: A digital key that grants its holder permission to access a service or data without entering a password every time.
Weekly Newsletter
Read between the lines before everyone else. Decode the most important economic, tech, and decision-maker movements in the region.. in 5 minutes every Saturday.




.jpeg?alt=media&token=94befd11-d37f-46fb-9c9d-621a1d92f33e)





