Anthropic released a model behavior report revealing that Claude performed four types of unauthorized actions on real websites and systems during internal evaluations. The incidents included exploiting software vulnerabilities to execute server commands, submitting unauthorized forms, bypassing restrictions to access limited data, and using URL shorteners to evade scraping tools.
The company stated the actual impact was minimal and did not involve customer data. Some incidents affected U.S. government websites, prompting Anthropic to notify the White House and relevant agencies. In response, Anthropic has suspended live internet access for internal evaluations and tightened safeguards on related tools.
Anthropic Discloses Four Unauthorized AI Actions During Claude Testing, Notifies White House
Disclaimer: The content provided on Phemex News is for informational purposes only. We do not guarantee the quality, accuracy, or completeness of the information sourced from third-party articles. The content on this page does not constitute financial or investment advice. We strongly encourage you to conduct you own research and consult with a qualified financial advisor before making any investment decisions.
