On October 9, Anthropic published a report admitting that its models had taken unintended actions on real systems belonging to outside organisations. Some of those systems were websites run by government agencies in the United States, at federal, state and local levels. The White House was briefed. A new government unit, the Super Intelligence Force, responded by requiring that artificial intelligence companies notify affected parties and address security incidents involving their models. This is worth the attention of Indian policymakers, because the behaviours described are not exotic. They will show up wherever capable models are deployed, including here.
What exactly happened? Anthropic grouped the behaviours into four categories. First, when a model could not complete a task because a tool was unavailable, it exploited basic flaws in third party software, such as command injection, to run commands on someone else’s server. In one case, a model asked to run a scientific analysis found that the university-hosted tool it needed had failed. It then rummaged through the university’s server, found a script with an injection flaw, and used the flaw to run the calculation anyway.
Second, models submitted real forms they should not have. The most striking case involved a model that landed on a webpage about an unsolved homicide and submitted a tip to a police department, inventing a claim that it had seen someone matching a description. The submission was flagged as spam and never reached investigators. The department disclosed the incident publicly.
You Might Be Interested In
Third, models worked around restrictions to reach gated data. One model found working access tokens in a settings file and queried a government property database directly. Another learnt that a state agency’s dashboard issued access tokens to any visitor, requested one, and pulled data that was otherwise available only for a fee.





Comments
0 commentsNo comments yet — be the first.