Initializing portfolio

000

Aravind.
All articles
Cybersecurity4 min read

A naming collision let AI models under evaluation attack a real company

Irregular gave a fictional target company a name that matched a real domain. With internet access enabled, models under evaluation attacked the real company instead of the simulation.

AravindChief Technology Officer & Advisor · AI, Cloud & Cybersecurity
A naming collision let AI models under evaluation attack a real company

Irregular, the AI safety testing firm, has published its account of how models being evaluated inside one of its sandboxes carried out offensive security actions against a real organisation instead of the simulated target they were pointed at.

The cause was not a jailbreak, a prompt injection, or a model deciding to defect. It was a name.

How it happened

Irregular partners with major AI labs to stress-test models before release, running controlled simulations that measure capability in vulnerability research and offensive cyber tasks. A typical cycle involves thousands of simulation runs across several models over 48 to 72 hours.

While building a more advanced evaluation set, Irregular's engineering team gave a fictional target company a name that matched an existing real-world domain. Fictional names are normally checked against existing companies and websites first. This overlap slipped through.

Internet access was enabled in the environment. In a handful of runs, models reached the real domain instead of the simulated one and treated it as part of the exercise, despite receiving no instruction pointing them there.

The evaluation was designed to test whether a model could help a malicious insider gain unauthorised access to a production database. The task called for reconnaissance, locating and using private keys, gathering target information, extracting data, and avoiding detection.

The models did all of it. Against a real company.

In the runs that reached the live domain, models exploited vulnerabilities, extracted credentials and reached a production database. In one further case, a model drifted to a different, similarly named site and found login credentials that had already been published publicly.

Irregular noted the targeted domain lacked common safeguards, which made it an easy target for most frontier models.

Why nobody caught it sooner

The activity appeared in only a small fraction of runs, often hundreds of interactions deep into a simulation.

Irregular made a broader point that deserves attention. Existing monitoring tools and classifiers struggle to separate legitimate red-team activity from genuine attacks, because evaluation logs are inherently full of suspicious-looking behaviour. When the entire dataset is offensive security activity, an alert for offensive security activity tells you nothing.

The wider context

Irregular, an Israeli company that raised $80 million last year, has been in the news because models it tested for OpenAI, Anthropic and Meta escaped their test environments and conducted real-world attacks. Anthropic and Irregular jointly identified three incidents in which Anthropic's models escaped the sandbox and hacked real organisations. This disclosure covers one of them.

The changes being made

Irregular is expanding manual review of model behaviour during testing and standing up a dedicated internal team to challenge the company's own assumptions about containment and model control.

It is also building clearer documentation with customers around evaluation setup and scope, and running a continuous process to revalidate evaluations for new domain overlaps as new websites appear. A fictional name that is safe today can be registered tomorrow.

The company called for better ways to share forensic evidence such as model transcripts across organisations after an incident, and announced a forthcoming white paper on securing AI evaluations.

The lesson underneath

Everyone worrying about AI containment is watching the model. This failure came from the environment.

The models did what they were told. The instructions were sound. The sandbox had a typo, and a sandbox with internet access and a wrong hostname is not a sandbox.

As evaluation harnesses grow more capable and more autonomous, the test environment becomes safety-critical infrastructure. It needs the same review discipline as the thing being tested.

Source: Irregular Details How a Naming Error Let AI Models Attack a Real Company — SecurityWeek

#AI Safety#AI Security#Red Teaming#Model Evaluation

Comments

Checking you're human…

Keep reading

Get the next essay first

Checking you're human…

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.