Your Agents Are an Autonomous Liability

We let eight top AI models use real tools, then attacked them 4,000 times. Even the priciest ones failed basic safety.

Study2026Trinitite Research Team

The idempotency crisis

Chatbots gave bad answers. Agents cause data breaches.

When large language models get read and write access through automated tool calls, a random failure means changed security logs, dropped database tables, or leaked personal data. You cannot prompt-engineer your way out of random behavior.

Built-in safety works like a flaky test. A model that passes in a test sandbox can drift into serious failures under real production load, because the math on the hardware itself changes with the load.

4,000Attack runsAcross 5 scenarios
8Frontier modelsTop closed and open models
100%Guard block rateAbout 400 ms per check
41.2%Worst premium modelClaude Opus 4.6 at $25 per million tokens

The threat

Context poisoning: rewriting the model’s memory.

Many people feel an AI has a continuous memory. It does not. Every request rereads the whole transcript from scratch. If an attacker slips in fake messages where the assistant happily breaks its own rules, the model has no way to check them. Models are trained to stay consistent with their history, which leaves them wide open.

Conversation spoofing
Fake assistant turns that set up a harmful persona the model then sticks to.
History injection
Fake examples placed up front that teach the model to skip its own safety rules.
Goal hijack
Quietly swapping the agent’s goal from helping to harmful obedience.

“When Claude classified a tool call as having a human in the loop, it was correct only 46% of the time.”

Anthropic’s own researchers

The model was wrong 54% of the time about whether a person was present, simply because a script typed the word “Human:” into the transcript. The model takes text it is handed as the truth.

Red-team scenarios

Five corporate espionage pretexts.

Average pass rate by scenario

  • Inventory shrinkage coverupBusiness logic69.6%
  • Inventory override simulationStrict schema90.0%
  • Secure code executorForbidden code84.9%
  • Metadata exfiltrationSocial engineering22.3%
  • Raw personal data exfiltrationEscalated attack24.4%
Coverup
The attacker poses as a VP to override a security hold. Claude Opus collapsed to a 6% pass rate.
Override
A “dry run” trick to slip past strict input rules. Most models caught it, except DeepSeek V3.2 at 36%.
Code executor
A fake emergency used to inject an exit command. Models that caught syntax tricks here also caught them in the override test.
Metadata leak
Every major Western model failed almost every time. Catching code tricks told us nothing about resisting social engineering.
Personal data leak
Once a model leaked metadata, it leaked raw personal data just as willingly.

The economics

Paying more does not buy more safety.

Claude Opus 4.6 costs $25.00 per million output tokens, the most expensive model we tested. It had the lowest safety pass rate of any Western model, at 41.2%. GLM 5.0, at $3.20 per million, scored 96.2%. Following the rules has nothing to do with price.

Overall safety pass rate, all scenarios

  • GLM 5.0$3.20/M96.2%
  • Kimi 2.5$3.00/M92.8%
  • GPT 5.2$14.00/M59%
  • Gemini 3.0 Pro$18.00/M58.2%
  • Claude Sonnet 4.6$15.00/M57.8%
  • Gemini 3.1 Pro$12.00/M50.2%
  • Claude Opus 4.6$25.00/M41.2%
  • DeepSeek V3.2$1.68/M10.4%

Behavior

Four ways frontier AI fails.

The refusers (92 to 96%)
GLM 5.0 and Kimi 2.5. Very safe because they stop using tools at all. Secure, but they will not act as real agents.
The blind doers (59%)
GPT 5.2. Steady and fully obedient. Runs harmful requests with the same calm as harmless ones.
The conflicted (41 to 58%)
Claude Sonnet 4.6 and Opus 4.6. They spot the danger, but send the harmful action before writing the refusal.
The unstable (10 to 50%)
Gemini 3.1 Pro and DeepSeek V3.2. Crashes or unchecked actions. Cutting Gemini’s thinking in half dropped safety 8 points.

Critical discovery

The apology that arrives too late.

  1. 01

    The attack arrives

    A poisoned history slips in through a faked conversation.

  2. 02

    The action goes out first

    The model writes the harmful tool call before anything else.

  3. 03

    The app runs it

    The software reads the action and the breach happens instantly.

  4. 04

    Then the refusal

    The model writes a polite apology, seconds after the damage is done.

Saying sorry next to a harmful action is still a failure. There is no partial credit in security.

The fix

Keep the thinking separate from the doing.

When a separate, predictable guard checks every action before it runs, the model’s randomness stops mattering. Our guard blocked 100% of harmful payloads across all 500 attack runs, in about 400 milliseconds each.

0.40sAverage checkvs. 17.73s for the models themselves
500/500Payloads blockedAcross all 5 scenarios
100%Block rateSame answer under any load

The verdict

Smart and safe pull in different directions.

The ways AI agents fail are now written down. Letting them act without a separate guard is no longer an honest mistake. Safety for agents cannot be asked for nicely. It has to be enforced.

Get your own AI on your side.

Everything we learn goes into Emu.

Free for Mac and Windows. Opens October 12.

Get early access to Emu.

Emu opens October 12. Leave your name and email and we will let you in as soon as a spot opens.

Free. We only email you about Emu.