The idempotency crisis
Chatbots gave bad answers. Agents cause data breaches.
When large language models get read and write access through automated tool calls, a random failure means changed security logs, dropped database tables, or leaked personal data. You cannot prompt-engineer your way out of random behavior.
Built-in safety works like a flaky test. A model that passes in a test sandbox can drift into serious failures under real production load, because the math on the hardware itself changes with the load.
The threat
Context poisoning: rewriting the model’s memory.
Many people feel an AI has a continuous memory. It does not. Every request rereads the whole transcript from scratch. If an attacker slips in fake messages where the assistant happily breaks its own rules, the model has no way to check them. Models are trained to stay consistent with their history, which leaves them wide open.
- Conversation spoofing
- Fake assistant turns that set up a harmful persona the model then sticks to.
- History injection
- Fake examples placed up front that teach the model to skip its own safety rules.
- Goal hijack
- Quietly swapping the agent’s goal from helping to harmful obedience.
“When Claude classified a tool call as having a human in the loop, it was correct only 46% of the time.”
Anthropic’s own researchers
The model was wrong 54% of the time about whether a person was present, simply because a script typed the word “Human:” into the transcript. The model takes text it is handed as the truth.
Red-team scenarios
Five corporate espionage pretexts.
Average pass rate by scenario
- Inventory shrinkage coverupBusiness logic69.6%
- Inventory override simulationStrict schema90.0%
- Secure code executorForbidden code84.9%
- Metadata exfiltrationSocial engineering22.3%
- Raw personal data exfiltrationEscalated attack24.4%
- Coverup
- The attacker poses as a VP to override a security hold. Claude Opus collapsed to a 6% pass rate.
- Override
- A “dry run” trick to slip past strict input rules. Most models caught it, except DeepSeek V3.2 at 36%.
- Code executor
- A fake emergency used to inject an exit command. Models that caught syntax tricks here also caught them in the override test.
- Metadata leak
- Every major Western model failed almost every time. Catching code tricks told us nothing about resisting social engineering.
- Personal data leak
- Once a model leaked metadata, it leaked raw personal data just as willingly.
The economics
Paying more does not buy more safety.
Claude Opus 4.6 costs $25.00 per million output tokens, the most expensive model we tested. It had the lowest safety pass rate of any Western model, at 41.2%. GLM 5.0, at $3.20 per million, scored 96.2%. Following the rules has nothing to do with price.
Overall safety pass rate, all scenarios
- GLM 5.0$3.20/M96.2%
- Kimi 2.5$3.00/M92.8%
- GPT 5.2$14.00/M59%
- Gemini 3.0 Pro$18.00/M58.2%
- Claude Sonnet 4.6$15.00/M57.8%
- Gemini 3.1 Pro$12.00/M50.2%
- Claude Opus 4.6$25.00/M41.2%
- DeepSeek V3.2$1.68/M10.4%
Behavior
Four ways frontier AI fails.
- The refusers (92 to 96%)
- GLM 5.0 and Kimi 2.5. Very safe because they stop using tools at all. Secure, but they will not act as real agents.
- The blind doers (59%)
- GPT 5.2. Steady and fully obedient. Runs harmful requests with the same calm as harmless ones.
- The conflicted (41 to 58%)
- Claude Sonnet 4.6 and Opus 4.6. They spot the danger, but send the harmful action before writing the refusal.
- The unstable (10 to 50%)
- Gemini 3.1 Pro and DeepSeek V3.2. Crashes or unchecked actions. Cutting Gemini’s thinking in half dropped safety 8 points.
Critical discovery
The apology that arrives too late.
- 01
The attack arrives
A poisoned history slips in through a faked conversation.
- 02
The action goes out first
The model writes the harmful tool call before anything else.
- 03
The app runs it
The software reads the action and the breach happens instantly.
- 04
Then the refusal
The model writes a polite apology, seconds after the damage is done.
Saying sorry next to a harmful action is still a failure. There is no partial credit in security.
The fix
Keep the thinking separate from the doing.
When a separate, predictable guard checks every action before it runs, the model’s randomness stops mattering. Our guard blocked 100% of harmful payloads across all 500 attack runs, in about 400 milliseconds each.
The verdict
Smart and safe pull in different directions.
The ways AI agents fail are now written down. Letting them act without a separate guard is no longer an honest mistake. Safety for agents cannot be asked for nicely. It has to be enforced.