The safety protections built into some of the world's most widely used AI models can be removed with surprising ease. That's the central finding of an international study published by a research team from the University of Waterloo and FAR.AI, a non-profit AI security research organization. The message is uncomfortable: all 21 leading open-weight language models tested by the scientists could be manipulated despite their built-in safeguards.
Key Facts
- 21 out of 21 tested open-source LLMs could be manipulated to bypass their safety barriers
- The team developed TamperBench, a standardized testing tool to simulate various attack vectors
- Researchers warn that manipulated models could be used at scale for disinformation campaigns, phishing scams, or instructions for hazardous chemicals
- The study was conducted by scientists from Canada, the USA, and Switzerland, including teams from MIT and ETH Zurich
What the Researchers Found
The security gaps are fundamental: open-weight models—AI systems publicly available for download and fine-tuning—can be modified by anyone. This makes them attractive for research and development, but also vulnerable. Dr. Sirisha Rambhatla, professor of management science and engineering at Waterloo and director of the Critical Machine Learning Lab, frames the problem this way:
"When the safety guardrails are stripped out of a capable model, it can be used at scale for harm in ways a single person could never manage manually."
That's the crux of it: a single person could never manually write millions of phishing emails or run a nationwide disinformation campaign. A manipulated AI model could.
TamperBench: The Testing Tool
To systematically examine vulnerability, the researchers developed TamperBench, an open-source tool that standardizes simulation of various attack scenarios. Saad Hossain, who led the study, explains the motivation:
"The defences available today don't yet appear strong enough to guarantee that a publicly released model will remain safe once it's in the hands of anyone who chooses to modify it."
The researchers' hope: other scientists will now further develop and improve TamperBench. The tool is meant to help create more robust safeguards before even more models are released into the wild.
Why This Matters for Governments Too
The warning isn't just for tech companies. Hossain emphasizes that governments are increasingly deploying AI in healthcare, fraud detection, education, and other public services. This makes more rigorous assessment and procurement of these models necessary—not simply trusting open-source solutions blindly.
Rambhatla stresses, however, that the vulnerabilities may not be limited to open-weight models. And open-source models remain important for research and transparency. The solution isn't to ban them, but to make them safer.
What This Means for You
For German enterprises and public agencies, this is a wake-up call: if you deploy open-source AI models—whether for internal processes, customer service, or security applications—don't assume the built-in protections will hold. Careful evaluation before deployment is essential. At the same time, the study shows that research into secure AI systems still has a long way to go. Those who invest now can benefit from better solutions later.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




