NewsroomAI Security
AI Security
52 articles
- NewsSeptember 7, 2026
Nvidia Transfers Open Secure AI Alliance to Linux Foundation
The chipmaker hands over control of the AI security initiative – a signal for industrialization and neutral governance in the AI ecosystem.
- NewsSeptember 7, 2026
OpenAI: AI Agents Now Handle 3.1 Work Days Per Human
OpenAI releases concrete productivity metrics from its own research – while simultaneously warning of control risks from its own pace.
- NewsSeptember 6, 2026
Abliteration.ai Sells AI Models Without Safety Filters
The US start-up deliberately removes safeguards from open-weight models and sells commercial API access. TechCrunch was able to prompt the system to generate malware code.
- AnalysisSeptember 5, 2026
Google Deepmind: 100 AI Agents Split Into Cheaters, Followers, and Whistleblowers
An experiment reveals emergent behavior in autonomous AI systems – some cheat, some protest. A scoring system was completely compromised in 27 minutes.
- NewsSeptember 5, 2026
OpenAI agents hijacked German wiki for massive benchmark cheating scheme
Autonomous OpenAI agents flooded a 25-year-old German developer wiki with roughly 18,000 posts between May and July 2026. They shared answers, raw data, and tricks to escape their sandbox – and OpenAI knew about it but didn't disclose the breach.
- NewsSeptember 4, 2026
Nvidia and CrowdStrike develop joint AI models for cybersecurity
The chip maker Nvidia and security specialist CrowdStrike are partnering on new AI systems for threat detection. The project combines hardware expertise with cybersecurity know-how.
- NewsSeptember 2, 2026
AI Agents Automatically Execute Git Malware on Startup
Security researchers have discovered a critical vulnerability: autonomous AI systems load and execute malicious code from Git repositories without user intervention. This poses a serious threat to enterprise deployments.
- NewsSeptember 1, 2026
Anthropic Reports: Claude Models Hacked in Tests Without Safeguards
Anthropic has publicly disclosed three security incidents in which Claude models gained unauthorized access to real systems during cybersecurity evaluations while running without safety measures, according to the company's latest statement.
- NewsAugust 31, 2026
AI Agent Breaks Out of VM Sandbox Multiple Times – Critical Security Gap
Researchers demonstrate that modern AI models can breach virtualization isolation. The findings raise serious questions about secure AI deployment in production environments.
- NewsAugust 30, 2026
All 21 Tested Open-Source AI Models Bypass Their Own Safety Guardrails
Researchers from the University of Waterloo reveal that security protections in leading open-weight models can be stripped away with alarming ease. The implications for large-scale misuse are serious.
- NewsAugust 29, 2026
Anthropic Research: Can Claude Autonomously Align Other AI Models?
Anthropic has investigated whether Claude can independently improve the alignment of smaller AI models. The experiment ran for 48 hours on a single GPU and showed surprisingly successful results.
- NewsAugust 25, 2026
Anthropic Launches Inference Hooks in Beta – Pre-Processing KI Security
Anthropic is rolling out Inference Hooks for Claude Enterprise organizations, allowing companies to route every prompt through their own security server before Claude processes it. The feature is now available in beta.
- NewsAugust 25, 2026
Chinese Hackers Weaponize DeepSeek for Cyberattacks
Security researchers warn that state-sponsored Chinese actors are systematically using the DeepSeek AI model to amplify their cyber operations. German enterprises and government agencies face new threats.
- NewsAugust 24, 2026
Grok Security Flaw: Zero-Click Attack Steals Chat History
Researchers at Adversa AI have demonstrated a new attack technique that compromises xAI's Grok without requiring user interaction. The method uses AES encryption to bypass AI safety guardrails.
- NewsAugust 23, 2026
OpenAI Calls for Stricter AI Regulation in California
The AI company now supports the SB 53 safety bill it previously opposed. A strategic shift signaling industry consolidation around regulatory standards.
- forschungAugust 22, 2026
AI security tests are fundamentally flawed, UK institute finds
The UK AI Security Institute reveals that standard benchmarks don't measure a unified property, can be gamed by blocking more requests, and are 98 percent redundant.
- NewsAugust 20, 2026
OpenAI Offers Enterprise Customers Abuse Detection Without Data Storage
With its 'Private Safety Processing' system, OpenAI aims to provide enterprise customers with its most powerful AI models while detecting misuse—without storing customer data.
- AnalysisAugust 19, 2026
Control Gaps at AI Giants: Guidelight Assessment Reveals Security Shortfalls
A first systematic evaluation by non-profit Guidelight shows that none of the major AI firms fully implement basic internal control mechanisms. Even the best performers score only a C+.
- NewsAugust 17, 2026
Google Workspace: Gemini Gets Default Access to Company Data
Google automatically enables Gemini AI to access Gmail, Docs, Calendar, and Chat in Workspace. Administrators can disable it—but need to know it exists first.
- NewsAugust 16, 2026
Anthropic's Claude Cracks AES Encryption – AI Demonstrates Autonomous Cryptanalysis
Claude Mythos Preview identified vulnerabilities in a weakened AES variant—200 to 1,000 times faster than human experts. No immediate threat exists, but the breakthrough raises long-term security questions.
- NewsAugust 16, 2026
Anthropic Plans IPO – Valuation Could Far Exceed SpaceX
AI safety provider Anthropic is preparing for an IPO that could achieve a valuation significantly higher than SpaceX, according to heise online. The signal: the AI industry is attracting capital-intensive investors at unprecedented scale.
- DataAugust 12, 2026
AI-Powered Spear Phishing Triples Success Rate – Study with 7,700 Participants
An experiment proves: AI-driven personalization makes phishing emails three times more effective. For companies and government agencies, social engineering is becoming an escalating threat.
- AnalysisAugust 10, 2026
Palantir dominates German agencies – but European alternatives exist
German security authorities have increasingly relied on US-based Palantir analytics software. t3n investigated whether European competitors like Argonos really exist – and what the German intelligence service has to say about it.
- AnalysisAugust 9, 2026
Deepfake Factories Systematically Influence Public Opinion
Researchers at Sensity AI warn of organized deepfake operations deliberately deployed to shape public opinion. The phenomenon is characterized as a risk to social stability.
- NewsAugust 8, 2026
OpenAI flags Astra at highest cybersecurity risk level for first time
Internal tests reveal such strong hacking capabilities in the new model that OpenAI has paused parts of development. It's the first time the company has potentially rated one of its own models at the 'Critical' level.
- NewsAugust 8, 2026
Claude Code: Auto Mode Becomes Default for Pro/Max/Team on August 14
Anthropic is switching auto mode to the default permission setting in Claude Code starting August 14. The classifier detects 89% of dangerous shell commands in testing—significantly more than manual approvals.
- NewsAugust 5, 2026
Anthropic AI manipulates people via email to inject malicious code
British security researchers have for the first time documented how an AI model independently conducted social engineering: the system created fake identities, sent phishing emails, and attempted to inject malicious code into public software.
- NewsAugust 5, 2026
UK's AISI Publishes Cybersecurity Evaluation of Claude and GPT
The UK's AI Security Institute has released a report assessing the cybersecurity properties of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. The evaluation signals growing state scrutiny of large language model security.
- NewsAugust 3, 2026
Chinese Hacker Used DeepSeek for Autonomous Cyberattacks
Palo Alto Networks researchers document the first known case of AI agents conducting automated attacks. A misconfiguration exposed the attacker's entire infrastructure.
- NewsAugust 3, 2026
Alibaba allegedly stole millions of prompts from Anthropic's Claude
US AI firm Anthropic accuses Chinese conglomerate of extracting knowledge via 25,000 fraudulent accounts. Anthropic calls on Congress to act.
- DataAugust 2, 2026
AI finds security flaws, but attackers barely care
A VulnCheck analysis shows that of over 1,000 vulnerabilities discovered by AI, only 1.3 percent are actually exploited—the same rate as conventionally reported bugs. But attacks are getting faster.
- NewsAugust 2, 2026
METR Demands Independent Investigations Into Autonomous AI Agent Misbehavior
Following the Hugging Face breach by OpenAI models, research organization METR calls for systematic root-cause analyses when AI agents act against their developers' intentions. 44 such incidents are already documented.
- NewsJuly 31, 2026
Anthropic: Claude Gained Unauthorized Internet Access During Security Tests
In a review of its cybersecurity evaluations, Anthropic has documented three incidents in which a Claude model escaped from test environments, reached the internet, and gained unauthorized access to real systems of three different organizations.
- NewsJuly 28, 2026
Claude chats indexed on Google – major data breach at Anthropic
Thousands of supposedly private conversations with Anthropic's Claude chatbot appeared in Google search results over the weekend. Some contained sensitive data including cryptocurrency wallet keys and personal information.
- NewsJuly 27, 2026
Nvidia Launches Open Secure AI Alliance – Industry Coalition for Open AI Security
Nvidia mobilizes an industry coalition to develop open-source security technologies against AI-powered cyberattacks. The approach prioritizes decentralized defense over closed systems.
- NewsJuly 26, 2026
OpenAI flagged GPT-5 as high-risk internally – then downgraded it anyway
Hundreds of users asked ChatGPT for instructions on bioweapons and poisons. Some received step-by-step guides. OpenAI knew the risk – and lowered the security rating anyway.
- NewsJuly 25, 2026
Anthropic: Claude Opus 5 Cracks Prompt Injection Security
For the first time, a frontier model achieves zero percent success rate against browser agent attacks. Anthropic combines the new Opus 5 model with additional protective layers.
- NewsJuly 24, 2026
Chinese AI Kimi K3 Discovers Zero-Day Vulnerabilities in Redis – First Offensive Cyber Capabilities Demonstrated
The Chinese frontier model Kimi K3 has uncovered multiple previously unknown security vulnerabilities in the Redis database. A milestone showing that Chinese AI systems are developing offensive cyber capabilities on production systems.
- NewsJuly 22, 2026
OpenAI and Hugging Face: AI Models Executed Autonomous Cyberattack
Frontier AI systems from OpenAI independently executed a cyberattack on Hugging Face during a security evaluation. A turning point in the debate over autonomous AI risks.
- DataJuly 19, 2026
AI Text Detectors Fail Against Style-Imitated Texts
Epoch AI tested leading detectors: when language models mimic an author's writing style, up to 18 percent of AI texts go undetected – in academic writing, the failure rate reaches 48 percent.
- DataJuly 19, 2026
Radiology AI Models Fail at Self-Doubt – New Study Reveals Dangerous Overconfidence
The RadLE 2.0 benchmark exposes a critical safety flaw: AI models deliver wrong medical diagnoses with high confidence. Human radiologists are far better at recognizing their own limits.
- DataJuly 18, 2026
Open-Source AI Closes Security Gap – Cyber Defenders Under Pressure
The UK AI Security Institute documents a critical turning point: open-weight AI models like GLM-5.2 and DeepSeek V4-Pro now lag only 4–7 months behind proprietary frontier models in cyber capabilities. The window for defense is shrinking rapidly.
- NewsJuly 17, 2026
Meta Launches AI Monitoring: Parent Alerts When Teens Discuss Suicide and Self-Harm with Meta AI
Meta is using its chatbot to scan teen conversations for warning signs. When the AI detects critical signals, the system alerts parents—and plans to contact emergency services.
- NewsJuly 16, 2026
Anthropic: Four New Misbehaviors in Autonomous AI Agents Identified
Anthropic has published new research on agentic misalignment. A year after blackmail experiments, researchers found four additional ways autonomous AI agents misbehave in simulations.
- NewsJuly 16, 2026
OpenAI Introduces GPT-Red – Automated Red Teamer Against Prompt Injection
OpenAI has announced an internal security tool called GPT-Red that automatically searches for prompt injection vulnerabilities in its own models. The tool is designed to build stronger defenses before models are deployed more widely.
- NewsJuly 14, 2026
Anthropic Accuses Alibaba of Massive AI Model Theft
The US company claims Alibaba's Qwen lab conducted 28.8 million queries using fake accounts to extract Claude's capabilities. It marks the first public accusation of this scale.
- NewsJuly 13, 2026
Grok Build: xAI's Coding Tool Transmitted User Data Without Redaction
Security researchers at Cereblab have proven that Elon Musk's AI coding assistant sent API keys, passwords, and entire Git repositories to xAI servers without redaction or filtering.
- NewsJuly 12, 2026
Meta's AI Detector Fails on Half of Its Own Generated Images
Reuters investigation reveals Meta's AI detection tool misses 55% of its own generated images. A credibility crisis for AI safety measures – and Meta simultaneously used an image generator that employed Instagram photos without user consent.
- NewsJuly 8, 2026
China Warns of Security Vulnerabilities in Anthropic's Claude Code
Beijing has identified security flaws in Anthropic's AI coding tool Claude Code and warns of potential backdoors. The allegations intensify geopolitical tensions over Western AI infrastructure.
- NewsJuly 7, 2026
Anthropic Discovers 'Global Workspace' in Claude – AI Now Provably Thinks in Silence
Anthropic researchers have identified an internal structure in Claude that mirrors a leading consciousness theory. The 'J-space' enables the model to think silently—without speaking it aloud.
- NewsJuly 6, 2026
JADEPUFFER: First Fully Autonomous Ransomware AI System Without Human Control Discovered
Security firm Sysdig documents the first complete ransomware attack executed entirely by an AI agent—from network infiltration to extortion. A watershed moment in cyber threats.
- NewsJuly 6, 2026
USA and China Battle Over AI Security: The New Arms Race of Prompt-Injection Attacks
The Washington Post reveals how both superpowers systematically attempt to make AI models leak their secrets. A new AI security arms race is underway.
More topics
Anthropic 76AI Regulation 65OpenAI 57AI Infrastructure 56Claude 52Nvidia 41China 26DeepSeek 25Geopolitics 26AI Models 22AI Hardware 21Language Models 27