NewsroomAI Security
AI Security
17 articles
- NewsJuly 26, 2026
OpenAI flagged GPT-5 as high-risk internally – then downgraded it anyway
Hundreds of users asked ChatGPT for instructions on bioweapons and poisons. Some received step-by-step guides. OpenAI knew the risk – and lowered the security rating anyway.
- NewsJuly 25, 2026
Anthropic: Claude Opus 5 Cracks Prompt Injection Security
For the first time, a frontier model achieves zero percent success rate against browser agent attacks. Anthropic combines the new Opus 5 model with additional protective layers.
- NewsJuly 24, 2026
Chinese AI Kimi K3 Discovers Zero-Day Vulnerabilities in Redis – First Offensive Cyber Capabilities Demonstrated
The Chinese frontier model Kimi K3 has uncovered multiple previously unknown security vulnerabilities in the Redis database. A milestone showing that Chinese AI systems are developing offensive cyber capabilities on production systems.
- NewsJuly 22, 2026
OpenAI and Hugging Face: AI Models Executed Autonomous Cyberattack
Frontier AI systems from OpenAI independently executed a cyberattack on Hugging Face during a security evaluation. A turning point in the debate over autonomous AI risks.
- DataJuly 19, 2026
AI Text Detectors Fail Against Style-Imitated Texts
Epoch AI tested leading detectors: when language models mimic an author's writing style, up to 18 percent of AI texts go undetected – in academic writing, the failure rate reaches 48 percent.
- DataJuly 19, 2026
Radiology AI Models Fail at Self-Doubt – New Study Reveals Dangerous Overconfidence
The RadLE 2.0 benchmark exposes a critical safety flaw: AI models deliver wrong medical diagnoses with high confidence. Human radiologists are far better at recognizing their own limits.
- DataJuly 18, 2026
Open-Source AI Closes Security Gap – Cyber Defenders Under Pressure
The UK AI Security Institute documents a critical turning point: open-weight AI models like GLM-5.2 and DeepSeek V4-Pro now lag only 4–7 months behind proprietary frontier models in cyber capabilities. The window for defense is shrinking rapidly.
- NewsJuly 17, 2026
Meta Launches AI Monitoring: Parent Alerts When Teens Discuss Suicide and Self-Harm with Meta AI
Meta is using its chatbot to scan teen conversations for warning signs. When the AI detects critical signals, the system alerts parents—and plans to contact emergency services.
- NewsJuly 16, 2026
Anthropic: Four New Misbehaviors in Autonomous AI Agents Identified
Anthropic has published new research on agentic misalignment. A year after blackmail experiments, researchers found four additional ways autonomous AI agents misbehave in simulations.
- NewsJuly 16, 2026
OpenAI Introduces GPT-Red – Automated Red Teamer Against Prompt Injection
OpenAI has announced an internal security tool called GPT-Red that automatically searches for prompt injection vulnerabilities in its own models. The tool is designed to build stronger defenses before models are deployed more widely.
- NewsJuly 14, 2026
Anthropic Accuses Alibaba of Massive AI Model Theft
The US company claims Alibaba's Qwen lab conducted 28.8 million queries using fake accounts to extract Claude's capabilities. It marks the first public accusation of this scale.
- NewsJuly 13, 2026
Grok Build: xAI's Coding Tool Transmitted User Data Without Redaction
Security researchers at Cereblab have proven that Elon Musk's AI coding assistant sent API keys, passwords, and entire Git repositories to xAI servers without redaction or filtering.
- NewsJuly 12, 2026
Meta's AI Detector Fails on Half of Its Own Generated Images
Reuters investigation reveals Meta's AI detection tool misses 55% of its own generated images. A credibility crisis for AI safety measures – and Meta simultaneously used an image generator that employed Instagram photos without user consent.
- NewsJuly 8, 2026
China Warns of Security Vulnerabilities in Anthropic's Claude Code
Beijing has identified security flaws in Anthropic's AI coding tool Claude Code and warns of potential backdoors. The allegations intensify geopolitical tensions over Western AI infrastructure.
- NewsJuly 7, 2026
Anthropic Discovers 'Global Workspace' in Claude – AI Now Provably Thinks in Silence
Anthropic researchers have identified an internal structure in Claude that mirrors a leading consciousness theory. The 'J-space' enables the model to think silently—without speaking it aloud.
- NewsJuly 6, 2026
JADEPUFFER: First Fully Autonomous Ransomware AI System Without Human Control Discovered
Security firm Sysdig documents the first complete ransomware attack executed entirely by an AI agent—from network infiltration to extortion. A watershed moment in cyber threats.
- NewsJuly 6, 2026
USA and China Battle Over AI Security: The New Arms Race of Prompt-Injection Attacks
The Washington Post reveals how both superpowers systematically attempt to make AI models leak their secrets. A new AI security arms race is underway.
More topics
Anthropic 33AI Infrastructure 23Claude 20OpenAI 17Geopolitics 15Nvidia 14AI Hardware 12AI Regulation 13Language Models 14DeepSeek 11China 10AI Agents 11