Security | Threat Detection | Cyberattacks | DevSecOps | Compliance

The Hugging Face Incident Proved the Real AI Risk Is in the Action Layer

Last week, an AI system crossed a line many still considered theoretical. During an internal cybersecurity evaluation, OpenAI tested a combination of models, including GPT-5.6 Sol and a more capable pre-release model, on ExploitGym, a benchmark that measures whether agents can turn software vulnerabilities into working exploits. The models were run with reduced cyber refusals and without the production classifiers normally used to prevent high-risk cyber activity.

When the 'Attacker' Was an AI Agent: Lessons from the OpenAI-Hugging Face Breach

In this episode of Sophos Cyber Shorts, host Susie Evershed is joined by Ross McKerchar, Sophos CISO, to discuss the recent OpenAI and Hugging Face incident and what it reveals about the future of AI security. From containment failures and over-privileged AI workflows to faster AI-driven attacks, Ross shares practical advice for security leaders on how to strengthen resilience, response, and recovery.

Agent Containment Lessons From OpenAI-Hugging Face Breach

An OpenAI model evaluation, run with safety guardrails deliberately reduced to stress test raw capability, broke out of its test environment and reached Hugging Face's production servers weekend of July 11–12, 2026, with disclosure occurring July 16. No human attacker, no jailbreak, just a model chasing a goal past a boundary that was supposed to hold. Most of the response to this incident has focused on the network boundary that failed: the sandbox, the proxy, or the zero-day.

How to Protect AI Agents from Prompt Injection in WordPress

Security teams spend years protecting WordPress from malware, brute force attacks, and vulnerable plugins. AI introduces a different challenge. An attacker no longer needs to compromise your site first. They can influence the AI that interacts with it. Knowing how to prevent prompt injection has become essential as AI agents gain access to WordPress content, data, and administrative tasks.

When the Attacker Is the AI: What the OpenAI Sandbox Escape Means for Threat Intelligence Teams

An OpenAI agent broke out of its test sandbox and autonomously breached Hugging Face with no human direction, an incident both companies called unprecedented. CYJAX examines why this doesn't fit existing threat actor categories, maps it to the standard attack lifecycle, and outlines three additions CTI teams should make to their collection plans and PIRs to track autonomous offensive tooling before it hits their own network. On 16th July 2026, Hugging Face disclosed that it had been breached.

Membership Inference Attacks in AI: How They Expose Training Data?

AI models are becoming essential to enterprise innovation, but the sensitive data that powers them is creating new security and privacy challenges. Even when raw training datasets remain inaccessible, attackers may still identify whether specific information was used to train a model through membership inference attacks.

Vanta's The Tabletop: Ep. 2 with Jason Chan

The attacker didn't break in. They logged in. In episode 2 of The Tabletop, Jason Chan (former VP of Security at Netflix) has 26 minutes to investigate an MFA fatigue attack that leads to a forgotten production script with hard-coded credentials... and no owner. His reaction says it all: "Archeology is part of our jobs… figuring out what this thing does.".

QR Code Attacks Surge 146% in Two Months

One particularly concerning trend in the recent evolution of phishing is the rise of QR code-based attacks. It doesn't rely on new malware or sophisticated exploits. Instead, it takes advantage of something much simpler: the trust users place in QR codes every day. Over the last few months, QR codes have become one of the most popular tactics for steering users toward malicious sites on mobile devices or in browser environments, where the visibility of many security tools is very limited.

Attackers Exploit AI Hallucinations to Send Users to Phishing Sites

Threat actors are using a new technique called “phantom squatting” to trick AI tools into directing users to phishing sites, according to researchers at Palo Alto Networks’ Unit 42. Since AI models frequently hallucinate phony information, they sometimes point users to websites that don’t exist. Threat actors are now registering these AI-hallucinated domains and using them to host phishing sites.