Topic

Security — AI news & analysis

AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›

Every Pick Right story tagged security — 12 articles, newest first. All news →

anthropic security

Anthropic just shipped its most restricted model to more customers — by taking away the prompt box

On 21 August 2026 Anthropic put Claude Mythos 5 behind Claude Security, its codebase vulnerability scanner, in public beta for Claude Enterprise. You cannot prompt the model. You get patches, alerts and CWE-tagged findings. That constraint is not a limitation Anthropic apologises for — it is the mechanism that made the release possible, and it is the clearest signal yet of how frontier capability will actually reach buyers. One question the announcement does not answer: what happens to your source code under the 30-day covered-model retention rule.

Read story →
microsoft copilot

Microsoft patched CoSnitch on Monday — but a patch can't un-poison an AI assistant's memory, and that's the part buyers keep missing

On 18 August 2026 Microsoft fixed CVE-2026-24301, a critical one-click flaw in Copilot Personal that Varonis reported in December 2025. The exploit chained auto-executing prompts, OAuth-connected Gmail and Drive access, and memory poisoning that survives password changes, session revocation and device re-enrollment. Microsoft says customers 'do not need to take any action.' For anyone whose memory was already poisoned, that is not true. Here's what actually happened and what to check.

Read story →
openai security

OpenAI stopped its largest training run — and put a number on what safety costs: about 20% of inference compute

On 18 August 2026 OpenAI disclosed a two-week pause on reinforcement-learning training for deployment-bound models, and said its largest planned frontier RL run remains on hold. The headline is the pause. The number buyers should write down is the monitoring overhead: roughly 20% of the inference compute being watched. Frontier capability is now gated by security engineering, not by compute — here's what that changes for roadmaps and budgets.

Read story →
openai anthropic

Both labs now agree single-request safety checks aren't enough for frontier models — only one is making you give up zero data retention for it

On 19 August 2026 OpenAI previewed Private Safety Processing: cross-session misuse detection that it says stays compatible with Zero Data Retention. It lands directly against Anthropic's covered-models policy, which requires 30-day retention for Mythos-class models and, per Anthropic's own documentation, makes them unavailable with ZDR and unusable under a HIPAA BAA. The two labs reach identical conclusions about the threat and opposite conclusions about the price. Here's what's actually shipping, what's still a preview, and how to choose.

Read story →
open-weights coding

Z.ai's GLM-5.3 pushes open-weights coding to the frontier's doorstep — and the model's cyber skills 'outgrew its training,' which is why you can't download it yet

Released 14 August 2026, GLM-5.3 is a post-training-only upgrade on the same 743B base as GLM-5.2 — and Z.ai says it is the strongest open-weights coding model it has measured, with Terminal-Bench 3.0 leaping from 4.6 to 28.3. But the headline is a vulnerability-discovery capability that scaled faster than the company expected, holding back the open weights for two weeks of safety hardening. Here is the grounded read for anyone choosing a coding model — and what the delay tells you about open-weights AI in 2026.

Read story →
anthropic claude-code

Claude Code just made 'auto' the default — a second AI now approves the commands you used to rubber-stamp, and it catches the dangerous ones you'd miss

From 14 August 2026, auto mode is the default in Claude Code for new sessions on Pro, Max and Team plans. A separate classifier model reviews every shell command, network call and agent message before it runs — replacing the permission prompt that Anthropic's own data shows you were approving 97% of the time. Here is exactly what is now allowed without asking, what is newly blocked, and how to decide whether to keep it on.

Read story →
anthropic claude-code

Claude Code now runs cloud sessions inside your own network — but your prompts still go to Anthropic, and that distinction is the whole story

On 13 August 2026 Anthropic opened a public beta of self-hosted environments for Claude Code, letting Team and Enterprise organizations run cloud agent sessions on infrastructure they control. Code checkouts, build artifacts, secrets and internal-service access stay in your network. But model inference still goes to api.anthropic.com, transcripts are stored, and you explicitly cannot route inference through Bedrock or Vertex. Here is what self-hosting actually gives you, what it does not, and how to tell whether it solves your problem or the wrong one.

Read story →
openai security

OpenAI built a model that writes exploits — and the interesting part is who is allowed to use it

GPT-5.6-Cyber completes 95% of exploit-chain, privilege-escalation and authentication-bypass requests, against 1.5% for the standard GPT-5.6 Sol it is built on. OpenAI is not selling it. Access runs through a vetted partner tier, Daybreak Red, and the model has already found a chainable zero-day in Chrome's V8 engine. The refusal rate was not a bug being fixed — it was a product decision.

Read story →
anthropic security

Anthropic proposes 'CVSS for AI jailbreaks' — a CJS-0 to CJS-4 severity scale, plus a HackerOne bounty on the restored Fable 5

On July 2, 2026, Anthropic — with Glasswing partners Amazon, Microsoft, and Google — proposed a Cyber Jailbreak Severity (CJS) scale grading AI jailbreaks from CJS-0 (Informational) to CJS-4 (Critical) on an exponential scale, across four axes: capability gain, breadth, ease of weaponization, and discoverability. The goal: a common language so AI developers and governments can talk about jailbreak risk in consistent terms. Anthropic also launched a HackerOne program for researchers to submit Fable 5 jailbreaks. It's the safety-governance response to the Fable 5 scramble and the Five Eyes cyber warning — the first standardized severity rubric for AI jailbreaks.

Read story →
regulation security

Why governments are suddenly gating frontier AI: the Five Eyes 'months, not years' cyber warning explained

On June 23, 2026, the Five Eyes cyber agencies (CISA, NSA, and the UK, Canada, Australia, and New Zealand equivalents) issued a joint statement warning that frontier AI models capable of overwhelming government and business defenses are 'months, not years' away. It's the security backdrop that explains the whole frontier-AI-regulation wave — the Fable 5 export-control shutdown, EO 14409's 'trusted partners,' the Mythos cyber-gating, ID verification. Here's what the warning actually says, why it's driving the government-gated regime, and what it means for the AI tools you use.

Read story →
anthropic security

Anthropic accuses Alibaba of the 'largest known distillation attack' on Claude — 25,000 fake accounts, 28.8 million exchanges

In a letter to the US Senate Banking Committee made public June 24, 2026, Anthropic accused Alibaba and its Qwen AI lab of 'brazenly' and 'illicitly' extracting Claude's capabilities — calling it the largest known distillation attack on the company. Anthropic says operators ran 28.8 million exchanges through roughly 25,000 fraudulent accounts between April 22 and June 5, targeting Claude's software-engineering and agentic-reasoning strengths. It follows February accusations against DeepSeek, Moonshot, and MiniMax. Here's what distillation is, why it matters for the tools you use, and the awkward connection to the Fable 5 export-control fight.

Read story →
anthropic security

Anthropic's Project Glasswing — Claude Mythos identifies 10,000+ critical vulnerabilities, kept restricted-access for safety

Anthropic's Project Glasswing update (May 26, 2026): Claude Mythos Preview — its unreleased frontier model — has identified more than 10,000 high- or critical-severity vulnerabilities in production software, including a 17-year-old FreeBSD remote-code-execution flaw (CVE-2026-4747). Partners include AWS, Apple, Broadcom, Cisco, Cloudflare, CrowdStrike, Google, JPMorgan, Linux Foundation, Microsoft, Mozilla, NVIDIA, Palo Alto Networks. Anthropic is keeping Mythos restricted-access, citing dual-use concerns and the absence of strong-enough safeguards across the industry.

Read story →