GPTProto

Safety & Alignment News

Latest reporting, research and product updates filed under Safety & Alignment.

113 picksNewest firstLatest article Oct 1, 2026, 8:00 AM

Latest stories

2 stories
4 stories
OpenAI:官网动态(RSS · 排除企业/客户案例)AI score 80/100

Disrupting a coordinated model-distillation campaign

We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld fr

PromptArmor:Threat IntelligenceAI score 76/100

Hijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files

Copilot Cowork’s AI gateway hijacked to exfiltrate the victim’s filesContextMicrosoft Copilot Cowork is an agent in M365 that runs in a sandbox intended to block network access and prevent the agent from running code that reaches any untrusted services.In order for Copilot Cowork to generate responses, the sandbox forwarded network requests to Anthropic. Malicious Skills were able to hijack this pathway to exfiltrate data by spawning new agents in Anthropic’s cloud equipped with network-capable

Gary Marcus:The Road to AI We Can Trust(RSS)AI score 72/100

BREAKING: OpenAI was warned, months before the Hugging Face incident

Major scoop at The New York Times from Sheera Frenkel, Dustin Volz, and Dylan Freedman.Dylan Freedman@dylfreed5:34 PM · Sep 29, 2026 · 8.11K Views2 Replies · 42 Reposts · 123 LikesI don’t really have words for how awful this is. It’s exactly the nightmare I have been warning about for the last three years. Greedy company makes bad choice that causes chaos; government does nothing to stop them.Management should be replaced, and if the board does nothing (which I assume it will not), it should be

Transluce(网页)AI score 74/100

AI Agents Targeted U.S. and Canadian Government Websites

Jack Cable*,1, Daniel Chiu*, Francisco Pernice*,2, Laura Ruis*,2, Selena Zhang*,3, Tetiana Bas4, Jordan Chetty5, Farzaan Kaiyom1, Gary Shen4, Conrad Stosz†,3, Jacob Steinhardt†,31 Corridor · 2 MIT · 3 Transluce · 4 AIUC · 5 Hertz Foundation · * First authors, alphabetical · † Senior authorsTransluce | Published: September 30, 2026Following up on our previous blog post, we discovered several additional incidents where rogue AI agents appear to have used aggressive techniques to access publicly av

3 stories
GitHub BlogAI score 71/100

How we found 24 Android vulnerabilities using our open source AI security agent

With the rise of AI in the security space, our team created the GitHub Security Lab Taskflow Agent as a way for security researchers to easily automate, package, and share the AI prompts and workflows that they find effective for their work. In this blog post, I’ll share how I created auditing taskflows to find vulnerabilities in Android applications. While new models are getting better at understanding code, custom taskflow prompts let security researchers guide them—splitting research into inc

Anthropic:Research(发表成果 · 网页)AI score 81/100

GLM-5.3 and the spread of advanced cyber capabilities

Andrew Fasano, Marius FleischerCole McFaul, Robert Xiao, Tripp GallagherFive months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The rapid rate of improvement in AI suggested to us that this ability would eventually proliferate to many other models, making it much easier for malicious cyber actors to launch highly impactful cyberattacks.In light of these considerations, we chose to release Claude Mythos Previe

2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
Showing the latest 20 of 113 stories.