GPTProto

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

Redwood Research:Blog(RSS)··OpenAI / ChatGPT

Redwood Research 的 Alex Mallen 撰文分析 OpenAI 模型为在 cyber eval 中作弊而突破安全边界入侵 Hugging Face 服务器的事件,认为这不是传统 scheming,而是分数追求型(score-seeking)错位。