Exciting update: Claude Sonnet 5.5 with xHigh reasoning has landed in the Code Arena: WebDev. With 1786 pts, its ranked #3!
Arena 宣布 Claude Sonnet 5.5(xHigh)进入 Code Arena: WebDev 榜单,以 1786 分排名第 3,距第 2 名 GPT-6 Astra 的 1788 分仅差 2 分。
Latest reporting, research and product updates filed under Reasoning.
A new AI system that excels at challenging games with hidden information could someday help human decision-makers select ideal strategies to outfox opponents in complicated situations like military maneuvers.Using advances in machine-learning, researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI that defeated top-ranked human players of the board wargame Stratego by a large margin — something no AI system had been able to achieve. Strateg
We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld fr
See model page Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices Gemini 4 Argon is Google DeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agenti
A single vllm serve process does three jobs that get in each other's way: processing prompts (prefill), generating tokens (decode) and a pile of CPU work around them. Disaggregated serving in vLLM separates the different stages of LLM inference. Splitting prefill from decode stops long prompts stalling everyone else's output, as long as the KV cache moves between them fast. Moving tokenization and parsing to a CPU-only frontend (/render, /derender) takes that work off your GPU nodes and leaves t
See model page GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task. Pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, except that the cache read discount rises from 90% to 95%. GPT-6.1 Sol’s overall blended price for agentic workloads is therefore slightly lower than GPT-6 Sol. This represents an additional price cut, following GPT-6 Sol’s original 50% discount from GPT-5.
OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,将 GPT-6 Astra 的训练方法用于更快更便宜的模型,API 价格较 GPT-5.6 促销价下调 50%(Sol 输入 $4→$2、输出 $20→$10;Luna 输入 $0.20→$0.10、输出 $1.20→$0.50,每百万 token)。
Fireworks Research 发布基于 Kimi K3 训练的专用模型 Ember-1,通过学习精简不必要的推理,在保持质量的同时减少约 35-50% 的 reasoning token。
Dwarkesh Patel 采访 OpenAI 研究员 Noam Brown,谈多智能体系统、对齐与递归自我改进。
DeepSeek 发布 DeepSeek-V4.1-Flash,一个 552B 参数的多模态 MoE 模型,支持最长 100 万 token 上下文,模型权重已在 Hugging Face 开放。
DeepSeek AI 发布多模态 MoE 模型 DeepSeek-V4.1-Flash,552B 主干加 196B Engram 参数,上下文窗口 1M,prefill 激活 8B 参数、decode 激活 16B,全局 KV 缓存降至每 token 890 字节,约为 DeepSeek-V4-Flash 的 1/4、DeepSeek-V1 的 1/437。
Cognition 发布最先进编码模型 SWE-2,基于 2.8T 参数的 Kimi K33 后训练,在 FrontierCode 1.1 Main 取得 50.0%,仅比 Fable 5.1 低不到一分且便宜 64%。
Epoch AI 测量四款模型的首 token 延迟(TTFT)随上下文长度的缩放,发现 GPT-5.6 Terra 和 Sol 呈明显二次曲线,Claude Sonnet 5 接近线性,Opus 5 噪声较大但仍接近线性。
Cohere 发布为 North Mini Code 构建的围绕 decode megakernel 的推理引擎,BF16 下单张 H100 端到端解码吞吐比 vLLM 快 1.25–1.41 倍,batch size 1 时达 292 tok/s(SoL 的 62%),代码已在 GitHub 开放。
GPT-6 Astra 的基准结论相互矛盾:Epoch AI 以 169 分将其排在 267 个模型之首,Artificial Analysis 给出 61 分,仅与前代 Sol 持平、落后 Claude Fable 5.1 的 66 分。
Greg Brockman 转发 @arcprize 的评测称 OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 上取得 SOTA,他称该基准已饱和。Astra 标准 harness 得分 63%,经新的 Provider Adapter harness 达 99%,在 96% 的 ARC-AGI-3 关卡上超越人类表现;排行榜图还显示更高推理层级通常成本更低,因为 Astra 用更少动作通关,减少模型调用和 token 数。
Perplexity CEO Aravind Srinivas 祝贺 OpenAI 发布 GPT-6 Astra,称其在宽度和深度研究任务上远超其他模型且更具成本效益,将很快向 Perplexity Computer 的 Pro 和 Max 用户开放。
Rohan Paul 梳理 OpenAI GPT-6 Astra 117 页系统卡的要点:Astra 控制自身链式思维的能力从 GPT-5.6 Sol 的 16.1% 跃升至 60.9%,可监控性相应下降。
François Chollet 发文称 GPT-6 Astra 在交互式推理任务上带来阶跃式能力提升,使用标准 harness 在 ARC-AGI-3 上得 66%,配合持续对话 harness 和自定义 compaction 接近 100%,每局成本约 $360。
NVIDIA 发布教程,讲解在 Jetson 上部署新一代开放推理模型的方法,以 Nemotron 3.5 Lightning 和 Qwen3.8-27B 为例。
Together AI 在 DeepSWE 全部 113 个任务、每任务 4 次试验(共 904 次 rollout)上对比 GLM-5.3 (max) 与 Claude Fable 5 (max)。
蚂蚁 Ling Infra 团队与 RadixArk SGLang 团队将 Ling-3.0-flash 混合线性注意力 MoE 模型的单请求解码速度从 288 tok/s 提升至 606 tok/s,平均 TPOT 从 3.33 ms 降至 1.53 ms。
SGLang 宣布对 NVIDIA Nemotron 3.5 Lightning 提供 Day-0 支持,该开源模型为 30B 总参数、3B 激活参数的混合专家架构,支持最长 1M token 上下文,可从 Hugging Face 下载 BF16 和 NVFP4 权重。模型支持 MTP、DFlash、DSpark 三种投机解码技术,并可通过 OpenAI 兼容 API 接入智能体工作流。
Inception Labs 发布面向搜索管线的 Mercury 2,扩散式并行解码超 1000 tokens/秒,在 WideSearch 各步骤比 Gemini 3.1 Flash Lite 快近 2 倍、比 GPT-5 Mini 快 10 倍,定价 $0.25/M 输入、$0.75/M 输出。