GPTProto

vLLM x AgentX: Optimizing for Real-World Agentic Serving

vLLM 官方博客(RSS)··DeepSeek

vLLM 团队发布针对智能体负载优化的博客,基于 SemiAnalysis 的公开 AgentX 基准展示成果:DeepSeek V4 Pro 在 GB300 上达到 83K total tokens/GPU-秒。