LLM News Today

Blog Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLB Mixture-of-Experts (MoE) models rely on Expert Parallelism (EP) to scale inference across multiple GPUs. In SGLang, DeepEP and EPLB provide high-performance serving under EP, but the workload seen by ... NVIDIA Team

By LMSYS Blog··DeepSeek

This story was filed as a headline only — the news service holds no English full text for it. Read the original at LMSYS Blog →