Blog Heterogeneous CPU + GPU EPD Disaggregation to Boost VLM Serving TL;DR We enabled heterogeneous Encode-Prefill-Decode (EPD) disaggregation via Dynamo and SGLang for Vision-Language Models (VLMs). By offloading vision encoding tasks to CPUs (the easiest-getting CPU... Intel & SGLang Team
This story was filed as a headline only — the news service holds no English full text for it. Read the original at LMSYS Blog →