Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090

unseenmars1 pts0 comments

SGLang on X: "The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:<br>- 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark<br>- 38.28 tok/s decode on DGX Spark<br>Qwen3.8-27B raises the bar again for what a small model" / X<br>Post

Log inSign up

Post

SGLang on X: "The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:<br>- 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark<br>- 38.28 tok/s decode on DGX Spark<br>Qwen3.8-27B raises the bar again for what a small model"

SGLang

@sgl_project

The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:<br>- 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark<br>- 38.28 tok/s decode on DGX Spark<br>Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks.

Long live the (small model) king! Run it locally with SGLang 👇

00:00

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Qwen

@Alibaba_Qwen

2h

We promised open weights for Qwen3.8. Now, time to meet them! 🎉

⚡ Qwen3.8-27B:<br>- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.<br>- 262K native context, easily extendable to 1M Show more

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">3:07 PM · Aug 14, 202677.5KViews

11<br>38<br>285<br>118

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">SGLang

@sgl_project

2h

Cookbook:

Qwen3.8-27B - SGLang Documentation

From docs.sglang.io

24<br>2.6K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Roberto Tomás C 🍉<br>@robertomasymas

2h

Haha yeah yeay! But its not “back” its the first model I know of to maintain number 1 slot until they replaced it themselves with a newer version 😃

1.4K

Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

SGLang@sgl_projectFollow<br>Run LLMs fast at any scale 🔗 https://t.co/F3u6wYESL0<br>Join our community https://t.co/fmlOfTOEec<br>For AI tech blogs & deep-dives 👉 @lmsysorg

Trending now

span empty before qwen3 sglang small

Related Articles