vLLM on X: "Kimi K3 is here: a 2.8T-parameter MoE with a 1M-token context and native vision, running on vLLM from day 0.
Here is our canonical deployment guide: architecture, kernels, recipes, and the flags to run it in production. https://t.co/SATHaHHYef" / X<br>Post
Log inSign up
Post
vLLM
@vllm_project
Kimi K3 is here: a 2.8T-parameter MoE with a 1M-token context and native vision, running on vLLM from day 0.
Here is our canonical deployment guide: architecture, kernels, recipes, and the flags to run it in production.<br>The Inference Engine Guide for K3 Deployment
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">5:07 PM · Jul 27, 202637KViews
14<br>54<br>375<br>264
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeff Steve<br>@JeffSte17327059
Jul 27
scan.co.uk/3xs/configurat…<br>do i need to buy 9 of these racks to achieve what you spoke about in the blog 60 users 400 tk/s decode?
svg]:size-5 text-body hover:bg-mix-current hover:bg-mix-amount-10 active:bg-mix-current active:bg-mix-amount-15 focus-visible:bg-mix-current focus-visible:bg-mix-amount-10 outline-current -m-2 shrink-0 cursor-pointer border-transparent p-0 text-body size-9 [&>svg]:size-[1.25em] [&>[data-engagement-icon]]:size-[1.25em] group-hover:bg-mix-current group-hover:bg-mix-amount-10" aria-label="View count" type="button" data-state="closed" href="/JeffSte17327059/status/2081818224883105964/quotes">630
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Ayush kumar
@Ayushkumar1808
Jul 27
What is the memory overhead of caching recurrent state per prefix?
svg]:size-5 text-body hover:bg-mix-current hover:bg-mix-amount-10 active:bg-mix-current active:bg-mix-amount-15 focus-visible:bg-mix-current focus-visible:bg-mix-amount-10 outline-current -m-2 shrink-0 cursor-pointer border-transparent p-0 text-body size-9 [&>svg]:size-[1.25em] [&>[data-engagement-icon]]:size-[1.25em] group-hover:bg-mix-current group-hover:bg-mix-amount-10" aria-label="View count" type="button" data-state="closed" href="/Ayushkumar1808/status/2081800403222499392/quotes">47
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">jeejeelee<br>@xiongmaogpu
Jul 28
cool
svg]:size-5 text-body hover:bg-mix-current hover:bg-mix-amount-10 active:bg-mix-current active:bg-mix-amount-15 focus-visible:bg-mix-current focus-visible:bg-mix-amount-10 outline-current -m-2 shrink-0 cursor-pointer border-transparent p-0 text-body size-9 [&>svg]:size-[1.25em] [&>[data-engagement-icon]]:size-[1.25em] group-hover:bg-mix-current group-hover:bg-mix-amount-10" aria-label="View count" type="button" data-state="closed" href="/xiongmaogpu/status/2081907911178309731/quotes">149
Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
vLLM@vllm_projectFollow<br>A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!
Trending now