Konstantin on X: "If you serve models with mla on sglang with --attention-backend flashinfer and you get radix cache hits above 8k your outputs are wrong right now.
DeepSeek V2/V3/R1/V3.1/V3.2, bailing, glm4moe_lite, Sarvam affected.
Fix: https://t.co/IRxPRx9Ojf" / X<br>Post
Log inSign up
Post
Konstantin on X: "If you serve models with mla on sglang with --attention-backend flashinfer and you get radix cache hits above 8k your outputs are wrong right now.
DeepSeek V2/V3/R1/V3.1/V3.2, bailing, glm4moe_lite, Sarvam affected.
Fix: https://t.co/IRxPRx9Ojf"
Konstantin
@advprop
If you serve models with mla on sglang with --attention-backend flashinfer and you get radix cache hits above 8k your outputs are wrong right now.
DeepSeek V2/V3/R1/V3.1/V3.2, bailing, glm4moe_lite, Sarvam affected.
Fix: github.com/sgl-project/sg…
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">5:49 PM · Aug 16, 202610Views
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Konstantin
@advprop
5m
flashinfer api proposal for clarity: github.com/flashinfer-ai/…
some context: there is merge v2 api in sglang and it does:<br>S = log(exp(s_a) + exp(s_b))<br>v = v_a·exp(s_a − S) + v_b·exp(s_b − S)
which is true if and only if s_a and s_b are natural logs. but this is not true for Show more
feat(prefill): add return_lse_base_on_e option to BatchPrefillWithRaggedKVCacheWrapper by advprop...
From github.com
Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
Konstantin@advpropFollow<br>head of applied research @whitecircle
Trending now