Shuo Yang on X: "Download: https://t.co/McSu3Uc1Cf
Code: https://t.co/fe3lnZU7xO
Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇" / X<br>Post
Log inSign up
Post
Shuo Yang on X: "Download: https://t.co/McSu3Uc1Cf
Code: https://t.co/fe3lnZU7xO
Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇"
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang
@Andy_ShuoYang
9h
Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization!
Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s
DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s
GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s Show more
00:00
84<br>160<br>1.1K<br>276K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang
@Andy_ShuoYang
9h
FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill
How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns.
More details in the technical report: arxiv.org/abs/2608.16157
12<br>42<br>320<br>35K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang
@Andy_ShuoYang
9h
FreeToken provides native GUI. No GGUF conversion. No building from source.
One-click install on Windows and Linux. FreeToken-desktop ships with agent harnesses built in — pick a model, pick an app, go.
95<br>11K
Shuo Yang
@Andy_ShuoYang
Download: flashml.ai
Code: github.com/FlashML-org/Fr…
Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇
flashml.ai<br>FreeToken — Bring Frontier to Edge<br>Download and run large language models on your own machine. Free for Windows & Linux.
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">5:42 PM · Aug 21, 20269.8KViews
79<br>12<br>171<br>192
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">tim su<br>@timbuksu
6h
Here's an interesting combination:<br>RTX 6000 Pro Blackwell + 32gb RAM
662
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">GopiNath<br>@gopinath9629
9h
256gb ddr4 ecc xeon e5 2680 v4 and 2x rtx306012gb and 2x rtx50608gb and via rpc rtx 3080 16gb
491
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Herbert West<br>@HerbertWest137
9h
Amazing! Rtx 4080 16GB, 32GB RAM.
437
Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
Shuo Yang@Andy_ShuoYangFollow<br>2nd year phd at Berkeley; Efficient ML System;
Trending now