Run frontier models on gaming GPUs

eob1 pts0 comments

Shuo Yang on X: "Download: https://t.co/McSu3Uc1Cf

Code: https://t.co/fe3lnZU7xO

Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇" / X<br>Post

Log inSign up

Post

Shuo Yang on X: "Download: https://t.co/McSu3Uc1Cf

Code: https://t.co/fe3lnZU7xO

Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇"

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang

@Andy_ShuoYang

9h

Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization!

Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s

DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s

GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s Show more

00:00

84<br>160<br>1.1K<br>276K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang

@Andy_ShuoYang

9h

FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill

How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns.

More details in the technical report: arxiv.org/abs/2608.16157

12<br>42<br>320<br>35K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Shuo Yang

@Andy_ShuoYang

9h

FreeToken provides native GUI. No GGUF conversion. No building from source.

One-click install on Windows and Linux. FreeToken-desktop ships with agent harnesses built in — pick a model, pick an app, go.

95<br>11K

Shuo Yang

@Andy_ShuoYang

Download: flashml.ai

Code: github.com/FlashML-org/Fr…

Reply with your GPU + RAM, and I'll tell you the biggest frontier model your machine can run 👇

flashml.ai<br>FreeToken — Bring Frontier to Edge<br>Download and run large language models on your own machine. Free for Windows & Linux.

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">5:42 PM · Aug 21, 20269.8KViews

79<br>12<br>171<br>192

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">tim su<br>@timbuksu

6h

Here's an interesting combination:<br>RTX 6000 Pro Blackwell + 32gb RAM

662

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">GopiNath<br>@gopinath9629

9h

256gb ddr4 ecc xeon e5 2680 v4 and 2x rtx306012gb and 2x rtx50608gb and via rpc rtx 3080 16gb

491

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Herbert West<br>@HerbertWest137

9h

Amazing! Rtx 4080 16GB, 32GB RAM.

437

Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

Shuo Yang@Andy_ShuoYangFollow<br>2nd year phd at Berkeley; Efficient ML System;

Trending now

span empty before shuo yang content

Related Articles