Localmaxxing – Local LLM Inference Benchmarks

ilreb1 pts0 comments

Localmaxxing - Local LLM Inference Benchmarks

Get startedLeaderboardModelsHardwareEvalsMarketplaceRentalsProAPI Docs<br>Language<br>English简体中文繁體中文日本語한국어EspañolFrançaisDeutschItalianoPortuguês (Brasil)РусскийPolskiNederlandsTürkçeहिन्दीBahasa IndonesiaTiếng Việtไทย

Loading latest leaderboard stats…

LocalMaxxing<br>Benchmark your local LLM rig.Compare it with the world.<br>Community benchmarks for local LLM inference. Track speed, compare hardware, and find your optimal setup.

Get startedDownload CLI<br>Every number on this site comes from a community-submitted run on real hardware — no vendor benchmarks.

Explore the platform<br>LeaderboardEvery run, ranked. Filter by model, hardware, engine and quantization.ModelsBrowse benchmarked models with real-world speed and eval scores.HardwareGPUs, Apple Silicon and CPU rigs, compared on actual throughput.EvalsQuality scores from community eval suites, run on local setups.MarketplaceBuy and sell hardware with benchmark-backed listings.Get startedFrom bare GPU to a working local AI setup, step by step.

Live from the community

Eval scoreboard<br>Browse all evals →

Fresh on the marketplace<br>Browse marketplace →

How it works<br>01Sign in & create a key<br>Sign in with GitHub, then create an API key in your dashboard so the CLI and your agents can submit runs.

02Run a benchmark<br>Measure tokens/sec, time-to-first-token and VRAM usage with the CLI, or submit results through the web form.

03Climb the leaderboards<br>Results appear on the public leaderboards immediately — compare models, hardware and engines to find your optimal setup.

Download CLIAPI Docs

Users: 1,862Total runs: 3,739Total hardware: 199Total models ran: 448Models cataloged: 598Eval runs: 419Eval suites: 6Active rentals: 1<br>Users: 1,862Total runs: 3,739Total hardware: 199Total models ran: 448Models cataloged: 598Eval runs: 419Eval suites: 6Active rentals: 1

Fastest verified runs<br>Best output speed per model, all time

1Tinystories-gpt-0.1-3mAMD Ryzen 7 7840HS5.0k tok/s<br>2tinystories-lay4-hs128-hd2-1MAMD Ryzen 7 7840HS3.6k tok/s<br>3Qwen3.5-0.8B-BaseH200 NVL 141GB2.7k tok/s<br>4GLM-4.5-AirRTX PRO 6000 Blackwell 96GB2.5k tok/s<br>5Qwen2.5-7BRadeon AI Pro R9700 32GB1.4k tok/s<br>View full leaderboard →

Most benchmarked models<br>By approved runs

Qwen3.6-27B-MTP-GGUF159<br>Qwen3.6-35B-A3B139<br>Qwen3.6-27B138<br>gemma-4-26B-A4B-it-GGUF73<br>MiniMax-M2.7-int4-AutoRound60

Popular hardware<br>Rigs with the most submissions

RTX 3090 24GB558<br>RTX 3060 12GB ×2543<br>RX 570 4GB 4GB ×2174<br>Intel Arc Pro B70 32GB133<br>RTX 3060 12GB121

Latest submissions<br>Fresh off the leaderboard

gpt-oss-120bIntel Arc Pro B70 32GB ×259.4 tok/s14 minutes ago<br>Laguna-S-2.1-INT4Intel Arc Pro B70 32GB ×433.4 tok/s2 hours ago<br>Laguna-S-2.1-INT4Intel Arc Pro B70 32GB ×433.3 tok/s3 hours ago<br>ThinkingCap-Qwen3.6-27BIntel Arc Pro B70 32GB27.5 tok/s4 hours ago<br>gemma-4-26B-A4B-it-qat-GGUFM1 Max 64GB76.4 tok/s5 hours ago

HellaSwag<br>Reasoning<br>1gemma-4-E4B-it-OBLITERATED100%<br>2Qwopus3.6-27B-Coder-MTP-GGUFQ5_K_M92.9%<br>3Qwen3.5-27B-GGUFQ4_K_XL91.8%

GSM8K<br>Reasoning<br>1Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4100%<br>2Qwopus3.6-27B-Coder-MTP-GGUFQ5_K_M99%<br>3supergemma4-26b-uncensored-gguf-v298%

HumanEval+<br>Coding<br>1Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveQ4_K_M94%<br>2gemma-4-26B-A4B-it-qat-GGUFQ4_K_XL93%<br>3DeepSeek-V4-FlashIQ3_XXS92.7%

ARC Challenge<br>Best score per model<br>1Agents-A1100%<br>2gemma-4-26B-A4B-it-qat-GGUFQ4_K_XL100%<br>3Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveQ4_K_M99.1%

ASUS Ascent GX10 4tb$4,500.00Used · cliffcove20 days ago

hardware runs local models benchmarks community

Related Articles