Get BaseRT · Base Compute
BaseRT<br>Meet BaseRT, the fastest runtime on Apple Silicon<br>$ curl -LsSf https://basecompute.co/install.sh | sh<br>Copy<br>Technical ReportsBaseRT docsGitHubDiscord
Benchmarks<br>Faster than MLX and llama.cpp<br>On Prefill, up to 6.4x vs llama.cpp and 3.9x vs MLX. Up to 1.33x on Decode.<br>BaseRTMLXLlama.cpp<br>Decode · tg128<br>Qwen3 0.6B · Q4
531
398+33%
386+37%
Llama 3.2 1B · Q4
342
298+15%
267+28%
Llama 3.2 3B · Q4
137
131+5%
120+14%
Prefill<br>Qwen3 30B-A3B · pp128 · Q4
2,478
639+288%
1,280+94%
Gemma 4 E2B · pp2048 · Q8
16,264
12,355+32%
2,547+539%
Tokens / sec · Apple M5 Pro
Coding agents<br>Run it with a<br>local coding agent<br>Serve a model with BaseRT, point your agent at it, and keep everything on your machine. No API keys, no data leaving your device.
# 1. Serve a model<br>basert serve basecompute/gemma-4-E4B-it
# 2. Install the coding-agent plugin<br>pi install git:github.com/basecompute/pi-basert
# 3. Run it — everything is set<br>pi
Models<br>Supported models<br>Qwen3<br>Qwen3.5<br>Qwen3.6<br>Llama 3.1<br>Llama 3.2<br>Gemma 3<br>Gemma 4<br>Mistral<br>Phi-3<br>Nomic BERT
Community<br>Join us on Discord<br>For engineers building on-device AI and working with open source models.<br>Join the Discord →
Melbourne & Berlin
Melbourne & Berlin