Fastest Inference Meta Muse Glimmer 30B on Apple

lukasonedge1 pts0 comments

Get BaseRT · Base Compute

BaseRT<br>Meet BaseRT, the fastest runtime on Apple Silicon<br>$ curl -LsSf https://basecompute.co/install.sh | sh<br>Copy<br>Technical ReportsBaseRT docsGitHubDiscord

Benchmarks<br>Faster than MLX and llama.cpp<br>On Prefill, up to 6.4x vs llama.cpp and 3.9x vs MLX. Up to 1.33x on Decode.<br>BaseRTMLXLlama.cpp<br>Decode · tg128<br>Qwen3 0.6B · Q4

531

398+33%

386+37%

Llama 3.2 1B · Q4

342

298+15%

267+28%

Llama 3.2 3B · Q4

137

131+5%

120+14%

Prefill<br>Qwen3 30B-A3B · pp128 · Q4

2,478

639+288%

1,280+94%

Gemma 4 E2B · pp2048 · Q8

16,264

12,355+32%

2,547+539%

Tokens / sec · Apple M5 Pro

Coding agents<br>Run it with a<br>local coding agent<br>Serve a model with BaseRT, point your agent at it, and keep everything on your machine. No API keys, no data leaving your device.

# 1. Serve a model<br>basert serve basecompute/gemma-4-E4B-it

# 2. Install the coding-agent plugin<br>pi install git:github.com/basecompute/pi-basert

# 3. Run it — everything is set<br>pi

Models<br>Supported models<br>Qwen3<br>Qwen3.5<br>Qwen3.6<br>Llama 3.1<br>Llama 3.2<br>Gemma 3<br>Gemma 4<br>Mistral<br>Phi-3<br>Nomic BERT

Community<br>Join us on Discord<br>For engineers building on-device AI and working with open source models.<br>Join the Discord →

Melbourne & Berlin

Melbourne & Berlin

basert llama qwen3 gemma apple basecompute

Related Articles