I made this after seeing someone posit the idea online yesterday over lunch then spent some time refining it. So far it s pretty impressive IMO! Right now I am running Qwen3-30B-A3B on my 24gb unified memory m4 MacBook Pro at 50 tok/sec and this should definitely not be working for such a large model on my middling hardware.Things are detailed in the README to get up and running and DESIGN.md has details on all the choices and such made along the way.