Release Picchio v1.0.0 · logxio/picchio · GitHub
//releases/show" data-turbo-transient="true" />
Skip to content
Search/
Sign in<br>Sign upAppearance settings
You signed in with another tab or window. Reload to refresh your session.<br>You signed out in another tab or window. Reload to refresh your session.<br>You switched accounts on another tab or window. Reload to refresh your session.
Dismiss alert
{{ message }}
logxio
picchio
Public
Notifications<br>You must be signed in to change notification settings
Fork
Star<br>17
Picchio v1.0.0
Latest
Latest
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
logxio
released this
22 Aug 21:24
v1.0.0
98fdc5d
I ran the same Qwen3.8-27B file on two NVIDIA cards. On an RTX 5090 with 32 GB, all 66 layers stayed on the GPU and decode reached 81.5 tok/s. On an RTX 4070 SUPER with 12 GB, 28 layers landed on the CPU and decode fell to 5.7 tok/s. Picchio caught the 14.3× gap.
Browse measured runs
Download picchio.pyz and run it with Python 3.9+ on macOS, Linux or Windows:
python picchio.pyz MODEL
Picchio reports GPU layer placement, prefill, decode, GPU work, memory, power and energy per token. Add --share bug-report for a paste-ready Ollama or llama.cpp Issue.
Assets
Loading
Uh oh!
There was an error while loading. Please reload this page.
-->
All reactions
You can’t perform that action at this time.