Run large language models at home, BitTorrent‑style

snorbleck1 pts0 comments

Petals – Run LLMs at home, BitTorrent-style

Petals

Run large language models at home, BitTorrent‑style

Generate text with Llama 3.1 (up to 405B), Mixtral (8x22B), Falcon (40B+) or BLOOM (176B)<br>and fine‑tune them for your tasks — using a consumer-grade GPU or Google Colab.

You load a part of the model, then join a<br>network<br>of people serving its other parts.<br>Single‑batch inference runs at up to 6 tokens/sec for Llama 2 (70B) and<br>up to 4 tokens/sec for Falcon (180B) — enough for<br>chatbots and interactive apps.

Beyond classic LLM APIs —<br>you can employ any fine-tuning and sampling methods, execute custom paths through the model, or see its hidden states.<br>You get the comforts of an API with the flexibility of PyTorch and 🤗 Transformers .

Thanks for subscribing!

We will email you only if we have really exciting updates.

Try now in Colab<br>Docs on GitHub

Top contributors right now:

Loading...<br>Oops, can't load the network status...

&bull;<br>and more

Contribute my GPU

Follow development in Discord or via email:

Subscribe

We send updates once a few months. No spam.

Submitting...

We sent you an email to confirm your address. Click it and you're in!

Featured on:

This project is a part of the BigScience research workshop.

home bittorrent style email large language

Related Articles