LFM2.5-Encoders for Fast Long-Context Inference on CPU

gmays1 pts0 comments

LFM2.5-Encoders for Fast Long-Context Inference on CPU

Log In<br>Sign Up

Back to Articles<br>a]:hidden">

LFM2.5-Encoders for Fast Long-Context Inference on CPU

Community Article Published<br>July 28, 2026

Upvote 65

+59

Fernando Fernandes Neto fernandofernandes Follow

LiquidAI

Edoardo Mosca EdoardoMosca Follow

LiquidAI

Maxime Labonne mlabonne Follow

LiquidAI

Leonie Monigatti iamleonie Follow

LiquidAI

Today, we release two new encoder models on Hugging Face: LFM2.5-Encoder-230M and LFM2.5-Encoder-350M . They match the quality of larger models but stay fast as inputs get longer. This means you can run document-scale jobs on the hardware you already have, even on CPU.

Here's what you get:

Strong for their size: match or beat larger encoders on GLUE, SuperGLUE, and multilingual tasks.

8,192-token context with latency that grows slowly as inputs get longer.

Fast on CPU : about 3.7× faster than ModernBERT-base at long context.

With these, you can build intent routers, policy linters, PII detectors, and text classifiers that run cheaply, all day. See the live demos below.

Why we built a general-purpose encoder

Last month we released LFM2.5-Retrievers, built for multilingual search. LFM2.5-Encoders come from the same family but serve a broader purpose. They're pre-trained with a masked-language objective, so you can fine-tune them for classification, token-level tasks, and search alike. Search is just one thing an encoder enables. That's why we built a general-purpose model instead of reusing the retrievers.

Encoders power many modern production NLP applications: classifiers, intent routers, safety filters. These jobs run all day, usually on CPU, on ever-longer inputs. BERT established this class of model, and recently ModernBERT pushed its accuracy, speed, and context further. LFM2.5-Encoders take the next step on the LFM2 architecture, where cost grows slowly as inputs grow.

How the encoders are built

We initialize the encoders from their respective LFM2 decoder backbones: LFM2.5-230M and LFM2.5-350M. Then we turn each causal decoder into a bidirectional encoder with a few changes:

Bidirectional attention mask: each token now sees the tokens on both sides, not just the ones before it.

Non-causal short convolutions: we pad them symmetrically so each token's convolution mixes in its neighbors on both sides.

Masked language modeling: we mask 30% of the tokens during training.

We train both models in two stages:

General language competence : a short-context masked-language objective on a large web corpus at a 1,024-token context.

Long-context adaptation : extending context to 8,192 tokens on the full data mix, strengthening factual, legal, and multilingual competence.

Benchmark Results

We fine-tune each model fully on every task and report the resulting score. Across the table, that's 14 models on 17 tasks pulled from GLUE, SuperGLUE, and multilingual classification.

We report the mean across five held-out seeds, so the numbers are stable run to run. The full framework and raw results are open-sourced.

LFM2.5-Encoder-350M ranks fourth of the 14 models. The three ahead of it are all larger, including a 3.5B model nearly 10 times its size. LFM2.5-Encoder-230M beats ModernBERT-base and every EuroBERT model, while being smaller than most of them. Both also score well above our own LFM2.5-Retrievers here.

Inference speed on CPU and GPU

Our encoders inherit the LFM2 backbone's fast inference. Since both our encoders and ModernBERT support an 8,192-token context, we measure speed across the full range.

Our encoders show their biggest edge on CPU. Here, LFM2.5-Encoder-230M is the fastest at every sequence length (even faster than the smaller ModernBERT-base for short inputs). With increasing input length, throughput decreases sharply for ModernBERT, while our LFM2.5-Encoders rise into the mid-range before tapering. At 8,192 tokens, ModernBERT-base takes over a minute and a half per forward pass versus about 28s for LFM2.5-Encoder-230M. This is about 3.7x faster. For developers, that means you can scan or classify a full contract, transcript, or long support thread in under 30 seconds on a laptop CPU.

On GPU, a similar pattern holds with a smaller margin: ModernBERT-base leads below ~1K tokens on the Apple GPU. Our encoders take the lead from about 2K tokens. This shows that for long inputs, LFM2.5-Encoders are the faster choice, and if you're running on CPU, dramatically so.

LFM2.5-Encoder demos

We built the demos below from fine-tuned LFM2.5-Encoders. Each one runs in a CPU-only Hugging Face space:

Zero-shot prompt routing : define your own routing lanes as free text. The model scores the whole prompt against every lane in one pass.

Zero-shot policy linting : check text against your company's rules, written as free text. It scores every token against every rule in one pass.

Spell checking : correct misspellings token by token.

PII detection : spot and remove 40 kinds of personal...

lfm2 encoders context encoder token modernbert

Related Articles