We tested our new Voice Isolation across 10 STT engines. Word errors dropped 46%

davitb1 pts0 comments

Krisp Voice Isolation 2.5: Cut STT Word Error Rate

AI Meeting Assistant

Back

AI Meeting Assistant

with #1 Noise Cancellation

Explore AI Meeting Assistant

AI Notetaker

AI Note Taker

Meeting Transcription

Meeting Recording

Meeting Summary

Real Time Voice AI

Noise Cancellation

Accent Conversion -<br>Speaker side

Accent Conversion -<br>Listener side

Use cases

IT Consulting

MSP

Call Center AI

Back

Call Center AI

AI that boosts call center productivity

Explore platform

Speech Assist

Accent Conversion

Real-time accent conversion for call center agents.

Voice Translation

Real-time AI voice translation for call center agents.

Noise Cancellation

Remove background noises, voices & echoes.

Agent Assist

Agent Assist

Real-time AI assistant for call center agents.

Speech Analytics

Speech Analytics

Call scoring, Compliance monitoring and more.

Developers

Back

Developers

with #1 AI Voice Models

Explore Developers

For Voice AI Agents

Voice Isolation

Isolate the primary speaker's voice

Turn-Taking

Improving turn-taking for AI

For Human-to-human Calls

Accent Conversion

Convert accents in calls

Noise Cancellation

Noise removal in calls

Voice Translation API<br>New

Real-time translation, self-serve

Customers

Pricing

Book a demo

Subscribe to get the latest insights weekly

Subscribed successfully

Subscribe

This form is protected by reCAPTCHA and the Google<br>Privacy<br>Policy and<br>Terms of<br>Service apply.

August 12, 2026<br>Voice Isolation 2.5: Built for STT, Not Just Human Ears

Written by Krisp Engineering Team

Max 9 min read

Share this post

Get Krisp for Free

Modern STT models don’t handle competing voices well

VIVA is a collection of real-time AI models for voice agents, including Voice Isolation (VI), turn-taking, and interruption prediction. The Voice Isolation model comes first in the chain, delivering clean audio by removing noise and background speech so the turn-taking model can tell when a caller has actually finished speaking. Those models run in production today across more than 1B+ minutes of voice-AI conversations every month.

Having achieved effective voice isolation in the VIVA models, we saw customers running VIVA in front of their STT as well, improving transcription accuracy on real-world calls, where background speech like a TV playing in the room would otherwise end up in the transcript.

That points to something modern speech-to-text still gets wrong. Background noise robustness in modern STT systems is largely a solved problem: feed today’s STT models fan noise or traffic, and they transcribe it well. Other people talking in the background is not transcribed well. When a second voice overlaps the primary speaker, word error rates degrade, and conventional denoising can’t help, because the interference is speech, not noise. Separating the primary speaker from competing voices is exactly what VIVA’s Voice Isolation does.

And the payoff is large.

🗣️ On calls with competing speakers, putting VIVA in front of the STT cuts average WER (word error rate) from around 36% to 11% across the STT models we tested, cutting the errors by roughly 70% . That is something that simply noise filtering can’t solve.

But as more teams run isolation ahead of an STT, the feedback and example recordings customers shared. It led us to a consistent insight:

People and STT models don’t listen the same way.

To isolate a primary speaker from overlapping conversation, VI sometimes has to aggressively remove secondary speech, and on the hardest segments, that removal can affect the primary voice too. To a human listener, the result is still perfectly usable. But an STT model, which isn’t trained on isolated-voice output, can read those affected moments as noise and drop words that were actually spoken.

The result wasn’t noise in the transcript; it was gaps in it. Deletions.

VI 2.5 , the Voice Isolation model within VIVA that we’re releasing today, is built to address that feedback. It isolates voice more surgically and shapes the output so STT models transcribe it more faithfully: fewer deletions on challenging segments and a lower WER across every model we tested. We’re shipping it in two sizes: the full VI 2.5, and a lite version for CPU-constrained and edge deployments.

Introducing Voice Isolation (VI) 2.5

VI 2.5 is our latest and most advanced general-purpose voice isolation model for conversational AI. It is STT-agnostic and lowers WER against no processing on every engine we tested, with its largest gains in reverberant conditions.

Four results define this release:

WER drops 46.4% on average. From 17.9% on untouched audio to 10.2% with Voice Isolation

2.5 in front of the STT, and the gain holds on every engine we tested.

69.7% fewer errors on competing speech. On calls with a competing speaker, the case that

decides whether a real call works, word error rate falls from 37% to 10.8% versus no

processing, the interference denoising can’t touch.

The clean-audio penalty is...

voice isolation noise speech models real

Related Articles