Krisp Voice Isolation 2.5: Cut STT Word Error Rate
AI Meeting Assistant
Back
AI Meeting Assistant
with #1 Noise Cancellation
Explore AI Meeting Assistant
AI Notetaker
AI Note Taker
Meeting Transcription
Meeting Recording
Meeting Summary
Real Time Voice AI
Noise Cancellation
Accent Conversion -<br>Speaker side
Accent Conversion -<br>Listener side
Use cases
IT Consulting
MSP
Call Center AI
Back
Call Center AI
AI that boosts call center productivity
Explore platform
Speech Assist
Accent Conversion
Real-time accent conversion for call center agents.
Voice Translation
Real-time AI voice translation for call center agents.
Noise Cancellation
Remove background noises, voices & echoes.
Agent Assist
Agent Assist
Real-time AI assistant for call center agents.
Speech Analytics
Speech Analytics
Call scoring, Compliance monitoring and more.
Developers
Back
Developers
with #1 AI Voice Models
Explore Developers
For Voice AI Agents
Voice Isolation
Isolate the primary speaker's voice
Turn-Taking
Improving turn-taking for AI
For Human-to-human Calls
Accent Conversion
Convert accents in calls
Noise Cancellation
Noise removal in calls
Voice Translation API<br>New
Real-time translation, self-serve
Customers
Pricing
Book a demo
Subscribe to get the latest insights weekly
Subscribed successfully
Subscribe
This form is protected by reCAPTCHA and the Google<br>Privacy<br>Policy and<br>Terms of<br>Service apply.
August 12, 2026<br>Voice Isolation 2.5: Built for STT, Not Just Human Ears
Written by Krisp Engineering Team
Max 9 min read
Share this post
Get Krisp for Free
Modern STT models don’t handle competing voices well
VIVA is a collection of real-time AI models for voice agents, including Voice Isolation (VI), turn-taking, and interruption prediction. The Voice Isolation model comes first in the chain, delivering clean audio by removing noise and background speech so the turn-taking model can tell when a caller has actually finished speaking. Those models run in production today across more than 1B+ minutes of voice-AI conversations every month.
Having achieved effective voice isolation in the VIVA models, we saw customers running VIVA in front of their STT as well, improving transcription accuracy on real-world calls, where background speech like a TV playing in the room would otherwise end up in the transcript.
That points to something modern speech-to-text still gets wrong. Background noise robustness in modern STT systems is largely a solved problem: feed today’s STT models fan noise or traffic, and they transcribe it well. Other people talking in the background is not transcribed well. When a second voice overlaps the primary speaker, word error rates degrade, and conventional denoising can’t help, because the interference is speech, not noise. Separating the primary speaker from competing voices is exactly what VIVA’s Voice Isolation does.
And the payoff is large.
🗣️ On calls with competing speakers, putting VIVA in front of the STT cuts average WER (word error rate) from around 36% to 11% across the STT models we tested, cutting the errors by roughly 70% . That is something that simply noise filtering can’t solve.
But as more teams run isolation ahead of an STT, the feedback and example recordings customers shared. It led us to a consistent insight:
People and STT models don’t listen the same way.
To isolate a primary speaker from overlapping conversation, VI sometimes has to aggressively remove secondary speech, and on the hardest segments, that removal can affect the primary voice too. To a human listener, the result is still perfectly usable. But an STT model, which isn’t trained on isolated-voice output, can read those affected moments as noise and drop words that were actually spoken.
The result wasn’t noise in the transcript; it was gaps in it. Deletions.
VI 2.5 , the Voice Isolation model within VIVA that we’re releasing today, is built to address that feedback. It isolates voice more surgically and shapes the output so STT models transcribe it more faithfully: fewer deletions on challenging segments and a lower WER across every model we tested. We’re shipping it in two sizes: the full VI 2.5, and a lite version for CPU-constrained and edge deployments.
Introducing Voice Isolation (VI) 2.5
VI 2.5 is our latest and most advanced general-purpose voice isolation model for conversational AI. It is STT-agnostic and lowers WER against no processing on every engine we tested, with its largest gains in reverberant conditions.
Four results define this release:
WER drops 46.4% on average. From 17.9% on untouched audio to 10.2% with Voice Isolation
2.5 in front of the STT, and the gain holds on every engine we tested.
69.7% fewer errors on competing speech. On calls with a competing speaker, the case that
decides whether a real call works, word error rate falls from 37% to 10.8% versus no
processing, the interference denoising can’t touch.
The clean-audio penalty is...