VoiceDuel — AI Voice Arena for Speech-to-Speech Models
VoiceDuel
Talk to any character.<br>Two voice models play it.
Character brief
Go live ⌘↵
Your microphone is only on during a conversation. Audio streams to the AI<br>voice providers being compared and is not recorded or stored — we keep the<br>vote and how long you talked, nothing else.
brief
edit
0:00
0:00
Switch channel
input — this should move while you speak
Which one pulled it off?
AChannel A<br>BChannel B
Channel A
Channel B
Run it again<br>new brief
Voice model arena standings
Live standings for the AI voice arena, fitted with Bradley-Terry over every vote cast. The ± is a parametric bootstrap, not a guess.
#ContenderProvider<br>ScoreW–L–T
VoiceDuel is an independent project. Models are compared blind and named only<br>after you vote. Theatre photograph: Teatro Grande, Pompeii.
About VoiceDuel
What is the VoiceDuel voice arena?
VoiceDuel is a blind AI voice arena for realtime<br>speech-to-speech models . You write a character brief — a hiring manager, an angry<br>customer, a medieval knight — and two different voice models play it. You<br>hold a live spoken conversation with each one, then say which performance<br>was better. Neither model is named until after you vote.
Which AI voice models does the arena compare?
Contenders are drawn from the current generation of realtime speech-to-speech<br>APIs, including OpenAI gpt-realtime , xAI Grok<br>Voice and Alibaba Qwen Omni Realtime . These are<br>true speech-to-speech models — audio in, audio out from a single model —<br>not a text-to-speech and speech-recognition pipeline stitched together.
What does the speech arena actually rank?
Most leaderboards rank vendors. VoiceDuel ranks configurations :<br>the same model with different turn-detection settings, voices and<br>instructions. How quickly an agent decides you have finished speaking, how<br>it handles being interrupted, and how well it stays in character matter<br>more to how a voice agent feels than which company trained it — and almost<br>nobody measures them in public.
How is the voice model leaderboard calculated?
Votes are fitted with a Bradley-Terry model rather than Elo,<br>because the full vote history is available and does not need an online<br>approximation. Each score carries a 95% confidence interval from a<br>parametric bootstrap, so you can see when a ranking is still noise.<br>Matchmaking favours pairs that have met least often, and slot order is<br>counterbalanced to cancel out the tendency to prefer whoever spoke last.
Is my audio recorded?
No. Your microphone is active only during a conversation. Audio is streamed<br>to the voice providers being compared for the duration of the call and is<br>not written to disk. What is stored is the vote itself, which two<br>configurations were paired, and how long each conversation lasted.