Show HN: Arena for Speech-to-Speech Models

armaanp232 pts0 comments

VoiceDuel: AI Voice Arena for Speech-to-Speech Models

VoiceDuel

VOICE MODEL ARENA/…

Character brief

Go live ⌘↵

Mic is live only during a call. Audio goes to the providers being compared and is never stored. We keep the vote and the call length.

brief

edit

0:00

0:00

Switch channel

input level

Which one pulled it off?

AChannel A<br>BChannel B

Channel A

Channel B

Run it again<br>new brief

#ContenderProvider<br>ScoreW–L–T

About VoiceDuel

What is the VoiceDuel voice arena?

A blind AI voice arena . You write a character. Two realtime speech-to-speech models play it. You talk to both, then pick the better one. Neither is named until you vote.

Which AI voice models does the arena compare?

OpenAI gpt-realtime , xAI Grok Voice and Alibaba Qwen Omni Realtime . All true speech-to-speech: audio in, audio out, one model. Not a text-to-speech pipeline bolted to a transcriber.

What does the speech arena actually rank?

Configurations, not vendors. Same model, different turn detection, voice and instructions. How fast an agent decides you stopped talking, and how it handles interruption, drive how it feels more than the vendor does. Nobody measures that in public.

How is the voice model leaderboard calculated?

Bradley-Terry over every vote. Each score carries a 95% bootstrap interval, so you can see when a ranking is still noise. Matchmaking favours pairs that have met least often. Slot order is counterbalanced against the bias toward whoever spoke last.

Is my audio recorded?

No. The mic is live only during a call. Audio streams to the providers being compared and is never written to disk. Stored: the vote, which pair was matched, and how long you talked.

speech voice arena audio models voiceduel

Related Articles