🔍 Read the full analysis: Open TTS Leaderboard Brings Scalable Evaluation To Multilingual Text-to-Speech on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has introduced an Open TTS Leaderboard that compares text-to-speech models using automated measures of speech accuracy, generation speed and speaker similarity. The project says its objective evaluations can run in hours, but the scores do not determine which voices sound most natural or which outputs listeners prefer.
Hugging Face has launched the Open TTS Leaderboard, a tool for comparing text-to-speech models on speech accuracy, inference speed and speaker similarity. The project is intended to make evaluations faster and more repeatable as the Hugging Face Hub hosted more than 8,000 TTS models as of September 30, 2026, according to the announcement; its automated scores are not a substitute for judging naturalness or listener preference.
The leaderboard reports word error rate and character error rate by comparing transcripts of generated speech with the text prompts used to produce it. Hugging Face says it uses Qwen3 automatic speech recognition for those transcript comparisons. For speed, the tool measures offline generation using inverse real-time factor on an H200 GPU, and streaming responsiveness using time-to-first-audio on both an H200 GPU and a CPU.
A separate speaker-similarity score is intended to assess voice identity preservation in cloning tasks. It compares WavLM embeddings from generated audio with embeddings from reference audio. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval; users can select other languages and switch to a voice-cloning view. The Open TTS Leaderboard offers a way to compare these evaluation measures across models.
Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among leading systems. These are results on the leaderboard’s chosen metrics, not an overall ranking of which systems sound best. The source material does not provide the full score tables or model-by-model evaluation sample sizes.
Faster Checks for Voice Models
A standardized, repeatable evaluation could help developers narrow down which open TTS systems to test for a particular language, voice-cloning task or latency requirement. The reported measures separate several practical considerations: intelligibility, generation speed and voice similarity. Teams building voice agents, for example, may pay close attention to time-to-first-audio, while a cloning application may prioritize similarity to a reference speaker.
Hugging Face says an evaluation using its objective metrics can take a couple of hours, compared with weeks for arena voting. That is the project’s estimate, not an independently verified performance comparison. Faster measurements may help address a gap in existing rankings: Hugging Face counted 16 open-weight models among 92 on Artificial Analysis as of September 30, 2026, and said Voice Arena showed a similar skew. It attributes the imbalance partly to the work required to host open models and commercial providers’ stronger incentives to seek placement.
The approach also makes tradeoffs easier to inspect rather than reducing every model to one preference score. A low error rate does not, by itself, establish that speech sounds expressive or natural. The leaderboard is best read as a set of technical indicators to guide further testing, alongside direct listening and any task-specific evaluation.
multilingual text-to-speech software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How TTS Comparisons Differ
Text-to-speech systems convert written prompts into spoken audio, and the growing number of model releases makes comparisons difficult to keep consistent. Existing references named by Hugging Face include TTS Arena v2, Artificial Analysis and Voice Arena. These use arena-style comparisons in which listeners hear outputs, vote on their preference and contribute to a ranking, often based on an Elo score calculated with a Bradley–Terry model.
Human voting captures preferences that automated measurements may miss, but it takes time to gather votes and requires the models to be hosted for testing. Hugging Face’s new system uses standardized datasets and automated measurements for much of its comparison process, while a separate “Listen” tab lets people hear outputs and submit preferences. The project says community votes may be incorporated into rankings later.
The measures have defined limits. Word and character error rates estimate how accurately speech can be transcribed by an ASR system; speaker similarity estimates how closely generated audio matches a reference voice. Neither directly measures naturalness, expressiveness or whether listeners prefer one voice. The announcement says Chinese, Japanese and Korean use character error rate. For languages beyond English and Chinese, it says Seed TTS Eval has no audio and scores come from CV3 Eval alone.
“The Open TTS Leaderboard does not replace human preference ranking.”
— Hugging Face
voice cloning speaker similarity tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Published Rankings
The supplied announcement does not include the full model list, detailed score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not show how closely the automated metrics align with listener judgments across languages, accents and speaking styles. Those omissions limit how much can be inferred from a model’s position or score alone.
Hugging Face says community votes may be added to the leaderboard as feedback accumulates, but has not specified a threshold or timetable. The announcement also does not say how frequently rankings will be refreshed, or how updates to models, datasets or evaluation methods will be tracked. Rankings may be useful for shortlisting systems, but the published information does not settle those questions.
As an affiliate, we earn on qualifying purchases.
Listening Tests and Future Votes
Users can go to the leaderboard’s “Listen” tab to select a language and dataset, compare generated audio and vote on outputs. The interface also lets users choose whether to compare voice cloning. Hugging Face asks participants to sign in with an account, saying this helps limit spam and bot submissions.
The project says it may use community preferences in future leaderboard results, but has not announced when or how votes would be incorporated. Readers and developers can expect the evaluation to become more informative as documentation, results and feedback are published; for now, the automated metrics should be considered alongside direct listening and the limits described in the announcement.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Open TTS Leaderboard measure?
It compares speech transcription error, generation speed and speaker similarity. The tool also lets users listen to generated audio and submit preferences.
Does the ranking show which voice sounds most natural?
No. Hugging Face says the objective metrics do not replace human preference ranking. The reported measures do not directly assess naturalness, expressiveness or listener preference.
Which models does Hugging Face identify as strong multilingual systems?
The announcement names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. It also lists several systems among leaders for English error rates; those results apply to the metrics and evaluation setup used by the leaderboard.
How long does an evaluation take?
Hugging Face says its objective-metric evaluation can take a couple of hours. The source compares that with weeks for arena voting, but does not provide an independent time study.
Will listener votes affect the rankings?
Hugging Face says community votes may be incorporated later. It has not specified when that would happen or what voting threshold or method it would use.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
