Start with what you make: emotional performances, local-audience content, your own voice or steady narration. The card scores below all use our internal video set.
Drama, entertainment and emotional performances
Nightingale
For creators and studios seeking premium delivery through emotional highs and lows. Nightingale ranks first in mean source-expression tracking on both tested languages.
Internal videos · Expression tracking: EN 0.45, KO 0.54. These scores measure vocal intensity, not overall quality.
Lectures, corporate videos and local audiences
Oriole
An accent that sounds natural to local audiences, based on listening evaluation. Choose Oriole for lectures and corporate videos that need a consistent voice from one sentence to the next.
Internal videos · EN voice consistency: 0.661, the highest mean. Korean consistency is led by Wren.
Your voice, interviews and high-volume production
Finch
For creators who want to keep their own voice when dubbing interviews or an entire channel. Finch had the highest mean speaker similarity in English and is designed for voice preservation and low-cost production.
Internal videos · EN speaker similarity: 0.564. Voice preservation and low cost are its focus; review English and Spanish accents before publishing.
Lectures, corporate narration and steady reading
Wren
For educators and businesses seeking steady narration. Wren suits lecture scripts, corporate introductions and read-aloud content, with the highest mean Korean voice consistency in the internal video set.
Internal videos · KO voice consistency: 0.713. EN predicted naturalness: 3.64/5, an automated score rather than a listening rating.
Dodo is the previous Expressive model and is included in the benchmark for context. Model availability and usage terms are shown in the product.
How we evaluate dubbing
Dubbing quality has several dimensions. We evaluate who the voice sounds like, how it follows the original expression, and how consistently it carries through a scene. We also examine translation, speech recognition error and timing.
01
Voice identity
Does the dubbed voice still resemble the original speaker?
Similarity to the original speaker
02
Expression
Does vocal intensity follow the performance in the source?
Arousal correlation · pitch analysis
03
Consistency
Does the generated voice stay coherent from line to line?
Consecutive-line voice similarity
Engineering benchmark · 22 September 2026
Explore the measurements.
Select a dataset, target language and metric to compare the models. Scores use the scale and evaluation conditions shown for each metric.
| Model | 00.250.50.751 | Mean |
|---|---|---|
| Wren | 0.341 | |
| Dodo | 0.353 | |
| Oriole | 0.323 | |
| Finch | 0.564 | |
| Nightingale | 0.414 |
Source: Perso engine-team benchmark, corrected 22 September 2026. Highlighted rows have the best mean for the selected metric; ties are highlighted together. This is not a claim of statistical significance or overall quality.
Finch has the highest mean on all four public read-speech targets and on internal English videos. On internal Korean videos, Nightingale scores 0.486 and Finch 0.483; the results are close.
Evaluation conditions
Public read-speech data: 13 language pairs, with Korean, English, Portuguese and Chinese sources and English, Spanish, Portuguese and Korean targets. Results are averaged by reader group. Internal videos: 29 dubbing pairs, 16 into English and 13 into Korean.
On FLEURS, Wren, Dodo, Oriole and Finch receive the transcript and timings and share a frozen translation. Nightingale runs speech recognition and translation itself. On FLEURS, Wren, Dodo, Oriole and Finch are scored from clean per-line clips; Nightingale output is voice-separated and aligned to source sentence windows. On internal videos, each model runs its own translation. Speech recognition error is evaluated on whole tracks for every model in that set.
Public read-speech data is licensed CC BY 4.0. Results cover selected languages and content, not every production scenario. No aggregate model ranking is calculated.
Card scores come from one internal video set. The explorer also includes public read-speech data. Listening evaluation provides separate evidence for accent quality.
13
FLEURS language pairs
Public read-speech data. Sources: Korean, English, Portuguese and Chinese. Targets: English, Spanish, Portuguese and Korean. 64 sentences per source language, grouped by reader. CC BY 4.0.
29
Internal video pairs
A separate set of customer-style videos, with 16 dubbing pairs into English and 13 into Korean. Evaluates content with tighter timing and more varied audio conditions.
What these results can tell you
The FLEURS evaluation supplies transcripts and timing to Wren, Dodo, Oriole and Finch, with one shared translation. Nightingale runs its own speech recognition and translation. These are pipeline results under documented conditions, not a controlled comparison of TTS alone.
Listening evaluation: in the September 2026 blind panel, Oriole had the lowest foreign-accent flag rate, 8%. This is a listening-panel finding, not an automated accent score. No single metric captures the whole listening experience.
Source: Perso engine-team benchmark, corrected 22 September 2026. Card scores: internal videos, 16 English and 13 Korean dubbing pairs. Small mean differences do not establish a reliable advantage. Check credit usage when using the product.
Does premium mean the highest score on every metric?
No. Nightingale is our premium engine offering. Different models lead different measurements. Voice identity, expression, translation and timing should be considered separately.
Why separate measurements from listening?
Automatic metrics examine specific signals. They do not fully capture flow, emotional nuance or whether a performance feels right for a scene. Listening remains essential.
Do the benchmark languages represent every supported language?
No. FLEURS results here cover four target languages; the internal video set covers English and Korean. Results should not be generalized to every language or content type.
Explore Perso Dubbing and choose the model that fits your next project.
Explore Perso Dubbing →