Perso Dubbing / Models

Meet our new premium dubbing engine.

Meet our new premium dubbing engine.

Nightingale

Nightingale

Nightingale

Nightingale is a premium AI dubbing engine for content where emotion matters. Compare model strengths and real-world video evaluations to choose the right dubbing model for your content.

Nightingale is a premium AI dubbing engine for content where emotion matters. Compare model strengths and real-world video evaluations to choose the right dubbing model for your content.

Find the right model
for your content.

Find the right model
for your content.

Start with what you make: emotional performances, local-audience content, your own voice or steady narration. The card scores below all use our internal video set.

Drama, entertainment and emotional performances

Nightingale

For creators and studios seeking premium delivery through emotional highs and lows. Nightingale ranks first in mean source-expression tracking on both tested languages.

Internal videos · Expression tracking: EN 0.45, KO 0.54. These scores measure vocal intensity, not overall quality.

Lectures, corporate videos and local audiences

Oriole

An accent that sounds natural to local audiences, based on listening evaluation. Choose Oriole for lectures and corporate videos that need a consistent voice from one sentence to the next.

Internal videos · EN voice consistency: 0.661, the highest mean. Korean consistency is led by Wren.

Your voice, interviews and high-volume production

Finch

For creators who want to keep their own voice when dubbing interviews or an entire channel. Finch had the highest mean speaker similarity in English and is designed for voice preservation and low-cost production.

Internal videos · EN speaker similarity: 0.564. Voice preservation and low cost are its focus; review English and Spanish accents before publishing.

Lectures, corporate narration and steady reading

Wren

For educators and businesses seeking steady narration. Wren suits lecture scripts, corporate introductions and read-aloud content, with the highest mean Korean voice consistency in the internal video set.

Internal videos · KO voice consistency: 0.713. EN predicted naturalness: 3.64/5, an automated score rather than a listening rating.

Dodo is the previous Expressive model and is included in the benchmark for context. Model availability and usage terms are shown in the product.

How we evaluate dubbing

Voice similarity.
Expression. Consistency.

Voice similarity.
Expression. Consistency.

Dubbing quality has several dimensions. We evaluate who the voice sounds like, how it follows the original expression, and how consistently it carries through a scene. We also examine translation, speech recognition error and timing.

01

Voice identity

Does the dubbed voice still resemble the original speaker?

Similarity to the original speaker

02

Expression

Does vocal intensity follow the performance in the source?

Arousal correlation · pitch analysis

03

Consistency

Does the generated voice stay coherent from line to line?

Consecutive-line voice similarity

Engineering benchmark · 22 September 2026

Explore the measurements.

Select a dataset, target language and metric to compare the models. Scores use the scale and evaluation conditions shown for each metric.

Model
00.250.50.751
Mean
Wren0.341
Dodo0.353
Oriole0.323
Finch0.564
Nightingale0.414

Source: Perso engine-team benchmark, corrected 22 September 2026. Highlighted rows have the best mean for the selected metric; ties are highlighted together. This is not a claim of statistical significance or overall quality.

Reading the result

Finch has the highest mean on all four public read-speech targets and on internal English videos. On internal Korean videos, Nightingale scores 0.486 and Finch 0.483; the results are close.

Evaluation conditions

Public read-speech data: 13 language pairs, with Korean, English, Portuguese and Chinese sources and English, Spanish, Portuguese and Korean targets. Results are averaged by reader group. Internal videos: 29 dubbing pairs, 16 into English and 13 into Korean.

On FLEURS, Wren, Dodo, Oriole and Finch receive the transcript and timings and share a frozen translation. Nightingale runs speech recognition and translation itself. On FLEURS, Wren, Dodo, Oriole and Finch are scored from clean per-line clips; Nightingale output is voice-separated and aligned to source sentence windows. On internal videos, each model runs its own translation. Speech recognition error is evaluated on whole tracks for every model in that set.

Public read-speech data is licensed CC BY 4.0. Results cover selected languages and content, not every production scenario. No aggregate model ranking is calculated.

Test data and evaluation methods

Test data and evaluation methods

Card scores come from one internal video set. The explorer also includes public read-speech data. Listening evaluation provides separate evidence for accent quality.

13

FLEURS language pairs

Public read-speech data. Sources: Korean, English, Portuguese and Chinese. Targets: English, Spanish, Portuguese and Korean. 64 sentences per source language, grouped by reader. CC BY 4.0.

29

Internal video pairs

A separate set of customer-style videos, with 16 dubbing pairs into English and 13 into Korean. Evaluates content with tighter timing and more varied audio conditions.

What these results can tell you

The FLEURS evaluation supplies transcripts and timing to Wren, Dodo, Oriole and Finch, with one shared translation. Nightingale runs its own speech recognition and translation. These are pipeline results under documented conditions, not a controlled comparison of TTS alone.

Listening evaluation: in the September 2026 blind panel, Oriole had the lowest foreign-accent flag rate, 8%. This is a listening-panel finding, not an automated accent score. No single metric captures the whole listening experience.

Source: Perso engine-team benchmark, corrected 22 September 2026. Card scores: internal videos, 16 English and 13 Korean dubbing pairs. Small mean differences do not establish a reliable advantage. Check credit usage when using the product.

Questions about dubbing quality

Questions about dubbing quality

Does premium mean the highest score on every metric?

No. Nightingale is our premium engine offering. Different models lead different measurements. Voice identity, expression, translation and timing should be considered separately.

Why separate measurements from listening?

Automatic metrics examine specific signals. They do not fully capture flow, emotional nuance or whether a performance feels right for a scene. Listening remains essential.

Do the benchmark languages represent every supported language?

No. FLEURS results here cover four target languages; the internal video set covers English and Korean. Results should not be generalized to every language or content type.

Choose a model.
Start dubbing.

Choose a model.
Start dubbing.

Explore Perso Dubbing and choose the model that fits your next project.

Explore Perso Dubbing →