Skip to content
/a/ FaceCue Performance Studio

Language Support

How language coverage is layered, what a transcript buys you, and what to expect outside the ten tuned languages.

Check Your Language

Three systems here have three different coverage lists and they do not line up, so start with the answer for the language you actually care about. Hover a chip for what it means.

Every language gets lip sync. What the list holds is the languages with at least one specialised tool behind them, whether that is a tuned recognizer, a transcript writer or a voice.

So the first chip reports which of the three lip sync routes you land on rather than whether you get one, and a language that is not in the list at all is still covered, through the universal route. The rest of this page is what those routes are.

Three Layers, Not a Limit

FaceCue does not have a hard limit on the languages it supports. It has three layers, and every language sits on one of them.

All three produce a bake, and a bake plays the same way whichever route made it. What changes is how closely the mouth matches what was actually said.

Ten Languages, Custom Tuned

These ten each have their own recognizer, their own refinements and extensive testing behind them. They are the accurate layer.

English, Greek, Spanish, German, Mandarin Chinese, French, Japanese, Russian, Korean and Portuguese (Brazil).

Tuning here is not cosmetic. Languages differ in which sounds carry meaning, which ones a listener reads off the mouth, and how they run into each other, so the rules that protect a readable shape are not the same rules in Greek as in Japanese. That per-language work is most of what separates these ten from the layer below.

Fifty More, Supported

Outside the ten, a multilingual decoder covers over fifty further languages with no per-language setup. It works with a transcript or without one.

It is the same underlying model the accurate layer starts from, so it is meaningfully better than reading the audio alone. What it does not have is the per-language work: the ten get refinements written for them specifically, and these are covered by rules generalised from those ten. Expect results that are comparable but not as exact.

Any Language at All

Below both, a universal path reads the audio itself and produces mouth shapes directly, without going through a language's sounds at all. Nothing is out of reach.

It is a less linguistically accurate route by construction, and results vary widely by language. All of them work.

With and Without a Transcript

Whichever layer you are on, a bake is more accurate when it has the script.

Hand it the words and the recognizer only has to work out when each one was said. Without them it works out the words as well, and a word it gets wrong is a set of mouth shapes it gets wrong. Both are real options. The first is simply more exact, because it starts with more to go on.

If you have audio and no script, FaceCue can often write one for you and bake from that, which gets you the exact route without having to type anything. Not every language it can bake is one it can write for, which the middle chip above tells you. Where it cannot, decoding free is what happens, and it still produces a performance.

Setting either of them up is covered in Recognizers.

Generating Speech, Voice Changer and Voice Clone

FaceCue can also generate speech, convert a voice onto another, or clone one from a few reference clips. That side is covered for twenty-three languages. It can also write a transcript from a recording, and that side covers fifty.

These tools run independently of the baker and can work alongside it, so a line can be written, spoken and baked without leaving the editor. They have their own chapter: Speech Production Tools.