The Rows¶
What each row of the Cue Clip Composer is telling you, and how to read the marks that appear on a cue. This page covers what the window shows you when it opens.
The window stacks the performance as rows, one stage per row, all sharing a timeline. Reading down the stack takes you from the sound at the top to the shapes on the face at the bottom.
Six rows are drawn when it opens, and a seventh is one toggle away. That is the working set, and it is the whole of this page. The Advanced and Expert toggles unfold a great many more, which is covered briefly at the end.
Audio Waveform¶
The sound itself, drawn across the clip.
It is the row you use to find your place. Everything else lines up against it, so a cue that looks wrong against the waveform usually is wrong, and a silence that has cues in it is worth a second look.
Words and Letters¶
What was said, and how it was broken up.
Words merges consecutive cues that belong to the same word into one span, so the row reads like the line of dialogue. Letters splits that down to the piece of spelling each cue came from, so night shows as n, i and ght sitting over the three cues that make it.
Together they answer the first question you have when something sounds wrong, which is whether the recognizer heard the right words at all. If the Words row says something the actor did not say, the problem is upstream of anything you can fix here, and the fix is a transcript. See Recognizers.
Visemes¶
The cues. This is the row you spend your time in.
Each block is one mouth shape, held for as long as the block is wide. The label is the shape's name, and the colour is that shape's colour, so a repeated sound is the same colour every time it comes round.
The row is a score, not the motion. A block says when a sound happens. It does not say how long the face spends on it, and the two are not the same thing. Shapes reach past their own block in both directions and run into each other, which is what the rounding mark is showing you when it trails back over the cues in front of a rounded vowel.
So blocks sit side by side because two sounds cannot happen at once, while the movement they produce overlaps almost all the time.
Gaps are ordinary. A pause has no cue over it, and neither does a stretch the recognizer could not place. What you will not find is two blocks on top of each other.
The Marks on a Cue¶
A cue carries more than its shape. Three things are drawn on the block itself, in three places that do not collide, so you can read all of them at once.
A stripe along the top edge is rounding. Lips round for some sounds and the mouth starts getting there early, so the strongest mark sits on the sound that is actually rounded and a fainter trail runs back over the cues it reaches.

A stripe along the bottom edge is secondary articulation, the smaller adjustments a sound picks up from what it is next to. It has a row of its own further down, and the stripe means the information survives that row being folded away.

A cross-hatch behind the label is reduction. Unstressed vowels in ordinary speech do not reach their full shape. They relax toward the schwa, the neutral vowel in the middle of the mouth that the a in about and the e in pencil both collapse into. The hatch marks a vowel that has gone that way, and it sits faintly behind the glyph so it never makes the label harder to read.

This Is the Difference Between Speech and Recitation
All three marks are FaceCue doing what a mouth does, not what a spelling says. Rounding arrives early because lips move ahead of the sound, reduction happens because unstressed vowels relax, and secondary articulation is a sound taking colour from its neighbours.
They are worth knowing about mostly so you recognise them when you see them, and so a cue that looks under-shaped reads as deliberate and not as a bug.
Reading a Cue¶
The marks tell you at a glance. Hovering a cue tells you in words, which is quicker than opening it when you only want to look. Hold Shift to keep it up while you move around, and let go to retract it.

This is where the three marks turn into words. Rounding and Reduction name what the stripe and the hatch are doing, so a shape you were not expecting can be traced back to the reason it is there, and the rest places the cue in its word and its syllable.
Cognitive Events¶
Moments where the character is thinking, marked as points on the timeline with an optional duration. They come from tags in the transcript, and you can place them yourself.
What a marked moment does today is in the eyes. Blinking slows, the way it does in someone concentrating, and the pupils widen and settle back as the moment passes. Both ease in and out over the event rather than switching at its edges, and both are tunable on the character's Eyes driver, including down to no response at all.
This row is the seventh, and it starts hidden. It sits behind Show ▸ Spans & Events along with the language spans row, because most clips carry neither and the two would otherwise sit empty above the output. Turn that category on when you are working with a clip that has them.
More Is Planned From These
A marked thinking moment is a general signal about the character's inner state, and the eyes are the first thing reading it. Others will follow, and looking away while thinking is the priority: it is the most recognisable thing a person does when they are searching for a word.
Events you author now carry forward, so a clip marked up today gets the additions when they land. See the roadmap.
Emotion Cues¶
The authored emotion timeline. Each cue draws as a trapezoid with five handles: where it starts, how fast it comes up, how long it holds, how fast it goes away, and where it ends.

Nothing here is inferred. Emotion is authored, whether it came from a tag in the transcript, from this row, from a Timeline track or from your own code. FaceCue does not guess how a character feels from the sound of their voice.
The row header carries Emotions and Expressions, which switch between the emotion as you authored it and the expressions it resolves into on this character. With both on, the authored shape and the expression shapes are drawn together, dotted and dashed, so you can see one becoming the other.

Hovering one shows the whole of it: its strength, how long it takes to rise, hold and fall, and the expressions it becomes with the weight each one carries. Those expressions run on their own timings inside the emotion, staggered in proportion, which is why an emotion arrives as a face doing several things slightly out of step and not as a single pose fading up.
This is the layer you would tune if an emotion reads wrongly on your character. The vocabulary of emotion names is fixed, and what each one looks like is yours. See Emotion.
Final Output¶
How much of each mouth shape lands on the face, drawn as stacked bands from nothing up to full.
This is the end of the chain and the row to read when you want to know what the audience will actually see. Everything above it is how FaceCue arrived here.
Unlike the rows above it, this one is computed instead of read from a file, so it needs a sweep behind it. Auto takes care of that, which is why the row is usually just there.
It is worth comparing against the Visemes row when something looks wrong. If a shape is in the cues but flat in the output, the cue is fine and something further down is holding it back. If it is missing from both, the problem is the bake.
What Advanced and Expert Add¶
Advanced and Expert unfold the stages in between, which is most of the window. They are worth reaching for when you want to know why a line looks the way it does, and worth folding away the rest of the time, because the overview is the thing you lose.
Those rows are documented where they live, in the window itself. Every row in the window carries a question mark on its right-hand side that explains that row, which is a better place for it than a page you would have to keep in step with the window.
Unfolding a Row Does Not Compute It
Like Final Output, these are drawn from a sweep instead of read from a file, so unfolding them shows nothing until one has run. Auto keeps that up to date as you change things.