A voice channel is a dedicated track for narration. Unlike a plain audio track, a voice channel knows its content is speech: it can be transcribed, edited word-by-word, cleaned up, and re-synthesised in a different voice — all from the right-side Transcribe Editor.
Add a voice channel
Click + Voice in the timeline toolbar (“Add voice channel”). A new voice track (V2 voice, etc.) appears above your video track.
Select the voice channel (click its header) to open the Transcribe Editor on the right. With a channel selected — but no clip — the panel shows the ACTIONS block; selecting a clip on the channel instead shows that clip’s Sound and Voice change controls.
Get audio onto the channel
There are four ways to put narration on a voice channel:
- Record — click the red Record button. After a 3-2-1 countdown the timeline plays from the playhead (muted) while your microphone records, so you can narrate to picture. Click Stop to finish; the take is dropped where the playhead started.
- Text to Speech — toggle the Text to Speech button, type the line you want spoken, pick a TTS model on the row beside the reference-voice tile (optional — a ✓ line confirms the picked file), and click Generate voice take. A new voice take is synthesised at the playhead.
- Add an audio file — drop an existing voice recording onto the channel.
- Extract from a screen recording — select a transcribed screen recording clip and click Extract voice with transcript (in its AI CAPTURED DATA panel). The spoken audio lands on a voice channel as an already-transcribed take and the video is silenced.
Each clip on the channel is a take — labelled Take 1, Take 2, … in timeline order.
Each take is a card with its own header strip: a numbered chip, a waveform of the take’s audio, its length, and a Re-transcribe icon that appears when you hover the card. The take the playhead is currently over is marked by its chip and waveform turning green — the card itself stays neutral.
Click a take’s header to fold it down to that strip — handy for parking takes you have finished with while you work on another. Folding lasts for as long as you stay in the panel; every take is open again next time you open it.
Transcribe a take
Transcription is on-demand, not automatic. In the TRANSCRIPT section, each take shows Transcribe this take; click it to convert the audio to text with per-word timing. Once transcribed the button becomes Re-transcribe — use it if you change the audio or the first pass was inaccurate.
After transcription the take’s words become individually selectable, and the take header shows the audio range of your current selection (e.g. 1.02–3.34s).
Edit by selecting words
Click a word, or drag across a range of words, then use the icons next to the selection’s time range:
- ✂️ Cut — “Cut audio range (silence the voice)”. Silences the selected words in place, leaving a quiet gap where they were.
- ⧉ Split — “Split out the range (remove it across the whole timeline)”. Removes the selected range and everything else on the timeline over that span, closing the gap.
- ✦ Regenerate — “Regenerate selected”. Opens a small panel to re-synthesise just the selected words (see below).
This is the fast way to fix a stumbled word or drop a sentence: select it, then cut, split, or regenerate.
Regenerate selected words
The Regenerate panel is pre-filled with the selected words. Edit the text if you like, choose a model, and optionally attach a reference voice (optional) audio file:
- No reference voice — uses the model’s default voice.
- With a reference voice — the new audio mimics that voice’s tone.
Click Regenerate and the freshly synthesised audio is spliced in at the start of your selection (“Regenerated voice spliced in at the selection start.”).
Check grammar
Click Check grammar in ACTIONS to scan every transcribed take for grammar issues and disfluencies. Results are shown per take:
- A suggestions card lists the issues found (e.g. “Use ‘their’ instead of ‘there’ → their”).
- In the transcript, filler words are dimmed and grammar issues are highlighted — hover for the detail.
- When a corrected version of the take is available, a Regenerate clean text button appears. It re-synthesises the whole take from the corrected text while cloning the take’s own voice, so you get the fixed wording in the original speaker’s voice.
Dismiss a card with its ×, or clear all highlights with Clear grammar in the transcript header.
Make captions
Every transcribed take has a Captions button that turns the transcript into timed subtitle lines on the timeline. See Captions.
Clip controls — Sound and Voice change
Select a voice clip on the timeline (rather than the channel header) to get its clip panel. A Voice Editor button at the top jumps back to the transcript view; below it one panel holds the clip’s controls:
- SOUND → Volume — 0–300% with an inline slider. 0% mutes the clip, 100% is the original level, above 100% amplifies.
- VOICE CHANGE — transform the whole clip into a different voice with a speech-to-speech model (default Chatterbox Speech-to-Speech). Attach a reference clip on the tile beside the model picker — a confirmation line under the row shows the picked file’s name and length. Click Generate voice to preview, then Replace voice to commit the new voice in place of the original (or Discard to drop the result). The reference voice you pick is remembered across clips and sessions, so you can apply the same voice take after take without re-selecting it.
Leave a Reply