Captions come from the audio on your timeline, with word-level timing, which is what makes the karaoke highlight advance word by word.
Generating them
Open the Auto Captions tab in the left panel and press Generate captions. It uses the first audio clip on the timeline, or the first video clip if there is no audio clip, not whatever you have selected. This costs 10 AI credits.
The transcription itself runs in your browser, which is why the button downloads a speech model the first time, around 150 MB, and caches it afterwards.
The free route
The Transcript panel in Text Mode transcribes clip by clip with Transcribe this clip, and its own label says Runs on your device, no credits used. Once a clip has a transcript, the Text Mode Captions panel turns it into caption clips with Add captions from transcript. That panel cannot transcribe on its own, so if it tells you to transcribe a clip first, that is why.
From a script instead
If you already have the words, paste them into the From script box and press Time & add captions. It spreads the words across the clip and spends no credits. The timing is estimated rather than heard, so it drifts on a long take.
Language
The in-browser model is English only. There is no language picker in the editor.
If the timing drifts
Transcription describes the audio it was given. Cut the timeline after generating captions and they no longer match, so regenerate them. Adding a second set on top of an existing one is what makes captions appear to flicker: the new set replaces the old by default for exactly that reason.
Fixing a misheard word
Select the words in the Transcript panel and press Correct. That rewrites the text without touching the timing or the footage. Cut in the same toolbar is different: it removes that footage from the timeline and closes the gap.
Correcting the transcript changes the source clip, not caption clips already on the timeline, so re-run the captions afterwards to see it.
Styling
There are 14 presets, including Clean, Karaoke, Word Box, Bold Yellow, Neon Pop, Hormozi and Gradient Pop. Karaoke is the default. Clicking a style card restyles captions that already exist.
Six presets carry their own highlight colour for the word currently being spoken, and that colour wins over anything you set elsewhere, including a gradient fill or a per-word colour. If a colour you chose is not showing on the spoken word, that is the reason. Change Active Word in the Inspector's caption section, or pick a preset without a built-in highlight.
What is not there
Vidmoat cannot import or export SRT or VTT files. Captions are text clips burned into the render, and there is no subtitle file to download.
If it fails
Unable to extract audio means the file has no audio track or your browser cannot decode its audio format. Re-export the file as H.264 with AAC audio and upload it again. A very large file skips browser analysis entirely and tells you so: playback and editing still work, only transcription and waveforms are skipped.