Appearance
Transcribe a video on your device
Transcription turns what people say into timed text that sits beside your storyboard. Every line is joined to the shot it was said in, so you can read the dialogue, search it, correct it, and take it out as subtitles or a script.
Speech recognition runs on your computer. The audio is read from the file already on your drive, and neither it nor the transcript is uploaded.
Free on every plan. StoryFolder downloads what it needs to recognise speech once, the first time. After that a transcription runs with the network off.
This is the local half of StoryFolder
Transcription is on-device. AI Autofill is not — it sends a frame per shot to a hosted model. Do not read one as a promise about the other; Local vs cloud draws the line.
Before you begin
- The video has finished importing and has an audio track.
- The first run downloads a model — 487 MB for the default — and needs the disk space for it. See Choose a model and language.
- Working offline soon? Download the model — and the speaker files, if you identify speakers — before you go.
- Transcription is published for macOS on Apple silicon, macOS on Intel, and Windows x64. On anything else the Transcript tab says so rather than offering a button that cannot work.
Transcribe a video
Most people never do this by hand. Automatic transcription is on out of the box, so a video is queued as soon as its import reaches Complete — a separate job from the import, so a transcription that fails can never fail your import.
To run one yourself:
- Open the video and go to the Transcript tab in the workflow bar, or press R.
- Choose Transcribe.
- Watch the phase line while it works.
The language, the model, and whether speakers are identified come from Preferences → Transcription. It is worth aiming those before a long run — the empty state links straight to them.
What happens while it runs
The progress line names the stage it is in:
| Stage | What it is doing |
|---|---|
| Queued | Waiting its turn |
| Downloading speech model | First run only, once per model |
| Preparing audio | Extracting audio from your file |
| Loading speech model | Reading the model into memory — larger models take longer |
| Transcribing | Recognising speech, producing timed words and lines |
| Identifying speakers | Grouping voices, if you asked for it |
Voice activity detection, on by default, keeps the recogniser off stretches with no speech in them. It is what stops timings drifting across a long musical passage or a silent cold open.
A video with music and no dialogue comes back as No speech found. That is a result about your audio, not a failed run.
Read and navigate the transcript
Click any word to move the player to it. Turn on Follow playback in the Display popover and the transcript scrolls itself as the video plays.
There are five view modes in the toolbar:
- By shot — the transcript laid against the shot grid. The one that matters for storyboard work.
- By speaker — grouped for screenplay-style reading.
- Continuous — paragraphs, for proofreading.
- Captions — one row per caption, for subtitle QC.
- Plain text — for drafting and pasting.
The Display popover also carries Speaker names, Timestamps and Highlight low confidence — a dotted underline under the words the model was least sure of. There is no percentage anywhere, because the underlying confidence is not calibrated well enough to print as one.
Correct a line
- Double-click the line.
- Type the correction. Enter adds a line break rather than saving.
- Click away to save. Esc abandons the edit instead.
Corrected lines are marked as edited, and the corrected text is what gets searched, copied and exported from then on. Word timings stay where they were, so playback still lands in the right place.
A correction improves this transcript. It does not teach StoryFolder a word for next time — there is no global vocabulary being learned from your edits.
Re-transcribe a video
Re-running is the answer when the language was wrong, the model was too small, or you want speakers identified on something transcribed without them.
Open Display → Re-transcribe… in the transcript toolbar.
Re-transcribing replaces the whole transcript
Every line and every timing is produced again from scratch, including lines you corrected by hand. StoryFolder counts your edited lines and tells you how many would be discarded — and if there are any, the first attempt is refused. The button then becomes Replace it anyway, so a re-run over your own work can only happen after you have read why.
Speakers you named or recoloured are carried over. Speakers that came from automatic detection are worked out again from scratch.
The confirmation lets you override the language and the task (transcribe or translate to English) for this run only, without touching your saved preference. Advanced settings adds the model, speaker detection and voice activity detection for the same single run. Anything left at its default changes nothing.
What a transcript makes possible
- Search inside it — with the search box in the transcript toolbar.
- Find spoken words from the library — library search reads transcript text alongside your metadata.
- Copy what is on screen — in whichever view mode and display settings you are using.
- Export subtitles, a script, a CSV or a Fountain file. See Search, copy, and export a transcript.
- Read the numbers — the Statistics tab carries lines, words, and each speaker's share.
- Ask a connected assistant about what was said. See Connect AI assistants.
Accuracy, and what to check
Speech recognition degrades on heavy accents, people talking over each other, noisy location sound, and specialist vocabulary. Product names and surnames are where you notice first.
Check the things a wrong word would cost you most: names, numbers, quotes you intend to publish, and the time ranges you are cutting to. On a long project, check the ends rather than the middle — the first and last minutes, the longest shot, one shot with music and no dialogue, and one where two people overlap.
Turn on Highlight low confidence and read the underlined words first.
Automatically detected speakers are groups of similar voice, not identified people. See Identify speakers.
Privacy, storage and offline use
Audio stays on your computer while it is transcribed, and so does the transcript StoryFolder writes. No third-party transcription account is involved.
The network is used once, at the start, to fetch what StoryFolder needs to recognise speech from its own servers. Every download is verified before it is used, so a file that arrived damaged is never run.
Preferences → Transcription shows what is taking up disk space — models, speech recognition, speaker detection — and offers Offload All Models to reclaim it. Offloading deletes downloaded models. It does not turn transcription off, and it does not touch a single transcript you already have: the model you have selected is downloaded again the next time something needs transcribing.
Related pages
- Choose a transcription model and language
- Identify speakers
- Search, copy, and export a transcript
- Troubleshoot transcription
- Library search · Local vs cloud