Appearance
Troubleshoot transcription
Transcription runs on your own computer, so when something goes wrong the causes are close to hand: what has been downloaded, your disk, your source audio, and the machine itself. Work down this table, then read the section that matches.
| What you see | Where to look |
|---|---|
| A download that never finishes | Downloads |
| Nothing happens when you transcribe | It will not start |
| No speech found | No speech |
| Words that are plainly wrong | Accuracy |
| Text that does not line up with the picture | Timing |
| The wrong people on the wrong lines | Speakers |
| A run that takes forever | Speed |
A download will not finish
- Check you are online, and check you have room for the file — the default model is 487 MB and Medium is 1.5 GB.
- Open Preferences → Transcription and start the download again from the model list.
- Watch the percentage. StoryFolder verifies the file before it will use it, so a download interrupted halfway is discarded rather than half-used — retrying is safe and starts clean.
Speaker detection and the file that lets StoryFolder skip silence download the same way, and the Transcription pane reports each with its own status and a Retry button.
If a download fails repeatedly, remove the model from the app's own list rather than deleting files by hand, then try once more. A corporate network that inspects traffic is worth ruling out — everything here comes from StoryFolder's own servers, and is verified before it is used.
Transcription does not start
- Is the import finished? Transcription is queued when the import reaches Complete, not before.
- Does the video have an audio track? A silent file has nothing to transcribe.
- Is transcription turned on? The Transcript tab says Transcription is turned off if the master switch in Preferences → Transcription is off, and offers to turn it back on.
- Is a model installed and selected? The first run downloads one; if that download never finished, start with the section above.
- Is the source file reachable? A video whose media has moved or is on a disconnected drive cannot be read. Relink it first.
- Is it available on this computer at all? Two messages mean there is nothing you can do — Transcription isn't available on this computer (it does not run on this kind of machine yet) and Transcription isn't ready in this build (it arrives in an update). Both say so plainly and offer no button, because there is nothing to press.
Write down the exact wording of any error before restarting the app.
StoryFolder reports no speech
No speech found in this video is a result, not a failure. Transcription ran, and what it heard was music, ambience or silence.
Check in this order:
- Is there speech in the video at all?
- Is it audible — not buried under music, not down at the noise floor?
- Is it on the channel you expect? A file whose dialogue sits on a track your player is not monitoring can sound empty.
If there is definitely speech, try naming the language instead of leaving detection to it, or a larger model. If it is a music video, leave it — re-running on footage with no dialogue will keep returning the same honest answer.
The transcript has many wrong words
- Choose the known language rather than Auto-detect.
- Move to a larger model, or to the English-only version if the audio is English. See Choose a model and language.
- Listen to the source. Clipping, heavy compression, wind, room echo and crosstalk cost more accuracy than any setting recovers.
- Correct names, numbers and specialist terms by hand — double-click the line, type, click away.
Corrections apply to this transcript only. StoryFolder does not learn a global vocabulary from your edits, so a name you fix in one project starts wrong in the next.
Turn on Highlight low confidence in the Display popover and read the underlined words first — it is the fastest way to spend correction time well.
Timing does not line up with playback
First, work out which kind of mismatch it is.
- Isolated — a line or two out of place, usually where speech overlaps or a long silence sits between lines. Correcting the text does not move timings; reassigning words does not either. Live with the small ones.
- Growing — the further into the video, the worse it gets. That is drift, and it usually means the recogniser ran across long non-dialogue stretches. Check Voice activity detection is on in Preferences → Transcription; it is on by default and it exists for exactly this.
- Wholesale — everything is offset. Check the video's media is still the same file that was transcribed. If it was relinked to a different edit, the words belong to the old one.
A re-transcribe is the fix of last resort
It replaces every line and timing, including your corrections. Read Re-transcribe a video before you start, and copy anything you cannot lose out first.
Systematic drift on ordinary dialogue is worth reporting — see what to send.
Speaker labels are wrong
Correcting labels by hand is almost always the smaller job. See Identify speakers for the three verbs:
- One person split across labels → merge.
- Two people under one label → add a speaker, then reassign their words.
- A misattributed passage → select it and assign it.
Overlapping speech and one-word turns are where automatic detection is weakest, and no re-run reliably improves them. Re-running detection also discards the automatic grouping you have been correcting — your named speakers survive a re-transcription, but the assignments made by detection are worked out again.
Processing is slow
The size of the model, the length of the video and the machine you are running on are what set the pace, in that order. Speaker identification adds another pass on top.
For a long file you mainly want to search, run Tiny or Base for a rough transcript, find the parts that matter, and re-run those videos at a larger size. Turning off speaker identification for the first pass saves time too.
There is no universal completion time worth publishing. Time a five-minute clip on your own machine and multiply.
Re-transcribing would replace manual work
What a re-run replaces: every line, every timing, and every correction you typed. StoryFolder counts your edited lines, tells you how many would go, and refuses the first attempt if there are any — you have to press Replace it anyway to proceed.
What it keeps: speakers you named or recoloured. Automatic speaker groups are worked out again from scratch.
Version history records transcript edits, including speaker changes, so a correction can be recovered. If the work is substantial, copy or export it before re-running rather than relying on that.
What to include with a support request
- StoryFolder version, operating system, and the processor family (Apple silicon, Intel, Windows x64).
- Free disk space.
- The video's duration, and whether it has an audio track.
- The model, the language, and whether speaker identification was on.
- The exact status or error, and the stage it stopped at.
- Your logs.
We do not need your footage, and we do not ask for it by default.