Skip to content

Choose a transcription model and language

StoryFolder ships with a sensible default and most people never change it. Move off it when you have a reason: you need results faster, you are short of disk space, you work in a language other than English, or a transcript came back weaker than the audio deserved.

Every model listed here runs on your computer. A bigger model is more work for your machine, never more data leaving it.

The quick choice

If you wantChooseDownload
The fastest rough pass, or the smallest downloadTiny77 MB
The default balanceSmall (multilingual)487 MB
The best recognition available hereMedium1.5 GB

Small (multilingual) is the default, and it is multilingual on purpose. The English-only and multilingual builds of a size are the same download, so defaulting to English-only would buy nothing and would cost anyone working in another language a second full-size download. small is also where the multilingual builds become genuinely usable across languages — the smaller ones are noticeably weaker than their English-only twins.

If everything you cut is in English, switch to the English-only version of the same size and keep the download you already have.

How model size changes the result

Larger models generally recognise more of what was said, especially where the audio is difficult. They cost you three things: download size, disk space, and processing time.

How long a transcription takes depends on the model, the length of the video and the machine you are running it on, so there is no useful number to publish here. Run one on a five-minute clip and you will know what your own machine does.

What a larger model will not fix: speech that is inaudible, clipped, or spoken over by someone else. If two people are talking at once, no model size resolves them into two clean lines.

English-only or multilingual

  • English-only — the (English) ones — for English speech. They are the stronger choice at every size, and markedly so at the smaller ones.
  • Multilingual for any other language, for material that switches language, and for automatic language detection.

The two versions of a size are the same download size. That does not make them interchangeable: an English-only model cannot transcribe French, and it refuses Translate to English outright, because translation needs a multilingual model.

Choose a language

The language setting lives in Preferences → Transcription → Language.

  • Auto-detect is the default and is right for most footage — a recording in one language it can hear clearly.
  • Pick the language when you know it. It is worth doing when the audio is quiet, when music opens the video, or when a mis-detection has already happened.
  • Mixed-language material is the hard case. Detection settles on one language for the run, so a video that switches halfway will be transcribed as though it did not.

The setting is a preference for every new run, and it can be overridden for a single run in the Re-transcribe dialog — along with Translate to English, which transcribes non-English audio straight into English.

Download and manage models

  1. Open Preferences (Cmd+P on Mac, Ctrl+P on Windows).
  2. Go to the Transcription pane.
  3. Open the Model list. Each model shows its size and whether it is installed.
  4. Download the one you want and wait for it to verify.
  5. Selecting it makes it the model every new transcription uses.

A download shows a live percentage. A model counts as installed only when the whole file has arrived intact — StoryFolder verifies it and will not run one it could not, so an interrupted download is discarded rather than half-used. Retry from the same list.

Download the model at startup is on by default, so your first import already has something to run against. Turn it off on a metered connection and the download waits until the first time you actually transcribe.

Deleting a model asks first and names the size you would have to download again. The model currently selected as your default cannot be deleted out from under itself.

Storage and offline planning

ModelLanguageDownload
Tiny (English) / Tiny (multilingual)English / 100 languages77 MB
Base (English) / Base (multilingual)English / 100 languages147 MB
Small (English) / Small (multilingual)English / 100 languages487 MB
Medium (English) / Medium (multilingual)English / 100 languages1.5 GB

There is no large model. Eight in all: four sizes, each in an English-only and a multilingual version.

Two more downloads are not models you choose, but they take disk:

Also downloadedWhat it doesDownload
Voice activity detectionSkips silence and music, which keeps timings from drifting~1 MB
Speaker detectionGroups voices into speakers, if you use it~60 MB

Both arrive on demand, the same way a model does.

Preferences → Transcription → Disk itemises models, speech recognition and speaker detection, with a total. Offload All Models deletes every downloaded model and reclaims the space; transcription stays on, and your selected model is downloaded again the next time something needs transcribing. No transcript you already have is affected — a transcript is text stored with the project, not something the model file holds.

Travelling, or working somewhere with restricted network? Download the model — and the speaker files, if you identify speakers — before you leave. After that a transcription runs with the network off.

When to try another model

  • The same words come back wrong in clearly recorded speech.
  • Proper names or specialist vocabulary are consistently mangled.
  • Heavy accents or noisy location sound.
  • You need a rough pass over hours of footage quickly — drop to Tiny or Base, then re-run the parts that matter at a larger size.

Re-transcribing replaces the whole transcript

Changing the model only affects new runs. To apply it to a video you have already transcribed you have to re-transcribe it, which replaces every line and every timing — including your hand corrections. Read Re-transcribe a video first.

Next step

Identify speakers