Skip to content

Transcription models

StoryFolder downloads a Whisper model the first time it needs one, then runs everything against that file locally. You choose which model in Preferences → Transcription → Model.

Which model does StoryFolder use by default?

Small (multilingual), shown in the app as 487 MB.

Two reasons it isn't the English-only build of the same size. The English-only and multilingual builds of a size are the same download, so defaulting to English would buy nothing and would cost anyone working in another language a second full-size download. And small is the size where the multilingual builds become usable across languages — the smaller multilingual models are noticeably weaker than their English-only twins.

If everything you cut is in English, switch to the English-only twin of the same size.

What models are available?

Eight: four sizes, each in an English-only and a multilingual build.

ModelSize shown in the app
Tiny (English) / Tiny (multilingual)77 MB
Base (English) / Base (multilingual)147 MB
Small (English) / Small (multilingual)487 MB
Medium (English) / Medium (multilingual)1.5 GB

Bigger models are more accurate; smaller ones are faster and take less disk. There's no large model.

How do I download, verify or remove a model?

Each model in the list has its own controls.

  • Download shows a live percentage. A model counts as installed only once the whole file has arrived intact.
  • Verify re-hashes the file on disk on demand.
  • Delete removes it. StoryFolder asks first and tells you the size you'd have to download again.

Model files come from StoryFolder's own servers and are checked against a known hash before they're used.

How much disk is transcription using?

Open Preferences → Storage. It itemises the transcription models, the recognition engine, the speaker-detection assets and the voice-activity model, and offers to free the space.

Can I use a different model for one video?

Yes. Open Display → Re-transcribe… in the Transcript tab and expand Advanced settings. Only models you've already downloaded are offered there, so a "just re-transcribe" click can never start a multi-gigabyte download by surprise.

Next step

Speakers