AI powered search

Search every word and image

Search by spoken words and visual content. See results across every video and every shot in your library.

A search box reading maya above two results, a shot tagged People Maya and a transcript line reading Where is she, Maya.
  • People matched a field you typed
  • Transcript matched a spoken line
  • Results group under the source .mov
  • Naming your own fields is Pro (/pricing)

Transcribe on your own machine

Transcription supports up to 100 languages and runs locally on your computer. Your footage is not uploaded to be transcribed.

The transcript sits under each shot

Every transcript line attaches to the shot whose time range it falls inside. By shot puts the thumbnails down the page with the spoken lines grouped beneath the shot they were said in, so the words stop being a separate document you keep in another window. It is the same join the search box uses, and the same join that carries into the CSV and the printed board. Four more reading modes are there when you want the text flat: By speaker, Continuous, Captions and Plain text.

The frame the line was spoken over

Shot 24500:13:36:05

I mean they're literally still out front talking to our neighbors.

Speaker 2Press play · 3.5s of the master's own audio
  • By shot
  • By speaker
  • Continuous
  • Captions
  • Plain text

AI Autofill

Search by visual content with AI Autofill

Intelligently extract the visual data that matters most for your workflow.

Customize the AI analysis to capture exactly what matters for your workflow, team, and studio.

More in choosing a field type and using your own categories.

Six cards, one per metadata field type, each labelled with the kind of value autofill returns for it.

Automatically index your entire video archive

Every shot detected and indexed by visual metadata. Every word transcribed. All searchable from one place.

The average StoryFolder customer indexes hundreds of videos and tens of thousands of shots according to a 2026 survey.

273
average number of videos
43,861
average number of shots

Find the shot, then cut elsewhere

Reach for Descript or Reduct when changing the text is meant to change the cut.

It does It does not
Transcribe locally on your computerTranscribe in the cloud
Index the shots and the words you importedCatalogue a drive it has never seen
Index by customizable AI vision analysis of every shotCatalogue your origial footage
Find the specific videos and shots matching your searchEdit video by editing text
Underline the words the model was unsure aboutPut a confidence percentage on them
Transcribe interviews and location soundHold accuracy through accents, overlap, noise or jargon

Export with the shot number

Five formats come out of the transcript: SRT, WebVTT, plain text, Fountain, and a CSV whose header reads start,end,shot,speaker,text. That shot column is the join leaving the app, so the file you hand a translator or a producer already knows which picture each line belongs to. The per-shot transcript prints under each card in the storyboard PDF as well, in full. None of these exports is paywalled.

  • Subtitles.srt
  • Subtitles.vtt
  • Transcript.txt
  • Transcript.csv
  • Script.fountain

Transcribe one interview

Sign in and download for macOS or Windows, then import one interview you have already delivered. The transcript queues itself. Type a single word from it into the Library box and watch which shots come back, and why. The b-roll reel beside it only needs a field or two typed on the shots that matter.

Sign in and download

macOS and Windows

FAQ

Does transcription upload my footage?
No. Whisper runs on the desktop as a local process, and your footage is not uploaded to be transcribed.
Is Library search fuzzy?
No. The Library and board box is substring matching across titles, your Text, Dropdown, Tag and Number fields and the transcript. Fuzzy matching lives in the transcript editor's own box.
Can search find a word the transcript got wrong?
In the transcript editor, yes, within bounds: tick tock finds TikTok, and a phrase can match across a line break. Short words get no allowance, so cute never reaches cube.
Can I search speaker names?
Not from the transcript editor's search box, which reads spoken words only. Use the By speaker view to read a named speaker's lines.
What happens on a silent clip?
Transcription reports No speech found in this video and offers Transcribe again. The shot stays findable through the fields typed against it.
Which languages can it transcribe?
The catalogue lists 100 languages plus auto-detect, and a separate task transcribes into English. The model that ships as the default is the multilingual one.
Which transcript exports are free?
All five: SRT, WebVTT, plain text, CSV and Fountain. The per-shot transcript in the storyboard PDF is free too.
Does it edit video from text?
No. Deleting a word in the transcript changes the transcript. For text-based video editing, use Descript or Reduct.