AI powered search
Search every word and image
Search by spoken words and visual content. See results across every video and every shot in your library.

- People matched a field you typed
- Transcript matched a spoken line
- Results group under the source .mov
- Naming your own fields is Pro (/pricing)
Transcribe on your own machine
Transcription supports up to 100 languages and runs locally on your computer. Your footage is not uploaded to be transcribed.
Read nextAbout local transcription
- Audiofrom the video you imported
- Local Whisper modela process on the desktop
- Transcriptqueued once the import finishes
- Joined to the shot

The transcript sits under each shot
Every transcript line attaches to the shot whose time range it falls inside. By shot puts the thumbnails down the page with the spoken lines grouped beneath the shot they were said in, so the words stop being a separate document you keep in another window. It is the same join the search box uses, and the same join that carries into the CSV and the printed board. Four more reading modes are there when you want the text flat: By speaker, Continuous, Captions and Plain text.

I mean they're literally still out front talking to our neighbors.
- By shot
- By speaker
- Continuous
- Captions
- Plain text
AI Autofill
Search by visual content with AI Autofill
Intelligently extract the visual data that matters most for your workflow.
Customize the AI analysis to capture exactly what matters for your workflow, team, and studio.
More in choosing a field type and using your own categories.

Powerful search features
Search transcripts, visual contents, notes and custom metadata for results across your library. Search for exact terms, similar matches, and filter by tag or other metadata.
Type a word from a field or a spoken line.
The cops just searched the apartment looking for you.
I mean they're literally still out front talking to our neighbors.
That's why I came in the back.
You're such a dummy.
Everything.
I couldn't you just wait it.
You only had six months left.
You know why, Maya?
What's the look?
Where is she, Maya?
At your uncle's.
What's up?
I didn't tell you when I came visit because I knew you'll get like this and I know you hate your uncle but he's been all right.
She's supposed to be here.
It's a lot, Benji.
I couldn't do it anymore.
I asked you to do this one thing for me.
Don't make this work.
I need your car keys.
My car keys?
You just didn't choose me again.
I just
go.
Just stand in the woods, get chickens, maybe go.
Love that.
You know I can't.
I see her.
All
right.
We're fine.
We've got this covered.
on.
Look at you're doing just fine.
Look at you.
Car, car, I'm taking your car.
Oh, no, no, no, are you, that's my baby, my
baby.
Stay up, stay
up.
Right there, copies.
My nephew.
Stay with your copies.
Stay with your
copies.
Don't
move.
Go
tell.
Hello, my sweet boy.
I have missed you so much.
I bet I missed you.
Your hair looks so long.
Just something.
My favorite.
I should keep it for later.
You're a good boy for coming to see me.
Anything to be here with you, mama.
Automatically index your entire video archive
Every shot detected and indexed by visual metadata. Every word transcribed. All searchable from one place.
The average StoryFolder customer indexes hundreds of videos and tens of thousands of shots according to a 2026 survey.
- 273
- 43,861
Find the shot, then cut elsewhere
Reach for Descript or Reduct when changing the text is meant to change the cut.
Read nextDescript for b-roll
| It does | It does not |
|---|---|
| Transcribe locally on your computer | Transcribe in the cloud |
| Index the shots and the words you imported | Catalogue a drive it has never seen |
| Index by customizable AI vision analysis of every shot | Catalogue your origial footage |
| Find the specific videos and shots matching your search | Edit video by editing text |
| Underline the words the model was unsure about | Put a confidence percentage on them |
| Transcribe interviews and location sound | Hold accuracy through accents, overlap, noise or jargon |
Export with the shot number
Five formats come out of the transcript: SRT, WebVTT, plain text, Fountain, and a CSV whose header reads start,end,shot,speaker,text. That shot column is the join leaving the app, so the file you hand a translator or a producer already knows which picture each line belongs to. The per-shot transcript prints under each card in the storyboard PDF as well, in full. None of these exports is paywalled.
Transcribe one interview
Sign in and download for macOS or Windows, then import one interview you have already delivered. The transcript queues itself. Type a single word from it into the Library box and watch which shots come back, and why. The b-roll reel beside it only needs a field or two typed on the shots that matter.
FAQ
Does transcription upload my footage?
Is Library search fuzzy?
Can search find a word the transcript got wrong?
tick tock finds TikTok, and a phrase can match across a line break. Short words get no allowance, so cute never reaches cube.Can I search speaker names?
What happens on a silent clip?
No speech found in this video and offers Transcribe again. The shot stays findable through the fields typed against it.