StoryFolder 0.4: AI Autofill, local transcription, and an assistant that can see your library

StoryFolder —

Last updated: 2 September 2026.

StoryFolder 0.4 is out for macOS and Windows. 0.4 adds AI Autofill, which fills the shot metadata fields you switch it on for; text transcription, and MCP (Model Context Protocol), which lets Claude Desktop, Claude Code or ChatGPT Codex read and edit your library through AI chat. New library features include folders, collections, quick filters and video-level metadata fields. 0.4 also introduces version history records the edits you make by hand, the ones an assistant makes, and the ones a paste makes, and lets you preview and roll each one back.

What changed in StoryFolder 0.4?

The StoryFolder shot inspector during an AI Autofill run: Shot Size, Interior/Exterior and Composition already filled on shot 245, Lighting filling, Motion and Content Type queued

AI Autofill looks at a representative frame per shot and automatically proposes a value for each shot metadata field you switched it on for. Every field ships with that switch off, and the Autofill action does not appear at all until at least one field has it on. Dropdown, Rating, Checkbox and Number answers are held to the domain you defined for that field. Tag and Text are deliberately open, so those two can come back with something you never listed. A run opens a modal that names the fields, the shots and the write mode and counts them before you commit; Only fill empty fields is the default and never touches a value someone typed. It is a Pro feature, enforced on the server, and it is a cloud call. Alongside it, AI Discover fields takes a description of the work you do and proposes fields for it: it suggests the schema and never fills a value, and anything it creates arrives with its Autofill switch off like every other field. Read how AI Autofill works and what it will not do.

A Run Rabbit frame with the audio waveform of the line spoken in it, the words spoken so far highlighted

Transcription runs whisper.cpp as a process on your own computer, on audio pulled from the file already on your drive. It is on by default: when an import reaches Complete, StoryFolder queues the transcription itself. The default model is small, the multilingual one, and there are eight to pick from across four sizes in English-only and multilingual pairs. Identify speakers is a separate switch, off by default and slower when on, and the speakers it finds can be renamed, merged and reassigned down to individual words. The catalogue holds 100 languages plus auto-detect, and a translate-to-English task sits beside plain transcription. You read the result in whichever of five modes suits the pass: By shot, By speaker, Continuous, Captions or Plain text. By shot is the one that changes how the app works, because from then on the words are part of the same index as the pictures: library search returns a shot on a transcript hit and says so, the assistant's search tool reaches the same text, and Statistics reports lines, words and speaker share. Transcription and its transcript and subtitle exports are free. Read what local transcription covers.

A chat with an assistant connected to StoryFolder over MCP: the question, the answer naming shots 239, 242, 243 and 249, and the four frames

Connect an AI assistant puts a local connector between StoryFolder and a desktop AI chat client. The protocol is MCP, the Model Context Protocol and includes sixteen tools that cover listing and reading projects, searching frames by metadata and transcript text, fetching frames as actual images, writing shot and video fields, creating fields, importing a video, queueing an export and following the job, and listing version history. Claude Desktop, Claude Code and ChatGPT Codex each have their own setup path in the app. A read-only mode hides every tool that writes. Read what an assistant can and cannot do with your library.

A folder of finished videos, a collection of shots from several videos, and a quick filter showing closeups from three films

The library gains collections, quick filters and video-level metadata fields. Folders are not new and are not claimed as new here; they file videos, and opening one narrows what you see to that folder and everything under it. A collection holds hand-picked shots from across different videos, and an open collection carries its own search box and filter bar. A quick filter saves the query, the filter chips and the card layout together as a destination you can jump back to. Video-level fields are new too: one flat list of fields whose values sit on the video rather than on a shot, which is where a client name or a job number belongs and where filtering by it filters whole projects. Autofill does not touch them. Searching or filtering from All items covers the whole library, filed videos included. One box searches titles, the fields you type into on a shot or on the video, and transcript text at once; a multi-word query needs every word to match, in any order; and each result says which layer it matched on. / or ⌘F, Ctrl+F on Windows, puts the cursor in that box. Read how folders, collections and quick filters fit together.

StoryFolder's Version history window: restore points on the right, the board as it was at the selected point on the left with the changed shots marked

Version history records one restore point per edit gesture, not one per shot, so a change written across forty shots is one entry and one undo. It covers the shot fields you type, video-level metadata, transcript corrections, titles and descriptions, folder moves, and the re-cut controls (sensitivity, minimum shot length, merges and forced cuts). ⌘Z, Ctrl+Z on Windows, takes back the last one the moment you regret it. Points minted by an assistant carry a badge saying so. Previewing a point never writes anything, and reverting captures the current state first, so the revert is itself reversible. Copy notes & metadata and Paste notes & metadata move values from one shot to a scope of others through a plan you confirm: which fields, which shots, and whether existing values are left alone. Read what version history actually captures.

Statistics counts your own fields. It buckets checkbox, dropdown, tag, rating and number values into a distribution per field, with an absence bucket for the shots that have nothing, and reports shot-duration mean, median, longest and shortest plus a transcript block. The tab opens on any plan; the charts are Pro.

The rest of the surface moved with it. Export is rebuilt around five outputs with a queue you can walk away from, and an export or a publish carries whatever search and filters you have active rather than the whole board. Import accepts a drop anywhere in the window, an "open with" from the operating system, and a URL. The account pane has an in-app plan picker, a one-time sign-in link instead of a key in a URL, self-serve account deletion and an analytics opt-out. There is an in-app What's New and a set of product tours.

Which AI work stays on your machine, and which does not?

Transcription and speaker identification run on your machine. AI Autofill does not: it encodes a frame from each shot and posts it to StoryFolder's server, which calls a hosted vision model. The video file itself is never uploaded for either.

Feature What leaves your machine Needs internet Plan
Transcription and Identify speakers Nothing from the audio Once, to fetch the engine and a model Free
AI Autofill One frame per shot, plus the video's title, description and video-level fields, plus your field names and instructions Every run Pro
Connect an AI assistant Whatever the assistant asks for and then sends to its own provider The connector itself runs on loopback Free

The model download is the part people get wrong in both directions. Transcription needs the network once, to fetch the whisper.cpp engine build for your platform and a model file. After that a transcription runs with the network off. That is a fact about where the audio went, not a certification, and StoryFolder makes no compliance claim.

Connect an AI assistant has two boundaries, and only one of them is local. The connector talks to the running app over loopback, so nothing about that hop leaves the machine. What the assistant then does with a shot note or a frame it asked for is between you and whoever makes your assistant, exactly as if you had pasted it into a chat box yourself. Read-only mode stops writing, not reading.

Two more things talk to StoryFolder's servers and always did: activating a licence, which is where your live plan comes back from, and usage telemetry, which carries counts and identifiers and never a title, a filename or a path. Telemetry has an opt-out in the account pane. Publishing a storyboard is a third, and an obvious one, since publishing is the act of putting it somewhere other people can open.

The privacy help page is the current authority on all of this.

What is free, what is Pro, and what needs an internet connection?

Transcription is free. AI Autofill is Pro. Custom metadata fields are Pro, and since Autofill fills custom fields, Pro is the practical floor for that whole workflow.

Free Pro
Transcription and Identify speakers AI Autofill
Connect an AI assistant, including writes Custom shot and video metadata fields
Collections and quick filters Spreadsheet / Shot List export
Version history, copy and paste metadata Statistics charts
Storyboard PDF, Images, Clips, Transcript and Subtitles export Password-protected share links

The export gate is narrower than people assume. Of the five outputs, only Spreadsheet / Shot List is gated. Storyboard, Images, Clips, and Transcript & Subtitles (SRT, VTT, TXT, CSV, Fountain) render on any plan. The free tier caps a board at 12 shots and a library at 3 boards.

Tiers and what each includes are on the pricing page.

What should you try first?

Pick the one that matches what is already on your machine.

  1. You have a library already. Open All items, type something you know is in a filed project, and confirm it comes back. Then build a filter you would rerun and choose Save as quick filter from the chip bar. The library page covers the rest.
  2. You have footage with people talking in it. Import one video and leave it alone. Transcription is on by default and queues itself once the import completes. Read it in By shot mode, then turn on Identify speakers in Settings → Transcription and re-transcribe if you need names. The transcription page covers models and correction.
  3. You log the same fields on every job. Create one or two shot fields in Settings → Shot Metadata, turn on the Autofill switch for those and nothing else, leave Autofill on import off, and run it on a handful of shots before anything larger. The Autofill page covers review and the failure modes.
  4. You already use a desktop assistant. Set it up from Settings → MCP Server or the Connect an AI assistant button, start read-only, ask it what is in your library, then turn writing on and make one small change. The assistant page covers the tools and attribution.
  5. You want a safety net first. Copy the metadata off one shot, paste it to a small selection, and open version history to see the entry it made. The version history page covers what is captured and what is not.

Installers for macOS and Windows are on the install page. An existing installation updates itself.

FAQ

Which platforms does StoryFolder 0.4 run on? macOS on Apple silicon, macOS on Intel, and Windows x64. There is no Linux build.

Is StoryFolder 0.4 a beta? The builds carry a beta label. Everything described here ships in them.

Is local transcription free? Yes. Transcription, Identify speakers, and the SRT, VTT, TXT, CSV and Fountain exports are all free. The engine and model download once from StoryFolder's servers.

Does AI Autofill upload my footage? No. It sends one frame per shot, plus the video's title, description and video-level fields and your own field instructions, to StoryFolder's server and on to a hosted vision model. The video file is not uploaded.

Can I undo an AI Autofill run? No. Autofill writes no restore point. Version history covers edits you make by hand, an assistant's writes and a paste, and the only guard on an Autofill run is Only fill empty fields, which is on by default.

Can ChatGPT connect to StoryFolder? Through Codex, OpenAI's assistant tool, in the Codex app, the CLI or the editor sidebar. The ChatGPT desktop chat window is not supported yet.

Does searching from All items cover every project? Yes. Searching or filtering from All items covers the whole library, including videos filed in folders; opening a folder narrows it to that folder and its subfolders.

A long shot board of which only the rows on screen are drawn

Is 0.4 faster than what came before? Yes, and noticeably so: the interface composites on the GPU, the shot board draws only the cards on screen, and on Apple silicon the app, the worker and the speech engines are all native builds (what makes 0.4 faster).

← All articles