Logging footage where nobody is scripted
StoryFolder —
Part of: How to search footage by what's in it and what's said
Last updated: 2 September 2026.
Nobody hands you a script for this footage. People walk into frame mid-sentence, talk over each other, and the same three people show up across forty hours of location audio wearing different roles depending on the day. A scripted interview gives you one clean speaker and a slugline. Reality and documentary give you neither, and most logging tools assume you have both. Someone on r/editors, describing what an NLE's project bin actually gives you for unscripted or documentary work: "NLEs are super far behind when it comes to logging and managing media in the project bin, and it makes selects far more annoying than they need to be, especially for unscripted or documentary work." (u/OliveBranchMLP, r/editors, 2025.)
Useful logging here needs a transcript and a speaker structure you can actually edit, rather than a wall of text carrying Speaker 1 labels nobody bothered to fix. StoryFolder separates speakers automatically, then lets you rename, merge, delete or reassign them. One thing to know before you build a habit around it: a re-transcribe rebuilds the speaker segments from scratch, so the naming pass is worth doing on the transcript you intend to keep.
New in 0.4. Speaker identification is new, free, and off by default; it runs on your machine alongside the transcription. How it works and what it gets wrong.
What makes reality footage harder to log than a scripted interview?
There's no script to check dialogue against, no single clean speaker track, and no NLE bin structure built for a cast that changes role by scene. One StoryFolder customer coined a name for the only workaround most people have without a real logging tool: "SCRUB v1", a full watch-everything pass, logged in the board title because there was nowhere better to put it. That's not a deliverable, it's a navigation verb somebody had to invent because nothing else was tracking it.
The honest version of the job: separate who's speaking, correct it where the machine got it wrong, and keep the correction around long enough to matter.
How are speakers separated and corrected?
Diarization runs locally as part of transcription, and it comes back with generic labels, Speaker 1, Speaker 2, and so on. The speaker engine, its default and its limits are described alongside the transcription itself. From there you can rename a speaker, give them a colour so a long transcript stays scannable by eye, merge two labels the diarizer split by mistake, delete one, or add a speaker that wasn't detected.
Correction goes finer than the whole segment, too. Drag-select a run of words inside a line and assign just that range to a different speaker, which is what you want for the moment someone talks over the person actually being interviewed. That reassignment splits the line into pieces at the word boundary you selected, and the split pieces lose their ability to be reverted to what the model originally wrote, since the line they came from no longer exists as one unit.
One thing diarization doesn't do: cleanly separate two people talking at once. When speech overlaps, the system resolves the segment to whichever speaker has the larger share of it, so you get one label on that segment.
What happens to your speaker work when you re-transcribe?
The transcript comes back with a fresh set of speaker segments, because diarization runs again over the audio from scratch. Plan the naming pass around that rather than around the previous transcript.
Re-transcribe a file, whether for a new model, a corrected language setting or any other reason, and the diarizer segments the audio again. What comes back is a new set of speaker segments rather than the old ones re-labelled, so a word-range reassignment belongs to the transcript you made it on.
The practical rule is the same one that applies to any pass you might redo: settle the model, the language and the speaker-detection switch first, run the transcribe you mean to keep, and do the naming and the word-range corrections after that.
Can the log be read and searched by speaker?
Yes. By speaker is one of five transcript reading modes, alongside By shot, Continuous, Captions and Plain text, and it groups every line under the person who said it. You can also narrow the transcript to specific speakers before running a search. That's a visibility filter you set first; transcript search itself reads spoken words, never speaker names.
Which exports keep speaker labels?
SRT and WebVTT carry speaker labels once a transcript has more than one speaker; VTT uses voice spans to do it. Fountain, built for screenplay-shaped text, emits proper character cues. The transcript CSV carries speaker alongside start time, end time, shot number and text on every row, which is the export worth reaching for if you're handing a reality log to someone who needs to sort or filter it later.
Where will diarization still need a human?
Heavy accents, people talking over each other, noisy location sound and specialist jargon all degrade what the model gets right. That's the same list that degrades transcription generally, and diarization inherits it. Low-confidence words get a dotted underline so you know where to look twice. It's off until you turn it on under Display, and it never shows a percentage, because the model's own confidence isn't calibrated well enough to print as one.
What does StoryFolder still not do for a reality workflow?
It doesn't edit video by editing text, doesn't round-trip picks back to an NLE, doesn't run a live remote review session, and it isn't a MAM or a drive catalogue. It also carries elapsed timecode rather than source timecode, with no reel, drop-frame or SMPTE frames field, so a shot list built from it reads to the second rather than the frame.
FAQ
Can StoryFolder rename detected speakers? Yes. Speakers can be added, renamed, recoloured, deleted or merged after diarization runs.
Can I reassign just part of a line to a different speaker? Yes. Drag-selecting a range of words splits the line at that point and assigns the selected range.
What does re-transcribing do to the speakers I already corrected? It runs diarization again, so the transcript comes back with a new set of speaker segments. Settle the model and the language first, then do the naming and the word-range corrections on the transcript you are keeping.
Can I search for a specific speaker's name? No. Search reads spoken words. Which speakers are visible is a separate filter you set first.
Which exports keep speaker labels? SRT and WebVTT once there is more than one speaker, Fountain as character cues, and the transcript CSV on every row.
Does StoryFolder separate two people talking at once? No. An overlapping segment is resolved to whichever speaker has the larger share of it.