Appearance
Identify speakers in a transcript
With Identify speakers on, StoryFolder groups similar-sounding voices and labels each line with the group it belongs to. It arrives as Speaker 1, Speaker 2, each with a colour. Turning those into names is your part, and it is the part that makes the transcript worth reading.
Speaker detection runs on your computer, like transcription itself, and is free on every plan.
Voice grouping is not identity recognition
StoryFolder detects that the voice changed and that these lines sound like each other. It does not know who anybody is, and it never will from the audio alone. A label is a group of similar voice until you name it.
Turn on speaker identification
- Open Preferences (Cmd+P on Mac, Ctrl+P on Windows).
- Go to the Transcription pane.
- Turn on Identify speakers. Every new transcription is labelled from then on.
Speaker detection needs about 60 MB of its own, downloaded the first time you use it, the same way a model is. If it is not available for your computer the switch says so rather than quietly doing nothing, and a transcript that cannot be split keeps a single speaker instead of failing.
Identification adds work on top of transcription, so a run with it on takes longer.
To label a video that is already transcribed, open Display → Re-transcribe… in the transcript toolbar, expand Advanced settings, and turn speaker detection on for that run.
Re-transcribing replaces the whole transcript
Every line and timing is produced again, including your hand corrections — StoryFolder counts them and refuses the first attempt if there are any. Speakers you have named or recoloured are carried over; speakers that came from automatic detection are worked out again from scratch.
What the automatic labels mean
Speaker 1 is a group of lines that sound like each other. Read the labels as a first pass, and expect these:
- One person split across several labels. A change of microphone, a phone call, a shout across a room.
- Two people combined into one label. Similar voices, or a conversation where they overlap constantly.
- Difficult material handled badly. Overlapping speech, one-word interjections, laughter, crosstalk and noisy location sound are where detection struggles most.
None of that is a malfunction to report. It is the shape of the problem, and the tools below are how it is finished.
Rename a speaker
- Open the speakers popover in the transcript toolbar, or use the Speakers panel in the inspector.
- Click the automatic label.
- Type the person's name.
The new name applies to every line assigned to that speaker, throughout this transcript — in the by-speaker view, in copied text, in exports and in the printed storyboard.
You can also change a speaker's colour, hide one from the view while you concentrate on another, and add a speaker that detection never found.
Reassign words to the right speaker
This is the fix for lines given to the wrong person, and for two people who ended up under one label.
- Drag across the words that belong to somebody else.
- A bar appears above the transcript naming what you selected.
- Choose the speaker it should belong to, or Unassigned.
The selection is what gets moved — StoryFolder splits at word boundaries rather than reassigning a whole line, so half a line can change hands. Select several lines at once when a whole passage was misattributed. Cancel on the bar clears the selection without changing anything.
If the voice is genuinely someone new, add a speaker first, then assign to it.
Merge duplicate speakers
Merge when one person was split across two labels.
- Open the speakers popover and choose … on the speaker you want to fold away.
- Choose Merge into…, then pick the speaker to keep.
Every line under the speaker you merged from moves to the one you kept, and the merged-from label is gone. It is one editing gesture on the transcript, so version history records it and it can be undone.
Delete or unassign a speaker
The four verbs do different things, and it is worth keeping them apart:
- Rename changes what the group is called. Nothing moves.
- Reassign moves the selected words to another speaker.
- Merge moves every line from one speaker to another and removes the empty label.
- Delete removes the label. Its lines stay exactly where they are and become Unassigned — no words are lost, and nothing about the transcript text changes. StoryFolder says so in the confirmation, with the line count.
All four are transcript edits, and all four are recorded in version history.
Re-running detection is the exception: that is a re-transcription, and it starts the automatic labels over.
Review speakers efficiently
Switch to the By speaker view and read each speaker's lines together — a misattributed line is far more obvious among a person's own words than it is in shot order.
Sample the beginning, middle and end of each speaker. Then check the places detection finds hardest: scene changes, off-camera voices, one-word answers, and any stretch where two people overlap.
Correct the transcript text first, names second. Searching for a name is only useful once the words are right.
Where speaker names appear
- In the transcript, when Speaker names is on in the Display popover.
- In copied text, following the same display settings.
- In SRT and WebVTT subtitles, and in the plain text, CSV and Fountain exports, once there is more than one speaker.
- In the printed storyboard, and in the Statistics tab, which shows each speaker's share of lines and words.
Plain text view deliberately drops speaker names and timestamps — it is there for prose you are about to paste somewhere else. See Search, copy, and export a transcript.
Troubleshooting
| What you see | What to do |
|---|---|
| One person appears as several speakers | Merge them into the one you want to keep |
| Several people appear as one speaker | Add the missing speaker, then reassign their words |
| No speakers at all | Check Identify speakers was on for the run, that its download finished, and that there is speech to split |
| Labels changed after re-transcribing | Expected. Named speakers carry over; automatic ones are worked out again |
Related pages
- Transcribe a video on your device
- Search, copy, and export a transcript
- Troubleshoot transcription
- Version history