Is AI video tagging any good?

StoryFolder —

Part of: How to log footage without tagging every clip by hand

Last updated: 2 September 2026.

You ran AI Autofill over a board. Some cells came back filled, some came back empty, and nothing on the board tells you which of those is a judgement and which is silence. So which of the filled ones are wrong, and how would you tell? Automatic tagging does one narrow job: a first pass on fields a frame can visibly support. A Dropdown, Rating, Checkbox or Number answer has to fit that field's own domain: one of your options, a whole number from 0 to 5, true or false, a finite number. A Tag or Text field has no domain at all, so anything can come back. What it will not do is let you skip the review. A valid-looking answer can still be the wrong answer, and nothing in the interface marks the difference.

Two things worth knowing before you run anything. StoryFolder does not stage a suggestion for you to accept or reject; a value it writes goes straight into the cell, the same way a value you typed would. And an autofill run creates no restore point, so there is no run-level undo. Those two facts, not an accuracy number, are what "is it good" actually depends on.

One practical gate: AI Autofill needs StoryFolder Pro, and so do the custom fields it writes into. A run sends one representative frame per shot to StoryFolder's server, which calls a third-party vision model. Transcription is the other case. That one runs on your own machine.

New in 0.4. Everything tested here ships in StoryFolder 0.4: the per-field switches, the domain check, the silent drop and the missing undo. The feature page carries the current defaults and the review method.

What does "good" mean for AI video tagging?

"Good" is three questions, and they have three different answers. Does the model answer inside the categories you defined? Yes for a Dropdown, Rating, Checkbox or Number: that is checked twice before anything is written. Is the answer factually right for the shot? No benchmark exists for this feature on this build, and none is claimed here. How expensive is a mistake to find? That depends on the field type you used.

The first is vocabulary validity, and it is the one that is proven and mechanical. A Dropdown, Rating, Checkbox or Number field is checked against its own domain twice, and an answer outside that domain is refused.

The second is factual correctness: is the answer the right one for this shot? That is a judgement call about the image, and there is no benchmark behind it. Nobody has measured a correctness rate for this feature on this build, and this page does not invent one.

The third is workflow recoverability: if the model is wrong, how expensive is it to find and fix that shot? That depends entirely on which field type you used it on, which is the part of this page worth your time.

Vocabulary validity is the part StoryFolder can promise. The other two are yours to judge, on your own footage, and the rest of this page is about how.

What happens when the AI chooses the wrong valid category?

It writes the value, and it looks exactly like something you typed.

Say you have a Content Type Dropdown with the options B-roll, Interview and GFX, and a run reaches a wide static shot of an empty room, the kind of establishing frame that could read as B-roll or as the setup for an interview that hasn't started. The model answers Interview. That is one of your three options, so it passes the server's domain check, and it passes the worker's second check before it is written. Both confirm the answer belongs to your vocabulary. A domain check tests membership. It says nothing about whether the member is right for this shot.

The value lands in the cell with no flag and no confidence score. StoryFolder does stamp which values came from autofill rather than a person, but nothing in the app displays that stamp: not on the board, not in the inspector, not in an export. So the only way to find the wrong Interview is the same way you would find a typo you made yourself: look at the shot. "Domain-constrained" is a promise about the vocabulary, and it is a weaker promise than "correct".

There is no undo, so the recovery is a filter, which is also the shape of reviewing an Autofill run at scale. On the board, hit + Filter, pick Content Type, pick is, and pick Interview, and the board narrows to the shots that came back with that answer. Scroll that set against the thumbnails and a wide empty-room shot sitting among real interviews stands out; correcting it is a click. You do this one option at a time (a filter shows you the shots holding a value, and there is no view that groups every value on the board side by side), which is why a field with three options is a two-minute review and a field with thirty is not. Leave Only fill empty fields on for any later run, so re-running never re-guesses over a value you just fixed.

What happens when the AI returns an invalid value?

It gets dropped, and the cell stays empty, with no indication that anything was attempted.

What happened What you see
A valid, wrong answer Written. Looks identical to a typed value.
An answer outside the field's domain Dropped silently. Cell stays empty.
The frame did not clearly support an answer Cell stays empty.
The model call failed for that shot Cell stays empty; the shot is marked failed internally, but that reason is not shown to you.

Four different situations, and from your chair three of them look exactly the same: an empty cell. Nothing distinguishes them (no toast, no badge, no per-field note), so "the model declined because the frame didn't support an answer", "the model answered something outside your options and got rejected" and "the call failed for this one shot" all arrive as silence.

That changes how you read a board after a run. An empty cell after autofill means exactly what an empty cell means before autofill: nobody has told you anything about that shot yet. It never means "the model checked and there was nothing there."

AI Autofill has no undo

That is literally true, and the write modal says so before you can run in overwrite mode.

A metadata edit you make in the app captures a restore point you can step back from. An autofill run does not. If a run goes badly, say a Dropdown option turns out to be ambiguous, or a field description steered the model wrong, there is no button that puts the board back.

The safety net is Only fill empty fields, on by default, which guarantees a run can never replace a value you already wrote, plus a small first run. Enable autofill for one or two fields, scope a run to a handful of shots or one video, and look at the result before you turn it loose on the next hundred boards. If the pattern of mistakes is one you can live with, say the model occasionally confuses static interview setups with B-roll, keep going. If it is not, turn the field's switch back off and log that one by hand.

Which fields produce the most reviewable results?

The field type decides how expensive a mistake is to find. Choose with that in mind, not only with "can a frame answer this".

Closed fields are the reviewable ones. Content Type, shot size, camera move, a B-ROLL checkbox, whether talent is visible, a 1–5 usability rating: each of these has a small, fixed set of possible answers, so a mistake is one of a handful of known wrong values, and you can pull every shot holding one of them onto the screen with a single filter. The exact mechanism behind that constraint is a separate page's job; the reviewability consequence is this page's.

Client, Job#, approval status, release-form status, and anything with a filesystem or contractual consequence are the expensive case. Nothing in the field stops a wrong answer from looking plausible. A Text field holding Job# has no domain at all, so a model can return a job number that is formatted correctly, sounds right, and belongs to nothing: no membership test to fail, and no filter that separates a real one from an invented one. Keep those typed or pasted.

Tag fields sit in between. They are open by design (a tag vocabulary has to be able to grow), so reviewing a tagged board means reading the tags that appeared, not filtering for the ones you expected.

What should a trustworthy test show?

Run one before you trust this on real client work, and run it on your own footage rather than taking anyone's word for it, including this page's.

Pick one video with a schema you actually plan to use, the same Content Type Dropdown and B-ROLL Checkbox you'd log with day to day. Run autofill on it with Only fill empty fields on, scoped to all shots. Then review it the way you would review it on a real job, which is by filter and not by scrolling. Take each closed field in turn, filter to each of its options, and look at that queue against the thumbnails. For the open Text and Tag fields, which have no queue to filter into, take a fixed sample across the beginning, middle and end plus the shots most likely to defeat a single frame, which are the dark ones, the graphics and the transitions. Count four things as you go: how many fields the model attempted, how many it wrote, how many it left blank, and how many of the written ones you had to correct. Keep that count with the date, the build you were on, and a short description of the footage; a number without those attached is not reusable evidence.

That is a test you can run in the time it takes to review the board you were going to review anyway. It will not produce a percentage that generalizes to different footage, different schemas, or a different day's model behavior. It will tell you whether this is worth turning on for the kind of shots you actually shoot.

Does one failed shot stop the run?

No. A shot whose frame can't be read, or whose model call fails outright, is marked failed and the run moves on to the next one.

Transient network and server errors retry on their own before a shot is given up on. Once a shot is marked failed, that's terminal for the run, but nothing else on the board is affected and the rest of the shots finish normally. The reason that shot failed is recorded internally and never shown to you, which is why its empty cell reads the same as every other empty cell.


FAQ

What image does the model actually look at? One representative frame per shot, plus the video's own title, description and video-level fields, and anything you typed into the run's optional context box.

Are the extra instructions I type saved for next time? No. The additional-context box is typed fresh every time you open the run modal and is not stored anywhere.

Does StoryFolder retry a failed shot automatically? Yes, for transient network or server errors. A shot the model genuinely couldn't answer is not retried.

Can I stop a run partway through? Yes. Stop autofill is cooperative and takes effect within a shot or two rather than instantly.

Can I tell why a cell is empty after a run? No. The reason is recorded internally but never shown in the app, so a declined answer, a rejected answer and a failed call all look the same.

How do I find the shots a run gave a particular value? Filter the board on that field — + Filter, the field, is, the value — and scroll the result against the thumbnails.

How do I protect values I've already corrected? Leave Only fill empty fields on. It's the default, and it means a rerun can never replace a value that already has something in it, corrected or not.

← All articles