Workflow6 min read

Why Two Reviewers Give the Same Video Different Notes

Reviewer disagreement isn't a people problem, it's a missing rubric and an unstated severity scale. How to calibrate a team so a cut gets the same answer whoever opens it.

By Editorial standards

Run an experiment that takes an hour and tells you more about your review process than any dashboard will. Take five videos you approved last month, strip the notes, and have two reviewers review them independently. Then compare.

Most teams doing this for the first time find agreement on which videos pass, and almost no agreement on the notes. Different issues flagged, different severity, different wording, and at least one case where reviewer A blocked something reviewer B didn't mention.

That gap is expensive in a specific way. It doesn't produce bad approvals often. It produces creators who cannot predict what will pass, which produces defensive over-delivery, extra rounds, and the belief on the creator side that review is a matter of taste and luck.

The four sources of disagreement

Reviewer variance is not carelessness. It comes from four things the process never told anyone.

No shared rubric. If the checklist lives in one reviewer's head and the other one inherited it verbally, they are running different checks. This is the common case on teams that grew from one reviewer to two without writing anything down.

No severity scale. Both reviewers notice the disclosure is at second 6 rather than second 5. One treats it as a blocker, one as a note. Neither is wrong, because nobody ever said which it was.

Unwritten standards. Half of any experienced reviewer's judgment is precedent: what this brand let through last quarter, what legal pushed back on once, which creator gets latitude. New reviewers cannot access any of it and will not know it exists.

Order and fatigue effects. The eleventh video of the afternoon gets a different review than the first, from the same person. This one is real, measurable, and mostly unfixable by trying harder.

Note that only the fourth is about the reviewer. The other three are documents that don't exist.

Build the rubric as checks with severities

A checklist that lists what to look at is half a rubric. The other half is what each finding means for the decision.

Three severities are enough:

  • Block. Cannot go live. Regulatory exposure, a forbidden claim, a missing or unusable disclosure, a competitor in frame.
  • Fix. Goes live once changed, but the change is required. Disclosure present but late, code missing, required phrase absent.
  • Note. Advisory, does not gate approval. Pacing, framing preference, a suggestion for next time.

The value is entirely in assigning them in advance, per check, in writing. Once that's done, two reviewers watching the same cut can still notice different things, but they can no longer disagree about what a given finding costs the creator.

One rule keeps the scale honest: a Note never blocks approval, and if a reviewer wants it to, it was misfiled and the rubric needs updating. Teams that let notes quietly gate approval end up with a rubric nobody trusts, and then everything is a blocker again.

If a check can't be assigned a severity in advance, it's probably not a check. It's creative direction, and it belongs in the brand pass rather than the compliance pass.

Calibrate on real cuts, not on the document

Writing the rubric aligns about half the variance. The rest only comes out with worked examples, because the disagreements that matter are about edge cases, and edge cases are not enumerable in advance.

Build a calibration set: ten to fifteen past videos that between them cover a clean pass, an obvious block, and, most importantly, six or seven genuinely ambiguous ones. Ambiguous is the point. A calibration set of easy cases proves nothing.

Then, once when a reviewer joins and once a quarter after:

  1. Everyone reviews the set independently, no discussion.
  2. Compare, finding by finding, and count the disagreements before anyone explains themselves.
  3. Discuss only the disagreements.
  4. Write the resolution into the rubric as a rule, with the example attached.

Step 4 is the one teams skip, and skipping it means the same debate recurs every quarter with a different outcome. The output of a calibration session is a document diff, not a shared feeling.

The examples are what make it work. "Disclosure must be conspicuous" aligns nobody. "This one passes, this one fails, here's the frame" aligns everyone, and it's how the unwritten precedent finally gets written.

Standardise the notes, not just the checks

Two reviewers can reach identical conclusions and still send feedback the creator experiences as inconsistent, because the format differs. Fix the shape of a note:

  • Timestamp. Where in the cut.
  • The rule. Which brief line or rubric check failed, quoted.
  • The required change. Specific enough to act on without a reply.
  • Severity. Block, fix, or note.

"Disclosure needs work" fails all four. "0:00 to 0:08, brief requires spoken disclosure in first 5s, add 'paid partnership with X' to the opening line, Block" fails none.

This is also where the reviewer's private uncertainty leaks into the creator's day. A reviewer who isn't sure whether something is a blocker tends to write ambiguously and let the creator decide how seriously to take it. That ambiguity is a large share of why cuts come back for a third round.

Measure agreement, because it decays

Alignment is not a state you reach. New reviewers arrive, campaign rules change, and precedent drifts.

Track one number: on the quarterly calibration set, the share of findings where reviewers agree on both the finding and its severity. Don't chase 100%. Something in the 80s on a deliberately hard set is a well-calibrated team, and a set on which everyone agrees completely is a set of easy videos.

Watch the direction rather than the level. A drop after a headcount change tells you the onboarding was verbal again.

Where the rubric pays for itself twice

A rubric written this way does two jobs. It aligns your reviewers, and it's the artifact that lets you raise capacity without hiring, because checks with fixed severities are checks that can be routed, delegated, or run automatically. A machine cannot apply an unwritten standard either, for the same reason a new hire can't.

It also feeds back upstream. Any check that keeps generating disagreement is usually a brief line that was never checkable. Reviewer disagreement is a decent detector for brief ambiguity, and fixing it there removes the finding entirely rather than standardising how you argue about it.

Run the same checks the same way every time

CherryBowl applies your rubric to every cut with fixed severities and timestamped evidence, so the mechanical findings are identical whoever is on the queue.

See the AI review a video

Or join the early-access waitlist.

or book a call

The takeaway

Two reviewers disagree because there's no written rubric, no severity scale, and no record of precedent. Write the checks with a block, fix, or note severity attached to each. Calibrate on a set of genuinely hard past videos, and write every resolution back into the document. Standardise the note format so consistent conclusions arrive as consistent feedback. Then measure agreement quarterly, because it decays quietly and shows up as creator confusion long before anyone names it.

Keep reading

More on workflow