Preparing a fair test

A test can only answer the question you built it to answer. What to hold steady, what to vary, and how to choose a length you can defend.

Updated

Two groups of strangers meet two versions cold, and whatever differs between those versions is what your answer will be about. That makes the preparation the study. Here is what keeps a comparison honest.

Change one thing

If version B has a new opening, a different voice, and a warmer grade, a win tells you the bundle won. It cannot tell you which part did the work, and you will carry all three changes forward without knowing which one you are paying for. Vary the one element under argument and hold the rest steady. The wizard's cards say the same thing in their own words: a talent test only isolates the talent when the talent is the only thing that changed.

Match length and framing

Two cuts of very different lengths are two different asks of a viewer, and the shorter one carries an advantage that has nothing to do with the work. Keep them close in length. Keep the aspect ratio, the captions, the audio treatment, the opening card, and the end card the same across both. Every difference you did not intend is a rival explanation for the result.

Finish both to the same polish

Testing a rough cut against a finished one answers a question nobody asked: people prefer finished. Temp music against a real mix, ungraded picture against graded, scratch voice against the final read. Any of those decides the test before a single person watches. If one version is not ready, wait for it. An unfair pair costs real money and buys back a result you already knew.

The audience testing setup wizard on the step where two versions are chosen
Screenshot: testing-wizard-1
The versions step. Whatever differs between these two is what your answer will be about.
The versions step. Whatever differs between these two is what your answer will be about.

Decide the primary question before you launch

One question decides the verdict, and it is frozen when the test launches. That is deliberate: it stops anyone (including you) from reading the results afterwards and promoting whichever number came out best. Pick the question you would defend in the room. Supporting questions still come back as diagnostics and the open answers still come back as quotes, so nothing is lost by committing early.

Choose your length with your eyes open

Participants are paid for their time, so the size of the ask sets the price. A short piece, a general audience, and a survey inside the first 4 questions sit at the base rate, roughly 4 credits per person. A long piece, or a long survey, raises the cost per person because it asks for more of someone's evening. A paid test starts at 25 credits, and the wizard quotes your exact number before anything runs. See What does a test cost?.

Length is a design choice as much as a budget one. If the argument is about the first ten seconds, test the cutdown you would actually run in a feed rather than bolting that question onto a six-minute film every participant has to sit through.

The checklist

  • One deliberate difference between the two versions.
  • Matched length, framing, and finish.
  • A primary question you would stand behind in public.
  • An audience that matches who the work is for. See Choosing an audience.
  • A sample size you can afford at that audience.

Questions

Can I test a rough cut at all?

Against another rough cut at the same stage, yes: the comparison is fair as long as both sides carry the same level of finish. What does not work is a rough cut against a finished one, because the audience will reward the finish and you will read it as a win for the edit.

Ready to try it?

Host video, collect frame-accurate review, and see who's watching. Free to start.

Start free