Reading your report

What each section of the report is for, what every verdict word means, and which people the scores are counted over.

Updated

Every report arrives in the same order and says only what the evidence supports. Reading one well is mostly knowing what each section is for, and what each verdict word costs to earn.

The order of the document

  1. Cover. The test, both versions with a frame from each, the audience, the date. The document stands on its own when you forward it.
  2. The Bottom Line. The verdict in plain sentences, and what it means for the decision in front of you. Most people read this and stop, which is why it is written to be enough on its own.
  3. In Their Words. Participant comments, grouped and attributed, with the full set one click away. Counts are stated honestly: five people who said the same thing are five people, not "viewers".
  4. By the Numbers. The score for each version, the gap between them, and the range around that gap. The quantitative record, for whoever wants to check the working.
  5. How This Was Run. The method in the open, including the participant flow.
The Bottom Line section of an audience test report with a plain-language verdict
Screenshot: testing-report-bottomline-1
The Bottom Line: the verdict in plain sentences, and what it means for the decision in front of you.
The Bottom Line: the verdict in plain sentences, and what it means for the decision in front of you.

The participant flow

The flow table shows, for each version, how many people were assigned, how many were exposed, how many finished the piece, how many answered, how many were voided and replaced, and how many were counted in the score. Reading down a column tells you where people left. It is there so the headline can be audited rather than believed.

The In Their Words section of a report showing attributed participant comments
Screenshot: testing-report-quotes-1
In Their Words: the audience speaking for itself, grouped and attributed.
In Their Words: the audience speaking for itself, grouped and attributed.

The verdict words

Each phrase is earned, and none of them is upgraded because a result came out close.

  • Outperformed. The strongest claim available, reserved for the primary question you set before launch. It needs three things at once: every group reached the sample committed at launch, dropout was balanced between the groups, and the gap is large enough to stand clear of chance. Miss one and the word is not used.
  • A directional advantage. One version is ahead and the evidence leans, but not far enough for the strongest claim, or the finding came from a supporting question rather than the primary one. Worth acting on when being wrong is cheap, worth a bigger test when it is not.
  • Too close to call. The versions landed near enough that naming a winner would be a guess. Still a finding: it says this choice will not move the needle much, so you can decide on craft, cost, or schedule. See Understanding your verdict.
  • Equivalent. Not a softer way of saying too close to call. Equivalence is reached only when the gap is small enough to rule out a meaningful difference in either direction, so it is a proven claim rather than a shrug.
  • N points higher in this sample. Descriptive. It describes the people who took this test and reaches no further. Read it as texture around the verdict, never as the verdict.

Who gets counted

Scores are computed over the people who experienced the whole piece and answered the primary question. Someone who abandoned the video halfway shows in the flow table and stays out of the score. That is stricter than counting everyone who clicked, and it is what keeps the two groups comparable.

What a report will not tell you

It will not tell you why: the design measures difference, not mechanism, so the words are where reasons live. It will not tell you what one person would have preferred with both versions in front of them, because nobody saw both. And it speaks to a first look at your work, not to what happens on the fifth viewing.

Questions

Is "equivalent" the same as "too close to call"?

No, and the difference matters. Too close to call means the test could not separate the versions. Equivalent means the gap was small enough to rule out a meaningful difference in either direction, which is a claim in its own right and the one that lets you ship the cheaper version with a straight face.

Ready to try it?

Host video, collect frame-accurate review, and see who's watching. Free to start.

Start free