
Guides
Why Does an A/B Testing Plan Template Need a Decision Log?
An A/B testing plan template needs a decision log because the numbers alone cannot tell you why a test won, what to do next, or which results to trust.
What to take away
- A decision log records the reason a test exists, the threshold that ends it, and the action taken afterward.
- The primary metric for most tests is conversion rate: completed goal actions divided by sessions in each variant.
- A common stopping rule is a 95 percent confidence level, which corresponds to a p-value below 0.05.
- Conversion rate cannot tell you why people converted or whether the gain will hold on another traffic source.
- Write the decision rule before launch, not after the data arrives.
A test without a written decision rule is a hobby. The log is what turns a split test schedule into an operating record you can defend in a budget meeting.
What to measure
Conversion rate is the share of sessions that complete the goal action you defined. If 4,000 sessions produce 200 signups, the rate is 5 percent. Record the denominator. A rate built on 400 sessions is not comparable to one built on 40,000.
Pair it with a secondary metric that guards against harm, such as revenue per session or refund rate. A variant that lifts signups while doubling cancellations has not won.
The log itself carries four fields per test: hypothesis, primary metric, minimum detectable effect, and the decision rule. Add a fifth for the date the rule was written. That date is the proof you did not move the goalposts.
If your team already maintains a campaign brief template, the decision log slots in beside it as the measurement half of the same discipline.
How to read it
| Signal | What it usually means | Next step |
|---|---|---|
| Lift with p below 0.05 | Result is unlikely under the null | Ship, then watch for two weeks |
| Flat result, wide interval | Underpowered test | Extend or raise traffic |
| Lift on primary, drop on guardrail | Trade-off, not a win | Decide which metric governs |
| No lift, tight interval | Real null | Log it and stop |
Read the confidence interval, not just the point estimate. A 12 percent lift with an interval of 1 to 23 percent is a different decision than the same lift with an interval of 9 to 15 percent.
US regulators do not set a significance threshold for private testing. The 95 percent convention comes from statistical practice, not law. What regulators do police is the claim you make afterward. If a test result becomes an advertised performance claim, the FTC advertising guidelines apply to how you substantiate it.
What it cannot tell you
Conversion rate is a ratio, and ratios hide their parts. A lift can come from a change in traffic mix rather than from your design. If a paid campaign started mid-test and brought cheaper, lower-intent clicks, the control group may look worse for reasons unrelated to the variant.
It also cannot tell you why. A button color test that wins tells you nothing about the mechanism. You will not know whether the gain came from contrast, position, or the wording above it.
Novelty is a real confound. Returning visitors react to change itself. A two-week test on a site with heavy repeat traffic may show a lift that fades by week five.
Finally, a single test on a single page does not generalize. A result on a desktop landing page may not hold on mobile. Pew Research Center data on mobile technology adoption shows how much of US traffic arrives on phones, which is why device splits belong in the log.
Attribution and its limits
Most teams run tests inside an analytics platform that assigns credit by last click. That model gives the variant on the final page credit for a decision shaped by an email sent three days earlier.
Record the attribution model in the log next to the result. A 6 percent lift under last-click may be a 2 percent lift under a multi-touch model. Both are true; they answer different questions.
The practical fix is to keep the test metric narrow and the reporting metric broad. Test on the page you changed. Report on revenue at the end of the quarter. Do not blend the two in one spreadsheet.
If you publish results externally, treat them like any other marketing claim. The FTC native advertising guide explains when sponsored or promotional content must be identified as such, and a case study built on test data falls under the same disclosure logic.
When to stop measuring and decide
Set the stopping rule before launch and write it down. A workable rule: run until the primary metric reaches a 95 percent confidence level, or until you hit a pre-set session cap, whichever comes first. If the cap arrives first, the result is inconclusive and the log should say so.
- Check that the sample reached the planned size.
- Confirm the guardrail metric did not degrade.
- Compare the observed lift against the minimum detectable effect you set.
- Record ship, kill, or iterate, with the reason in one sentence.
- Archive the raw numbers so a later reader can recheck the call.
A useful discipline is to schedule the review meeting when the test launches, not when it ends. That removes the temptation to keep a test running until it produces a pleasing number.
For teams building the surrounding paperwork, the marketing report templates show how test outcomes fold into monthly reporting without turning into a wall of charts. And if a test produces a result worth telling publicly, the standards in marketing case studies explain what evidence a reader will expect before they believe the number.
Common questions
What confidence level should a US team use?
Ninety-five percent is the common default, matching a p-value below 0.05. Some teams use 90 percent for early directional reads. Pick one and apply it consistently, because switching thresholds between tests makes the log useless.
How long should a split test run?
At least one full business cycle, usually two weeks, and long enough to reach the planned sample. Stopping early because the graph looks good is the most common way teams fool themselves.
Does a decision log need legal review?
Not for the test itself. It needs review when a result becomes a public claim, since substantiation and disclosure rules then apply to the wording you publish.
What goes in the log when a test loses?
Everything. A documented null result saves the next planner from rerunning the same idea, and it is the only record that shows your testing program is producing knowledge rather than activity.






