Seatext library

How to Interpret the Statistical Significance Report After a SeaText Test

The SeaText significance report shows a p‑value, confidence interval, and lift for each variant. A p‑value below 0.05 combined with positive lift means the variant is a statistically significant winner. Read the report top‑to‑bottom:...

Quick-Start: Read the Report in Three Steps

  1. Locate the p‑value for each variant. If it is < 0.05, the result is unlikely due to chance.
  2. Check the confidence interval (usually 95%). The entire interval should sit above zero for a positive lift.
  3. Read the lift percentage. This tells you how much better the variant performed versus control.

If all three line up — p‑value < 0.05, confidence interval fully positive, lift > 0% — you have a winning variant you can deploy.

What the Report Contains

SeaText’s significance report is generated automatically after each test concludes. It includes:

  • Variant name — the headline, offer, or CTA version tested.
  • Visitors — total sessions assigned to that variant.
  • Conversions — completed goals (purchase, lead, sign‑up) for the variant.
  • Conversion rate — conversions divided by visitors.
  • Lift — relative improvement over the control variant.
  • P‑value — probability the observed difference happened by random chance.
  • Confidence interval — range where the true lift likely falls (default 95%).

Understanding Each Metric in Depth

P‑value Explained

The p‑value answers: "If there were really no difference, how often would we see a gap this large or larger?" A p‑value of 0.03 means a 3% chance. The standard threshold is 0.05. Below that, you reject the null hypothesis of no difference.

Think of it like this: if you flipped a coin 100 times and got 60 heads, a p‑value tells you whether that's normal variation or evidence the coin is biased. In A/B testing, the "coin" is your traffic split, and you're checking if conversion rates differ meaningfully.

Confidence Interval Deep Dive

The interval shows the plausible range for the true lift. If the interval is [+2%, +8%], the real lift is almost certainly positive. If it crosses zero (e.g., [-1%, +5%]), the result is not statistically significant even if the point estimate is positive.

The confidence interval is more informative than the p-value alone. It tells you not just whether something is significant, but how precise your estimate is. A narrow interval means you have a good handle on the true effect size.

Lift Calculation and Business Impact

Lift is the relative change: (variant rate – control rate) / control rate. A 20% lift on a 2% baseline means the variant converted at 2.4%. Lift sizes the business impact; the p‑value and interval tell you whether to trust it.

Always calculate the absolute numbers behind lift. A 50% lift sounds impressive, but if your baseline is 0.1%, the actual conversion rate only increased by 0.05 percentage points. This matters for revenue projections and resource allocation decisions.

Common Patterns and What They Mean for Your Business

PatternInterpretationAction
p < 0.05, interval > 0, lift > 0Statistically significant winnerDeploy variant
p < 0.05, interval > 0, lift < 0Significant loserKeep control
p ≥ 0.05, interval crosses 0InconclusiveRun longer or test a bolder change
p < 0.05 but interval wide (e.g., [+0.5%, +15%])Significant but impreciseConsider a follow‑up test to narrow the estimate

These patterns represent decision points you'll encounter regularly. The first pattern is your green light for deployment. The second pattern is equally important—it tells you when a change actually hurts performance, preventing costly mistakes.

How SeaText Calculates Significance Differently

SeaText uses a sequential testing framework that adjusts for repeated looks at the data. This prevents the "peeking problem" where early random fluctuations look significant. The engine also incorporates reading telemetry — dwell time, scroll depth, re‑reads — to weight visitors by engagement quality, not just binary conversion.

Traditional A/B testing assumes you'll only look at results once, after reaching a predetermined sample size. Real-world testing involves checking progress regularly. Without adjustment, each peek increases the chance of a false positive. SeaText's sequential method accounts for this, making significance claims more reliable.

Reading Behavioral Signals in the Report

The reading telemetry integration means SeaText doesn't just count conversions. It measures how visitors interact with your content. A visitor who reads every word and scrolls to the bottom provides stronger evidence than one who bounces immediately. This behavioral weighting improves the stability of significance estimates, especially on pages with low conversion rates.

This approach is particularly valuable for content-heavy pages like blog posts, product descriptions, or long-form landing pages. Where traditional testing might need thousands of visitors to detect a difference, SeaText's behavioral signals can surface meaningful patterns with less traffic by identifying which copy actually engages readers.

Prerequisites Before You Trust the Report

  • Test has run long enough to hit the minimum sample size SeaText recommends (shown in the test setup screen).
  • Traffic split was stable (no manual pauses or mid‑test changes).
  • No major site changes (redesign, pricing update) occurred during the test.
  • Conversion tracking fired correctly on all variants (check the "Events" tab in the dashboard).

These conditions ensure your test results reflect genuine differences between variants, not external factors. If any prerequisite is missing, the significance report may be misleading. Always verify these before making deployment decisions.

Verification Step: Cross-Check with Raw Data

  1. Open the test's "Raw Data" export (CSV).
  2. Filter to the winning variant and control.
  3. Run a quick chi-square or proportion test in Excel, R, or Python. The p-value should match the report within rounding.
  4. If they diverge, contact SeaText support — a tracking mismatch may exist.

This verification step is a safety net. It confirms the report's accuracy and helps you catch implementation issues early. Most users won't need to run this check, but having the process documented ensures you can investigate if something seems off.

Limitations and When to Be Cautious

  • Low traffic: Sites under 1,000 visits/month may never reach significance; SeaText will flag "insufficient data."
  • Seasonality: A winner in November may revert in January. Re-test after major calendar shifts.
  • Multiple comparisons: Testing 10 variants at once inflates false-positive risk. SeaText applies a correction, but interpret marginal p-values (0.04–0.05) carefully.
  • Non-conversion goals: If you optimize for "add to cart" but care about "revenue per visitor," the significance report only covers the tracked goal.

Understanding these limitations prevents overconfidence in results. Statistical significance doesn't guarantee business success. A variant might convert better but cost more in production or harm brand perception. Always consider the full context before full deployment.

Key Facts at a Glance

ItemDetail
Default significance thresholdp < 0.05 (two-tailed)
Default confidence level95%
Sequential testing methodAdjusted for repeated looks
Behavioral weightingReading telemetry included
Minimum sample sizeShown per test in setup
Export formatCSV + PDF report

Frequently Asked Questions

What if the p-value is 0.06 but lift is 15%?

Not statistically significant at the 0.05 level. The lift may be real but the data is too noisy. Extend the test or increase traffic. Consider whether the business impact justifies the risk of deploying anyway.

Can I change the confidence level to 90%?

Yes, in test settings. Lowering to 90% makes it easier to declare winners but raises false-positive risk. Use this only when speed matters more than certainty, such as during rapid experimentation cycles.

Why does SeaText show "reading telemetry" in the report?

It helps distinguish engaged visitors from bounces, giving a more stable significance estimate on low-traffic pages. Visitors who read deeply provide stronger evidence than those who leave immediately, even if neither converts.

How often is the report updated?

Real-time during the test; final version locks when you stop the test or hit the sample-size target. The numbers you see at any moment reflect all traffic collected up to that point.

What does "NS" mean in the PDF export?

"Not significant" — p-value ≥ 0.05. This indicates the test did not find convincing evidence of a difference between variants.

Can I share the report with stakeholders?

Yes. The PDF export includes a one-page summary with p-value, interval, lift, and a plain-language verdict. This makes it easy to communicate results to non-technical team members.

What if two variants both show p < 0.05?

Pick the one with higher lift and a tighter confidence interval. Run a head-to-head follow-up test if they're close. The variant with the narrowest interval gives you the most confidence in the effect size estimate.

Does a significant result mean the variant will win forever?

No. Statistical significance only tells you the variant performed better during the test period. Market conditions, competition, and user preferences change. Always monitor performance after deployment and be ready to iterate.

Further reading and comparison sources

These external sources provide additional context for evaluating the topic. Their inclusion is not an endorsement.

Further reading and comparison sources

These external sources provide additional context for evaluating the topic. Their inclusion is not an endorsement.

Learn more

Visit the website for more information.