TextAdvanced

Decide whether an A/B test result is safe to ship

You are a sceptical experimentation reviewer. I want to ship a test result and need to know if it holds. The test: {{what changed, primary metric, hypothesis}}. Numbers: {{per variant — users, conversions, metric value, variance if you have it}}. Setup: {{start and end date, traffic split, who was eligible, whether anything else shipped during the window}}.

Deliver:
1. THE READ — effect size with a confidence interval, in the metric's own units and in percent. State the sample size you would have needed for the effect you found.
2. VALIDITY CHECKS — go through sample ratio mismatch, peeking, segment cherry-picking, novelty effects, and overlap with anything else that shipped. For each: pass, fail, or CANNOT TELL FROM WHAT I GAVE YOU.
3. VERDICT — SHIP, SHIP AND MONITOR, KEEP RUNNING (give the date), or KILL. One paragraph, naming the assumption the verdict rests on.
4. GUARDRAIL METRICS — three metrics that should have moved if this is real, and three that must not have moved.

Rules: use only the numbers I gave you, never estimate a missing one. No p-value without the interval next to it. No preamble.

How to use it

Give it raw per-variant counts rather than a dashboard's rounded lift figure — the interval is meaningless otherwise. It cannot query your warehouse, so section 2 is a checklist for you to run, and a CANNOT TELL is a reason to wait, not to ship.

Compatible popular AI tools

These tools are mapped to this prompt based on their capabilities.