Nothing to hand? Load the — a clean win with a guardrail worth a second look — the , a null result the team is about to act on that the test was never large enough to produce, or the , where a headline win sits on top of an assignment that never happened. All three replay a saved run for free. Or press to watch the checker name its findings with no model call at all.
The statistics are arithmetic, so they happen here
Two proportions, a pooled standard error for the p-value and an unpooled one for the interval, a delta-method interval on the relative lift, Welch's test when the metric is continuous. None of that needs a model and none of it should cost anything, so all of it runs in this tab before you sign in. The one thing worth saying about the arithmetic is that the pooled and unpooled standard errors are both computed, because using one where the other belongs is the most common way a homemade calculator is quietly wrong.
The useful question is whether the test could have answered itself
A p-value alone cannot tell you whether a null means “no effect” or “no visibility.” So the engine computes the sample size the declared MDE required, the smallest effect the exposure you actually collected could have detected, and the power the design achieved against the effect it saw — and separates the two kinds of null. When the test could not see the effect it was designed for, it says so and solves: at the observed direction, this many per arm would settle it, which is this many more than you have, about this many more days at the rate you were running.
Most broken experiments are broken before the p-value
The split is checked against the plan with a chi-square at the 0.001 convention rather than 0.05, because this check runs on every test and a 5% threshold would condemn one healthy experiment in twenty. The run length is checked against whole weekly cycles. Every declared interim look is priced, with the Pocock equal-boundary constant where it applies. The guardrail family gets a Holm step-down correction, so a five-metric dashboard cannot manufacture a significant regression by sheer count. A sample-ratio failure stops the readout: nothing below a broken assignment can be read.
The metered pass is judgement, and it is held to the measurement
What a model is for: whether the hypothesis was testable in the first place, what a flat metric means for the product, the novelty effect or segment heterogeneity or unit mismatch the arithmetic cannot see, and the experiment that should run next. It returns exactly one reading per measured metric, and the app counts them — a metric read twice fails as loudly as one left out. Each reading is checked against the measured significance, its stated decision is set beside the rule's rather than merged into it, and a metric the engine itself could not test is excluded from the partition so the pass is never blamed for a gap the browser already found.
What does Experiment Desk compute that a significance calculator does not?
Whether the test could have answered its own question, and what it would take if it could not. A calculator returns a p-value; this returns the p-value, the interval, the power the design actually achieved against the effect it saw, the sample size its own declared MDE required, and the smallest effect the exposure collected could have detected. When the answer is inconclusive it solves for the missing driver rather than reporting a gap: at the observed direction, 17,565 per arm would settle it, which is 13,475 more than were collected, about 27 more days at the rate this test was running. It also runs the checks a calculator has no inputs for - the sample-ratio chi-square, the whole-week duration check, the cost of every interim look, and a Holm correction across the guardrail family.
Do I have to sign in?
No. The parsing, every statistical test, the intervals, the power and sample-size arithmetic, the sample-ratio check, the duration and peeking checks, the Holm correction, the decision rule, the complete readout, the results CSV, the measurement JSON and every validity check run in this tab with no account and nothing charged. Only the judgement pass is metered, and the three bundled examples replay a saved run for free.
Why is a null result treated as two different outcomes?
Because it is two different outcomes. A test that was large enough to find the effect it was designed for and did not find it has learned something: stop. A test that could only ever have detected an effect five times larger than the one it was looking for has learned nothing about the question it asked, and calling that a null is the single most expensive mistake in experimentation. The engine separates them by comparing the smallest detectable effect against the declared MDE, and says which one this is.
What is checked about the AI pass?
Every measured metric must get exactly one reading - a metric read twice fails as loudly as one left out, so the partition is counted rather than sampled. Each reading is compared against the measured significance, so calling a p = 0.18 result a win fails. The pass's stated decision is set beside the one the rule reaches from the same facts, and where they disagree both are shown rather than merged. A metric the engine itself could not test is excluded from the partition, so the pass is never blamed for a gap the browser already reported.
What does a run cost, and what if my balance is low?
A worst-case amount is reserved and only what the run uses is charged. Before you submit, the estimate is checked against your balance and the answer is shown at the run button rather than after it: signed out, it says to sign in and why; under the model's minimum, it names the exact shortfall and offers a top-up; and a reply cut short by a low balance says so rather than presenting itself as a complete readout.
Where does the readout go afterwards?
Out of the browser, with its audit attached. The readout downloads as markdown for the decision doc, every metric, arm, design figure, validity check and every line of the pass's verification as a CSV for a spreadsheet, and the whole measurement as JSON. Readouts are saved to your account, so the re-read after an extended test starts from the last one with the numbers back in the boxes, and any past readout can be set against what is on screen as a delta on the p-value, the effect and the decision.
A derived work of @phuryn/ab-test-analysis: its method — validate the setup before reading the result, check sample size and duration and randomization, compute significance and the interval, check the guardrails, then translate all of it into a ship / extend / stop decision — is what this app implements and measures. The skill asks a model to do the statistics; this computes them and holds the model to them.