In a proper experiment you assign subjects randomly and the assignment guarantees the groups are alike. On a live website you cannot do that: the pages you are fixing are the pages that have the defect, and that is not a random property. Deep pages have thin content more often. Old templates lack structured data more often.
So the honest framing is a quasi-experiment. You are not claiming the groups were assigned at random. You are claiming they are comparable on the things that matter, that they were selected before the outcome was known, and that whatever moved the market moved both.
Choosing the set
- Comparable demandSimilar impression volume over the 28 days before the change. A control set of pages nobody visits cannot tell you anything about pages people do.
- Comparable position in the siteSimilar crawl depth and, where possible, the same section. A homepage does not control for a blog post.
- Currently healthyIndexable, returning 200s, not already mid-migration. A control that is broken in its own way is measuring its own recovery.
- Genuinely untouchedIf the same deploy also changed the control pages, there is no control. This is the failure mode that produces the most confident wrong answers.
- FrozenChosen at the moment of the change and never recomputed. A set that is re-selected later is being selected partly on the outcome, which quietly guarantees the result you were hoping for.
Reading the result
Compute the change in the metric for the treated pages and the change for the control pages over the same window, then subtract. That difference of differences is your effect. If both rose by the same amount, your effect is zero and the rise was the market.
Two thresholds then decide the verdict: how large a difference counts as an effect, and how large counts as decisive enough to call early. Both are judgement calls, and anyone who tells you their thresholds are derived from first principles is selling something. Ours are written down as invented, with a note to validate them against real data before they carry weight.
The most common invalid control set is 'the rest of the site'. On a sitewide template change, the rest of the site was also changed. On a seasonal business, the rest of the site has a different season. Comparability is a claim you have to be able to defend, not a default.
Is this the same as an SEO split test?
It is the practical cousin. Split testing usually means randomising pages within a template and serving variants; a frozen control set works after the fact on a change that has already been made, which is the situation an agency is actually in.
What if there is no adequate control set?
Say so and do not issue a verdict. On a small site or a sitewide change this is common. Recording 'landed, not measurable' is honest and still useful, because the dates are in the ledger when someone asks later.
Try it on your own site, free
No card, and the free tier does not expire — the first verdict takes about six weeks to arrive.