New: the nine ways developer tools go invisible in AI answers.
citon

Measurement

Whether any of this can be measured at all, what a number has to survive before it means anything, and why a single visibility score is noise.

Short answer

What does Measurement cover?

Whether any of this can be measured at all, what a number has to survive before it means anything, and why a single visibility score is noise.

Why a level reading is not a measurement

Most writing about AI visibility reports a level. A score, a share of voice, a rank, read once and pasted into a slide. The trouble starts when you read it twice.

Twelve money queries were run five times each, on one model, on the same day, with nothing changed between runs. Seven of the twelve came back with a different outcome. A single reading of where a company stands is not a weak measurement. It is a coin toss with a decimal point on it.

What survives the noise

A difference between two arms survives where a level does not. Across 20,000 random splits with no work applied at all, the null difference centred on zero, which is what makes a treated group readable against a held out control group when one number on its own is not.

Sample size then decides whether that difference means anything. A twelve by five design has a minimum detectable lift of 51.1 percentage points, which is another way of saying it can detect nothing worth paying for. Forty queries by forty samples brings that to 9.9.

The posts collected here run along both edges of that. Some take apart the tools people buy to answer this question, what they charge, and the point at which a tool on its own stops being able to say whether a rising number was you or the model changing its mind. The rest set out method: what a query has to survive before it counts, why the split between the two arms is written down before any baseline is taken, and which assumptions in that design are still unproven rather than quietly buried.

Prompt Tracking, How Many Prompts and How Many Repeats Before a Change Is Real

Every page ranking for prompt tracking says how to pick prompts. None says how many, or how many repeats, before a weekly change stops being noise.

44 min

How to Track AI Visibility, and How to Read What It Tells You

The trackers solved collection and left interpretation alone. The record schema, the sampling trade, the null distribution, and whether your design could see a win.

41 min

AirOps Alternatives: Content-AEO vs Pure Measurement

Nine real AirOps alternatives, split into content production and pure measurement, plus the buyer split most vendor comparison pages miss entirely.

25 min

Best AI Visibility Tools (and Where a Tool Alone Runs Out)

Seven AI-visibility tools compared on what they actually measure and what they cost, plus the one question none of them can answer on their own.

24 min

Profound Alternatives: What Teams Actually Switch To (and Why)

Five real alternatives to Profound compared on what they track and what they cost, built from the actual switching threads where teams explain why they left.

23 min

Profound Pricing: What an AI Visibility Tool Costs

Profound publishes two tiers and hides the third. What Starter and Growth gate, and what Enterprise costs per brand per country, per one practitioner's report.

23 min

Why a single AI visibility score is noise

Twelve money queries, five identical repeats, one model, one day. Seven changed their answer, which breaks every before-and-after published in this category.

24 min