New: the nine ways developer tools go invisible in AI answers.
citon

Answer engine optimization

The answer to your category question is already written. Today it names somebody else.

We measure where you stand in live AI answers, show you the exact pages the models cite when your buyers ask, and go and get you onto them. Every engagement runs against a held-out control group, so at day ninety you get a causal number rather than a chart that went up.

The answer they already getSample

“which rate limiting api should we use?”

1The incumbentCited
2A second vendorCited
3A third vendorCited
yourapi.exampleNot on the list
  • Live-answer measurement
  • Money-query mapping
  • Cited-source extraction
  • Owned answer pages
  • Off-site placement
  • Review-site presence
  • Agent readiness
  • Causal-lift reporting

91.5%

of AI citations point off-site

901 citations, 60 answers, 82 distinct hosts. Our own measurement.

7 of 12

queries flipped outcome across identical repeats

One model, five repeats each, same day, nothing changed.

9.9pp

smallest lift the instrument can detect

At 40 queries by 40 samples. Anything smaller is unproven.

The short version

The whole service, in one paragraph a model can quote.

Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.

Checked against the capsule on this page58 words
Defines the term in sentence oneYes
Between 40 and 60 words58
No pronoun pointing outside itselfNone
No marketing framingNone

The count is measured from the rendered string, so this pane cannot disagree with the paragraph beside it.

Short answer

What is answer engine optimization?

Answer engine optimization is the practice of earning a place in the answers AI assistants generate, rather than in a ranked list of links. Citon measures where a company stands in live AI answers, identifies the pages models cite when buyers ask category questions, earns placement on those pages, and runs every engagement against a held-out control group.

Why the score is the free part

A visibility score is a level reading, and a level reading flips on repeat.

Twelve money queries were run five times each on one model, on the same day, with nothing changed between runs. Seven of the twelve flipped outcome. Any single read of where you stand is therefore a coin flip wearing the clothes of a metric, which is why the score is handed over free and the charge sits somewhere else.

What survives the noise is a difference. Across 20,000 random splits with no work applied at all, the null difference centred on zero, so unbiased noise cancels between a treated arm and a matched control. That finding is what the whole engagement rests on, and it is evidenced rather than argued.

20,000 random splits, no work appliedOur data
zerocontrol aheadtreated ahead
Mean 0.0016SD 0.215Splits 20,000

Unbiased noise cancels between a treated arm and a matched control. That is what the whole engagement rests on, and it is evidenced rather than argued.

What the instrument does

Repeats
5 identical runs before one query counts
Depth
40 queries by 40 samples at baseline
Floor
9.9pp, the smallest lift it can detect
Arms
Treated, plus a held-out control

Who this is written for

Written to the company losing the answer, not to the one winning it.

The citation gap concentrates in funded challengers rather than in the company the category is named after. In a narrow category the best-in-class question and the head-to-head question land on the same one or two names, so the definer already wins both answers and has nothing to buy here.

This page is for you if

  • An incumbent takes the category answer, and you are named only when a buyer already types your name.
  • You are better on the axis your buyers actually test, and the answer has not caught up.
  • You are funded, seed to Series C, with an outside marketing budget producing work nobody can attribute.
  • Your category question returns a shortlist, and you are not on it.

It is not for you if

  • You are the name the category gets described with, and both the category answer and the head-to-head answer already return you.
  • You want the score rather than the work behind it. The score is the free part.
  • Your buying runs through procurement, pilots and design-in cycles rather than through an assistant.
  • You want a lift number faster than ninety days, which is shorter than a control group takes to say anything.

Ask the category question with no brand name in it, five times, on two engines. If an assistant names you every time, keep your money.

A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.

Two questions, one nameSample

best rate limiting api

The incumbentBoth times

incumbent vs alternatives

The incumbentBoth times
The gap is in the company losing both answers, not the one winning them.

A definer already holds both slots, so there is nothing here to buy. That is the qualifying step, and it runs before a call rather than on one.

The four disciplines

How a query becomes a citation, and how we prove we caused it.

One team runs the measurement, the diagnosis, the on-site and off-site work and the report. The alternative is four vendors who each own a slice and none of whom own the outcome.

12 queries, 5 identical repeatsSample
best rate limiting apiStable
api gateway for startupsFlipped
cheapest webhook apiStable
how to handle 429sFlipped
webhook delivery at scaleFlipped
rate limit vs quotaStable
idempotency key apiFlipped
best api observabilityStable
token bucket libraryFlipped
api key rotation toolsFlipped
usage based billing apiStable
burst traffic protectionFlipped

7 of 12 queries changed outcome across identical repeats, same model, same day, nothing altered between runs. 60 of 60 calls succeeded, so this is the answer moving rather than the harness failing.

01

Find out where you actually stand

Money queries are built and disambiguated first, then run at depth against live answers rather than against a source proxy. A query counts once it has survived five identical repeats, because a single read flips: seven of our first twelve did. The treated and control split is written down before any baseline is taken.

  • Money-query mapping
  • Live-answer sampling
  • Five identical repeats
  • Pre-registered split
Why one read is not a reading

Twelve queries, five identical repeats, same day, nothing changed: seven flipped. All 60 calls succeeded, so this is the answer moving rather than the harness failing.

Our own measurement, step zero, 60 of 60 calls

Agent readiness, run end to endMode 08
DiscoverFound in the registrypass
InstallOne command, no manual steppass
AuthenticateKey flow needs a browserfail
Finish a taskNever reachedblocked

Mode 08 of nine, unreadable to an agent. For a CLI or an MCP server the run stopping here and the answer not naming you are the same commercial event.

02

Name the failure mode, and the pages behind it

Going invisible is not one problem. The audit of 64 companies saturated at nine failure modes, and yours is named rather than guessed at. Alongside it comes the cited-source list, a crawlability and citability audit, and an agent-readiness pass asking whether an agent can discover, install, authenticate and finish a real task.

  • Nine-mode classification
  • Cited-source extraction
  • Citability audit
  • Agent-readiness pass
Where the taxonomy came from

Sixty-four companies were mapped and the failure modes saturated at nine. Not one of them published a control group or a causal-lift number.

Our own audit, 64 companies mapped

The same hosts, ranked twiceSample

By domain rating

1GitHub
2Stack Overflow
3Reddit
4G2

By citation depth

Reddit9
Hacker News7
GitHub6
G24

Across 901 citations, 91.5 percent pointed off-site over 82 distinct hosts, so no one high-authority domain carries a category. Our own measurement.

03

Go and get onto the pages the answers cite

Across 901 citations we counted, 91.5 percent pointed off-site over 82 distinct hosts, so most of the month is spent on somebody else's property. Roundups, review sites, community threads and developer surfaces, targeted by citation depth rather than by domain rating, with owned answer pages carrying the half a model can quote directly.

  • Off-site placements
  • Owned answer pages
  • Review-site presence
  • Developer surfaces
Why depth beats domain rating

91.5 percent of 901 citations pointed off-site over 82 hosts, so no single high-authority domain carries a category. Depth beats rating.

Our own measurement, 82 distinct hosts

Difference in differencesIllustration
Treated armDay 0 to 90
31%55%+24pp
Held out armDay 0 to 90
30%32%+2pp
Caused by the work22ppThreshold set day 0

24 minus 2. The threshold this is read against is written down before any work starts, so it cannot be moved afterwards.

04

Report how much of the movement we caused

Work happens on a treated arm. A matched arm is held out and deliberately left alone, so the difference between the two differences at day ninety is how much of the movement we caused.

  • Held-out control arm
  • Threshold written first
  • Difference-in-differences
  • Reported at every tier
Why the difference survives the noise

Across 20,000 random splits with no work applied to either arm, the null difference centred on zero to within two parts in a thousand. Unbiased noise cancels between the arms, so what is left at day ninety is attributable rather than ambient.

Our own measurement, 20,000 permutation splits

How the instrument was tested

An instrument gets tested before it gets trusted. Here is the test.

Twelve money queries, five identical repeats each on one search-capable model, sixty of sixty calls returned. Then 20,000 random splits with no work applied to either arm, to find out whether a difference could be read out of pure noise. It could not: the null centred on zero. That is the finding the whole engagement rests on, and it is evidenced rather than argued.

measurement calls succeeded at step zero

60 of 60

measurement calls succeeded at step zero

Twelve money queries, five repeats each, one search-capable model.

random splits with no work applied

20,000

random splits with no work applied

The null difference centred on zero, so unbiased noise cancels.

smallest lift the design can detect

9.9pp

smallest lift the design can detect

At 40 queries by 40 samples. Anything smaller is unproven.

Design, drawnIllustration
Causal lift+24pp
Treated ControlDay 0 to 90

The chart is a drawing of the design rather than a result, and your engagement produces its own, read against a threshold written down before any work starts. A before-and-after reading with nothing held out is a different and weaker thing: seven of our first twelve queries flipped outcome across identical repeats on the same day, so a level read once is a coin flip wearing the clothes of a metric.

What ships

What we ship in an answer-engine engagement.

Counted per month rather than described, because a plan with no numbers in it is a plan nobody can hold anyone to, including us.

Cluster one, validated

best rate limiting apiIn
5 repeats, stable
api gateway for startupsIn
5 repeats, stable
cheapest webhook apiOut
5 repeats, flipped

Money queries, disambiguated and validated

20 to 30 across one cluster at the entry tier, 60 to 90 across three above it. A query counts once it has survived five identical repeats.

Shipped this month

6 of 6
Capsule first

What is rate limiting for production APIs?

Rate limiting caps how many requests a key may make in a window.

Owned answer pages

6 a month at the entry tier, 12 above it. Capsule first, structured for extraction rather than for browsing, which is a different craft from writing to rank.

Placed where the answers pull from

news.ycombinator.comDepth 7
github.comDepth 6
reddit.comDepth 9

Off-site placements on hosts already cited

6 a month at the entry tier, 15 above it, targeted by citation depth rather than by domain rating, because the deepest page about your specific thing usually wins.

Outside corroboration

2Category listings, kept current
g2 and capterra

Review-site presence

G2 and Capterra, because a model reaches for outside corroboration the moment two tools claim the same capability.

Four stages, end to end

Quarterly
Discover and installPassed
AuthenticateFailed
Finish a real taskBlocked

Agent-readiness pass

Can an agent discover, install, authenticate and finish a real task. Audited at the start, then re-checked quarterly.

Reported at every tier

Illustration
Arms2ThresholdDay 0Floor9.9pp

Causal-lift report at day ninety

Difference-in-differences against the held-out control arm, measured against a threshold written down before any work started so it cannot be moved afterwards.

Price
From $6,000 a month
Term
90-day minimum on every tier
Before that
A free Gap Report, walked through on a call
Report
The causal-lift number, at every tier

Start here, free

The Gap Report: 12 queries, 5 repeats each, 3 engines.

Delivered in five working days, and walked through live rather than emailed as a PDF. You get your failure mode named, the hosts cited on your queries ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.

Apply now

The rest of the family

One play, run in one niche, all the way down.

The other pages are not other services. They are this method pointed at a narrower set of buyers, which is why the instrument and the control arm are the same on all of them.

What the family shares
One measurement, one held-out control arm
Answer engineReddit and communityThree categories

Every page in this family is the same method pointed at a narrower set of buyers, which is why the pricing does not change per page.

Measuring lift, answer engine workA chart going up is not evidence that the work you paid for did it.Read it →

Questions

The five we get asked about this service.

Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.

Rendered against the graph6 of 6
Question
Question
Question
Question
Question
Question

The first row is the answer capsule, which is the same string in the page and in the graph, so the two cannot disagree.

How is this different from an AI visibility tool?

A tool reports a score. That part is handed over free, because monitoring is commoditised and it starts near a hundred dollars a month. What you pay for is the work across the surfaces models cite, plus a held-out control group that says how much of the movement the work caused.

Can AI citations be measured at all, given the answers move?

The level cannot be measured reliably. Twelve queries run five times each on one model produced seven flips. The difference between a treated arm and a matched control still measures cleanly, because unbiased noise cancels: across 20,000 random splits with no work applied, the null difference centred on zero.

What is the difference between AEO, GEO and AI SEO?

Very little in practice. Answer engine optimization, generative engine optimization and AI search optimization all name the same job: being present in a generated answer rather than in a ranked list of links. The naming argument matters less than whether the work is measured against something held out.

What is in the free Gap Report?

Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.

How quickly will we see something?

Diagnosis lands inside the first three weeks. The causal-lift report comes at day ninety, which is why the entry tier carries a ninety-day minimum. Anybody promising a measured citation lift faster than that is not running a control group, and without one there is no way to settle the argument in either direction.

Free gap report

When the model answers,
be the one it names

Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.

12 queries · 5 repeats each · 3 engines · 5 working days

Free · Walked through live · Five working days

Gap ReportSample
12 queries · 5 repeats each

Queries we run

best rate limiting api
your-api alternativesYou, 1 of 12
cheapest webhook api
api gateway for startups

Cited instead of you

reddit.com38%
g2.com26%
news.ycombinator.com21%

91.5% of citations point off-site

Your failure mode, named
Hosts ranked by citation depth
The pre-registered baseline