Query set
20 to 30 install-intent queries
Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.
For CLIs and MCP servers
Command-line tools, MCP servers, agent-callable APIs and the manifests that describe them. This is the sharpest version of the problem in the whole market, because the model that recommends your server is the same model that decides whether to call it.
Which MCP server should I install for Postgres?
Named in the answer
91.5%
of AI citations point off-site
901 citations, 60 answers, 82 hosts. Our own measurement.
7 of 12
queries flipped outcome across identical repeats
One model, five repeats each. A single read is a coin flip.
9.9pp
minimum lift the instrument can detect
At 40 queries by 40 samples. Anything smaller is unproven.
The short version
Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.
What is rate limiting for production APIs?
Rate limiting caps how many requests a key may make in a window, so a burst from one client cannot starve the rest.1
Short answer
Citon is an AI citation agency for command-line tools and MCP servers. The work measures whether assistants name a given server when developers ask what to install, extracts the registries, lists and threads those answers cite, and earns placement across them. For agent-callable tools, being named and being called are the same event.
Why this is the sharpest version of the problem
Everywhere else in this market those are two budgets. A model that names your server in an answer is a model that can install it and call it in the next turn, so one piece of work buys the recommendation and the invocation together.
That is why the same sentence has to serve two readers. The developer reads it to decide what to install. The agent reads it to decide what to call. Written once, correctly, it does both.
Page one ranking
What the model read
Rank 1 appears in 0 of the 3 pages read
Who this is written for
The citation gap concentrates in funded challengers, not in the server everybody already reaches for. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer already wins both answers and has nothing to buy here.
This page is for you if
It is not for you if
Ask an assistant which server to install for your integration, five times, on two engines. If yours comes back every time, there is nothing here worth paying for.
A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.
What that test returns
6 of 10 answers
The queries that decide it
Install-intent queries are the money queries here. They are asked by a developer, and increasingly by an agent acting for one, and they are answered from third-party lists.
The last row is the one that matters most and the one nobody optimises for, because it has no product name in it and no keyword tool reports it.
Which MCP server should I install for Postgres?
For read-only Postgres access, Your Server1 is the one most consistently recommended. It exposes schema introspection and parameterised queries.2
Sources
Where the answer comes from
The pages an answer about a CLI or an MCP server is actually assembled from. The last one is unique to this vertical, and it is the reason the two halves collapse into one.
Branded-win, generic-invisible
Ranked by citation depth
The canonical index a model consults before it recommends anything installable. Placement here is a citation and a discovery path at once.
The one owned surface a model quotes verbatim. It has to answer what the tool is, what it needs and what one command does, above the fold.
Where a tool gets vouched for by somebody with no stake in it, which is the corroboration models weight most heavily.
Targeted by topical depth rather than domain rating. In this vertical the deepest page is often a personal blog with fifty readers and a citation on every engine.
The manifest an agent reads at selection time. Same words, second reader, and the only place on the internet where copy is executed rather than read.
Not a marketing blog post
What actually ships
One team runs the query map, the registry and list work, the manifest pass and the measurement, because splitting them across vendors puts the recommendation and the invocation in two different budgets again.
Query set
Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.
Listings
Registry entry
npx yourtool.example
Listed
Awesome list row
Pull request open
Both are tracked to a state, including the one still open. A listing nobody re-checks is a listing that quietly disappears.
Manifest
Written for a human
A fast, friendly way to work with your data.
Written to be selected
Reads a table and returns the rows matching a filter. Use when the question asks for records by a field value.
The docs page is never read at selection time. The description is the whole interface, and most tools ship the first version.
Off-site
Month one, ranked by citation depth rather than by domain rating. 91.5 percent of the citations we counted pointed off-site, so this is where the month goes.
Validated
one money query, nothing changed between runs
7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.
Baseline
40 queries1,600 reads per timepoint
The design states what it would miss. A 12 by 5 pilot had a floor of 51.1pp, which is why it is not the design.
Arms
Matched on baseline citation share, row for row. An unmatched pair of arms produces a number that cannot be defended in either direction.
Report
24 minus 2. The threshold this is read against is written down before any work starts, so it cannot be moved afterwards.
Start here, free
Delivered in five working days, and walked through live rather than emailed as a PDF. You get your failure mode named, the hosts cited on your queries ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.
The diagnosis
Going invisible is not one problem. The audit of 64 companies saturated at nine modes, and these three are where command-line and agent-callable tools land.
You are named as a component inside somebody else's setup guide, never as the subject of the question about what to install. Endemic to servers that ship as part of a larger stack.
A rename, a fork, or a name shared with an unrelated package splits your equity across two entities the model never merges. Registries make this worse, because both entries stay live.
An adjacent category eats the install question outright, and the framing goes with it. The phrasing for agent tooling has already turned over twice.
The other six are in the full taxonomy. Your Gap Report names which one is yours, which is the part a visibility score cannot see.
What we report
The outcome here is a share of the answers on install-intent queries, held over time. The number at day ninety is how much of that movement we caused, not how much movement there was.
Nothing has been done to either arm at the moment the division is recorded. That is what makes the day 90 comparison a comparison rather than a story told afterwards.
Your money queries are split into a treated arm and a held-out control arm, and the split is written down before anything starts.
A line drawn after the reading is a description of the reading. This one is signed and dated on day 0, and the reading is taken against it on day 90.
The pass threshold is declared in writing first, so it cannot be moved afterwards.
Across 20,000 random splits with no work applied, the difference between two halves centred on zero: mean 0.0016, standard deviation 0.215. Our own measurement.
Both arms are measured at full sample for the pre-period.
Nothing is published, placed or posted against these queries for the whole engagement. They are the only reason the other number can be read as caused rather than as coincident.
Giving up half the query set is the expensive part of this method and the part that cannot be faked afterwards.
Work happens on the treated arm. The control arm is deliberately left alone.
At day ninety both arms are re-measured on the same channel, and the difference between the differences is the number you get.
The chart above is an illustration of the design, not a result. The instrument exists and has passed its own kill test, and it has not yet produced a causal-lift figure for a client. Anyone showing you one today is showing you a level reading, which flips on repeat.
The rest of the family
Questions
Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.
Because the reader is the same system. A model that names your server in an answer is the model that decides whether to call it in the next turn, and both decisions run off the same third-party pages plus the same tool descriptions. Every other vertical has to fund those two outcomes separately.
A registry entry makes you findable, not cited. Across 901 citations we counted, 91.5 percent pointed off-site across 82 distinct hosts, so an answer about what to install is assembled from lists, threads and comparisons rather than from one index. Being present is the floor, not the outcome.
The level cannot be measured reliably. We ran twelve queries five times each and seven flipped outcome. The difference between a treated arm and a held-out control still measures cleanly, because unbiased noise cancels. Our design detects a 9.9 percentage point lift at 40 queries by 40 samples.
Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.
Yes, and they are treated as product surface rather than as copy. A tool name and description are what an agent reads at selection time, so they are written to be unambiguous against neighbouring tools first and readable second. That pass is part of the work, not an add-on.
Free gap report
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 3 engines · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site