New: the nine ways developer tools go invisible in AI answers.
citon

For CLIs and MCP servers

For a CLI or an MCP server, getting cited and getting chosen are one job.

Command-line tools, MCP servers, agent-callable APIs and the manifests that describe them. This is the sharpest version of the problem in the whole market, because the model that recommends your server is the same model that decides whether to call it.

Which MCP server should I install for Postgres?

Named in the answer

1incumbent.example
2roundup.example
3guides.example
cut
younot named
  • MCP servers
  • Command-line tools
  • Agent-callable APIs
  • Tool manifests
  • Local dev tooling
  • Registry listings
  • Terminal-first products
  • Agent frameworks integrations

91.5%

of AI citations point off-site

901 citations, 60 answers, 82 hosts. Our own measurement.

7 of 12

queries flipped outcome across identical repeats

One model, five repeats each. A single read is a coin flip.

9.9pp

minimum lift the instrument can detect

At 40 queries by 40 samples. Anything smaller is unproven.

The short version

The whole offer, in one paragraph a model can quote.

Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.

Your page54 words

What is rate limiting for production APIs?

AQuoted back, attributed

Rate limiting caps how many requests a key may make in a window, so a burst from one client cannot starve the rest.1

Short answer

What does an AI citation agency do for a CLI or MCP server?

Citon is an AI citation agency for command-line tools and MCP servers. The work measures whether assistants name a given server when developers ask what to install, extracts the registries, lists and threads those answers cite, and earns placement across them. For agent-callable tools, being named and being called are the same event.

Why this is the sharpest version of the problem

Being cited in an answer and being chosen by an agent are the same commercial event.

Everywhere else in this market those are two budgets. A model that names your server in an answer is a model that can install it and call it in the next turn, so one piece of work buys the recommendation and the invocation together.

That is why the same sentence has to serve two readers. The developer reads it to decide what to install. The agent reads it to decide what to call. Written once, correctly, it does both.

Page one ranking

1yourapi.example
2competitor.example
3roundup.example
4guides.example

What the model read

reddit.com
g2.com
news.ycombinator.com
nothing else

Rank 1 appears in 0 of the 3 pages read

Who this is written for

Written to the company losing the answer, not to the one winning it.

The citation gap concentrates in funded challengers, not in the server everybody already reaches for. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer already wins both answers and has nothing to buy here.

This page is for you if

  • An established server takes the answer for your integration, and yours is named only after a buyer has been told your name elsewhere.
  • You are listed in a registry and almost never cited, so discovery is happening without you in the room.
  • You are funded, seed to Series C, with an outside marketing budget aimed at a landing page nobody in this path visits.
  • Your tool descriptions were written for a human reading docs, and an agent is now the thing reading them.

It is not for you if

  • Yours is the server the category is described with, and both the category answer and the head-to-head answer already return it.
  • You want the visibility number on its own. That part is free and needs no retainer.
  • The tool is internal, or unreleased, with no public question being asked about it yet.
  • You need proof of lift inside a quarter, which is shorter than a control group can speak.

Ask an assistant which server to install for your integration, five times, on two engines. If yours comes back every time, there is nothing here worth paying for.

A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.

What that test returns

Citation rate

6 of 10 answers

The queries that decide it

The answer is written before your buyer reaches your site.

Install-intent queries are the money queries here. They are asked by a developer, and increasingly by an agent acting for one, and they are answered from third-party lists.

best mcp server for postgresincumbent
mcp servers for github issuesregistry
yourcli.example alternativesyou
cli for tailing kubernetes logsincumbent
which mcp server should i installunclaimed

The last row is the one that matters most and the one nobody optimises for, because it has no product name in it and no keyword tool reports it.

AAssistant···

Which MCP server should I install for Postgres?

For read-only Postgres access, Your Server1 is the one most consistently recommended. It exposes schema introspection and parameterised queries.2

Sources

reddit.comg2.comnews.ycombinator.com

Where the answer comes from

Five surfaces decide it, and four of them are not yours.

The pages an answer about a CLI or an MCP server is actually assembled from. The last one is unique to this vertical, and it is the reason the two halves collapse into one.

Cited sourcesMode 01
reddit.com38%
g2.com26%
news.ycombinator.com21%
yourapi.example15%

Branded-win, generic-invisible

Ranked by citation depth

01

Registries and awesome-lists

The canonical index a model consults before it recommends anything installable. Placement here is a citation and a discovery path at once.

02

Your README and install block

The one owned surface a model quotes verbatim. It has to answer what the tool is, what it needs and what one command does, above the fold.

03

Community threads

Where a tool gets vouched for by somebody with no stake in it, which is the corroboration models weight most heavily.

04

Roundups and comparisons

Targeted by topical depth rather than domain rating. In this vertical the deepest page is often a personal blog with fifty readers and a citation on every engine.

05

Tool names and descriptions

The manifest an agent reads at selection time. Same words, second reader, and the only place on the internet where copy is executed rather than read.

Not a marketing blog post

What actually ships

The work, and the thing that proves the work moved the number.

One team runs the query map, the registry and list work, the manifest pass and the measurement, because splitting them across vendors puts the recommendation and the invocation in two different budgets again.

Query set

20 to 30 install-intent queries

One cluster, mappedSample
01Category questionIncumbent
02Head to headIncumbent
03Constraint questionUnclaimed
04Task questionUnclaimed
26 queries, one cluster

Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.

Listings

Registry and awesome-list placement, tracked

Where an agent looksSample
AAgent resolving a tool for the task

Registry entry

npx yourtool.example

Listed

Awesome list row

Pull request open

Both are tracked to a state, including the one still open. A listing nobody re-checks is a listing that quietly disappears.

Manifest

Tool names and descriptions written for selection

One capability, two descriptionsSample

Written for a human

A fast, friendly way to work with your data.

Written to be selected

Reads a table and returns the rows matching a filter. Use when the question asks for records by a field value.

Selected on the description alone

The docs page is never read at selection time. The description is the whole interface, and most tools ship the first version.

Off-site

6 placements a month on hosts already cited

Six placementsSample
Reddit
Hacker News
G2
Stack Overflow

Month one, ranked by citation depth rather than by domain rating. 91.5 percent of the citations we counted pointed off-site, so this is where the month goes.

Validated

5 identical repeats before one query counts

One query, five identical runsSample

one money query, nothing changed between runs

Named in 3 of 52 flips

7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.

Baseline

40 queries by 40 samples, 9.9pp floor

40 queries by 40 samplesOur design

40 queries1,600 reads per timepoint

not detectable9.9pp and up

The design states what it would miss. A 12 by 5 pilot had a floor of 51.1pp, which is why it is not the design.

Arms

Treated, plus a held-out control group

Two arms, matchedSample
Treated 15Held out 15
34%34%
21%21%
12%12%
6%6%
Registered before either arm is touched

Matched on baseline citation share, row for row. An unmatched pair of arms produces a number that cannot be defended in either direction.

Report

Difference-in-differences at day 90

Difference in differencesIllustration
Treated armDay 0 to 90
31%55%+24pp
Held out armDay 0 to 90
30%32%+2pp
Caused by the work22ppThreshold set day 0

24 minus 2. The threshold this is read against is written down before any work starts, so it cannot be moved afterwards.

Start here, free

The Gap Report: 12 queries, 5 repeats each, 3 engines.

Delivered in five working days, and walked through live rather than emailed as a PDF. You get your failure mode named, the hosts cited on your queries ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.

Apply now

The diagnosis

Three of the nine modes account for most of this vertical.

Going invisible is not one problem. The audit of 64 companies saturated at nine modes, and these three are where command-line and agent-callable tools land.

Mode 04

Default by inclusion

You are named as a component inside somebody else's setup guide, never as the subject of the question about what to install. Endemic to servers that ship as part of a larger stack.

Mode 06

Identity orphaned

A rename, a fork, or a name shared with an unrelated package splits your equity across two entities the model never merges. Registries make this worse, because both entries stay live.

Mode 09

Category-query absorption

An adjacent category eats the install question outright, and the framing goes with it. The phrasing for agent tooling has already turned over twice.

The other six are in the full taxonomy. Your Gap Report names which one is yours, which is the part a visibility score cannot see.

Failure modes9 in the taxonomy
010203040506070809
Default by inclusionYours

What we report

Causal lift is the number. Not a score, and not activity.

The outcome here is a share of the answers on install-intent queries, held over time. The number at day ninety is how much of that movement we caused, not how much movement there was.

  1. The order, drawnSample
    30 validated queriestreated 15held out 15validatedsplit, day 0work starts
    Which queries sit in which arm is written down at the split and not reopened.

    Nothing has been done to either arm at the moment the division is recorded. That is what makes the day 90 comparison a comparison rather than a story told afterwards.

    01

    Split

    Your money queries are split into a treated arm and a held-out control arm, and the split is written down before anything starts.

  2. The line, set on day 0Illustration
    0102030declared +12ppday 90reading dueday 0day 90
    Signed before work startsCannot be moved

    A line drawn after the reading is a description of the reading. This one is signed and dated on day 0, and the reading is taken against it on day 90.

    02

    Threshold

    The pass threshold is declared in writing first, so it cannot be moved afterwards.

  3. Pre-period, both arms at full depthIllustration
    020406031.2%Treated30.8%Held out
    Gap at day 00.4ppInside the band

    Across 20,000 random splits with no work applied, the difference between two halves centred on zero: mean 0.0016, standard deviation 0.215. Our own measurement.

    03

    Baseline

    Both arms are measured at full sample for the pre-period.

  4. Where the month landsSample
    Treated arm22 items, month 1
    Answer pagePlacementMentionAnswer pagePlacementMentionand 16 more
    Held-out arm0 items, 90 days

    Nothing is published, placed or posted against these queries for the whole engagement. They are the only reason the other number can be read as caused rather than as coincident.

    Giving up half the query set is the expensive part of this method and the part that cannot be faked afterwards.

    04

    Treatment

    Work happens on the treated arm. The control arm is deliberately left alone.

  5. Design, drawnIllustration
    Causal lift+24pp
    Treated ControlDay 0 to 90
    05

    Readout

    At day ninety both arms are re-measured on the same channel, and the difference between the differences is the number you get.

The chart above is an illustration of the design, not a result. The instrument exists and has passed its own kill test, and it has not yet produced a causal-lift figure for a client. Anyone showing you one today is showing you a level reading, which flips on repeat.

Questions

The five we get asked in this category.

Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.

Why do citation work and agent adoption collapse into one job here?

Because the reader is the same system. A model that names your server in an answer is the model that decides whether to call it in the next turn, and both decisions run off the same third-party pages plus the same tool descriptions. Every other vertical has to fund those two outcomes separately.

We are already in the main registry. Why are we not being recommended?

A registry entry makes you findable, not cited. Across 901 citations we counted, 91.5 percent pointed off-site across 82 distinct hosts, so an answer about what to install is assembled from lists, threads and comparisons rather than from one index. Being present is the floor, not the outcome.

Can this be measured when the answers change every time?

The level cannot be measured reliably. We ran twelve queries five times each and seven flipped outcome. The difference between a treated arm and a held-out control still measures cleanly, because unbiased noise cancels. Our design detects a 9.9 percentage point lift at 40 queries by 40 samples.

What is in the free Gap Report?

Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.

Do you write the tool descriptions as well?

Yes, and they are treated as product surface rather than as copy. A tool name and description are what an agent reads at selection time, so they are written to be unambiguous against neighbouring tools first and readable second. That pass is part of the work, not an add-on.

Free gap report

When the model answers,
be the one it names

Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.

12 queries · 5 repeats each · 3 engines · 5 working days

Free · Walked through live · Five working days

Gap ReportSample
12 queries · 5 repeats each

Queries we run

best rate limiting api
your-api alternativesYou, 1 of 12
cheapest webhook api
api gateway for startups

Cited instead of you

reddit.com38%
g2.com26%
news.ycombinator.com21%

91.5% of citations point off-site

Your failure mode, named
Hosts ranked by citation depth
The pre-registered baseline