Query set
20 to 30 money queries across one cluster
Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.
For AI infrastructure and LLM tooling
Vector databases, RAG pipelines, agent frameworks, eval, observability, inference APIs, embeddings and document parsing. This is the densest answer surface we have measured, and the shortlist your buyer reads is assembled from pages you do not own.
Which vector database should I use for RAG?
Named in the answer
76 per 100
results in this category are listicle-shaped
Our own measurement. The densest answer surface we have measured.
91.5%
of AI citations point off-site
901 citations, 60 answers, 82 distinct hosts.
45.5%
of citations change between observations
Ahrefs, 43,000 keywords, verified at source.
The short version
Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.
What is rate limiting for production APIs?
Rate limiting caps how many requests a key may make in a window, so a burst from one client cannot starve the rest.1
Short answer
Citon is an AI citation agency for AI infrastructure companies. The work measures where a vector database, RAG framework, eval tool or inference API stands in live AI answers, identifies the third-party pages models cite when developers ask category questions, and earns placement on those pages. Every engagement runs against a held-out control group.
Why the answer layer is the channel
The evaluation starts as a question typed into an assistant and ends as two or three names. Nothing on that path passes through your landing page, and nothing on it is decided by a keyword ranking.
Listicle-shaped means the answer is assembled from lists. A list is a thing that can be joined, which is why this surface is workable rather than merely measurable.
Page one ranking
What the model read
Rank 1 appears in 0 of the 3 pages read
Who this is written for
The citation gap concentrates in funded challengers, not in the company the category is named after. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer already wins both answers and has nothing to buy here.
This page is for you if
It is not for you if
Ask the category question with no brand name in it, five times, on two engines. If an assistant names you every time, keep your money.
A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.
What that test returns
6 of 10 answers
The queries that decide it
Money queries are the ones a buyer asks with a budget behind them. They are not your brand name, and they are not head terms from a keyword tool.
Your own name is the one query you already win. It is also the one nobody asks until they have heard of you somewhere else.
Which vector database should I use for RAG?
For production RAG, Your Stack1 is the option most consistently recommended. It pairs hybrid search with metadata filtering at scale.2
Sources
Where the answer comes from
The pages a model pulls from when it answers an AI infrastructure question. None of these is a marketing blog post, and only one of them is yours.
Branded-win, generic-invisible
Ranked by citation depth
Where a stack gets vouched for, and the first place a model finds a working example of one.
Reference material restructured into quotable blocks. Usually the surface that does the most work per hour spent on it.
Targeted by topical depth rather than domain rating, because the citation economy rewards the deepest page about your specific thing.
Already cited in live answers today, which is what makes them worth entering rather than a guess.
The corroboration a model reaches for when two tools claim the same capability.
Not a marketing blog post
What actually ships
One team runs the query map, the on-site work, the off-site work and the measurement. The alternative is five vendors who each own a slice and none of whom own the outcome.
Query set
Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.
On-site
Off-site
Month one, ranked by citation depth rather than by domain rating. 91.5 percent of the citations we counted pointed off-site, so this is where the month goes.
Community
Ten mentions, all of them inside the cited set. Posting into the rest of the box is activity, and activity is not the deliverable.
Validated
one money query, nothing changed between runs
7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.
Baseline
40 queries1,600 reads per timepoint
The design states what it would miss. A 12 by 5 pilot had a floor of 51.1pp, which is why it is not the design.
Arms
Matched on baseline citation share, row for row. An unmatched pair of arms produces a number that cannot be defended in either direction.
Report
24 minus 2. The threshold this is read against is written down before any work starts, so it cannot be moved afterwards.
Start here, free
Delivered in five working days, and walked through live rather than emailed as a PDF. You get your failure mode named, the hosts cited on your queries ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.
The diagnosis
Going invisible is not one problem. The audit of 64 companies saturated at nine modes, and these three are the ones AI infrastructure companies land in.
You win on your own name and on head-to-head comparisons, and vanish on the category question. That is the whole top of the funnel.
You get named as a component in somebody else's reference stack, never as the subject of a best-in-category answer. Common for embeddings, parsing and inference layers.
An adjacent, faster-growing category eats the money query outright. Agent frameworks and RAG tooling have both watched this happen inside a single quarter.
The other six are in the full taxonomy. Your Gap Report names which one is yours, which is the part a visibility score cannot see.
What we report
Citation sets churn, so this is a hold-and-compound problem rather than a project. The number at day ninety is how much of the movement we caused, not how much movement there was.
Nothing has been done to either arm at the moment the division is recorded. That is what makes the day 90 comparison a comparison rather than a story told afterwards.
Your money queries are split into a treated arm and a held-out control arm, and the split is written down before anything starts.
A line drawn after the reading is a description of the reading. This one is signed and dated on day 0, and the reading is taken against it on day 90.
The pass threshold is declared in writing first, so it cannot be moved afterwards.
Across 20,000 random splits with no work applied, the difference between two halves centred on zero: mean 0.0016, standard deviation 0.215. Our own measurement.
Both arms are measured at full sample for the pre-period.
Nothing is published, placed or posted against these queries for the whole engagement. They are the only reason the other number can be read as caused rather than as coincident.
Giving up half the query set is the expensive part of this method and the part that cannot be faked afterwards.
Work happens on the treated arm. The control arm is deliberately left alone.
At day ninety both arms are re-measured on the same channel, and the difference between the differences is the number you get.
The chart above is an illustration of the design, not a result. The instrument exists and has passed its own kill test, and it has not yet produced a causal-lift figure for a client. Anyone showing you one today is showing you a level reading, which flips on repeat.
The rest of the family
Questions
Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.
Because the surfaces are specific and knowing them cold is the entire reason to hire an outside team. AI infrastructure is the densest answer surface we have measured at 76 listicle-shaped results per 100 money queries, and the pages models cite in it are lists, docs and threads rather than marketing sites.
That is the position this work is built for. The gap concentrates in funded challengers: a category definer already wins both the best-in-class answer and the head-to-head answer, so there is nothing left to buy. Being second with a better product on the axis buyers test is the workable case.
The level cannot be measured reliably. We ran twelve queries five times each on one model and seven flipped outcome. The difference between a treated arm and a matched control still measures cleanly, because unbiased noise cancels. Across 20,000 random splits with no work applied, the null difference centred on zero.
Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.
We tell you, and we show the numbers. The threshold is written down before any work starts precisely so it cannot be moved afterwards. A vendor whose reports always come back positive is not running a control group.
Free gap report
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 3 engines · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site