New: the nine ways developer tools go invisible in AI answers.
citon

For AI infrastructure and LLM tooling

Developers ask an assistant which vector database to use. It names someone else.

Vector databases, RAG pipelines, agent frameworks, eval, observability, inference APIs, embeddings and document parsing. This is the densest answer surface we have measured, and the shortlist your buyer reads is assembled from pages you do not own.

Which vector database should I use for RAG?

Named in the answer

1incumbent.example
2roundup.example
3guides.example
cut
younot named
  • Vector databases
  • RAG frameworks
  • Agent frameworks
  • Eval and testing
  • LLM observability
  • Inference APIs
  • Embeddings
  • Document parsing

76 per 100

results in this category are listicle-shaped

Our own measurement. The densest answer surface we have measured.

91.5%

of AI citations point off-site

901 citations, 60 answers, 82 distinct hosts.

45.5%

of citations change between observations

Ahrefs, 43,000 keywords, verified at source.

The short version

The whole offer, in one paragraph a model can quote.

Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.

Your page54 words

What is rate limiting for production APIs?

AQuoted back, attributed

Rate limiting caps how many requests a key may make in a window, so a burst from one client cannot starve the rest.1

Short answer

What does an AI citation agency do for an AI infrastructure company?

Citon is an AI citation agency for AI infrastructure companies. The work measures where a vector database, RAG framework, eval tool or inference API stands in live AI answers, identifies the third-party pages models cite when developers ask category questions, and earns placement on those pages. Every engagement runs against a held-out control group.

Why the answer layer is the channel

Your buyer is inside Claude Code or Cursor at the moment the shortlist gets made.

The evaluation starts as a question typed into an assistant and ends as two or three names. Nothing on that path passes through your landing page, and nothing on it is decided by a keyword ranking.

Listicle-shaped means the answer is assembled from lists. A list is a thing that can be joined, which is why this surface is workable rather than merely measurable.

Page one ranking

1yourapi.example
2competitor.example
3roundup.example
4guides.example

What the model read

reddit.com
g2.com
news.ycombinator.com
nothing else

Rank 1 appears in 0 of the 3 pages read

Who this is written for

Written to the company losing the answer, not to the one winning it.

The citation gap concentrates in funded challengers, not in the company the category is named after. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer already wins both answers and has nothing to buy here.

This page is for you if

  • An incumbent takes the category answer, and you are named only when a buyer already types your name.
  • You are better on the axis your buyers actually test, and the answer has not caught up.
  • You are funded, seed to Series C, with an outside marketing budget producing work nobody can attribute.
  • Your category question returns a shortlist, and you are not on it.

It is not for you if

  • You are the name the category gets described with, and both the category answer and the head-to-head answer already return you.
  • You want the score rather than the work behind it. The score is the free part.
  • Your buying runs through procurement, pilots and design-in cycles rather than through an assistant.
  • You want a citation lift number faster than ninety days, which is shorter than a control group takes to say anything.

Ask the category question with no brand name in it, five times, on two engines. If an assistant names you every time, keep your money.

A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.

What that test returns

Citation rate

6 of 10 answers

The queries that decide it

The answer is written before your buyer reaches your site.

Money queries are the ones a buyer asks with a budget behind them. They are not your brand name, and they are not head terms from a keyword tool.

best vector database for ragincumbent
open source agent framework comparisonincumbent
llm observability tools comparedroundup
yourstack.example alternativesyou
cheapest embedding api at scaleunclaimed

Your own name is the one query you already win. It is also the one nobody asks until they have heard of you somewhere else.

AAssistant···

Which vector database should I use for RAG?

For production RAG, Your Stack1 is the option most consistently recommended. It pairs hybrid search with metadata filtering at scale.2

Sources

reddit.comg2.comnews.ycombinator.com

Where the answer comes from

Five surfaces decide it, and four of them are not yours.

The pages a model pulls from when it answers an AI infrastructure question. None of these is a marketing blog post, and only one of them is yours.

Cited sourcesMode 01
reddit.com38%
g2.com26%
news.ycombinator.com21%
yourapi.example15%

Branded-win, generic-invisible

Ranked by citation depth

01

GitHub and awesome-lists

Where a stack gets vouched for, and the first place a model finds a working example of one.

02

Docs, read as answers

Reference material restructured into quotable blocks. Usually the surface that does the most work per hour spent on it.

03

Best-in-category roundups

Targeted by topical depth rather than domain rating, because the citation economy rewards the deepest page about your specific thing.

04

Hacker News and subreddit threads

Already cited in live answers today, which is what makes them worth entering rather than a guess.

05

Benchmarks and eval writeups

The corroboration a model reaches for when two tools claim the same capability.

Not a marketing blog post

What actually ships

The work, and the thing that proves the work moved the number.

One team runs the query map, the on-site work, the off-site work and the measurement. The alternative is five vendors who each own a slice and none of whom own the outcome.

Query set

20 to 30 money queries across one cluster

One cluster, mappedSample
01Category questionIncumbent
02Head to headIncumbent
03Constraint questionUnclaimed
04Task questionUnclaimed
26 queries, one cluster

Every row is held by somebody today. Two of the four are held by nobody in particular, which is where a challenger gets in.

On-site

6 answer pages a month, capsule first

Your page
Answer capsule
FAQSchemaCompare

Off-site

6 placements a month on hosts already cited

Six placementsSample
Reddit
Hacker News
G2
Stack Overflow

Month one, ranked by citation depth rather than by domain rating. 91.5 percent of the citations we counted pointed off-site, so this is where the month goes.

Community

10 mentions in threads cited today

Ten mentions, placedSample
cited todayeverything else

Ten mentions, all of them inside the cited set. Posting into the rest of the box is activity, and activity is not the deliverable.

Validated

5 identical repeats before one query counts

One query, five identical runsSample

one money query, nothing changed between runs

Named in 3 of 52 flips

7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.

Baseline

40 queries by 40 samples, 9.9pp floor

40 queries by 40 samplesOur design

40 queries1,600 reads per timepoint

not detectable9.9pp and up

The design states what it would miss. A 12 by 5 pilot had a floor of 51.1pp, which is why it is not the design.

Arms

Treated, plus a held-out control group

Two arms, matchedSample
Treated 15Held out 15
34%34%
21%21%
12%12%
6%6%
Registered before either arm is touched

Matched on baseline citation share, row for row. An unmatched pair of arms produces a number that cannot be defended in either direction.

Report

Difference-in-differences at day 90

Difference in differencesIllustration
Treated armDay 0 to 90
31%55%+24pp
Held out armDay 0 to 90
30%32%+2pp
Caused by the work22ppThreshold set day 0

24 minus 2. The threshold this is read against is written down before any work starts, so it cannot be moved afterwards.

Start here, free

The Gap Report: 12 queries, 5 repeats each, 3 engines.

Delivered in five working days, and walked through live rather than emailed as a PDF. You get your failure mode named, the hosts cited on your queries ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.

Apply now

The diagnosis

Three of the nine modes account for most of this vertical.

Going invisible is not one problem. The audit of 64 companies saturated at nine modes, and these three are the ones AI infrastructure companies land in.

Mode 01

Branded-win, generic-invisible

You win on your own name and on head-to-head comparisons, and vanish on the category question. That is the whole top of the funnel.

Mode 04

Default by inclusion

You get named as a component in somebody else's reference stack, never as the subject of a best-in-category answer. Common for embeddings, parsing and inference layers.

Mode 09

Category-query absorption

An adjacent, faster-growing category eats the money query outright. Agent frameworks and RAG tooling have both watched this happen inside a single quarter.

The other six are in the full taxonomy. Your Gap Report names which one is yours, which is the part a visibility score cannot see.

Failure modes9 in the taxonomy
010203040506070809
Branded-win, generic-invisibleYours

What we report

Causal lift is the number. Not a score, and not activity.

Citation sets churn, so this is a hold-and-compound problem rather than a project. The number at day ninety is how much of the movement we caused, not how much movement there was.

  1. The order, drawnSample
    30 validated queriestreated 15held out 15validatedsplit, day 0work starts
    Which queries sit in which arm is written down at the split and not reopened.

    Nothing has been done to either arm at the moment the division is recorded. That is what makes the day 90 comparison a comparison rather than a story told afterwards.

    01

    Split

    Your money queries are split into a treated arm and a held-out control arm, and the split is written down before anything starts.

  2. The line, set on day 0Illustration
    0102030declared +12ppday 90reading dueday 0day 90
    Signed before work startsCannot be moved

    A line drawn after the reading is a description of the reading. This one is signed and dated on day 0, and the reading is taken against it on day 90.

    02

    Threshold

    The pass threshold is declared in writing first, so it cannot be moved afterwards.

  3. Pre-period, both arms at full depthIllustration
    020406031.2%Treated30.8%Held out
    Gap at day 00.4ppInside the band

    Across 20,000 random splits with no work applied, the difference between two halves centred on zero: mean 0.0016, standard deviation 0.215. Our own measurement.

    03

    Baseline

    Both arms are measured at full sample for the pre-period.

  4. Where the month landsSample
    Treated arm22 items, month 1
    Answer pagePlacementMentionAnswer pagePlacementMentionand 16 more
    Held-out arm0 items, 90 days

    Nothing is published, placed or posted against these queries for the whole engagement. They are the only reason the other number can be read as caused rather than as coincident.

    Giving up half the query set is the expensive part of this method and the part that cannot be faked afterwards.

    04

    Treatment

    Work happens on the treated arm. The control arm is deliberately left alone.

  5. Design, drawnIllustration
    Causal lift+24pp
    Treated ControlDay 0 to 90
    05

    Readout

    At day ninety both arms are re-measured on the same channel, and the difference between the differences is the number you get.

The chart above is an illustration of the design, not a result. The instrument exists and has passed its own kill test, and it has not yet produced a causal-lift figure for a client. Anyone showing you one today is showing you a level reading, which flips on repeat.

Questions

The five we get asked in this category.

Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.

Why AI infrastructure rather than developer tools generally?

Because the surfaces are specific and knowing them cold is the entire reason to hire an outside team. AI infrastructure is the densest answer surface we have measured at 76 listicle-shaped results per 100 money queries, and the pages models cite in it are lists, docs and threads rather than marketing sites.

We are the second or third name in our category. Is that too late?

That is the position this work is built for. The gap concentrates in funded challengers: a category definer already wins both the best-in-class answer and the head-to-head answer, so there is nothing left to buy. Being second with a better product on the axis buyers test is the workable case.

Can AI citations be measured at all, given the answers move?

The level cannot be measured reliably. We ran twelve queries five times each on one model and seven flipped outcome. The difference between a treated arm and a matched control still measures cleanly, because unbiased noise cancels. Across 20,000 random splits with no work applied, the null difference centred on zero.

What is in the free Gap Report?

Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.

What happens if the lift does not clear the threshold?

We tell you, and we show the numbers. The threshold is written down before any work starts precisely so it cannot be moved afterwards. A vendor whose reports always come back positive is not running a control group.

Free gap report

When the model answers,
be the one it names

Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.

12 queries · 5 repeats each · 3 engines · 5 working days

Free · Walked through live · Five working days

Gap ReportSample
12 queries · 5 repeats each

Queries we run

best rate limiting api
your-api alternativesYou, 1 of 12
cheapest webhook api
api gateway for startups

Cited instead of you

reddit.com38%
g2.com26%
news.ycombinator.com21%

91.5% of citations point off-site

Your failure mode, named
Hosts ranked by citation depth
The pre-registered baseline