Skip to content
← All articles

Benchmarking AI Visibility: 6 Sites, 3 Platforms, 10 Checks

A citable reference for Citability's AI visibility and citation measurements: the six-site benchmark from April 2026, a fixed twelve-prompt citation panel from September 2026, and the sample limits, definitions and denominators behind every number.

Chudi Nnorukam||7 min read

This page is the reference record for what Citability has measured about AI visibility and AI citation, and for what those measurements can and cannot support. If you want to quote a number from it, each section tells you the instrument, the date, the sample and the denominator you need to quote it responsibly.

Two datasets live here. The first is the original six-site benchmark from April 2026. The second is a fixed citation panel from September 2026. They use different instruments and are never added together.

Definitions used on this page#

  • Visible (mentioned): the AI answer names the brand or its topic, with or without a link.
  • Cited: the AI answer links to a URL on the domain as a source. In the September panel this means the domain appeared in a source URL returned by the engine itself, not a link read from the answer text.
  • Unknown: the query errored or was not run. An unknown cell is left out of the denominator and is never counted as zero.
  • Interval: where a rate appears with a range, the range is a 95 percent Wilson score interval for that count and sample size.

Dataset 1: six-site benchmark, measured 2026-04-07#

Instrument: the AI Visibility Readiness (AVR) framework, with ten automated infrastructure checks per domain and roughly 20 topic queries per domain across ChatGPT Search, Perplexity and Claude.

Sample: six domains picked by hand to span domain authority 28 to 97. The sample was not drawn at random and it does not represent any population of websites.

Machine-readable copy: avr-benchmark.json. The table below renders straight from that file, so the figures on this page cannot drift from the data.

DomainDomain AuthorityAI VisibilityAI CitationNotes
ahrefs.com92100%5%Foundation-ready: always mentioned, rarely the cited source
chudi.dev2825%0%Foundation-strong infrastructure, no topic citations at baseline
semrush.com91PartialPartialFoundation-ready; strong schema; only partially scored
reddit.com97UntestedUntestedFailed basic infrastructure checks; cited via training data, not structure
medium.com95UntestedUntestedFailed basic infrastructure checks
x.com96UntestedUntestedFailed basic infrastructure checks

Source: /avr-benchmark.json (point-in-time baseline, measured 2026-04-07). Visibility = mentioned in an AI answer; citation = linked as the source.

Denominators you need to know. The query count per domain was recorded as "roughly 20", not as an exact number, and per-query records from the April run are not published. So "5 percent cited" for Ahrefs means about one scored query in twenty. It is not a precise rate, and no interval is given for it because the exact denominator is not on record.

Partial and Untested. Semrush was partly scored and has no rate on record. Reddit, Medium and X failed the infrastructure checks and were not scored on AI platforms at all, so they carry no visibility or citation figure. Treat these cells as missing, not as low.

Infrastructure checks#

SiteDArobots.txtsitemapAnswer-firstFreshnessJSON-LDMeta descCanonicalHTTPSHeadingsOG tagsPass / Partial / Fail
ahrefs.com92PassPassPassPassPassPassPassPassPassPass10 / 0 / 0
semrush.com91PassPassPassPassPassPassPassPassPassPass10 / 0 / 0
chudi.dev28PassPassPassPassPassPassPassPassPassPass10 / 0 / 0
reddit.com97PartialPassFailFailFailPassPassPassPassPass6 / 1 / 3
medium.com95PassPartialFailPartialFailPassPassPassPassPass6 / 2 / 2
x.com96PassPartialFailFailFailPartialPassPassFailPass4 / 2 / 4

An earlier version of this table showed a single score out of ten (7, 7 and 5 for the last three rows). Those scores do not follow from the cells under one consistent rule: Medium and X fit a rule that counts Partial as half a point, and Reddit does not. The April run recorded no definition of Partial for a single check. This version shows the raw counts instead, so you can apply your own rule.

What Dataset 1 supports#

  • Supported: on 2026-04-07 Ahrefs was mentioned on every scored query and cited on about one in twenty. For a well-known brand, being named and being cited were far apart.
  • Supported: the three highest-authority domains failed more infrastructure checks than the three lower-authority ones in this sample.
  • Not supported: that any single check causes or predicts citation. Only two domains have a full citation figure. No correlation can be estimated from two points, and this page no longer claims one.
  • Not supported: that Reddit, Medium or X are cited mainly through training data. That remains a hypothesis. They were not scored on AI platforms, so the benchmark holds no evidence either way.

Dataset 2: fixed citation panel, measured 2026-09-08#

What it measures: whether AI engines cite citability.dev on twelve fixed prompts, next to a control brand (Profound, tryprofound.com) asked the same prompts. This is a first-party panel about one domain. It is a tracking instrument for Citability itself, not a field benchmark of the market.

Instruments: four, each a different system, each read once on the same day. Perplexity (sonar) and Gemini (gemini-2.5-flash) were read through their APIs. Claude and ChatGPT were read through their subscription command line tools with web search. Because the four columns are four instruments, a difference between columns is not a trend and the columns are never summed.

Prompts: twelve, frozen before the run. Four are about the brand, its founder or a named comparison (prompts 1, 2, 3, 10). The other eight are category, method or research questions, such as "What tools track whether ChatGPT cites my website?"

Citability, cited / measured#

InstrumentCitedMeasuredUnknown95% interval
Perplexity API61200.25 to 0.75
Gemini API31110.10 to 0.57
Claude, subscription51200.19 to 0.68
ChatGPT, subscription41200.14 to 0.61

Control (Profound), cited / measured#

InstrumentCitedMeasured95% interval
Perplexity API4120.14 to 0.61
Gemini API1100.02 to 0.40
Claude, subscription1120.01 to 0.35
ChatGPT, subscription3120.09 to 0.53

What Dataset 2 supports#

  • Supported: every instrument that measured prompts 1, 2, 3 and 10 cited citability.dev on them, apart from Gemini on prompt 3. Engines cite the domain when asked about it by name.
  • Supported: on prompts 5, 6, 8, 11 and 12 all four instruments returned zero citations for citability.dev. Prompt 9 returned zero on three instruments and was unknown on Gemini. When a buyer asks which tools exist, this domain was not cited on 2026-09-08.
  • Not supported: that Citability is cited more than Profound. Every pair of intervals overlaps. Twelve prompts cannot separate the two.
  • Not supported: any change over time. Each instrument was read once. The Perplexity column will not be rerun, so later readings will come from the other three instruments only.

How big a sample a change needs#

The same limit applies to your own measurements. Citability's internal analysis of its own ten-cell baseline grid (five queries across two engines) found that ordinary run-to-run variation on ten cells produces a noise floor of about 37 points on a 100-point score, and 45 points on the axis that rests on only four cells. Roughly 100 cells brings that floor down to about 12 points. A before-and-after claim built on a dozen queries is usually inside the noise, and that includes the two datasets on this page.

Each of these is its own dataset, with its own instrument and date. They are listed here so you can find them, not merged into the figures above.

  • Robots.txt census: robots.txt files of 1,207 marketing and SEO agency domains, published 2026-08-18.
  • Which domains AI cites most: Bing Webmaster Tools AI Performance for citability.dev, read 2026-09-03, 14 cited pages and 39,838 citations. Microsoft describes that report as a sample of citation activity, not a complete record.
  • Citation share benchmarks: reference ranges for citation share and citation rate.

Cite or reuse this data#

You can quote any figure on this page if you name the dataset, its measurement date and this URL. A suggested form:

Citability. "Benchmarking AI Visibility: 6 Sites, 3 Platforms, 10 Checks."
Dataset 1 measured 2026-04-07; Dataset 2 measured 2026-09-08.
https://citability.dev/blog/benchmarking-ai-visibility-6-sites

The Dataset 1 figures are available as JSON. The check definitions and evidence tiers are on the methodology page. To run the same infrastructure checks on your own site, start a free scan. For the author's view on why entity authority limits these numbers, see Entity Mention Graph and AI Visibility on chudi.dev.

Topics:benchmarks·ai-visibility·data

Chudi Nnorukam

AI-Visible Web Architect

Builds chudi.dev and citability.dev. Authored the AI Visibility Readiness Framework. Contributor at freeCodeCamp /news.

chudi.dev|Published |Updated

Check your AI visibility

Free scan. No account required. Results in 10 seconds.

Start Free Scan