The AI Visibility Study
Search is being answered, not just listed. This is an ongoing measurement of how often real small and mid-sized businesses are named when an AI assistant answers a buying-intent question about their category — and what their websites are doing, or failing to do, to be findable at all.
- 94.3% of audited sites are never named in an AI answer
- 2.1% of 5407 assistant answers named the business asked about
- 9.1% block at least one AI crawler in robots.txt
What the AI Visibility Index measures
The Index tracks two different things, because "being visible to AI" is two different problems and sites tend to fail at them independently.
Answer presence
When an assistant is asked the kind of question a customer actually asks — the best provider of a category, in a place — is this business one of the ones it names? Every audited business is queried against ChatGPT, Claude, Gemini and Perplexity, and the businesses named in each answer are matched against the audited one.
Machine readability
What an assistant's crawler finds when it arrives: whether robots.txt lets it in, whether there is a sitemap, whether the homepage states in machine-readable form what the business is and where it operates. These are measured first-party, by our own crawler, on the site itself.
Presence is the outcome; readability is the input a site owner controls. A site can be perfectly marked up and still go unnamed, which is why both are published side by side rather than combined into a single score. The audit that produces these signals is the same free scan anyone can run.
Key findings
-
94.3%
of audited sites never surface in an AI answer
Across 5,407 assistant answers to buying-intent questions, only 10 of 175 sites were named even once.
-
3.2%
Gemini names a local business most often
Gemini named the audited business in 3.2% of its answers, ahead of every other assistant in the panel — the widest spread between two assistants asked the same questions on the same days.
-
9.1%
block at least one AI crawler in robots.txt
38 of 418 sites disallow an AI user-agent outright, most of them without distinguishing between crawlers that train models and agents that fetch a page to answer a live question.
-
37
sites block GPTBot; only four block OAI-SearchBot
Blocking is aimed at training crawlers, not at retrieval. GPTBot is disallowed by 37 sites and ClaudeBot by 35, while the agents that actually fetch pages to answer a question — OAI-SearchBot (4) and Perplexity-User (3) — are left open. Sites blocking the whole group may be giving up citations they meant to keep.
-
55.0%
publish any schema.org structured data
230 of 418 sites carry structured data on their homepage, leaving the rest to be understood from prose alone.
-
19.1%
publish LocalBusiness schema
The markup that states a business's name, address and category in machine-readable form is the least-adopted signal measured — present on 80 sites — even though most of this corpus is local businesses.
The data
Each assistant was asked the same questions about the same businesses on the same days. ChatGPT answered 1213 of its queries; the 187 that failed at the provider are excluded from its denominator rather than counted as absences.
Training crawlers harvest pages to build models; retrieval agents fetch a page because someone asked a question right now. Blocking the first costs a site nothing in today's answers. Blocking the second makes it uncitable — and the gap between the two columns is far smaller than that distinction deserves.
Methodology
The dataset
Every figure comes from audits run through Website Auditor's free scanner between 2026-03-30 and 2026-07-31: 510 stored audits across 441 distinct domains, in 37 detected sectors. Aggregates are computed over Website Auditor's reports, the analytics view that excludes our own parent company auditing its own properties, so internal traffic cannot inflate a share. The exact aggregate is committed as a SQL function and re-run unchanged for each edition, so consecutive editions are comparable rather than separately judged.
How each number is calculated
- One domain, one vote. Both pillars use the most recent audit per domain. Counting audits instead would give a site scanned nine times nine times the influence over every share.
- Answer presence covers 2026-07-23 to 2026-07-31: 175 sites and 5407 assistant answers. A site counts as present if it was named in at least one answer from at least one assistant. Of 176 sites audited in that window, one produced no assistant answers at all and is excluded rather than counted as absent.
- Machine readability covers 2026-05-16 to 2026-07-31 across 418 sites — a longer window and a larger sample, because these signals are read by our own crawler and depend on no third-party provider.
- Failed queries are excluded, not counted as absences. A provider timeout is a fact about the provider, not evidence that a business is invisible. Per-assistant failure counts are published above so a smaller denominator is visible rather than inferred.
- Percentages are derived, never stored. The published dataset holds integer counts only; every share on this page is computed from two of them at render time, so a headline cannot drift away from the data behind it.
Assumptions and limitations
These are the caveats that decide whether a figure here is safe to quote. They are published with the same weight as the findings.
- The corpus is self-selected, not a random sample of the web: every site in it is one whose owner chose to run a free audit, which skews toward small and mid-sized businesses actively working on their web presence. Treat the shares as descriptive of that population, not of all websites.
- Cross-assistant figures start on 2026-07-23 because earlier audits fanned a single Perplexity answer across the ChatGPT, Claude and Gemini cards. Those earlier rows are real answers, but they are Perplexity's, so comparing assistants over them would compare one provider with itself.
- ChatGPT answered 1,213 of its 1,400 panel queries; 187 failed at the provider and are excluded from its denominator rather than counted as absences. Its rate therefore rests on a smaller sample than the others.
- Appearance is measured by matching the audited business against the businesses named in an assistant's answer. Name matching is fuzzy, and a business identified only by its domain — roughly one in five — is harder to match, so presence is more likely under-counted than over-counted.
- Assistants are non-deterministic and their indexes move. These figures describe the answers given during the stated window, and re-running the same queries later will not reproduce them exactly.
- Crawl signals are read from the homepage, robots.txt and sitemap.xml only. A site carrying structured data on inner pages but not its homepage counts as having none.
Citation and reuse
Figures are free to reproduce with attribution and a link to this page. For the underlying methodology, or a cut of the data for a specific sector, email support@website-auditor.io.
Update history
This page is the permanent home of the Index. Each new edition updates the figures above and adds a dated row here rather than moving to a new URL, so links and citations keep resolving to current data.
-
First edition. Cross-assistant presence measured over 2026-07-23 to 2026-07-31, the first period in which ChatGPT, Claude, Gemini and Perplexity were each queried through their own provider; crawl signals measured over 418 sites audited since 2026-05-16.
Where does your site sit in this data?
The Index is built from the same free audit anyone can run. It checks the AI readiness signals measured above and queries the four assistants for your business by name — no signup, results in about a minute.
Run a free AI visibility auditAlready know the drill? Start from the homepage scanner, or read what the audit checks first.