Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

SpikeyCoder1 pts0 comments

AI Visibility Index: how often small businesses appear in AI answers | Website Auditor

here never loaded — CSP<br>style-src-elem is 'self' — so the site silently ran on fallback<br>fonts. Preload the latin subsets: they are referenced from base.css,<br>so without this the browser cannot discover them until that<br>stylesheet parses. -->

Skip to main content

Our Terms and Privacy Policy were updated on July 27, 2026. Read the Terms and the Privacy Policy.

The AI Visibility Study

Search is becoming an answer, not just a list of links. This is an ongoing<br>measurement of how often real small and mid-sized businesses are named when<br>an AI assistant answers a buying-intent question about their category, and of<br>what their own websites do or fail to do to earn that mention.

Built from 531 website audits across<br>458 domains and 38 detected sectors.<br>The corpus runs to 2026-08-02, the date of its most recent audit.<br>Free to cite and quote with attribution.

94.8%<br>of audited sites are never named in an AI answer

1.9%<br>of 5,978 assistant answers named the business asked about

8.9%<br>block at least one AI crawler in robots.txt

What the AI Visibility Index measures

The Index tracks two things separately, because "being visible to AI" is<br>really two problems with two different causes.

Answer presence

A customer asks an assistant for the best provider of a category in a place,<br>and the assistant answers with a handful of names. This pillar asks whether<br>the audited business is one of those names. Every audited business is queried<br>against ChatGPT, Claude, Gemini and Perplexity, and each answer is matched<br>against the audited business.

Machine readability

This pillar is what an assistant's crawler finds on arrival: whether<br>robots.txt lets it in, whether the site publishes a sitemap, whether the<br>homepage states the business's identity and location in machine-readable<br>form. Every one of these signals is first-party, read by our own crawler<br>directly from the site.

Presence is the outcome; readability is the input under a site owner's control.<br>A site can be perfectly marked up and still go unnamed, which is why both are<br>published side by side rather than combined into a single score. The signals<br>come from the same free scan anyone can run on their own site.

Key findings

94.8%

of audited sites never surface in an AI answer

Only 10 of 193 sites were named even once across 5,978 assistant answers to buying-intent questions.

2.9%

Gemini has the widest reach in the panel

Gemini named the audited business in 2.9% of its answers, ahead of every other assistant in the panel. All four were asked the same questions about the same businesses on the same days, so the gap is a difference between assistants rather than a difference between questions.

8.9%

block at least one AI crawler in robots.txt

39 of 436 sites disallow an AI user-agent outright, most of them without distinguishing between crawlers that train models and agents that fetch a page to answer a live question.

38

sites block GPTBot; only 4 block OAI-SearchBot

Blocking is aimed at training crawlers rather than at retrieval agents. GPTBot is disallowed by 38 sites and ClaudeBot by 36, while the agents that actually fetch a page to answer a live question, OAI-SearchBot and Perplexity-User, are disallowed by 4 sites and 3 sites. A site that blocks the whole group pays for it in lost citations.

54.6%

publish any schema.org structured data

238 of 436 sites carry structured data on their homepage. The rest leave an assistant nothing to read but prose.

19.3%

publish LocalBusiness schema

The markup that states a business's name, address and category in machine-readable form is the least-adopted signal measured here, present on 84 sites. Most of this corpus is local businesses.

The data

Share of answers naming the audited business

2026-07-23 to 2026-08-02 &bull;<br>193 sites &bull;<br>5,978 answers

Gemini

2.9%<br>8 of 193 sites

ChatGPT

1.7%<br>5 of 193 sites

Claude

1.6%<br>5 of 193 sites

Perplexity

1.6%<br>8 of 193 sites

Every assistant was asked the same questions about the same businesses on the<br>same days, so the differences below are differences between assistants.<br>ChatGPT answered 1,352 of its 1,544 queries;<br>the 192 that failed at the provider are excluded from<br>its denominator rather than counted as absences.

AI crawlers blocked in robots.txt

Share of 436 audited sites disallowing each user-agent

CCBot

8.7%<br>38 sites - training

GPTBot

8.7%<br>38 sites - training

Bytespider

8.5%<br>37 sites - training

ClaudeBot

8.3%<br>36 sites - training

Google-Extended

8.3%<br>36 sites - training

anthropic-ai

1.4%<br>6 sites - training

PerplexityBot

1.1%<br>5 sites - retrieval

ChatGPT-User

0.9%<br>4 sites - retrieval

OAI-SearchBot

0.9%<br>4 sites - retrieval

Perplexity-User

0.7%<br>3 sites - retrieval

Training crawlers harvest pages to build models; retrieval agents fetch a page<br>in response to a live question. Blocking the first costs a site nothing in<br>today's answers. Blocking the second makes the...

sites answers assistant site audited business

Related Articles