The ReadableByAI Index — YC Fall 2025 vs. Established SaaS, Raw-HTML Readability | ReadableByAIMenu ▾<br>THE READABLEBYAI INDEX<br>A quarter of YC Fall 2025 startups are blank pages to AI crawlers.<br>25.5% of 145 YC Fall 2025 homepages ship an empty client-rendered shell to GPTBot, ClaudeBot and PerplexityBot — no headline, no product description, nothing but a . Among 455 established SaaS companies measured the same way, it's 2.9%. That's a 8.9x gap. Measured 8 August 2026.<br>What was measured, and what wasn't<br>Every homepage in this Index was fetched twice: once with a baseline browser user-agent, and once each with the twelve crawler user-agents (eleven AI-vendor crawlers plus bingbot as a non-AI control) ReadableByAI probes, all from the same datacenter IP. The readability numbers below — visible word count, rendering classification, robots.txt contents — come from the baseline fetch, and they are identity-independent : whether a page ships its content in raw HTML or hides it behind a client-side render is true for every visitor, bot or browser, and needs no further confirmation.<br>What we deliberately did not publish is bot-specific access results — which user-agents got a 200, a 403, or a challenge page. A datacenter IP claiming to be ClaudeBot is an unverified probe, not the real crawler; vendors authenticate their bots by published IP range, so a challenge to our probe cannot distinguish a genuine block from a WAF simply doubting an impostor. Only the site's own server logs settle that, and we don't have them. So this Index reports what any fetch can prove — content and permission — and leaves access claims out entirely.<br>One further limit, stated plainly because it is the strongest objection to this method: a site can serve different HTML to different requesters. The common implementation keys on the user-agent string, and this Index detects it — every company named below returned the same content to every crawler identity tested , verified by comparing the response returned to each identity against the baseline (four were byte-for-byte identical across all thirteen fetches; the other two varied by under one percent in size, consistent with per-request timestamps). The four companies that did serve crawlers a materially different page are reported separately rather than counted as failures. What an outside probe cannot rule out is a site that varies its content by verified crawler IP range rather than by user-agent. That is rare, and nothing in this dataset suggests it, but it cannot be disproven from outside — which is the honest reason server logs matter and probes alone are never the last word.<br>That limit is not a flaw in this Index; it is the boundary of what any outside probe can establish, ours included. Resolving it needs the one record we do not have: the target's own server logs, with every hit checked against each vendor's published crawler IP ranges. That separates verified crawlers from impostors, shows which pages they actually retrieved, and surfaces the fetches they abandoned midway — none of which is visible from the outside, and all of which is what the paid audit reads. A roadmap item, not yet built and therefore not sold: dual-origin probing from residential and datacenter networks simultaneously, which would expose IP-sensitive bot management without logs.
The comparison<br>145 of 147 YC Fall 2025 companies and 455 of 514 established SaaS companies returned a clean 200 to the baseline fetch and are counted below. The rest are excluded from every percentage on this page, not silently dropped: 2 YC domains (2 non-200 responses) and 59 SaaS domains (51 non-200, 8 unreachable).<br>PopulationClean baselineCSR_SHELLSSR_THINSSR_FULLMedian wordsrobots.txtllms.txtYC Fall 2025145 / 14737 (25.5%)14 (9.7%)94 (64.8%)630111 (76.6%)48 (33.1%)Established SaaS455 / 51413 (2.9%)11 (2.4%)431 (94.7%)1239446 (98.0%)247 (54.3%)<br>CSR_SHELL = under 150 visible words in raw HTML. SSR_THIN = 150–399. SSR_FULL = 400+. Median visible words: 630 for YC vs. 1239 for SaaS — the typical SaaS homepage ships roughly double the readable content of the typical YC Fall 2025 homepage, before either one is judged by whether it's reachable at all.<br>The comparison isn't one-directional. 4 companies in the established-SaaS set serve AI crawlers more content than they served our baseline fetch — bot-aware, dynamic rendering that detects a non-browser request and responds with fuller markup instead of a thinner one. The two largest gaps we measured: a payroll platform serving roughly 220x the byte size to a crawler that it serves to a plain browser fetch, and an ML-ops platform at roughly 141x (both public companies; unnamed here because per-company detail belongs to the named-examples policy above). Some teams have already solved this deliberately — it's evidence this is a solvable engineering problem, not an unavoidable one.
The 145 YC Fall 2025 homepages, anonymized<br>These are the 145 Y Combinator Fall 2025 companies that returned a clean...