47.66° N 117.43° W Summit · Profitable Growth
I Ran Six AI Search Audits. Crawler Access Was Almost Never the Problem.
Quick Take
I audited how AI assistants answer buyer questions for six businesses, 3,623 graded answers in all. Every AI search crawler I tested could read every site. In five of the six, what held the business back was corroboration: what other sites say about it, and whether that’s true. The sixth was missing the page format engines quote.
What I measured
Between late September and early October I ran the same audit on six sites. The mix was ecommerce stores in home goods, baby gear and tools, a law firm, a medical practice, and my own site.
Client figures in this post are reported only in aggregate. Per-client results will be published only with each client’s written permission. My own site’s numbers are mine, so those appear in full.
For each business I built a panel of 48 to 59 buyer questions from real demand: converting search terms, Search Console queries and customer phrasing. I locked each panel and recorded its SHA-256 hash before the first run, so I can’t quietly swap out a question that makes a result look bad. The re-tests reuse the same file, byte for byte.
The engines were ChatGPT, Claude, Gemini, Perplexity, Google AI Mode, Google AI Overviews and Microsoft Copilot, though not every engine ran in every audit. Quotas, CAPTCHAs I didn’t solve and a Microsoft terms wall got in the way. The plan was three runs of every question on every engine, because the same question asked twice gets different answers. Quotas cut that on several engines, and each report says where.
A separate verifier pass re-ran a sample in fresh sessions. Anything that didn’t reproduce got dropped or downgraded.
Every rate below carries its raw count and a 95% Wilson interval. A percentage built from a few hundred answers that change run to run doesn’t mean much without its range.
What I expected vs what I found
I went in testing two ideas. Most AI search advice says to check crawler access first. My notes from earlier audits said access was rarely the problem and corroboration usually was. I tested both instead of assuming either.
Access was the one that had already bitten me. In June my Cloudflare firewall was returning 403 errors to five AI crawlers, and nothing in my analytics showed it. I fixed it then.
This time access was clean everywhere. I fetched each home page and a set of money pages with the major AI crawlers’ user agents. All 546 fetches came back 200. All six sites graded A on access.
One caveat. Those fetches sent the crawlers’ user-agent strings from a residential IP, not from the crawlers’ own IP ranges. A firewall rule keyed to verified bot IPs could behave differently, and I haven’t tested that.
The pattern: engines read you, then ask someone else
I grade each audit on six layers, from L1 (can the crawler get in) to L6 (can you measure the traffic). L5 is corroboration: anything outside your own site that confirms you exist and are good. In five of the six audits, L5 was the binding constraint.
Across the five client audits, the share of non-brand answers that named the business ran from 4.4% (17 of 388, 95% CI 2.8% to 6.9%) to 22.9% (181 of 791, 95% CI 20.1% to 25.9%). Those are the questions a buyer asks without typing the business name, which is where new customers come from. Even at the top of that range, the business went unnamed in more than three answers out of four.
My own site sat below the whole range: 1.0% (2 of 203, 95% CI 0.3% to 3.5%). When someone typed my name, engines found me 85.4% of the time (41 of 48, 95% CI 72.8% to 92.8%). They could read me fine. When someone asked who to hire, they reached for marketplaces, directories and roundup articles.
The reason was plain once I looked. I had two independent referring domains. The engines drew on 81 slots across roundups and results pages for my buyer questions, and I held none of them. Both of my non-brand mentions came on the same question, which matched one of my pages almost word for word. ChatGPT’s held up on a re-run. Google AI Mode’s didn’t. That’s about as far as on-site work goes alone.
The client audits showed the same mechanism. Engines cited a business’s own pages to explain how something works, then named someone else when the question turned to who to buy from. For that answer they trusted directories, review sites, editorial roundups and forums.
The exception
One audit didn’t fit. That business did well on questions about itself and close to nothing on comparison questions.
The pages winning those comparisons weren’t from bigger or more trusted sites. They shared a format: a verdict up top, a comparison table, a recent date and a named author. That business had no live page in that format. Its binding layer was L4, content format, and the audit found off-site trust wasn’t what stopped those answers.
That changed my thinking. Trust signals decide questions about who you are. Comparison questions go to whoever has the page in the right format.
The accuracy finding
Getting named is half the job. The other half is what the engine says once it names you.
Three of the client audits counted wrong facts per answer. Across those three, between 16.4% (36 of 219, 95% CI 12.1% to 21.9%) and 31.3% (81 of 259, 95% CI 25.9% to 37.2%) of answers that named the business got at least one fact wrong. The audits drew the base a little differently, so read that as a range, not a ranking. The other two counted errors in ways that don’t convert to a per-answer rate, so I left them out.
The errors had sources. Most traced to a stale directory profile, an old listing, a profile nobody maintained anymore, or the business’s own pages saying two different things. Plenty of reviews and listings didn’t protect a business either, because corroboration only helps when it’s right.
My own site had its share. Claude told people I don’t have a Google Business Profile. I do, with reviews, and the claim reproduced on a re-run.
I would rank accuracy ahead of getting recommended, for two reasons. A wrong fact in a brand answer reaches the person who already knows your name and is checking you out, which is the person closest to buying. And accuracy is the faster fix, because most errors trace to a page someone can correct or claim. That’s why brand fact-error rate is the 30-day metric in my re-test plan, and non-brand mention rate is the 90-day one.
What this means if you own a business
- Check access once, then move on. Confirm robots.txt allows the AI crawlers and fetch a few pages with their user agents. In six audits it came back clean every time.
- Ask the engines about you by name. Read the answers for wrong facts, then find where each one came from. It’s usually a directory or an old listing. Fix the ones you control first.
- Find the sources engines cite for your buyer questions. Get onto the real ones. Don’t seed reviews. The FTC’s fake-review rule covers that.
- Put your name in the passage engines lift. If your page gets cited and your name doesn’t, the paragraph that gets quoted doesn’t say who you are.
- Publish the format for comparison questions. Verdict first, a table, a date, a named author.
If you’re rebuilding soon, read my notes on migrating a site that AI assistants already cite.
What I don’t know yet
- No re-tests yet. The first are due in late October. These are first baselines, not before-and-after numbers. I wrote down the success rules on October 4, before any re-test: a change only counts when the 95% intervals don’t overlap.
- Engines answer differently run to run. On my own audit, the verifier’s fresh sessions matched 21 of the 23 sampled runs it could complete (91.3%, 95% CI 73.2% to 97.6%). One of the two misses was a recommendation of me that vanished on the re-run.
- Location. Most runs came from my location in Washington. I wrote the city into local questions, and the reports flag answers that localized to me anyway.
- Signed-in leakage. On several audits my own signed-in assistant accounts carried my context into some runs. Those runs were excluded or flagged.
- Coverage. My own site got 251 of 721 planned runs before quotas stopped me. I ran 225 more after I had started fixing things, and left them out of the baseline.
- Sample. Six sites isn’t a market study. Five are clients, picked because they’re my clients, and one is me.
If you want this run on your own business, the method and what you get back are on the AI answer audit page. If the fix turns out to be the site and its listings, that’s what my service-business websites build covers.
More reading
-
The Fields You Most Want to Change Are the Ones That Vanish
A Google Ads mutate can return success and change nothing. The documented way to build a field mask drops any value equal to its type's default, and off is a default.
-
A 200 Is Not Proof (The Redirect That Eats Your gclid)
A redirect can return a correct 301 to a correct page and still strip the gclid off every ad click that passes through it. The click lands, the conversion comes back attached to nothing. Here is the check.
Want a review like this on your account?
Want this kind of review
on your account?
Thirty minutes on the phone. Same person on the call as on the work. Walk out with a clear set of next steps.