/

What was measured, and when

Three cohorts, each probed the same way: real businesses with real websites, two questions per business in a customer's own words, two engines, three answers per question per engine, every answer stored in full.

CohortDateBusinessesStored answers
Nine US wedding markets21 July to 11 August 20269 city cohorts5,904
Phoenix, AZ home services12 August 202662744
Utrecht province, Netherlands19 August 202670840

The engines were the API models behind ChatGPT (OpenAI's gpt-4o-mini) and Claude (Anthropic's Claude Haiku 4.5) in all three. Both are cheap consumer-tier models, and neither is a paid or search-connected tier, so nothing here describes what someone on a different model sees.

The intermediary rate, three times over

CohortAnswersNamed an intermediaryNamed the business asked about
Wedding, who-to-hire questions2,2321,407 (63%)6
Wedding, where-to-marry questions3,672383 (10%)131
Phoenix home services744465 (63%)35
Utrecht installers840794 (95%)0

Read the wedding rows against each other first, because they are the same corpus split by question type. Asked where to get married, the answers name places: the directory rate is 10 percent and 131 answers named the venue in question. Asked who to hire, the directory rate is 63 percent and six answers named the business in question. The question decides the shape of the answer, not the market.

Then read the two non-wedding rows. Phoenix produced the same 63 percent against a completely different trade in a different state. Utrecht produced 95 percent and a flat zero, in a different language and a different country. Three measurements, three markets, one direction.

The Phoenix ratio

Phoenix is the sharpest version of it because both halves were counted in the same corpus. Across 744 stored answers, a lead-generation marketplace was named in 465 and the 62 probed businesses were named, between them, in 35. That is 13.3 answers naming a marketplace for every one naming a local company.

Named in the Phoenix corpusAnswersShare of 744
Yelp46162.0%
Angi / Angie's List23832.0%
HomeAdvisor9312.5%
Thumbtack263.5%
At least one of the four46562.5%
Any of the 62 probed businesses354.7%

Search engines and accreditation bodies were counted separately and deliberately left out of the 465: Google appeared in 444 answers and the Better Business Bureau in 148, and a trade does not buy a lead from either. Counting them would have made the headline number look bigger than the claim deserves.

Both narrow numbers are narrow on purpose. The 35 counts only businesses whose hits a human read and confirmed the answers genuinely name: one cohort member trades under an ordinary noun that matches 90 answers using the word in its everyday sense, and seven more matched companies with near-identical names under different ownership. Counting those would have claimed visibility that the text does not support. The 30-of-62 homepage figure below is a floor for the same reason, not a ceiling.

The part that is not about bad websites

The obvious explanation is that these are neglected websites. In Phoenix it is not.

Readiness, on the same auditAverage score
Phoenix home services (62 sites, 12 August 2026)80.2
Nine wedding cohorts71.9 to 75.1
Utrecht installers (70 sites, 19 August 2026)75.9

The Phoenix cohort scored higher than any wedding cohort measured, with a minimum of 45 and a maximum of 100 across 62 sites and no incomplete samples. These are agency-built, well instrumented sites. And 30 of those 62 advertise a lead-generation directory on their own homepage, as a profile link, a seal, or an award claim in visible prose. Across all 199 Phoenix prospects the same scan returned 97, which is the same rate to within half a point.

So the picture is a trade that builds a competent website, pays a marketplace for leads, prints that marketplace's badge on its own homepage, and is still not the name the assistant gives the homeowner who asks.

Where the Dutch cohort differs, and where it does not

Utrecht produced the most complete absence measured: zero of 70 businesses named, in any of 840 answers, on either engine. But the reasons are not the American ones, and three of them run the other way.

  • Cloudflare settings that keep AI crawlers out, the most common silent block found in the US sweeps, appeared on none of the 70.
  • Sites serving their text only after JavaScript runs: three of 70.
  • Contact details as readable text: 69 of 70 publish a phone number, an email address and usually a street address plainly on the page.

What is missing instead is self-description. None of the 70 passes the JSON-LD check: 13 publish none at all, and the other 57 publish something that does not describe the business. Sixty-two of 70 have no llms.txt. Those are the sites of businesses that are not hard to read, they just never say what they are.

One number in the Dutch answers is worth its own line. Yelp, which barely operates in the Netherlands, appeared in 358 answers, more than twice as often as the Dutch lead platform that Dutch tradespeople actually pay per lead. The assistant answers a Dutch question with an American reflex.

What this does not establish

Every number above is a count, on a stated date, from stored text that can be re-read. None of it is a forecast, and four limits are worth stating plainly.

  • Stored absence is not invisibility. In Phoenix, 61 of 62 businesses appeared in none of their own 12 answers, but a full-depth re-read of the whole corpus found three of them named elsewhere in it. The defensible absent set is 44 of 62, not 61, and that is the number we use.
  • One city each, one date each. Phoenix is one home-services market on one day. A second home-services city has not been measured, so nothing here is a statement about home services generally, or about Phoenix at any other time.
  • Correlation, not causation. That businesses advertise on directories and that answers name directories are two separate measurements over the same market. Nothing here shows one causes the other.
  • The badge rate is a live-page reading, taken over homepages only on 12 August 2026, so it drifts as those sites change. The answer counts read stored text and are stable.

The readiness half is the half a business controls, and it moves in days. How the score is produced sets out every rule behind it, and the free audit runs it on any address you give it, quoting the line behind every deduction.