XYZ BIMCollective · Portugal Research
xyzbim.eu / pesquisa / visibilidade-ia · panel · round 0 · 2026-09-05

AI visibility — who the answer engines can read in construction

A recurring public panel: ten fixed sector questions, put to the answer engines (ChatGPT, Claude, Gemini, Perplexity, Copilot) and to the sources they read. Round 0 measures the retrieval layer — being crawlable, citable and structured — of 11 domains that answer those questions. Later rounds add the Bing SERP and the five assistants, signed in.

domains audited · fixed questions11 · 10 sources the answer engines read
outside Common Crawl (window 2026-21 to 2026-34)1 xyzbim.eu
true llms.txt3 of 11 two clean soft-404s on verification
blocking AI crawlers0 of 11 robots.txt of all domains
Verdict — the only source invisible to the corpus is the house itself — and the fault is not technical. Of the eleven sources the answer engines read for construction questions in Portugal, ten are in
Common Crawl, several with thousands of captures. The one that is not — zero captures in every collection tested
up to 5 September 2026, including one from 2024 — is the only one that publishes a true llms.txt (one of three in
the whole set), robots.txt
open to all AI crawlers, a sitemap with 77 URLs and structured data. The door is in order; the problem
is that nobody links to it from outside. No third-party mentions, no listicles, no directories — and without
inbound links there is no entry into the crawl.

1 · The five signals, the eleven domains

Audited by script on 2026-09-05, against public endpoints. Common Crawl is the open corpus that feeds training and retrieval; llms.txt is the front page for models; robots.txt says who may come in; the sitemap and the JSON-LD say what there is to read and in what form. The house's row is highlighted — this panel measures us too.

domainCommon Crawlllms.txtrobots (AI crawlers)sitemapJSON-LD
dre.pt
PT institutional
4000+no (soft-404)open
thefuture3d.com
listicle cited by assistants
4000+yesopen17
ted.europa.eu
EU institutional
3000+noopen2124
ipq.pt
PT institutional
2110nono robots.txt (404)5
base.gov.pt
PT institutional
2102+nono robots.txt (404)
impic.pt
PT institutional
1110nounreadable (403)
builtcolab.pt
PT institutional
783no (soft-404)open
oasrs.org
PT institutional
675nono robots.txt (404)
clutch.co
business directory
645yesopen592
truecadd.com
listicle cited by assistants
280noopen3481
xyzbim.eu
the house
0yesopen771

Captures in Common Crawl, by domain

Same window as the four recent collections (2026-21 to 2026-34). Counts cut off at 1000 captures appear as “1000+”; linear scale up to 1000. The house is the red bar.

thefuture3d.com
4000+
dre.pt
4000+
ted.europa.eu
3000+
ipq.pt
2110
base.gov.pt
2102+
impic.pt
1110
builtcolab.pt
783
oasrs.org
675
clutch.co
645
truecadd.com
280
xyzbim.eu
0

2 · Readings from round 0

The corpus already speaks institutional Portuguese. The Diário da República (the official gazette; 4000+ captures in the window), IPQ, BASE
and IMPIC are in Common Crawl in bulk — by age and by links, not by tidiness: IPQ's sitemap
declares 5 URLs, most of the institutional sites have no sitemap at all and none announce entities in JSON-LD
(round 0, verified on 2026-09-05) [measured].
When an answer engine answers a Portuguese regulatory question, the raw material exists in the corpus;
what these sources lack is form (structure, metadata), not presence.

The house has the form and lacks the presence — the exact inverse. It is the round's most useful finding: the two kinds
of failure are independent and call for different work. For the institutional sites, fix the form. For the
house, fix the presence — get referenced from pages that are already in the crawl (directories,
listicles, company profiles), because Common Crawl has no submission form: entry is through links.

llms.txt is rare and sometimes an accident. Three of the eleven domains have a true llms.txt. Two others
answer `200 OK` on `/llms.txt` with an HTML error page — a naive checker would count “has llms.txt”.
The method lesson: every boolean signal in this panel was confirmed by a second content check, not
by HTTP status code.

No domain blocks AI crawlers. The eleven robots.txt (four of them nonexistent, one unreadable)
allow GPTBot, ClaudeBot, PerplexityBot, CCBot and company. In 2026, in this sample, blocking models
is no longer the norm — whoever is outside the corpus is outside for not being reached, not for refusing to be read.

3 · How we measure — and what this panel is not

The panel. Ten fixed questions of the construction sector in Portugal, kept stable across rounds:

Round 0 measured the retrieval layer — the objective conditions for a domain to be read and cited —
in five signals: captures in Common Crawl (CDX index, four recent collections, domain and subdomains,
count cut off at 1000 captures per collection) [measured], `llms.txt` (existence and content verified, not just the
HTTP status code), `robots.txt` (rules for ten AI crawlers), `sitemap.xml` (declared URLs) and the JSON-LD of the
homepage (announced structured entities).

What this panel is not. It does not measure whether the assistants cite or recommend anyone — that is round 1 and
requires signing in to each assistant. It does not score or award prizes: it records verifiable signals.

Two method hazards on record, for whoever replicates. First: the Common Crawl index returns
spurious `404`s when a shard is under load — the first pass of this round said that dre.pt and
thefuture3d.com were absent, and was wrong; the published counts only accept a zero after a
second confirming call. Second: Bing serves plausible organic results, but from a different query to
clients without a session — for the same question, four attempts returned results about a bridge in
Dublin, a BBQ restaurant, a hockey game and Reddit pages. The Bing SERP is measured with a human
session; any tool that says otherwise is measuring noise — or is not saying how it measures.

4 · Round 1

Round 1 adds the two layers that require a human session: the Bing SERP for the ten questions (which
domains are returned) and the five assistants — ChatGPT, Claude, Gemini, Perplexity and Copilot — with the
same questions, recording who is mentioned and who is cited with a link. Monthly cadence; the questions do not
change, so the rounds stay comparable. If you want one of your domains audited in the next round,
write to [email protected] — the retrieval audit is free and takes minutes.