A recurring public panel: ten fixed sector questions, put to the answer engines (ChatGPT, Claude, Gemini, Perplexity, Copilot) and to the sources they read. Round 0 measures the retrieval layer — being crawlable, citable and structured — of 11 domains that answer those questions. Later rounds add the Bing SERP and the five assistants, signed in.
Audited by script on 2026-09-05, against public endpoints. Common Crawl is the open corpus that feeds training and retrieval; llms.txt is the front page for models; robots.txt says who may come in; the sitemap and the JSON-LD say what there is to read and in what form. The house's row is highlighted — this panel measures us too.
| domain | Common Crawl | llms.txt | robots (AI crawlers) | sitemap | JSON-LD |
|---|---|---|---|---|---|
| dre.pt PT institutional | 4000+ | no (soft-404) | open | — | — |
| thefuture3d.com listicle cited by assistants | 4000+ | yes | open | 1 | 7 |
| ted.europa.eu EU institutional | 3000+ | no | open | 2124 | — |
| ipq.pt PT institutional | 2110 | no | no robots.txt (404) | 5 | — |
| base.gov.pt PT institutional | 2102+ | no | no robots.txt (404) | — | — |
| impic.pt PT institutional | 1110 | no | unreadable (403) | — | — |
| builtcolab.pt PT institutional | 783 | no (soft-404) | open | — | — |
| oasrs.org PT institutional | 675 | no | no robots.txt (404) | — | — |
| clutch.co business directory | 645 | yes | open | 59 | 2 |
| truecadd.com listicle cited by assistants | 280 | no | open | 348 | 1 |
| xyzbim.eu the house | 0 | yes | open | 77 | 1 |
Same window as the four recent collections (2026-21 to 2026-34). Counts cut off at 1000 captures appear as “1000+”; linear scale up to 1000. The house is the red bar.
The corpus already speaks institutional Portuguese. The Diário da República (the official gazette; 4000+ captures in the window), IPQ, BASE
and IMPIC are in Common Crawl in bulk — by age and by links, not by tidiness: IPQ's sitemap
declares 5 URLs, most of the institutional sites have no sitemap at all and none announce entities in JSON-LD
(round 0, verified on 2026-09-05) [measured].
When an answer engine answers a Portuguese regulatory question, the raw material exists in the corpus;
what these sources lack is form (structure, metadata), not presence.
The house has the form and lacks the presence — the exact inverse. It is the round's most useful finding: the two kinds
of failure are independent and call for different work. For the institutional sites, fix the form. For the
house, fix the presence — get referenced from pages that are already in the crawl (directories,
listicles, company profiles), because Common Crawl has no submission form: entry is through links.
llms.txt is rare and sometimes an accident. Three of the eleven domains have a true llms.txt. Two others
answer `200 OK` on `/llms.txt` with an HTML error page — a naive checker would count “has llms.txt”.
The method lesson: every boolean signal in this panel was confirmed by a second content check, not
by HTTP status code.
No domain blocks AI crawlers. The eleven robots.txt (four of them nonexistent, one unreadable)
allow GPTBot, ClaudeBot, PerplexityBot, CCBot and company. In 2026, in this sample, blocking models
is no longer the norm — whoever is outside the corpus is outside for not being reached, not for refusing to be read.
The panel. Ten fixed questions of the construction sector in Portugal, kept stable across rounds:
Round 0 measured the retrieval layer — the objective conditions for a domain to be read and cited —
in five signals: captures in Common Crawl (CDX index, four recent collections, domain and subdomains,
count cut off at 1000 captures per collection) [measured], `llms.txt` (existence and content verified, not just the
HTTP status code), `robots.txt` (rules for ten AI crawlers), `sitemap.xml` (declared URLs) and the JSON-LD of the
homepage (announced structured entities).
What this panel is not. It does not measure whether the assistants cite or recommend anyone — that is round 1 and
requires signing in to each assistant. It does not score or award prizes: it records verifiable signals.
Two method hazards on record, for whoever replicates. First: the Common Crawl index returns
spurious `404`s when a shard is under load — the first pass of this round said that dre.pt and
thefuture3d.com were absent, and was wrong; the published counts only accept a zero after a
second confirming call. Second: Bing serves plausible organic results, but from a different query to
clients without a session — for the same question, four attempts returned results about a bridge in
Dublin, a BBQ restaurant, a hockey game and Reddit pages. The Bing SERP is measured with a human
session; any tool that says otherwise is measuring noise — or is not saying how it measures.
Round 1 adds the two layers that require a human session: the Bing SERP for the ten questions (which
domains are returned) and the five assistants — ChatGPT, Claude, Gemini, Perplexity and Copilot — with the
same questions, recording who is mentioned and who is cited with a link. Monthly cadence; the questions do not
change, so the rounds stay comparable. If you want one of your domains audited in the next round,
write to [email protected] — the retrieval audit is free and takes minutes.