Proofwire
← All writing

Do press releases show up in AI answers?

· 7 min read

We cannot tell you whether an assistant will cite your release, and neither can anyone selling you a number for it. What can be measured is the layer underneath: whether a press page is retrievable by a machine at all. We read the robots.txt of every outlet in our catalogue, and the results are unhelpful to anyone selling AI visibility — several of the best-known mastheads block the major AI crawlers from the whole domain, and one fences off the exact path a release lands on.

Two different claims get sold as one in this industry, and only one of them can be checked.

The first is that a press release produces pages which exist at stable URLs on domains a machine can read. That is measurable. We measured it, and the results include some that are bad for anyone selling AI visibility.

The second is that an AI assistant will cite your company because of it. We have no evidence for that, we cannot generate any that would survive scrutiny, and we are not going to imply otherwise. Competing distributors sell "AI Search Visibility" and "Cited in Google AI Overviews" as line items. We do not, and the rest of this article is why.

What we can show: which pages a machine is allowed to read at all

Every publisher states, in a file at the root of its domain, which automated agents may fetch which paths. The format is a published standard — RFC 9309 — and the rules are read per named agent. We read those files for all 26 outlets in our catalogue on 28 and 29 August 2026. This is the most useful thing anyone can tell you about press releases and AI assistants, and almost nobody tells you it.

OutletWhat its robots.txt saysConsequence for a compliant assistant
AP NewsDisallow: /press-release/* for all agents, plus a whole-domain block on several named AI crawlersPress releases are off-limits, by name, to everyone
MarketWatchDisallow: / for all agents, under a Dow Jones notice prohibiting automated collectionWhole domain off-limits
ReutersAllowlist: roughly 60 named agents permitted, everything else disallowedOff-limits unless the assistant is on the list
USA TodayWhole-domain block on several named AI crawlers; the permissive carve-outs in the file belong to two other companies' agent groups onlyOff-limits to some assistants, permitted to others
Business InsiderWhole-domain block on named AI crawlersOff-limits to those agents
Yahoo FinanceWhole-domain block on named AI crawlersOff-limits to those agents
CoinMarketCap/headlines/* disallowed for everyone; the academy and community paths allowedIts saleable press path is off-limits; its other paths are not
Benzingarobots.txt itself returns 403 behind a bot challengePolicy unreadable — we could not determine it
The remaining 22No relevant restriction foundRetrievable

Read that top row again, because it is the one that matters commercially. AP News is the outlet this industry sells hardest, and AP fences its press-release section off from automated retrieval by name, for every agent. We saw that enforced in production: a search query scoped to apnews.com through one major assistant's search API came back as an error saying the domain is not accessible to its user agent. Not "no results". Refused.

It goes further than the assistants. When we searched for the full text of one specific AP-syndicated release, the text turned up on Morningstar, Yahoo Finance, MarketScreener, 3BL Media, CSRWire and the issuer's own newsroom — and not on apnews.com at all. The AP copy of a release is, on this evidence, the copy least likely to be found by anything.

One publisher goes as far as keeping its paid content out of its own news index: USA Today adds a specific directive excluding its sponsor-story path from Google News, aimed at Googlebot-News, the separate token Google documents for its news crawler. That is a publisher telling a search engine not to treat the thing you bought as news.

The limits of that evidence, stated plainly

A robots.txt file is an instruction that well-behaved crawlers follow. It is strong evidence about permission, not proof about every system in existence:

  1. It governs live retrieval. It says nothing about whether a model absorbed a page during training, which is unobservable from outside.
  2. Compliance varies, and we cannot audit it. Google says as much about the format: the instructions in a robots.txt file cannot enforce crawler behaviour — it is up to the crawler to obey them. We can only show that at least one production system enforced it, which we did.
  3. It changes. A domain that is open today can close tomorrow, without notice.
  4. It is per-agent. USA Today's file is permissive to one company's crawlers and closed to another's, so "can an assistant read this" has different answers for different assistants.

That is still a far better basis for a purchasing decision than a promise about AI citations. It is checkable by you, today, in a browser, for free.

A page that stops existing cannot be cited by anything

Retrievability is worthless if the URL dies, and press pages die more often than the sales material suggests.

Digital Journal has deleted an entire tier of its press archive. We checked its section indexes wire by wire: the paths for ACCESS Newswire, GlobeNewswire and 24-7 Press Release returned 200, while five reseller-wire sections returned HTTP 410 Gone. A 404 means not found; 410 is the status a publisher returns to tell search engines a page is permanently removed and should be de-indexed. Individual articles behave the same way — reseller-wire URLs still sitting in search results return 410 when opened.

We also found one MSN placement that is still findable in search and renders "This story is unavailable", and one MarketWatch URL still cited in a competitor's live case study that returns "Page Not Found". A third outlet changed the published slug on a delivered release, so the URL in the customer's report now redirects rather than resolving.

Against that, plenty of pages last. A release published on Street Insider on 26 February 2026 was still live when we opened it six months later. CoinMarketCap articles from June and July were still live in August. Persistence is real; it is just not universal, and nobody can promise it. We do not sell permanent links and we will not, because we have watched them stop being permanent.

Archiving does not rescue this either. We asked the Wayback Machine's availability API for the fifteen specific real placements in our verification study: two of the fifteen had any snapshot at all.

What we cannot prove

We cannot tell you an assistant will cite you, and we cannot tell you how likely it is. Four reasons, none of which we can engineer around:

  1. We have no citation data. We do not see assistants' query logs or retrieval traces. Nobody outside those companies does. Any "we got you cited" claim in this market is either an anecdote or a screenshot.
  2. Retrievability is a precondition, not a cause. Everything measured above establishes that a page can be read. Whether a system chooses to surface it — and attribute it — depends on ranking behaviour none of these companies documents, and which they change without notice.
  3. The answers are not reproducible. Ask the same question twice and you can get two different sets of citations. A method that gives different results on Tuesday is not a method you can sell a monthly report against.
  4. Duplicate copies compete with each other. Your release is published verbatim at many URLs on the same day; we watched one submission appear on two major finance sites within hours. A retrieval system that surfaces one of them has no obligation to surface yours, or to name your company rather than the wire.

There is a fifth reason that is really about us: a claim we cannot verify after the fact is a claim we cannot stand behind when a customer asks. We can re-open a URL in twelve months and tell you whether your release is still on it — a sample report shows the state of every row on the day it was written. We cannot re-run an AI answer from March and show you what it said.

Why the screenshots are not evidence

You will be shown screenshots of an assistant naming a client. Treat them the way you would treat a screenshot of a stock chart. There is no timestamp you can trust, no way to know how many attempts produced it, no way to know whether the prompt named the company outright, and no way to reproduce it. It is not falsifiable, which means it is not proof — of anything, in either direction. We have not run that test and we are not going to publish one, because a test we ran on ourselves is not evidence either.

The claim we are willing to make

A press release buys you published pages at real URLs, carrying your announcement in your words, on domains that people and machines can find. Where the publisher permits automated retrieval — which is the majority of our catalogue, and specifically not AP News, MarketWatch or Reuters — those pages are readable by an assistant that goes looking.

That is a precondition for being cited. It is not a promise of citation, it is not a ranking, and it is not a number we will put on an invoice.

If someone quotes you a figure for AI visibility, ask them one question: what would have to be true for that number to be wrong, and how did you check? We have never heard an answer.

Where these figures came from

  • robots.txt files read directly on 28 and 29 August 2026 for apnews.com, usatoday.com, marketwatch.com, businessinsider.com, finance.yahoo.com, reuters.com, coinmarketcap.com, benzinga.com and the rest of our 26-outlet catalogue.
  • A query scoped to apnews.com through Anthropic’s search API returned 400 with the message that the domain is not accessible to its user agent — a production system enforcing the directive, observed 28 August 2026.
  • A text search for one specific AP-syndicated release surfaced the same text on Morningstar, Yahoo Finance, MarketScreener, 3BL Media, CSRWire and the issuer’s own newsroom, and not on apnews.com.
  • Persistence and decay, checked 28 August 2026: a Street Insider release from 26 February 2026 still live six months later; CoinMarketCap articles from June and July 2026 still live; one MSN placement still findable in search but rendering "This story is unavailable"; one MarketWatch URL still cited in a live case study returning "Page Not Found".
  • Digital Journal section indexes returning HTTP 410 Gone for five reseller-wire paths (getnews, indnewswire, vehement-media, pr-distribution, binary-news-network) while access-newswire, globenewswire and 24-7-press-release returned 200, checked 28 August 2026.
  • Wayback Machine availability API, queried for the 15 specific real placements in our verification study: 2 of 15 held any snapshot: https://archive.org/help/wayback_api.php
  • RFC 9309, Robots Exclusion Protocol (Internet Standards Track, September 2022) — the standard the files above are written against, including per-user-agent group matching: https://www.rfc-editor.org/rfc/rfc9309.html
  • Google Search Central, "Introduction to robots.txt" — robots.txt "is not a mechanism for keeping a web page out of Google", and its instructions cannot enforce crawler behaviour: https://developers.google.com/search/docs/crawling-indexing/robots/intro
  • Google Search Central, "Google crawlers" — Googlebot-News listed as a robots.txt user-agent token distinct from Googlebot: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
  • RFC 9110, HTTP Semantics, section 15.5.11 — the 410 (Gone) status code: https://www.rfc-editor.org/rfc/rfc9110.html#name-410-gone

We distribute press releases to 300+ outlets, then open every published link and report what actually went live.

View packages