How to tell if your press release was actually published
· 7 min read
A send receipt confirms that a distribution network accepted your release and transmitted it. It does not confirm that a page exists, that the URL resolves to your release rather than to an unrelated article, that the name found on the page is really yours, that the section will not be deleted, or that the URL still works next month. Each of those is a failure we reproduced against live pages in August 2026.
We buy distribution wholesale, so we receive the same document a customer receives. It is worth describing precisely, because almost everything sold as proof of coverage in this category is built from it.
What a send receipt actually contains
When a distribution network confirms that a release went out, the payload carries some or all of the following: an identifier for the submission, a timestamp, the state the job reached in the network's own lifecycle, the circuits or lists it was pushed to, and a set of URLs the network says picked it up. That last part is called a pickup file.
Every one of those facts is true and useful. Notice what each of them is a fact about.
- The timestamp is a fact about the network's queue.
- The lifecycle state is a fact about the network's own workflow. In our supplier's vocabulary, "distributed" is a state that comes after "published" — it describes where the job got to internally, not what an outlet decided.
- The pickup file is a fact about what the network was told, or expected, or observed at the moment it generated the file.
None of them is a fact about a page that exists right now with your release on it. That is the whole distinction, and it is not an accusation about anyone's honesty. A receipt is an accurate record of a send. It is simply not a record of a publication, because the sender is not the publisher.
Six things a receipt cannot see
Each of these is a failure we reproduced against live pages in August 2026, not a hypothetical.
A URL that resolves but is not your release. Requesting a never-existing article ID on a major national daily returned HTTP 200 and 556 KB of a real, unrelated article. Reproduced twice. Any check that treats a 200 as confirmation counts that as a placement. Google has a name for the pattern — a soft 404, a success code served over an error — and it is common enough that Search Console reports it as its own category.
A page that confirms itself. Press-release URLs carry the company name in the slug, and many sites write the requested URL back into the page as a canonical tag and an og:url property. A checker that fetches a URL and then searches the page for the company name can find the name it just supplied. The check proves only that it asked.
A name that matches by accident. We ran real, confirmed-live placement pages against decoy company names. 14 of 36 tests matched: "Apple", "Market" and "Insider" were each found on pages with no connection to them, picked up from navigation and ticker rails. A company called Market Financial gets confirmations off page furniture.
A section the publisher deletes. One outlet segments its press-release archive by originating wire. Five of those sections, all belonging to low-cost reseller wires, returned 410 Gone — the status that tells search engines a page is permanently removed, and one Google acts on by dropping the URL from its index. The three sections that survived all belonged to wires with editorial gatekeeping and higher prices. Individual articles behaved the same way. Which is to say the survival odds of your page are partly a fact about which network carried it, decided before you bought.
A URL that quietly stops resolving. A published slug on one sector blog was changed after delivery, so the originally indexed URL now redirects. One aggregator URL still sitting in the search index rendered "This story is unavailable" when opened.
A removal nobody can detect. Three outlets in our own catalogue return 403 to automated clients. On those, a pulled page and a blocked check look identical from the outside. Deletion there is only ever discovered by a person opening the page.
Why checking is not simply a better crawler
The honest reason the receipt is the standard deliverable is that verification is a labour cost rather than a software cost, and it does not fall with scale.
We measured our own catalogue. Roughly 46% of the placements a customer buys by name can be confirmed by machine. The rest — about 3.1 placements per order — need a person to open the page. Not because the software is bad, but because of policy: AP News disallows its press-release path to all crawlers in robots.txt, whose semantics RFC 9309 standardises; MarketWatch disallows everything and publishes a notice that automated collection is prohibited without express written permission from Dow Jones; Street Insider blocks automated clients at the firewall while its robots file permits the path. Note that the two mechanisms are unrelated — Google's own robots.txt documentation is careful that the file is a request to crawlers, "not a mechanism for keeping a web page out of Google", and it certainly is not a firewall.
At published rates for verified remote assistants — $6.24 to $7.99 an hour across three marketplaces — a three-to-four minute check works out at $1.02 to $1.66 per order, or 16 to 22 hours a month at 105 orders. Repeating that check every month for a year would take it to roughly 90 to 120 hours a month, which is the arithmetic that decided we would not sell ongoing monitoring at all.
Those are small numbers in absolute terms and structural ones in practice: they are recurring human hours, they scale linearly with orders, and no amount of engineering removes them, because the constraint is another company's access policy. That is a sufficient explanation for why receipts are the default. It does not require anyone to be acting in bad faith.
The uncomfortable half, for us: in our flagship package, three of the nine named outlets can never be machine-checked — AP News and MarketWatch disallow it in robots.txt, Street Insider blocks it at the firewall. A person opens those. Benzinga used to make it four; it answered a checker on one of three passes, and it is in no package of ours now. We would rather publish that than sell around it.
What a coverage report has to say, per row
The unit is the row, and the row needs five things.
| Field | Why it is there |
|---|---|
| The URL | A named page, not an outlet logo or a circuit name |
| The date it was opened | A confirmation is a statement about a moment. Without a date it is a claim |
| Who confirmed it | Opened by machine, opened by a person, or reported by the network and not independently confirmed. These are three different things and must never share a label |
| The link attribute | Read off the live page, because it varies by release and by content type on the same outlet — and because Google's definitions of nofollow, sponsored and ugc are what the value actually means |
| The outcome, with a reason | Live, not found, or could not be checked — and when it could not be checked, why: robots disallowed, bot challenge, timeout, wrong content type, name not found |
Two rules matter more than the fields. A check that was blocked is never reported as a missing placement — a 403 is a fact about the checker, not about the page, and converting it into a failure would tell a customer their release was not published when the truth is that nobody looked. And a 200 is never reported as live on its own, because of the first failure mode above.
The corollary is that a good report is partly a list of things nobody could confirm. A report showing only wins is not a report; it is a selection. A sample report is published with those rows filled in, including the ones a person had to open and the ones we could only take from the network's pickup file. None of the five fields needs our software to check, either — the ten-minute version you can run yourself is deliberately published.
Five questions worth asking any distributor
You can ask these before you buy, and the answers are more informative than any sample.
- Which of these URLs did you open yourselves, and on what date?
- What does the report say for a placement you could not open — and does that row look different from a placement that was confirmed?
- What is the link attribute on each page, and did you read it off the page or off a rate card?
- What happens when a page is removed a month from now? Do I hear about it, and how?
- Which outlets in this package block automated checking, and what do you do about those?
The last one is the tell. Every distributor selling AP News is selling an outlet whose robots.txt disallows its press-release path to all crawlers. There is a good answer to that question — a person opens it and the row says so — and there is an answer that avoids it.
What none of this can promise is permanence, a ranking, or that any given outlet chooses to publish. Those are decisions made by publishers and search engines, and the list of things distribution simply cannot do is worth reading before you decide how much of it to buy. What can be promised is that you are told, per URL, what was confirmed, by whom, and when.
Where these figures came from
- RFC 9309, Robots Exclusion Protocol — what a robots.txt Disallow rule does and does not bind: rfc-editor.org/rfc/rfc9309.html
- RFC 9110, HTTP Semantics — the definitions of 403 Forbidden and 410 Gone: rfc-editor.org/rfc/rfc9110.html
- Google Search Central, "Introduction to robots.txt" — robots.txt is "not a mechanism for keeping a web page out of Google": developers.google.com/search/docs/crawling-indexing/robots/intro
- Google Search Central, "How HTTP status codes affect Google Search" — 4xx handling, index removal and the soft-404 category: developers.google.com/search/docs/crawling-indexing/http-network-errors
- Google Search Central, "How to specify a canonical URL" — canonical tags and og:url as page-supplied self-description: developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls
- Google Search Central, "Qualify your outbound links to Google" — the definitions of nofollow, sponsored and ugc: developers.google.com/search/docs/crawling-indexing/qualify-outbound-links
- Supplier fulfilment payloads and lifecycle vocabulary: our own integration research, August 2026 — "distributed" is a supplier lifecycle state that follows "published".
- Soft-404 reproduction (HTTP 200 with an unrelated article for a non-existent article ID), reproduced twice on 28 August 2026.
- Decoy company-name testing against confirmed-live placement pages: 14 of 36 tests matched, 28 August 2026.
- Verification feasibility measurements, 28 August 2026: 46% of named placements machine-confirmable, 3.13 human escalations per order, $1.02–$1.66 per order at published VA rates of $6.24–$7.99 per hour.
- robots.txt for apnews.com and marketwatch.com (including the Dow Jones notice on automated collection), and Street Insider WAF behaviour, read 28 August 2026.
- Digital Journal section index status codes (410 Gone on five reseller-wire paths, 200 on two gatekeeping wires) and an MSN URL rendering "This story is unavailable", checked 28 August 2026.
We distribute press releases to 300+ outlets, then open every published link and report what actually went live.
View packages