Discussion starter: a digital PR question, posted with an answer from Niraj Raut to open the thread. If you have dealt with this on a site, reply with what you saw, especially where it differs.
- Digital PR
- Indexing
Digital PR links from news sites often sit on pages that later get noindexed or pruned. How much link decay have you measured after 12 months?
Disclosure: Niraj Raut, who posted this answer, runs the SEO consultancy linked at the end of it.
1 reply
-
adminAdmin
I can’t hand you a clean “X% lost after 12 months” figure, and I would be wary of anyone who quotes one without showing how they decided a link was lost. I could not find a public benchmark that tracks digital PR placements on news sites over 12 months with a transparent method. What does exist says attrition starts early: Pew Research found that 8% of pages that existed in 2023 were already inaccessible by October 2023, and about one in five pages from its 2021 sample were gone two years later. Those figures cover web pages in general. They don’t cover news articles specifically, and they don’t measure links.
The more useful point is that “link decay” covers several different events. A link can stop counting while the article still loads, because the page has dropped out of Google’s index or the link was stripped or qualified. A link can also show as “lost” in a third-party tool while the article is perfectly healthy. Your 12-month number will depend far more on your placement mix (original articles versus syndicated copies, which publisher groups, which sections) and on your definition of “lost” than on news sites as a category.
What the public data measures, and why none of it is your 12-month number
Source What it measured Relevant finding Why it doesn’t answer the question Pew Research Center (2024) About 1 million pages sampled from Common Crawl snapshots between 2013 and 2023, checked in October 2023. A page only counted as gone if it returned an error code showing the page or host no longer exists. 38% of 2013 pages inaccessible; 8% of 2023 pages; roughly one in five 2021 pages after two years. Separately, 23% of news pages had at least one broken outbound link. Measures page death only, not noindex, link removal or rel changes. The news finding is about links on news pages pointing elsewhere, not news pages disappearing. Ahrefs link rot study (2022, updated 2024) Links pointing to about 2 million domains, tracked from 2013 onward in Ahrefs’ own crawl data. 74.5% of links counted as lost. Of lost links: 34.2% removed from pages that still exist, 47.7% on pages dropped from Ahrefs’ index, 5.99% redirects, 0.82% canonical changes, 0.73% noindex. Cumulative over roughly nine years with no breakdown by age. “Dropped from index” means Ahrefs’ index, not Google’s. Covers all link types, not PR placements. Two lessons carry over. First, links removed from pages that still load are a large share of the loss in the Ahrefs data, so a “does the page still return 200?” check will undercount decay. Second, the biggest bucket is a crawler-index state, which Ahrefs says can mean a page couldn’t be crawled or indexed or that the domain no longer exists. A tool’s “lost” count mixes real deaths with the tool’s own crawl decisions.
The five states a PR link can be in a year later
State What usually happened What Google has said How to detect it 1. Live, indexed, followed Nothing changed Can still be ignored; being on an indexed page is not a guarantee it counts Crawl plus an index check 2. Live but qualified Publisher added nofollow, sponsored or ugc, often through a sitewide policy Since 2019 these attributes are treated as hints, not blocks Crawl and read the rel attribute 3. Live, link removed Content refresh, correction, legal complaint, link policy Nothing specific; the link simply no longer exists Crawl for any link to your domain 4. Live but out of Google’s index Noindexed during pruning, syndicated copy kept out of the index, or crawled and not indexed Long-term noindex leads Google to drop the page and stop following its links; for pages that aren’t indexed, “I wouldn’t assume the links do anything” Meta robots, X-Robots-Tag, canonical, plus an index proxy 5. Gone or redirected 404 or 410, redirect to a section front or homepage, domain lapsed The linking page no longer exists in a usable form Status code and redirect target Important distinction:
“Lost in a tool” and “lost to Google” are different measurements. State 4 is the one most reports miss: the article loads, the link is there, it is followed, and your crawler reports it as healthy. If Google has dropped that page from its index, John Mueller’s advice is not to assume the link does anything. He also said links on indexed pages can be ignored, so a surviving link is a ceiling on value, not proof of value.
Syndicated copies carry most of the decay from day one
In May 2023 Google changed its documentation: it no longer recommends cross-domain canonicals for syndicated content, and says partners who want to avoid duplication should block the republished copy from being indexed. Many regional newspaper networks and news aggregators republish the same story across dozens of mastheads. If a campaign report counts every copy as a link, it is counting pages that, under Google’s own guidance, may be built to stay out of the index. Those links don’t decay over 12 months. Many of them were never on pages Google was likely to use.
Hypothetical example:
A campaign report lists 55 linking pages: 9 original articles on different publishers and 46 syndicated copies across two regional networks. At 12 months a crawl shows 49 still load with the link, which looks like 11% decay. Now suppose 40 of the syndicated copies carry noindex or can’t be found in Google, and 2 of the 9 originals had the link removed during an article update. The number that matters (independent, indexable, followed placements) went from 9 to 7, a 22% loss. Same campaign, two very different decay figures. These numbers are illustrative, not measured.
Noindexing and pruning: what is documented and what is guesswork
Google’s position on noindex and links comes from Mueller in 2017: in the short term Google may still follow links on a noindexed page, but once the noindex has been there long enough, the page is removed completely “and then we won’t follow the links anyway”. When asked how long that takes, he said it depends, partly on crawl frequency. For old news articles that get recrawled rarely, that transition could plausibly be slow, but nobody outside Google can time it precisely.
Pruning is a publisher business decision, not a Google requirement. When CNET deleted thousands of older articles in 2023 (using pageviews, backlinks and time since last update to decide whether to redirect, repurpose or deprecate), Google’s Search Liaison responded: “Are you deleting content from your site because you somehow believe Google doesn’t like ‘old’ content? That’s not a thing!” He added that removal might help crawling on a massive site. Publishers still prune for crawl, storage, legal and commercial reasons. Once the news cycle ends, most PR placements become exactly the low-traffic archive pages those criteria pick out. That is my reasoning, not a measured rate.
What nobody has shown well is whether a link’s value lingers after it disappears. Some practitioners describe rankings holding for a while after links are lost. I haven’t seen a controlled test that separates that from normal ranking noise or from slow recrawling of the linking page, so I would not bank on it in either direction.
How I would measure your own decay curve
- Build a placement ledger on day one. For each linking URL record the publisher and publisher group, original or syndicated, section, date found, anchor, target URL, rel values, meta robots, X-Robots-Tag, canonical target and status code. Save an archived copy as evidence of the original state.
- Split originals from syndicated copies before you report anything. Report copies as reach, not as links.
- Recrawl the cohort on a fixed schedule, for example at 30, 90, 180 and 365 days, using a crawler in list mode with custom extraction for links to your domain and their rel values. Check rendered HTML as well as raw HTML, because some news templates inject body content with JavaScript.
- Estimate Google-side status with proxies, and label them as proxies. URL Inspection only works on properties you own, and Mueller has said there is no API for testing whether a page is indexed. An exact-match search for the article headline is a rough check. Search Console’s Links report helps, but Google says it shows a sample and may omit URLs such as non-indexed pages, so a disappearance there is a weak signal, not proof.
- Report survival by cohort, publisher group and state rather than one blended percentage.
- Before chasing removed links, find out why they were removed. A correction or legal edit needs a different conversation from a routine content refresh.
Fetch the linking URL (raw and rendered)↓4xx, 5xx, or redirected somewhere else? Yes: state 5↓Link to your domain still present? No: state 3↓Noindex, canonical to another URL, or not findable in Google? Yes: state 4↓nofollow, sponsored or ugc added since placement? Yes: state 2↓Otherwise state 1: live and followed (still not proof it is counted)Which patterns should change what you do
Pattern at 12 months Likely explanation What I would change Most loss sits in syndicated copies Copies noindexed, canonicalised or pruned by the network Stop counting copies in link KPIs and pricing; report them as reach Loss concentrated in one publisher group CMS migration, archive policy or section closure Weight that group lower when targeting; ask about archive policy before relying on it Links removed from articles that still load Edits, corrections, tightened outbound link policy Reclaim quickly, and favour hooks where your link is the source of the data or tool, which makes it harder to remove Links changed to nofollow or sponsored Publisher-wide link policy Treat as a hint, not zero, but don’t value it like a followed editorial link Originals mostly survive but target pages didn’t move Links were not the constraint on those pages Re-examine whether PR is the right lever for those targets before the next campaign Watch for this:
Survival is not impact. A link that survives 12 months could still be counted. It doesn’t tell you the link moved rankings. If you use decay figures to justify spend, pair them with ranking and traffic movement on the target pages, and be honest that the attribution is loose.
If this were my campaign, I would stop reporting “links landed” and start reporting “independent, indexable, followed placements surviving at 6 and 12 months”, segmented by publisher group, with syndicated copies as a separate reach metric. After two or three cohorts you will have your own decay curve, and it will be more useful than any industry average because it reflects your publishers and your hooks. When people in this thread do share numbers, ask for the method alongside them: whether copies were counted, which tool defined “lost”, and whether noindexed pages were treated as alive.
Need help with this on your own site?
Niraj Raut works with businesses in Nepal, Australia, the UK, Europe and the US on technical SEO, ecommerce SEO, local SEO and AI search.
Add a reply Cancel reply
Replies are open to approved members. Log in or apply to join.