Discussion starter: a technical SEO question, posted with an answer from Niraj Raut to open the thread. If you have dealt with this on a site, reply with what you saw, especially where it differs.
- Crawling
- Indexing
- Technical SEO
Logs show Googlebot spending most of its hits on expired listings on a 2M-URL classifieds site. Would you 410 them in bulk, noindex them, or let them age out?
Disclosure: Niraj Raut, who posted this answer, runs the SEO consultancy linked at the end of it.
1 reply
-
adminAdmin
The short answer: for listings that are permanently gone, return a 410 (or 404) at the server, and build that into the listing lifecycle rather than doing a one-off purge. Noindex is the wrong tool for a crawl problem, because Google has to fetch a page to see the noindex. “Let them age out” only works if the expired pages already return a 4xx. If they return 200 with an “ad expired” notice, they are soft 404s, and Google’s own documentation says those keep getting crawled.
The main qualification is that a high share of hits on dead URLs is not automatically a problem. Googlebot rechecks URLs it already knows about, and Google has said that it does this with spare capacity. On a classifieds site the decision should depend on two things more than on the raw share of hits: what status code expired listings return today, and how quickly new listings get crawled.
First, work out whether those hits are waste or leftovers
These two situations look identical in a pie chart of Googlebot hits, but they need opposite responses.
Expired listings return 200
This is real waste. Google’s crawl budget guide says soft 404 pages “will continue to be crawled, and waste your budget”. Every recrawl of a dead ad is a fetch that could have gone to a new one.
Expired listings already return 404 or 410
This is mostly housekeeping. John Mueller has described Google rechecking known 404 and 410 URLs “when we have nothing better to do on this website”, and said it is not costing crawl capacity. If new listings are crawled promptly, there may be nothing to fix.
What I would check before changing anything
- Verify the bot. Classifieds sites attract heavy scraping, and scrapers often spoof the Googlebot user agent. Filter logs to requests that pass reverse DNS or match Google’s published IP ranges, which moved to a new location in March 2026. An unverified “Googlebot” share can be badly inflated.
- Join logs to your listings database. For each hit, record listing status (active or expired), days since expiry, the status code served, response time, whether the URL is in a sitemap, and whether it has internal links pointing at it.
- Measure what you actually care about. Median time from listing publication to first verified Googlebot request, the share of active listings crawled in the last 7 days, and how many active listings sit in “Discovered - currently not indexed”.
- Cross-check the Crawl Stats report. Its breakdowns by response and by purpose (discovery means a URL crawled for the first time, refresh means a recrawl) tell you whether new URLs are getting discovery crawls while old ones soak up refresh crawls.
- Find where Googlebot keeps finding expired URLs. Common sources are sitemaps that are never pruned, “related listings” and “recently viewed” modules, seller profile pages listing past ads, and archive pagination.
- Check whether any expired URLs still earn anything. Pull organic clicks from Search Console and referring domains for the expired set. On most classifieds sites this is a small tail, but it decides which URLs get a different treatment.
What you find Likely interpretation Action Expired URLs return 200 and new listings take days to get a first crawl Dead pages are competing with live inventory for crawl High priority: move expired listings to 410 and prune discovery sources Expired URLs return 200 but new listings are crawled within hours Index quality issue more than a crawl issue Still move to 410, but with less urgency Expired URLs already return 4xx and new listings are crawled quickly Normal rechecking of known URLs Leave it; fix nothing that is not broken Expired URLs return 4xx but new listings are still slow The bottleneck is elsewhere Check server response time, parameter and facet URLs, and link depth to new listings How each option behaves once Googlebot arrives
A useful way to picture the difference: a 410 is a house marked “moved out, no forwarding address”. The postie still walks past occasionally but stops knocking. A noindex page is a house where someone opens the door every time and says “please don’t write about us”. The message is respected, but the postie has to knock to hear it, every time. Technically, Google has to request a URL and receive the response before it can act on either signal, but a 4xx tells it there is nothing there, while a 200 with noindex tells it there is a live page it should keep checking.
Option Crawling Indexing Main risk 410 or 404 Crawl frequency for those URLs gradually decreases; Google says a 404 is “a strong signal not to crawl that URL again” Removed from the index A bug that serves 410 to active listings Noindex on a 200 page Still requested; Google’s crawl budget guide says not to use noindex for this reason Dropped once the tag is seen Crawl waste continues indefinitely 200 “ad expired” page (ageing out) Soft 404s continue to be crawled Usually dropped as soft 404, slowly and inconsistently The status quo you are trying to escape 301 to the homepage or a loosely related category Redirect is followed each time For homepage redirects, Mueller has said Google mostly treats them as soft 404s Confuses users with no SEO upside robots.txt disallow Blocked URLs “stay part of your crawl queue much longer” Google cannot see a 410 or noindex behind the block Needs a separate URL pattern for expired ads, which most sites lack On 404 versus 410: Google’s documentation says all 4xx codes except 429 are treated the same, and Mueller has said the difference is “so minimal that I can’t think of any time I’d prefer one over the other for SEO purposes”. Pick 410 for listings that are gone for good and 404 if the same URL can come back when a seller relists, mainly because it keeps your own reporting honest.
Two details trip people up. First, 4xx responses “have no effect on crawl rate”, per the same documentation. A bulk 410 does not make Googlebot back off your host, it lowers demand for those specific URLs, and any freed crawl only goes to active listings if Google has reason to want them. Second, Google’s JavaScript documentation now says rendering might be skipped for non-200 pages. So the body of a 410 page is for users, not Google, and the expired state must be decided on the server. A single-page app that flips to “expired” in JavaScript after returning 200 gives Google the worst of both.
The approach I would take: expiry as part of the listing lifecycle
Listing goes live: 200, in an active-listings sitemap with accurate lastmod,unavailable_afterset to the expiry date if you know it↓Listing expires: remove it from sitemaps and from related-listing and seller-profile modules the same day↓Can the same URL be relisted soon? If yes, keep a short grace period on 200 before it returns 404↓Does it earn organic clicks or have referring domains? If yes, 301 only to a genuinely equivalent page (same model and location), otherwise treat as below↓Everything else: 410, with a useful body showing similar active listings and a search boxThe
unavailable_afterrobots rule is underused on classifieds. It tells Google not to show the page after a given date, and Google’s documentation says Googlebot “will decrease the crawl rate of the URL considerably” after that date. Setting it at publication means you are not waiting for a batch job to catch up.If part of the site is a jobs board, the rules differ slightly. Google’s JobPosting documentation lists a past
validThroughdate, a 404 or 410, or removing the markup as valid ways to expire a posting, and recommends the Indexing API for job URLs. The Indexing API is limited to job postings and livestream pages, so it is not an option for general listings.Hypothetical example:
A classifieds site has 2 million known URLs: 350,000 active listings and 1.65 million expired ones. Expired ads return 200 with a “this ad has expired” template and remain in paginated seller profiles. Verified Googlebot requests split 70/30 in favour of expired URLs, the average listing lives 21 days, and the median new listing waits 3 days for its first crawl. Here the fix is urgent: a meaningful slice of each listing’s short life passes before Google has seen it. Change one fact, so that expired URLs already return 404 and new listings are crawled within hours, and the same 70/30 split is harmless.
Rolling out a bulk 410 without breaking live inventory
- Start with the safest cohort. Listings expired for months, with no organic clicks and no referring domains. Expand in stages once monitoring shows no active listing ever returned a 4xx.
- Derive status from the database, server side. Add an automated check that samples active listing IDs daily and alerts on any non-200 response. This is the failure that costs revenue.
- Watch CDN caching. A long cache lifetime on a 410 can keep serving “gone” after a seller relists. Keep 4xx cache times short for URLs that can come back.
- Fix discovery sources at the same time. A 410 on a URL that is still linked from 50,000 seller profile pages will keep being rediscovered and rechecked.
- Expect a long tail. Google says it “won’t forget a URL that it knows about”. Some requests to 410 URLs will continue for months, and that is normal.
Watch for this:
Do not judge success by hits on expired URLs reaching zero. Judge it by active-listing crawl freshness and indexing. Chasing zero leads to robots.txt blocks, which keep URLs in the crawl queue longer and hide the 410 from Google.
How to measure it, and what would change the plan
Crawl allocation is a site-level effect, so it cannot be split-tested cleanly. You can split the expired cohort by listing ID (say odd IDs get 410, even IDs stay as they are) to measure how fast per-URL recrawl rates decay, but the benefit to new listings has to be read from a before-and-after comparison. Mark the deployment dates with Search Console’s custom annotations and note any core updates or infrastructure changes in the same window.
- Primary metrics: median time to first crawl for new listings, share of active listings crawled in 7 days, and share of verified Googlebot requests reaching active listings.
- Secondary metrics: indexed active listings, organic clicks to listing and category pages, and organic enquiries or seller contacts, since traffic that does not produce replies is not the goal.
- Timeframe: allow 8 to 12 weeks, because recrawl rates for known URLs decay gradually rather than switching off.
Two results would change the plan. If a noticeable group of expired URLs earns clicks for model or location searches, the demand is real and deserves an evergreen page (a model-in-city page, for example) that those URLs can redirect to. And if time to first crawl for new listings does not improve after the change, the constraint is not the dead inventory. Look at server response times, facet and parameter URLs, and how many clicks deep new listings sit.
If this were my site, I would 410 the long-dead, zero-value cohort first, stop feeding expired URLs into sitemaps and internal modules on the day they expire, set
unavailable_afteron new listings, and keep noindex only for pages that must stay live for users. Then I would judge the work entirely by how quickly new listings get crawled and indexed.Need help with this on your own site?
Niraj Raut works with businesses in Nepal, Australia, the UK, Europe and the US on technical SEO, ecommerce SEO, local SEO and AI search.
Add a reply Cancel reply
Replies are open to approved members. Log in or apply to join.