Discussion starter: an SEO measurement question, posted with an answer from Niraj Raut to open the thread. If you have dealt with this on a site, reply with what you saw, especially where it differs.
- Search Console
- SEO Analytics
Search Console’s branded queries filter: have you checked its classification against your own brand regex, and how far off was it?
Disclosure: Niraj Raut, who posted this answer, runs the SEO consultancy linked at the end of it.
1 reply
-
adminAdmin
Expect the headline numbers to disagree, sometimes by a factor of two or more. Most of that gap usually has little to do with Google misclassifying queries. It comes from the two methods counting different populations of queries and using different definitions of “brand”. Once you strip those two things out, the row-level disagreement that is left is the only part that measures the classifier’s accuracy, and it is usually much smaller than the headline gap.
I would not put much weight on anyone’s single “it was X% off” figure, because the result depends on how distinctive the brand name is, how many product and people names the business owns, how much traffic sits in anonymised queries, and how carefully the regex was written. What is useful is a way to split the gap so the number you end up with actually means something.
What Google has documented about the classifier
Google announced the filter in November 2025 and expanded it to all eligible sites in March 2026. It describes an internal AI-assisted system that treats a query as branded when it includes the brand name in different languages, misspellings of it, or references to products or services unique to the brand. The Search Console help page adds limits that matter before you compare anything:
- It is not available for sites with a low number of impressions, and Google has also said it is not available for sub-properties such as a
/blog/URL-prefix property. - The classified history starts in March 2025, so you cannot use it to restate earlier years.
- Google states that “some queries might be incorrectly identified as branded or non-branded” and that the classification does not affect ranking.
You also cannot tune it. Asked whether site owners can add or suggest brand terms, John Mueller said there is no customisation at the moment and he was not aware of plans for it. So when you find a disagreement, the only thing you can change is your own regex, or your own written definition of brand.
The three gaps hiding inside “how far off was it”
Most comparisons put one number (brand share of clicks from the filter) next to another (brand share of clicks from a regex run over exported query rows). That single comparison mixes three different things, and only one of them is an error.
Gap What causes it How to detect it Is it a classifier error? Population gap Your regex only sees query rows Search Console displays. Google classifies internally and may be counting queries you never see. The sum test below No. The two methods have different denominators. Definition gap Google’s “brand” includes typos, other languages and unique products. Your regex may include founders or domain-name queries, or leave out product lines. Row-level review of disagreements No. It is a policy choice you should make explicitly. Classification error The classifier misses a term both definitions agree is brand, or tags a generic query as brand. Row-level review of disagreements Yes The population gap, and a quick test for it
Google’s help page says anonymised queries (queries too rare to show, withheld for privacy) are “included in chart totals, unless a query filter is applied (for example, ‘queries containing’ or ‘queries not containing’ a given string)”. Your regex is a query filter. Whether you apply it in the interface or run it over exported rows, the anonymised long tail drops out.
What Google does not document is whether the branded filter behaves like a regex filter in this respect, or classifies anonymised queries internally. Practitioner write-ups disagree. Spotibo’s March 2026 analysis states that the filter covers anonymised clicks, while Patrick Stox’s guide says branded and non-branded clicks will not add up to total clicks. You can settle it for your own property in a couple of minutes:
- Note total clicks for a 28-day window with no filters applied.
- Apply the branded filter and note clicks.
- Switch to non-branded and note clicks.
- If branded plus non-branded roughly equals the unfiltered total, the filter is classifying traffic your regex can never see. If the sum falls short by roughly your anonymised share, both methods are working from similar populations.
If the sums match, that alone can explain a dramatic headline gap. My reasoning: people search a brand name in fairly repetitive ways, so branded queries tend to clear the privacy threshold and appear as visible rows, while the rare, long, varied queries that get anonymised lean non-branded. A regex run only on visible rows would then overstate brand share. Spotibo’s comparison of ten projects over two weeks is consistent with that: the regex method overstated brand share by more than 50% in six of the ten, and one project showed 52.5% brand by regex against 18.9% by the filter. Ten projects from one agency is a small sample, so treat it as a plausible mechanism, not a benchmark for your site.
The definition gap is where the useful work is
Once the populations match, most remaining disagreements come down to what counts as “brand”. These are the categories I would check first:
- Product and sub-brand names. Google’s definition includes unique products or services. Many brand regexes only cover the company name.
- Misspellings, spacing and transliteration. “brandname”, “brand name”, “brand-name”, and non-Latin spellings in other markets. The classifier is meant to catch these. Many regexes do not.
- People. Founders, executives and authors. Ann Smarty’s review found the filter tagged a founder’s name as branded but missed the title of the founder’s book, and it also included some queries for unrelated executives, products and competitor names.
- Brand names that are ordinary words. If the brand is also a dictionary word, a regex over-matches every query containing that word. A contextual classifier may not. This is one place where the filter could plausibly beat a regex.
- Third-party brands you sell. For retailers and marketplaces, check how manufacturer brand queries are treated. They should not count as your brand in either method.
- Comparison queries. “yourbrand vs competitor” matches a simple regex. Check how the filter treats it, and decide whether it belongs in brand reporting at all, since it often behaves more like consideration-stage demand than navigation.
How I would run the comparison properly
- Pick a clean window and compare clicks first. After Google stopped supporting the
&num=100parameter in September 2025, many sites saw desktop impressions drop, while clicks were affected far less, according to Brodie Clark’s analysis. Brand share of impressions across that date is not a like-for-like series. - Run the sum test so you know whether you are comparing like populations.
- Export the query table twice, once under the branded filter and once under non-branded. The interface export stops at 1,000 rows, so on larger sites split the export by country, device or directory to reach more of the tail.
- Compute both shares on the same rows. Take the filter’s share as branded export clicks divided by the clicks in both exports combined, and run your regex over exactly those rows. This removes the population gap from the comparison.
- Build a two-by-two matrix weighted by clicks, not by query count. A thousand one-click typos matter less than one navigational query with thousands of clicks.
- Read the disagreement cells by hand, highest clicks first, and label each row as a definition choice or a genuine error.
Matrix cell What it usually contains What to do Both say branded Core brand and navigational queries Nothing. This is your agreement baseline. Both say non-branded Generic demand Nothing Google branded, regex not Typos, other languages, product names, people Mostly regex gaps, so add them. Some are definition choices (a founder’s name, say). A few are errors. Regex branded, Google not A brand word used generically, third-party brands, newer products Google has not yet associated with you Tighten the regex if it over-matches. Otherwise record it as a known classifier miss. Hypothetical example:
A B2B software site gets 46% brand share of clicks from its regex and 24% from the filter’s headline figure. The sum test shows branded plus non-branded equals the unfiltered total, so the filter includes anonymised traffic the regex never saw. Computed on the same exported rows, the two methods now agree much more closely, and most of the remaining disagreement sits in “Google branded, regex not”: misspellings and two product names. A handful of rows are genuine errors. The accurate summary is “our regex and the classifier mostly agree on visible branded clicks, and the headline gap was a denominator problem”, not “Google is off by half”. (All figures invented for illustration.)
Which one to trust for which job
Job Search Console’s filter Your own regex Headline brand share including the long tail Better, if the sum test shows it covers anonymised queries Likely to overstate brand share when the anonymised tail is large Long-running KPI trend Weaker: history only from March 2025, and you cannot inspect or freeze the model Better: stable, versioned and documented Finding brand variants you missed Strong for typos, languages and products Only as good as your list Brand names that are common words Plausibly better, because it is contextual Over-matches unless written carefully Sub-properties, low-volume sites, joins with other data Not available for sub-properties or low-impression sites. Spotibo reported no API access in March 2026, so check the current API reference before building a pipeline on it. Portable to the API, BigQuery exports, paid search and rank tracking Several brands on one domain You cannot choose which brands count You define each brand group Watch for this:
Switching reporting from a regex to the filter mid-series creates a fake trend break. If you switch, report both side by side for an overlap period and mark the date with a Search Console custom annotation. Also remember the model is Google’s and can change. If brand share steps up or down with no matching movement in other brand-demand signals (brand campaign impressions in paid search, direct traffic, Google Trends), consider a classification change before telling stakeholders brand demand moved.
What I would report as the accuracy number
If this were my reporting, I would stop quoting the headline brand-share gap as the classifier’s error, because it mostly measures denominators. The accuracy figure worth sharing is the click-weighted agreement on the same visible rows, plus a short list of the disagreement types behind the rest. I would use the filter for the headline split (once the sum test confirms what it covers) and as a free discovery tool for brand variants my regex misses. I would keep the regex as the definition of record for trend KPIs and anything that has to join with other data, and I would re-run the two-by-two matrix each quarter, because a model you cannot inspect needs periodic checking against a definition you can.
Need help with this on your own site?
Niraj Raut works with businesses in Nepal, Australia, the UK, Europe and the US on technical SEO, ecommerce SEO, local SEO and AI search.
- It is not available for sites with a low number of impressions, and Google has also said it is not available for sub-properties such as a
Add a reply Cancel reply
Replies are open to approved members. Log in or apply to join.