Book a strategy call Book a call

Methodology v2.8

The method.

Published in full, including the weights and the parts that are a judgement call. A measurement nobody can argue with is not a measurement, it is an assertion.

01What the index measures

Australians find brands in four different places now, and a brand can be dominant in one and invisible in the next. The index measures all four on the same scale, every month, and publishes the whole field rather than a shortlist.

It measures discovery, not quality. Nothing here says a brand is better, cheaper or more trustworthy than another. It says how findable it is when someone goes looking for what it sells.

The unit that gets scored is a board: one subcategory, one fixed keyword set, every brand in it. Categories are navigation. There are 59 live boards across 9 categories in the August 2026 edition.

02The four channels and their weights

The presence score blends four channels on fixed weights:

Search discoverability
45%
AI visibility
30%
Social discoverability
15%
Brand demand
10%

Why the weights sit where they do

Search discoverability 45%

Organic ranking is the largest discovery surface in this market and the foundation the others stand on. Assistants retrieve from the web, so ranking for a category’s buying terms is a precondition for being cited when someone asks about them.

It carries five points more than it did before methodology 2.3. They came across from brand demand.

AI visibility 30%

An AI answer to a commercial question is not content about a shortlist. It is the shortlist, and being named is being considered. The authority that earns a citation also earns the organic rank, so the two channels compound rather than duplicate.

Answers vary between runs. That is shown in every brand’s edition history rather than smoothed away, and it is not handled with a lower weight. Weight says how much a surface matters, not how cleanly it can be measured.

Social discoverability 15%

A genuine third surface, and the one most likely to surface a challenger, because social rewards new entrants in a way branded search volume never does. It mines the brands named inside ranking content rather than ranking accounts, so a category dominated by creators still reports which brands those creators name.

Reddit joined YouTube, TikTok and Instagram at methodology 2.4 on measurement, not intuition. Across 60 published brands it cut the channel’s zero rate from 47% to 32%, and its correlation with AI citation was −0.20, so it reads an independent signal. IKEA scored zero for mattresses on every video platform and 68.4 on Reddit.

The weight stays at 15% because this is the least mature measurement of the four. Instagram cannot be geo-filtered and covers Reels only, and a repeated identical query returns roughly a fifth different results. Those limits are disclosed here rather than priced into the weight.

Brand demand 10%

The most reliable channel and the least earned. It records existing awareness, correlates with size, and at any meaningful weight turns the index into a ranking of who is biggest, which is the thing this methodology exists to avoid.

What a channel does to a ranking is set by its spread, not its level. On the first four-channel run brand demand’s standard deviation was 24.5 against roughly 14 for every other channel, so at a nominal 15% it was driving 22% of the ranking in the published top 20. Cut to 10% at methodology 2.3, it drives 15%, which is what the weight was always meant to mean.

Calibrated on one category, checked on another

The same weights apply to every board and every edition. They were derived from mattresses and then checked against online furniture, a deliberately different category, before being kept. The table shows how much of each board's ranking spread every channel actually drove across the published top 20, against the weight it was meant to carry.

Channel Weight Mattresses Online furniture
Search discoverability 45% 40.9% 48.7%
AI visibility 30% 25.5% 25.4%
Social discoverability 15% 18.4% 11.6%
Brand demand 10% 15.2% 14.3%

Brand demand runs slightly hot in both and is left there. Two boards are a thin basis for a second cut, and the channel exists as an anchor so that an obscure brand cannot top a table on citations alone. Social moved most between the two boards. That is a real difference between categories, and it belongs in the channel scores, which are published unweighted, rather than in a weight no reader ever sees.

A presence score of 72 means the same thing in retail as it does in travel, which is what lets boards and editions be compared at all. Every channel score is published unweighted, so anyone can re-weight and reproduce a different composite.

Read the full rationale as it appears in the data export

Search Discoverability carries the most weight because organic ranking is both the largest discovery surface in this market and the foundation the others are built on: assistants retrieve from the web, so ranking for a category's buying terms is a precondition for being cited when someone asks about them. AI Visibility is second because an AI answer to a commercial question is not content about a shortlist, it is the shortlist — being named is being considered — and because the authority that earns a citation also earns the organic rank, so the two legs compound rather than duplicate. It is measured across seven engines in a depth no other ranking covers.

AI answers do vary between runs, which is why the index publishes monthly and carries every brand's full per-edition history rather than treating one reading as settled: the volatility is shown rather than smoothed away, and it is not handled by a lower weight, because weight is meant to express how much a surface matters and not how cleanly we can measure it. Social Discoverability holds 15%. It is a genuine third surface rather than a proxy for the other two, and the leg most likely to surface a challenger, because social rewards new entrants in a way branded search volume never does.

It mines the entities named inside ranking content rather than ranking accounts, so a category dominated by creators still reports which brands those creators name. 4 - YouTube, TikTok, Instagram and Reddit - having run on the three video platforms alone until 2026-08-09. 20, so it reads an independent signal rather than restating one.

It is not a better platform than the three it joined - alone it scores more zeros than they do - it is blind to different brands, which is the only reason to add a fourth. 4 on Reddit. The leg is still held at 15% rather than higher because its measurement is the least mature of the four legs: Instagram results cannot be geo-filtered and cover Reels only, and repeat calls to an identical query return roughly a fifth different results.

Those limits are disclosed rather than quietly priced into a weight. 3 on measurement rather than opinion. It is the most reliable of the four and the least earned: it records existing awareness and correlates with size, and at any meaningful weight it quietly turns the index into a ranking of who is biggest, which is the thing this methodology exists to avoid.

Its influence had to be measured rather than assumed, and the right measure is spread and not level: a leg's median shifts every brand's score together and moves nobody up or down, while its standard deviation is what separates them. 5 against roughly 14 for every other leg, so at a nominal 15% it was driving 22% of the ranking across the published top 20 and 24% across the top 50 - the exact failure this paragraph warns about, happening at the weight that was supposed to prevent it. At 10% it drives 15%, which is what the weight was always meant to mean.

The five points went to Search Discoverability, which was running slightly under its intent at 35%. The same weights apply to every edition, and the calibration is uniform rather than per-category. It was derived from mattresses and then checked against online furniture, a deliberately different category, before being kept.

4% against 30. 3% against 10, and is left there rather than cut again: its spread is itself category-dependent, two subcategories is a thin basis for a second adjustment, and the leg exists as a sanity anchor so that an obscure brand cannot top the table on citations alone. 6% in furniture, which is the genuine category difference this methodology says belongs in the leg scores rather than in a weight - social search matters more to a mattress buyer than to someone buying a dining table, and the unweighted leg reports that instead of hiding it.

A Presence Score of 72 means the same thing in retail as it does in travel, which is what allows editions to be compared at all. Where categories genuinely differ, and they do — social search matters far more in travel than in finance — that difference belongs in the leg scores, which are published unweighted, rather than hidden inside a weight no reader ever sees. 0, published so it can be argued with; every leg score is reported unweighted so any reader can re-weight and reproduce a different composite.

The weights are the same for every board and every brand, published here before anyone asks, and unchanged within a methodology version. A weight that moved to suit a result would make the whole index worthless.

03How each channel is read

Each channel is scored out of 100 on its own terms. What a channel excludes is stated as plainly as what it counts, because that is where most disagreements about a number actually start.

Search discoverability

Counts. Where the brand ranks across the board keyword set.

Does not count. Paid placement of any kind. An ad above the organic result is not the brand ranking, and a board that counted it would rank budgets.

AI visibility

Counts. How often it is named across the questions put to seven engines.

Does not count. Whether the answer was flattering. Being named in a comparison that ranks you third is still being named, and sentiment is not part of the score.

Social discoverability

Counts. Whether it surfaces against searches made inside the platforms.

Does not count. Follower counts, engagement and posting volume. This measures whether a brand surfaces against a search made inside the platform, not how large or busy its account is.

Brand demand

Counts. How many people search for the brand by name.

Does not count. Anything but the brand name itself. Category searches belong to the search channel, so demand cannot be inflated by ranking for generic terms.

The engines

chatgpt
30%
copilot
5%
gemini
20%
google ai mode
10%
google ai overviews
30%
perplexity
5%

0, not a measured usage share. AI Overviews carries the widest reach because it appears inside mainstream Google Search rather than requiring a separate assistant. ChatGPT is weighted equally as the dominant standalone assistant.

Gemini follows on Google and Android default placement. AI Mode, Perplexity and Copilot are weighted lower on smaller Australian audiences. Every per-engine score is published unweighted so any reader can re-weight and reproduce a different composite.

claude answers from pre-training only, with no retrieval step, so it never enters the composite. It is reported separately as pre-training recall. The gap between a brand's live-web visibility and its pre-training recall separates brands the models know from brands the models can currently find.

04How the presence score is built

Each channel's score out of 100 is multiplied by its weight, and the four results are added. That is the whole calculation. A brand strong on one channel alone cannot climb far, which is the point of weighting four of them.

Where a channel could not be measured it shows n/m, never zero, and its weight is shared proportionally across the channels that did run, so the weights behind any published score always total 100. Contributions are rounded to one decimal place and the rounding residue is carried into the largest contribution, so the parts always sum to the score printed on the row.

Inside the AI channel

citation
40%
inclusion
60%
portability
0%
prominence
0%

Two questions carry the score: does the assistant name you, and does it link to you. Inclusion leads because being named is the precondition for everything else. Citation is weighted heavily because a linked mention is worth far more than a name-drop and because it is the component that separates brands the rest cannot: measured on the mattresses edition, Emma Sleep is named on 48% of prompts and linked on 7%, while Bedshed is named on 14% and linked on 80%.

Prominence and Portability are still measured and published unweighted, but they no longer score. 88, so the two were largely one signal taking 65% of the leg between them, and it can only be measured where the answer is a ranked list, which was 62% of mentions. The remaining 38% earned nothing for a reason that is a property of the answer's format rather than the brand.

It is also unstable on small samples, scoring one brand 78 off seven positioned mentions. Portability, the share of prompts where every live engine named the brand, was zero for 156 of 160 discovered brands and so deducted an almost identical amount from almost everyone rather than ranking anyone; it is retained as a published statistic because 'only four brands in the category are named by all six engines' is a finding worth reporting, just not a scoring component. 2 against the first full-fidelity run rather than chosen in advance.

A worked example, using a real row from a real board, sits on the index home page .

05How a board is built

The brand set is not chosen. Every entity named in any captured answer, cited as a link in any answer, or holding a top-10 organic position for any category keyword is scored — with no inclusion floor, so a brand named once counts. Entities are canonicalised and labelled by type (retailer, marketplace, manufacturer, publisher, comparison, government) but never excluded on type: an assistant answering a shopping question by pointing at a forum is a finding, not noise. The top 100 by Presence Score are recorded per subcategory and the top 20 shown by default; the full discovered set is warehoused.

Sources: entity names in captured AI answers (ranked list items and prose); domains cited as links inside AI answers; top-10 organic domains for the subcategory's category keywords.

Top 100 recorded per board, top 20 shown by default. Replaces the frozen Power Retail Top 500 selection used in methodology 1.0. That version measured eight pre-chosen brands per subcategory and could not see a challenger, a marketplace or a publisher taking the answer.

Nobody is chosen and nobody is excluded on type. Retailers, marketplaces, manufacturers, publishers, comparison sites and government pages are all labelled and all scored. An assistant answering a shopping question by pointing at a forum is the finding, not noise to be cleaned away.

Platforms we search inside are the exception, and they leave for a different reason: once we run a query inside YouTube or Reddit, every result there is hosted by it, so including them would let an instrument top its own reading.

The question set is fixed per board: 25 questions put to 7 engines, plus a separate set of category keywords for the search channel. The two are held apart on purpose. An assistant answers a conversational question and Google ranks a keyword; the conversational forms have little or no search volume, so reusing them would score every brand against results nobody visits.

06How movement is calculated

Movement is the change in a brand's rank and score against the previous edition of the same board. It is a column, not a separate URL: monthly editions overwrite the same canonical path so link equity consolidates.

Editions computed a different way are never compared. When the methodology changes, every brand's number moves for a reason that has nothing to do with the brand, and publishing that as movement would be indefensible. So an edition is only compared with the most recent prior edition that shares its methodology version, and where none exists the movement column shows a dash rather than a number.

That is why movement is blank across the whole index today. The method has changed between every edition published so far, most recently to v2.8, so no two consecutive runs have yet been directly comparable. Movement appears once a version holds for two editions.

Read direction off a brand's full history rather than one step. AI answers vary between runs and small gaps between adjacent brands are noise.

07Terms

Presence score
The published number, out of 100. Four channel scores, weighted and added.
Board
One subcategory with a fixed keyword set. The unit that is actually scored, and the only level at which brands compete against each other.
Category
A group of boards. Navigation, not a market. Nothing is scored at category level, and a category page's figures are always board figures with their board named.
Edition
One monthly publication of the index, identified by its run date.
Channel
One of the four things measured. Called a leg in the exported data.
Inclusion
The share of a board's questions whose answers name the brand.
Citation
The share of answers that link the brand's own site as a source.
Discovered / recorded / published
How many brands the run found, how many it kept, and how many the board shows by default. The gap between them is itself a finding about the category.
n/m
Not measured this edition. Never a zero, and never treated as one.

08What the index cannot tell you

Stated plainly and before anyone has to go looking, because an index that hides its gaps is not a methodology, it is marketing.

Citation coverage
Measured on google ai overviews, chatgpt, gemini, google ai mode, perplexity, copilot. Citation is measured on every weighted engine as of methodology 2.1. Engines report sources two different ways and the difference is invisible in the answer text: ChatGPT, Perplexity, AI Mode and Copilot quote source links inline, while Gemini returns them as grounding chunks and Google AI Overviews returns a references array. Both of the latter were parsed and discarded, so scanning the answer text found no links and every brand recorded a structural zero across 50% of the engine weight. They were excluded from the component in methodology 1.0 to stop that misrepresenting the result; the adapters now capture the structured sources, so the exclusion is gone. Gemini reports the publisher domain rather than a usable URL, because its grounding uri is a Vertex redirect carrying no publisher identity, so the domain is what gets recorded. Citation scores are not comparable with editions published under methodology 1.0.
Ambiguous brand names
Some brand names are ordinary words. Matching is anchored to word boundaries and is case-insensitive, so "freedom of choice" is not counted as a mention of Freedom Furniture, but a genuinely ambiguous sentence could still be miscounted either way.
Run-to-run variance
AI answers vary between runs for the same question, and the variance is not small: two identical runs put one brand at 83.8 and then 70.0. The headline is always the single-month reading, and every brand carries its full history alongside it. Read a brand's direction off that history rather than one month-on-month step, and treat small gaps between adjacent brands as noise.
Search location skew
The Australian SERP sample our search data comes from carries some regional bias. It is reliable for which national retailers appear, less so as a precise national rank.
Being named is not being liked
Social discoverability counts whether a brand is named, not whether it is recommended, so a brand can score well here on a thread of complaints. That is deliberate: the index measures discovery, and being argued about is still being discovered. We measured it before deciding. Across 259 mentions on two boards, 118 were positive and 17 negative, and sentiment ran slightly against how much a brand is discussed rather than with it, so scoring by sentiment would have rewarded the brands nobody talks about.
Video is read to a depth
The video platforms read each result's caption plus, for the top ten results per question, what the video actually says. About a third of videos publish no transcript, so some are read more deeply than others. Reading only captions was worse: it missed brands that are named out loud and never typed.
The question set is ours
A different reasonable set would produce different numbers. That is why ours is published in full.

09Corrections and challenges

Can a brand be removed from a board?
No. The field is what the run found, and a table anyone could edit themselves out of would measure willingness to complain rather than discovery.
Can a brand pay to be included, or to move up?
No. There is no paid placement, no submission fee and no way to influence a score. The method is the same for every brand on a board, including our own clients and our competitors.
Can a brand ask for its category to be scored?
Yes, and it costs nothing. Categories are added whole, so every brand in one gets measured at the same time on the same keyword set whether they asked or not. Ask for a category .
What if a number looks wrong?
Tell us and we will re-run the scoring over the stored answers for that board. Every raw answer is kept alongside the scores precisely so a disputed ranking can be re-checked against what the engines actually said, rather than re-queried months later to get different ones. Scoring is deterministic, and the order brands appear in our configuration has no effect on their scores: that is enforced by an automated test, because it is the thing that would quietly invalidate the whole exercise.
How often does it recalculate?
Monthly, versioned by run date, with the methodology version tracked separately so a scoring change is visible as a scoring change rather than disguised as movement. Current edition: au-retail-2026, run 2026-08-16.

Back to the index