Guide for marketers
To track a brand in ChatGPT, write a fixed set of the questions your buyers actually ask, run each one several times on each engine on a fixed schedule, and record for every answer whether the brand was named, in what position, which sources were cited and in what tone. The result is a named-rate you can compare month to month. Everything else in this guide is detail on doing that without fooling yourself.
This is a method a competent marketer can run by hand with a spreadsheet and an afternoon each month. It is also the method any credible tool automates, so understanding it is the only way to judge whether a tool measures anything real.
Four surfaces matter: ChatGPT, Perplexity, Google Gemini and Google AI Overviews. They behave differently enough that a reading from one does not transfer to the others.
AI Overviews deserves separate attention. It reaches roughly 2 billion users a month and appears on roughly 25-30% of informational searches, which makes it the widest-reaching of the four. It is also the only genuinely location-aware surface. In one of our own audit runs, the query "best hair transplant clinic" issued from Lisbon returned an AI Overview naming local clinics, while the identical query from London returned none. Measuring AI Overviews from the wrong location produces a reading that describes somebody else's market.
The prompt set is the whole experiment. Questions invented at a desk produce a flattering, useless picture; questions taken from sales calls, support tickets, site search logs and the "people also ask" boxes in your category produce a picture you can act on. Cover five intents.
| Intent | What it tells you | Example shape |
|---|---|---|
| Branded | What the engines already believe about you, including anything wrong | "What is [brand], the [category] company?" |
| Category | Whether you are named at all when nobody mentions you | "Best [category] providers for [segment]" |
| Alternatives-to | Which competitive set the engines place you in | "Alternatives to [competitor] for [use case]" |
| How-to-choose | Which criteria the engines teach buyers to apply | "How do I choose a [category] provider?" |
| Comparison | How you are characterised next to a named rival | "[Brand] vs [competitor] for [segment]" |
Every prompt must carry its category, without exception. Brand names collide with the world, and an engine that cannot tell which sense you mean will answer the wrong question confidently. In one of our own measured incidents, a sportswear question about "Puma" returned advice about the Ruby web server of the same name. "Is Puma good?" is not a measurable prompt; "Is Puma a good brand for running shoes?" is.
Write the prompts in the language and phrasing your buyers use. A stiff translated question often returns no AI Overview where natural native phrasing returns a full one, so a translated prompt set can make a healthy market look empty.
These figures come from audits we run ourselves. If you would rather have your own numbers than ours, apply for a free audit on one brand — real prompts, real answers, across all four engines.
Three rules make the difference between a measurement and an anecdote.
Cost is not the obstacle. Google AI Overview queries cost about $0.002 each through SERP data providers, so systematic measurement is genuinely cheap. What the work actually costs is execution and judgment.
Keep one row per answer rather than one per prompt, so that variance between runs stays visible in the data.
| Field | Why it matters |
|---|---|
| Prompt, engine, run number, date, country | Makes the row reproducible and the trend comparable |
| Was the brand named (yes or no) | The headline metric; everything else is secondary |
| Position in the answer | First named brand and fifth named brand are not the same outcome |
| Every other brand named | Gives you the competitive set the engine believes in |
| Every cited URL | Tells you which pages the engine trusts in your category |
| Tone of the mention | Recommended, listed neutrally, or named with a caveat |
| Any factual error about the brand | These are usually the fastest thing to fix |
Copy the full answer text into the sheet as well. The raw answer is the evidence behind every number you will later present, and a number nobody can trace back to an answer is not worth presenting.
Four readings come out of the sheet.
Named-rate is the share of all answers in which the brand appeared. Read it per engine and per intent, never as a single blended figure. A brand that is named reliably on branded prompts and rarely on category prompts has a specific, solvable problem, and a blended score would hide it.
Average position is where in the answer the brand appears when it is named. Movement here usually precedes movement in named-rate.
Share of voice is your named-rate against the named-rates of the other brands the engines return. This is the number that tells you whether you are in the consideration set the engines are building for buyers.
Own-versus-third-party citations is the reading most people skip and the one that decides strategy. In our own audit of a global consumer brand covering 232 answers across three engines with every citation logged, 92% of cited sources were third-party pages: 100% on Gemini, 98% on Perplexity and 80% on ChatGPT. Of the 196 answers that named the brand, 157 cited no page on the brand's own site. Publishing more pages on your own domain is therefore a minority lever.
Set the scoreboard before you start. Across those same 232 answers, not one contained a contact address. AI answers hand the user a brand name rather than a link, so the value appears as branded search volume and direct visits rather than as AI referral traffic in analytics. A team measuring only referral traffic from ChatGPT will conclude that nothing happened.
Manual tracking works well for one brand, one country and one language. It stops being viable once you add markets, product lines or a competitor set you want scored the same way, because the number of rows multiplies with each one. It also stops being viable when you need AI Overviews from several locations, since that requires a data source rather than a browser.
At that point the options are to narrow the scope deliberately, to automate collection through a SERP data provider, or to hand the measurement to a provider who already runs the pipeline. Whichever you choose, the requirements do not change: fixed prompt set, several runs per prompt, every citation logged, and raw answers available as evidence.
The sheet gives you a priority order rather than a to-do list.
No amount of this work guarantees placement. Correct records, citable content and accurate third-party coverage raise the probability that an engine names you; they do not control the output. Anyone offering a guarantee is describing a mechanism that does not exist.
Monthly, with several runs per prompt on each occasion. Checking more often mostly measures the engines' own variance, since repeated identical prompts return different brands 40-60% of the time. Checking less often makes it hard to tell whether a change followed something you did.
The engines sample from many possible responses and draw on sources that change between runs. That is why a single check is unreliable and why any figure worth reporting is a rate across multiple runs rather than a single observation.
Mostly by being accurately represented on pages other people own. In our own audit of 232 answers, 92% of cited sources were third-party pages, and 157 of the 196 answers that named the brand did not cite the brand's own site. Own-site work still matters for correcting facts and for supplying clear, quotable explanations, but the majority of the leverage sits off-site.
One audit of a global consumer brand: 232 answers across three engines, every cited link classified. Measured by BrandsNode, August 2026.
Third-party pagesThe brand's own site
Only from that market. Location decides whether an AI Overview appears at all, and phrasing matters as much as place: a stiff translated question often returns nothing where natural native phrasing returns a full answer. Use a data source that lets you set the country, or have someone in the market run the queries.
Yes, but measure it on the right scoreboard. Across 232 logged answers in our own audit, none contained a contact address, so the effect appears as branded search, direct visits and shortlist inclusion rather than as referral sessions. Track those alongside the named-rate.
We will run a full audit on one brand using the method described here: an agreed prompt set asked across ChatGPT, Perplexity, Gemini and Google AI Overviews, several runs each, with every answer and cited source logged. You receive the prompt set, the named-rate and position, the citation log and a prioritised fix list — in your own branding if you are an agency.
Apply for a free AI visibility auditOne brand, no cost, no card, no obligation. We reply with the audit or with an honest reason it would not tell you anything useful.