How to Track Brand Mentions in ChatGPT in 2026: The Agency Method
ChatGPT has no Search Console - no impressions, no analytics, no export. Here is the method agencies use to track brand mentions across ChatGPT and the other answer engines without trusting a black-box score, and why "citations" and "mentions" are two different numbers.
TL;DR: ChatGPT gives you no analytics, so tracking brand mentions means running your own prompt set, reading the answers, and recording two separate booleans per answer - was the brand named, and was its domain cited. Store the raw answer behind every number. Do not trust a single "visibility score": it merges two signals that fail for different reasons, and it cannot be checked. The method below is what makes the number defensible to a client.
Millions of buyers now ask ChatGPT for recommendations before they ever touch Google. If your client is not in that answer - or is named but points at a competitor - they lose the deal upstream, invisibly. So tracking it matters. The problem is that ChatGPT ships with no Search Console: no impressions, no query report, no export. You have to build the measurement yourself.
Why the "visibility score" is the wrong tool
Most trackers hand you one number: "62% visibility." It feels like progress and tells you nothing. Sixty-two percent of what - named, cited, or both, blurred together? You cannot act on it, and you certainly cannot defend it when a client asks what changed. We have watched a client's citation count climb for a quarter while the brand's actual name never made it into a single answer the reader saw - the number went up and the deals still went to whoever the engine named instead. That is the trap of a merged score: it rewards the signal that feels good over the one that wins the deal. The fix is to stop merging signals.
The two signals you actually track
Every answer gives you two independent facts. Keep them apart:
- Mention - is the brand named in the answer the reader sees? Yes or no.
- Citation - is the brand's own domain in the sources under that answer? Yes or no.
The engine can name a brand without linking it (it "knows" the name but pulls facts from a review site) or cite the brand's page without ever saying the name (the content answered the question and the brand got dropped - a conversion gap). One score cannot hold both. Two booleans can.
The method
- Build a prompt set. 20-50 real buyer questions in the client's category - the things prospects actually ask, not brand-name lookups. Question-shaped queries surface the truth; brand lookups flatter everyone.
- Run each prompt in ChatGPT deliberately. Fresh session, note whether web search was on, which model, which country - these change the answer, so record them as part of the run.
- Capture the exact answer and its sources. Copy the verbatim text and the cited URLs. Not a summary - the artifact.
- Mark the two booleans: named (Y/N), own domain cited (Y/N).
- Store the raw answer behind the number. This is the step trackers skip and the step that makes your report survive scrutiny. A month later, "you moved from source gap to stronghold on this prompt" only means something if you can reopen both answers.
- Diff against your own history, on a schedule. The baseline is the client's past runs, not an invented target. Flag a change only when it moves beyond run-to-run noise.
Step 1: build the prompt set
The prompt set is the whole experiment. Get it wrong and every number after it is noise. Build five categories, five prompts each, for twenty-five prompts total - broad enough to be meaningful, small enough that you will actually re-run it every week.
| Prompt category | Question it answers | What to record |
|---|---|---|
| Category discovery | Does the engine know the brand exists in this category at all? | Named, cited, and which brands were named instead |
| Comparison and alternatives | Who does the engine put the brand up against? | Named, cited, and the competitor set the engine chose |
| Buying criteria | Does the brand appear when the buyer asks how to choose? | Named, cited, and which source the criteria came from |
| Problem-first | Does the brand surface for the pain, not the product name? | Named, cited, and whether the answer stays generic |
| Brand-direct | What does the engine say when asked about the brand by name? | Named, cited, and whether any claim is factually wrong |
How you know you did it right: read the twenty-five prompts back and ask whether a real prospect would type them. If more than a handful contain the brand name, the set is flattering, not diagnostic.
Common mistake: pasting a keyword-tool export into the set. Search keywords are short and fragmentary; the questions people put to ChatGPT are long, conversational, and usually carry a constraint the keyword never held. Ask the client's sales team for the actual wording of the questions they field on calls. If you want a starting set and a matching scorecard without building one from scratch, the free AI visibility checker for ChatGPT, Perplexity and Gemini generates both.
Step 2: run each prompt in a clean session
Open a new chat for every prompt. Do not run twenty-five questions down one thread, because the model carries your earlier turns forward and answer five is contaminated by answer four.
Turn off personalization and memory before you start, and record which model you used, whether web search was active, and which country you were in. Every one of those changes the answer, so all four belong in the run record next to the result. If your account has months of history discussing this client, use a logged-out session or a separate browser profile - memory is the single most common reason two people get different answers to the same prompt and each concludes the other made an error.
How you know you did it right: ask the same prompt twice in two fresh sessions. The answers should differ in wording but land on a similar set of brands. If they name completely different brands, that prompt is high-variance and needs more samples before you report anything from it.
Common mistake: treating your own logged-in ChatGPT as the neutral observer. It is the least neutral instance available to you, because it has been absorbing your interest in this client for weeks.
Step 3: record two booleans and store the raw answer
For each answer, record exactly two facts: was the brand named, and was the brand's own domain among the cited sources. Two columns, yes or no. Then paste the full answer text and the list of cited URLs into a third and fourth column, with tracking parameters stripped.
Resist every urge to add a quality rating at this stage. "Positively mentioned, 7 out of 10" is a judgement you will not reproduce next month, and the moment a client questions it you are defending a feeling instead of a record. The booleans reproduce. The distinction between the two columns is set out in mention vs citation, and the entire method rests on keeping them apart.
How you know you did it right: hand your sheet to a colleague with your booleans hidden and ask them to re-derive the two columns from the raw answers alone. If they agree with you on every row, your recording is objective rather than personal.
Common mistake: storing a screenshot instead of the text. Screenshots cannot be searched, diffed, or quoted in a report, and in three months nobody will be able to tell which sentence changed.
Step 4: re-run on a fixed cadence and diff against your own history
One run is a snapshot and proves almost nothing. Pick a cadence - weekly for most brands, daily only for the handful of prompts where a single answer decides a large deal - and hold it. The baseline is always the client's own previous runs, never an industry average or an invented target.
When you diff, compare booleans first and prose second. A prompt that flipped from named to not-named is a flag. A prompt where the same brands appear in a different order is usually sampling noise, and reporting it as movement will cost you credibility the first time it reverses on its own.
How you know you did it right: your flags survive a second look a week later. If most of last week's flags quietly un-flag themselves, your threshold is too sensitive and you are reporting weather, not climate.
Common mistake: editing the prompt set between runs. The moment a prompt changes, its history resets. Keep a frozen core set for trending and a separate exploratory set for new questions.
A worked example: one prompt, two booleans
Here is the method on a single prompt, so the abstraction turns concrete. Say the client is a project-management SaaS and the buyer question is "what's the best project management tool for a small agency?"
You run it in ChatGPT with web search on, US, and capture the verbatim answer. The answer recommends four tools by name and shows a row of source chips underneath. Now you read the two booleans against the client:
| What you read in the stored answer | Boolean | Quadrant |
|---|---|---|
| Client is named in the prose ("…and Acme for smaller teams") | Mentioned = Yes | Stronghold - defend it |
| A source chip links acme.com | Cited = Yes |
Change one fact and the quadrant flips. If the answer names Acme but every source chip points at a review directory and a competitor's comparison page, that is a source gap - the reader hears the name and clicks through to someone else. If the answer never says "Acme" but one of the source chips is an acme.com blog post the model paraphrased, that is a conversion gap - your content did the work and the brand vanished. Same prompt, same engine, completely different fix - which is exactly why one merged "visibility %" is useless and the two booleans are not.
The discipline is to record all three artifacts for that run - the verbatim answer, the exact cited URLs, and the metadata (model, country, search on/off) - so that a month later you can reopen this precise answer and prove which way it moved. Repeat for every prompt in the set, and the pile of runs becomes a quadrant map you can hand a client.
Do it across all six engines, not just ChatGPT
ChatGPT is one surface. Buyers also land in Perplexity, Gemini, Google AI Overviews, Google AI Mode, and Copilot, and the same brand can be a stronghold in one and whitespace in another. Each engine surfaces answers and sources differently, so each has its own tracking guide linked above. Track the same prompt set across all of them and sort every answer into a quadrant - the source gap vs conversion gap framework turns the raw booleans into fixed moves. For where GEO fits next to AEO and classic SEO, see GEO vs AEO vs SEO. The engine that behaves least like ChatGPT is Perplexity, which prints numbered sources on nearly every answer: what changes there is covered in Perplexity visibility tracking. Google AI Mode is the other outlier, because it fans one question into several sub-searches and leaves nothing to hold a rank in - the instruments built for it are compared in best Google AI Mode tracking tools.
A worked example: one illustrative brand, one week
The following is an illustrative walkthrough, not a real client. The numbers are invented to show the shape of the output, not to report a result.
Take a fictional project-management tool called Northbeam Tasks. We build twenty-five prompts across the five categories, run each once in ChatGPT in a clean session, and fill in two columns.
Out of twenty-five answers, Northbeam is named in nine and its own domain is cited in four. Three of those four citations sit inside answers where the brand was never named at all - the engine used a Northbeam comparison page to assemble its criteria list, then recommended two competitors. That is a conversion gap, and it is invisible to anyone tracking a single blended score, because the citations read as a win.
Six of the nine mentions cluster in the brand-direct category, which is the least valuable bucket - people asking about Northbeam by name already know Northbeam. In category discovery, the bucket that decides whether a stranger ever hears the name, it is named zero times out of five.
The work order writes itself. The comparison page is already trusted enough to be cited, so it needs the brand named in its own conclusions rather than only inside a table cell. And the category-discovery whitespace needs entry content before any of the citation work matters. None of that comes out of a score. All of it comes out of two boolean columns and the stored answers sitting behind them.
When monitoring ChatGPT brand mentions does not work
This is the most defensible method available, which is not the same as reliable in every case. Be honest with clients about four limits, in the first report rather than the last.
Single samples are noisy. Generative answers are stochastic. One run of one prompt tells you what happened once. Anything you report from a single sample can reverse on the next run with nothing having changed on the client's site.
Answers drift for reasons you cannot see. Model updates, index refreshes, region, time of day, and whether web search fired all move the result. A drop between two weekly runs may reflect a silent model change rather than anything about the brand.
Personalization and memory contaminate results. A logged-in account that has discussed this brand carries that forward. Clean sessions and disabled memory reduce it; nothing eliminates it, and no browser-based method reproduces what a stranger in another country sees.
Manual runs do not scale to change detection. Twenty-five prompts by hand is an afternoon. Twenty-five prompts across six engines every week across eight clients is over a thousand runs a week, and the failure mode is not effort - it is that one skipped week erases the trend line you were selling.
If any of these limits is load-bearing for the claim you want to make, put it in the report. A method that states its error bars persuades better than one that pretends it has none.
Make it a client-provable habit
Running this by hand once is an audit. Running it on a schedule, per client, with every answer stored, is a monitored service - and the thing that renews a retainer, because you can show proof of movement instead of asserting it. If you would rather buy the pipeline than build it, we reviewed the ChatGPT brand monitoring tools on the one axis that decides it - whether the tool stores the answer behind the number. Jincove runs exactly this method across the six engines with an isolated workspace per client and evidence behind every number. See how it compares to the score-first tools: Jincove vs Profound and Jincove vs Otterly.
See it on your own brand first
Request a free, human-run audit: send one URL and an email, and we hand-run ChatGPT, Perplexity, and Gemini and reply with the exact answers, the sources, and whether each named or only cited the brand - within two business days, no card, no account. It is the fastest way to see the two-boolean method on a brand you care about.
Related blogs
Related Post
Expand your knowledge with these hand-picked posts.

Best AI Visibility Tools for Agencies in 2026: An Honest Shortlist
Most "best GEO tool" roundups rank on engine count. Agencies buy on a different axis - client roster, white-label, and evidence you can hand a client. Here is the honest shortlist, what each tool is actually best at, and the three questions that decide it.
Gan Liu

How to Track Brand Mentions in Perplexity in 2026: The Agency Method
Perplexity shows numbered citations and a Sources list under every answer. The agency method for reading brand mention vs citation off it - no black-box score.
Gan Liu

How to Track Brand Mentions in Gemini in 2026: The Agency Method
Gemini sometimes grounds on Google Search and shows sources, sometimes answers from the model alone. Here is the agency method for tracking brand mentions in Gemini with two booleans and stored evidence.
Gan Liu
Start with the free audit
Send us one brand. We’ll run it through the engines and send back the answers, citations, and sources — so the first report your client sees is already backed by evidence.
Free audit: ChatGPT, Perplexity and Gemini, run by hand. Paid work covers all six engines.