The GEO Audit Checklist for Agencies in 2026: From First Audit to Retainer
A repeatable GEO audit you can run across a whole client book - how to scope prompts with the client, run six engines, sort answers into quadrants, and turn the one-off audit into a monitored retainer with a report the client can check.

TL;DR: A GEO audit for an agency is not a site scan or a schema score. It is a repeatable operation: pick the buyer questions the client actually gets asked, run them through the answer engines, mark two booleans per answer (named, cited), sort each into a quadrant, and hand back a report where every cell opens to the stored answer. Do it once and it is a pitch. Put it on a schedule and it is a retainer.
The source gap vs conversion gap framework explains the scoring. This piece is the operations layer on top of it - how to run that framework across a client roster without it becoming a pile of ad-hoc screenshots.
Before you audit: scope with the client
The audit is only as good as the prompt set. Get it from the client, not from a keyword tool.
- Pull 20-50 real buyer questions. Ask the client what prospects actually type or say - question-shaped queries ("best X for Y", "is Z worth it", "alternatives to..."), not brand-name lookups. Brand lookups flatter everyone and prove nothing.
- Name the competitors. You want to know when an engine recommends someone else. List the three or four the client loses deals to.
- Agree what "win" means per prompt. For some queries, being named is the goal. For others, being the cited source is what drives the click. Decide before you run, so the report is not arguing with itself later.
The checklist
Run this the same way every time, on every client:
- Lock the prompt set (20-50 question-shaped queries, agreed with the client).
- Run each prompt across the engines - ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode, Copilot - and capture the exact answer text and the cited sources, not a summary.
- Mark two booleans per answer: brand named (Y/N), client's own domain cited (Y/N). Keep them separate - they fail for different reasons.
- Drop each answer into its quadrant: stronghold, conversion gap, source gap, or whitespace. The two gap quadrants are defined term by term in the glossary; this checklist only tells you where to put the answer.
- Assign the fixed move per quadrant: defend and monitor, brandable content, reclaim the cited source, or triage-and-enter.
- Store every raw answer. Each number in the report must open back to the verbatim answer and its sources, with tracking stripped from the URLs. This is the part that survives a client meeting.
- Note the competitors that won. Where an engine cited someone else, record the exact URL. That list is your client's content roadmap.
That is the whole audit. No black-box score, no "you are at 40% visibility." A quadrant, a move, and a link to the evidence for every prompt.
The checklist as a table you can hand a junior
Same seven steps, written so someone else can run them without you in the room. Every row carries a pass condition and the ticket to raise when it fails - that is what makes it an audit rather than a walkthrough.
| Check | How to check it | Passes when | Ticket if it fails |
|---|---|---|---|
| Prompt set locked | Read every prompt back; count how many contain the brand name | 20-50 question-shaped prompts, agreed in writing, under a fifth brand-direct | Re-scope with the client's sales team before running anything |
| Competitors named | Confirm the three or four accounts the client loses deals to | A named list exists and the client agrees with it | Book 20 minutes with the client; do not guess from a keyword tool |
| Win condition per prompt | Each prompt is tagged: named is the win, or cited is the win | Every prompt carries exactly one tag, decided before the run | Tag them now - deciding after the run is grading your own homework |
| Runs captured | Open three stored answers at random and inspect the record | Verbatim answer text plus cited URLs, tracking stripped | Re-run the affected prompts; a summary is not a record |
| Two booleans marked | Have a second person re-derive the booleans from the raw answers | Both people agree on every row | Rewrite the marking rule until it is reproducible by anyone |
| Quadrant assigned | Every answer sits in stronghold, conversion gap, source gap or whitespace | No answer is unassigned or sitting in two quadrants | Fix the boolean pair first - an unassignable row means the data is wrong |
| Competitor citations logged | For every answer citing someone else, the exact URL is recorded | A deduplicated URL list exists, sorted by frequency | Re-open the stored answers and extract them; this list is the content roadmap |
Prompt set: why it decides everything, and how it fails
Every downstream number inherits the prompt set's bias. A set built from a keyword export measures search demand. A set built from the client's sales calls measures the questions that actually precede a purchase. Only the second one produces findings a client acts on.
The most common failure looks like this: thirty prompts, twenty-two of which contain the client's brand name. The report comes back showing strong visibility, the client is pleased, and nothing about the account changes - because the audit only measured what happens after someone already knows the brand. Cap brand-direct prompts at a fifth of the set and push the rest into category discovery, comparison, buying criteria, and problem-first questions.
Two booleans: why one number always breaks
Named and cited are independent facts. They fail for different reasons and they are fixed by different work, so merging them into a single percentage destroys the only actionable information in the audit. A brand can be named constantly while its own domain is never cited, or its content can be cited repeatedly while the brand is never named - the conversion gap, and the most common finding in a first audit.
The most common failure looks like this: an analyst adds a sentiment or quality rating alongside the booleans to make the report feel richer. Next month a different analyst rates the same answers differently, the trend line moves for no reason anyone can explain, and the client notices. Keep the two booleans objective and put judgement in the written recommendation, where it belongs and where it is labelled as judgement.
Evidence storage: the step that survives the client meeting
The stored raw answer is what separates an audit from an assertion. When a client asks why they are paying you, the defensible artefact is not a chart. It is the answer ChatGPT gave in March next to the answer it gave in June, with the change visible in both.
The most common failure looks like this: the team stores screenshots. Screenshots cannot be searched, diffed, or quoted, and six months later nobody can tell which sentence changed. Store the answer text and the cited URLs with tracking parameters stripped, and keep the screenshot as a supplement rather than as the record.
Competitor citations: the roadmap hiding in your audit
Every answer that cited someone other than your client points at a page a machine already decided was the better source. Collected across a whole audit and sorted by frequency, that list is the most concrete content brief an agency can produce, and it costs nothing extra because you already have the data sitting in your sheet.
The most common failure looks like this: the competitor URLs are read once during the audit and never extracted. The audit ships, the client asks what to write next, and the team goes back to a keyword tool - having already had the answer and thrown it away.
The 60-minute version of the GEO audit checklist
Sometimes you need something credible before a pitch tomorrow morning, not a full audit. Six steps, one hour, one engine.
- Ten prompts, not thirty (10 min). Two from each category. Skip the client interview, pull the questions from their own sales page, and flag the set as provisional.
- ChatGPT only (20 min). One engine, one clean session per prompt. Coverage across six engines is what the paid audit adds; one engine is enough to prove the method works.
- Two booleans, nothing else. Named and cited. No ratings, no sentiment, no score.
- Paste every answer into one document (5 min). Unformatted is fine. The point is that the evidence exists and can be reopened.
- Sort into four buckets (15 min). Stronghold, conversion gap, source gap, whitespace. Count each one.
- Write three sentences (10 min). The biggest bucket, the single most surprising answer, and the one fix you would ship first.
That is a defensible mini-audit as long as you say out loud what it skipped: one engine, ten prompts, one sample each, no trend line. Our AI visibility checker will generate the prompt set and a matching scorecard so the first ten minutes cost you nothing.
What a GEO audit checklist can and cannot prove
An audit is measurement, not a guarantee, and being specific about that boundary is what stops a retainer ending in an argument twelve months later.
What it proves. What a given engine answered, for a given prompt, at a given time, from a given region - and, re-run on a schedule, whether that answer changed. That is a real, checkable, dated fact, and it is more than most marketing reporting can honestly claim.
What it cannot prove. It cannot prove causation: if a brand appears in more answers after you shipped content, the audit shows a correlation between a fix and a change, not a mechanism, because the engine may have updated independently that week. It cannot prove what a different user sees, since personalization, memory, and region all move the answer. It cannot promise a ranking, because nobody controls what a generative model says. And it cannot attribute revenue on its own - being cited and being clicked are different events, and the second one lives in the client's analytics, not in ours.
Say all of that in the first report rather than the last. Agencies that scope the claim honestly keep retainers longer than agencies that promise a number will go up, which is the same reason we publish the white-label GEO reports guide with our own shipping status stated plainly rather than implied. And if a term in your report is doing work you cannot define on demand, check it against the GEO glossary before it reaches a client.
Turn the audit into a retainer
The one-off audit is the pitch. The retainer is the same prompt set, re-run on a schedule, diffed against the client's own history:
- Set a cadence. Weekly for most, down to daily or every few hours for high-stakes queries. The baseline is always the client's own past runs, never an invented target.
- Flag only real movement. A prompt that named the client last month and names a competitor this month is a flag. Sampling noise is not. Every flag opens to the answer that changed.
- Report monthly. One evidence report per client per month: what moved, into which quadrant, and the fix you shipped or recommend. Proof of movement is the thing that renews a retainer.
What to package and charge
Price the service line on runs, not vibes. A run is one prompt on one engine. A first audit of 30 prompts across six engines is 180 runs; monthly monitoring re-runs a tighter high-value set. If you use Jincove underneath, pricing is usage-based with a credit pool shared across your whole roster, so your unit cost drops as you add clients - which is exactly the margin math a service line needs. Full detail on packaging it as a retainer is in GEO for agencies, and the pricing arithmetic behind it is in what to charge for GEO services.
Run your first one free
Want the checklist run on a real client before you build the process yourself? Request a free, human-run audit: send one client URL and an email, and we hand-run ChatGPT, Perplexity, and Gemini, then reply with the exact answers, sources, and quadrants within two business days. No card, no account. Use it as the sample audit in your next GEO pitch.
Related blogs
Related Post
Expand your knowledge with these hand-picked posts.

GEO for Agencies: How to Offer AI Visibility Monitoring as a Service
How SEO, PR, and brand agencies can package AI visibility monitoring as a retainer: a re-runnable evidence report where every mention and citation opens to the exact AI answer it came from.
Gan Liu

White-Label GEO Reports: The Agency Structure
Build a white-label GEO report that survives the client asking how you know. Here is the section order, the two numbers that matter, and what to leave out.
Gan Liu

What to Charge for GEO Services: Pricing and Packaging for Agencies
How to price and package GEO as an agency service line - the two real cost lines, three packaging shapes with the arithmetic behind them, what the client actually receives, and the four situations where you should decline the engagement.
Gan Liu
Start with the free audit
Send us one brand. We’ll run it through the engines and send back the answers, citations, and sources — so the first report your client sees is already backed by evidence.
Free audit: ChatGPT, Perplexity and Gemini, run by hand. Paid work covers all six engines.