The GEO Audit Checklist for Agencies in 2026: From First Audit to Retainer

A repeatable GEO audit you can run across a whole client book - how to scope prompts with the client, run six engines, sort answers into quadrants, and turn the one-off audit into a monitored retainer with a report the client can check.

G
Written byGan Liu
Read Time8 min
Posted onSeptember 1, 2026
The GEO Audit Checklist for Agencies in 2026: From First Audit to Retainer

TL;DR: A GEO audit for an agency is not a site scan or a schema score. It is a repeatable operation: pick the buyer questions the client actually gets asked, run them through the answer engines, mark two booleans per answer (named, cited), sort each into a quadrant, and hand back a report where every cell opens to the stored answer. Do it once and it is a pitch. Put it on a schedule and it is a retainer.

The source gap vs conversion gap framework explains the scoring. This piece is the operations layer on top of it - how to run that framework across a client roster without it becoming a pile of ad-hoc screenshots.

Before you audit: scope with the client

The audit is only as good as the prompt set. Get it from the client, not from a keyword tool.

  • Pull 20-50 real buyer questions. Ask the client what prospects actually type or say - question-shaped queries ("best X for Y", "is Z worth it", "alternatives to..."), not brand-name lookups. Brand lookups flatter everyone and prove nothing.
  • Name the competitors. You want to know when an engine recommends someone else. List the three or four the client loses deals to.
  • Agree what "win" means per prompt. For some queries, being named is the goal. For others, being the cited source is what drives the click. Decide before you run, so the report is not arguing with itself later.

The checklist

Run this the same way every time, on every client:

  1. Lock the prompt set (20-50 question-shaped queries, agreed with the client).
  2. Run each prompt across the engines - ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode, Copilot - and capture the exact answer text and the cited sources, not a summary.
  3. Mark two booleans per answer: brand named (Y/N), client's own domain cited (Y/N). Keep them separate - they fail for different reasons.
  4. Drop each answer into its quadrant: stronghold, conversion gap, source gap, or whitespace. The two gap quadrants are defined term by term in the glossary; this checklist only tells you where to put the answer.
  5. Assign the fixed move per quadrant: defend and monitor, brandable content, reclaim the cited source, or triage-and-enter.
  6. Store every raw answer. Each number in the report must open back to the verbatim answer and its sources, with tracking stripped from the URLs. This is the part that survives a client meeting.
  7. Note the competitors that won. Where an engine cited someone else, record the exact URL. That list is your client's content roadmap.

That is the whole audit. No black-box score, no "you are at 40% visibility." A quadrant, a move, and a link to the evidence for every prompt.

The checklist as a table you can hand a junior

Same seven steps, written so someone else can run them without you in the room. Every row carries a pass condition and the ticket to raise when it fails - that is what makes it an audit rather than a walkthrough.

CheckHow to check itPasses whenTicket if it fails
Prompt set lockedRead every prompt back; count how many contain the brand name20-50 question-shaped prompts, agreed in writing, under a fifth brand-directRe-scope with the client's sales team before running anything
Competitors namedConfirm the three or four accounts the client loses deals toA named list exists and the client agrees with itBook 20 minutes with the client; do not guess from a keyword tool
Win condition per promptEach prompt is tagged: named is the win, or cited is the winEvery prompt carries exactly one tag, decided before the runTag them now - deciding after the run is grading your own homework
Runs capturedOpen three stored answers at random and inspect the recordVerbatim answer text plus cited URLs, tracking strippedRe-run the affected prompts; a summary is not a record
Two booleans markedHave a second person re-derive the booleans from the raw answersBoth people agree on every rowRewrite the marking rule until it is reproducible by anyone
Quadrant assignedEvery answer sits in stronghold, conversion gap, source gap or whitespaceNo answer is unassigned or sitting in two quadrantsFix the boolean pair first - an unassignable row means the data is wrong
Competitor citations loggedFor every answer citing someone else, the exact URL is recordedA deduplicated URL list exists, sorted by frequencyRe-open the stored answers and extract them; this list is the content roadmap

Prompt set: why it decides everything, and how it fails

Every downstream number inherits the prompt set's bias. A set built from a keyword export measures search demand. A set built from the client's sales calls measures the questions that actually precede a purchase. Only the second one produces findings a client acts on.

The most common failure looks like this: thirty prompts, twenty-two of which contain the client's brand name. The report comes back showing strong visibility, the client is pleased, and nothing about the account changes - because the audit only measured what happens after someone already knows the brand. Cap brand-direct prompts at a fifth of the set and push the rest into category discovery, comparison, buying criteria, and problem-first questions.

Two booleans: why one number always breaks

Named and cited are independent facts. They fail for different reasons and they are fixed by different work, so merging them into a single percentage destroys the only actionable information in the audit. A brand can be named constantly while its own domain is never cited, or its content can be cited repeatedly while the brand is never named - the conversion gap, and the most common finding in a first audit.

The most common failure looks like this: an analyst adds a sentiment or quality rating alongside the booleans to make the report feel richer. Next month a different analyst rates the same answers differently, the trend line moves for no reason anyone can explain, and the client notices. Keep the two booleans objective and put judgement in the written recommendation, where it belongs and where it is labelled as judgement.

Evidence storage: the step that survives the client meeting

The stored raw answer is what separates an audit from an assertion. When a client asks why they are paying you, the defensible artefact is not a chart. It is the answer ChatGPT gave in March next to the answer it gave in June, with the change visible in both.

The most common failure looks like this: the team stores screenshots. Screenshots cannot be searched, diffed, or quoted, and six months later nobody can tell which sentence changed. Store the answer text and the cited URLs with tracking parameters stripped, and keep the screenshot as a supplement rather than as the record.

Competitor citations: the roadmap hiding in your audit

Every answer that cited someone other than your client points at a page a machine already decided was the better source. Collected across a whole audit and sorted by frequency, that list is the most concrete content brief an agency can produce, and it costs nothing extra because you already have the data sitting in your sheet.

The most common failure looks like this: the competitor URLs are read once during the audit and never extracted. The audit ships, the client asks what to write next, and the team goes back to a keyword tool - having already had the answer and thrown it away.

The 60-minute version of the GEO audit checklist

Sometimes you need something credible before a pitch tomorrow morning, not a full audit. Six steps, one hour, one engine.

  1. Ten prompts, not thirty (10 min). Two from each category. Skip the client interview, pull the questions from their own sales page, and flag the set as provisional.
  2. ChatGPT only (20 min). One engine, one clean session per prompt. Coverage across six engines is what the paid audit adds; one engine is enough to prove the method works.
  3. Two booleans, nothing else. Named and cited. No ratings, no sentiment, no score.
  4. Paste every answer into one document (5 min). Unformatted is fine. The point is that the evidence exists and can be reopened.
  5. Sort into four buckets (15 min). Stronghold, conversion gap, source gap, whitespace. Count each one.
  6. Write three sentences (10 min). The biggest bucket, the single most surprising answer, and the one fix you would ship first.

That is a defensible mini-audit as long as you say out loud what it skipped: one engine, ten prompts, one sample each, no trend line. Our AI visibility checker will generate the prompt set and a matching scorecard so the first ten minutes cost you nothing.

What a GEO audit checklist can and cannot prove

An audit is measurement, not a guarantee, and being specific about that boundary is what stops a retainer ending in an argument twelve months later.

What it proves. What a given engine answered, for a given prompt, at a given time, from a given region - and, re-run on a schedule, whether that answer changed. That is a real, checkable, dated fact, and it is more than most marketing reporting can honestly claim.

What it cannot prove. It cannot prove causation: if a brand appears in more answers after you shipped content, the audit shows a correlation between a fix and a change, not a mechanism, because the engine may have updated independently that week. It cannot prove what a different user sees, since personalization, memory, and region all move the answer. It cannot promise a ranking, because nobody controls what a generative model says. And it cannot attribute revenue on its own - being cited and being clicked are different events, and the second one lives in the client's analytics, not in ours.

Say all of that in the first report rather than the last. Agencies that scope the claim honestly keep retainers longer than agencies that promise a number will go up, which is the same reason we publish the white-label GEO reports guide with our own shipping status stated plainly rather than implied. And if a term in your report is doing work you cannot define on demand, check it against the GEO glossary before it reaches a client.

Turn the audit into a retainer

The one-off audit is the pitch. The retainer is the same prompt set, re-run on a schedule, diffed against the client's own history:

  • Set a cadence. Weekly for most, down to daily or every few hours for high-stakes queries. The baseline is always the client's own past runs, never an invented target.
  • Flag only real movement. A prompt that named the client last month and names a competitor this month is a flag. Sampling noise is not. Every flag opens to the answer that changed.
  • Report monthly. One evidence report per client per month: what moved, into which quadrant, and the fix you shipped or recommend. Proof of movement is the thing that renews a retainer.

What to package and charge

Price the service line on runs, not vibes. A run is one prompt on one engine. A first audit of 30 prompts across six engines is 180 runs; monthly monitoring re-runs a tighter high-value set. If you use Jincove underneath, pricing is usage-based with a credit pool shared across your whole roster, so your unit cost drops as you add clients - which is exactly the margin math a service line needs. Full detail on packaging it as a retainer is in GEO for agencies, and the pricing arithmetic behind it is in what to charge for GEO services.

Run your first one free

Want the checklist run on a real client before you build the process yourself? Request a free, human-run audit: send one client URL and an email, and we hand-run ChatGPT, Perplexity, and Gemini, then reply with the exact answers, sources, and quadrants within two business days. No card, no account. Use it as the sample audit in your next GEO pitch.

Start with the free audit

Send us one brand. We’ll run it through the engines and send back the answers, citations, and sources — so the first report your client sees is already backed by evidence.

Free audit: ChatGPT, Perplexity and Gemini, run by hand. Paid work covers all six engines.