How to Choose an AI Visibility Provider: 12 Questions to Ask First
Twelve questions to put to any AI visibility vendor before you sign - what a good answer sounds like, and what a vague one is hiding. Including the four where our own answer is not good enough yet.

TL;DR: Feature-comparison tables do not separate AI visibility vendors, because every vendor claims every feature. What separates them is how they answer twelve specific questions about method, evidence, isolation and billing. This page gives you the twelve, what a good answer sounds like, and what a vague answer is actually hiding. Jincove answers all twelve at the bottom, including the four where our answer today is "we do not do that yet."
There are already two shortlists on this site. This is not a third. Three pages, three jobs:
- Best AI visibility tools for agencies ranks the category for an agency buyer - roster economics, isolation, deliverability.
- ChatGPT brand monitoring tools narrows to one engine and compares seven products on it.
- This page ranks nothing. It is the interview script you run against whichever vendors those two put on your shortlist, ours included.
It exists because every vendor in this category ships a dashboard, claims six engines, and shows a number going up. The marketing pages are close to interchangeable. The differences live one level down - in how the number was produced and who eats the cost when a run fails - and none of that is on a feature grid.
How to use these twelve questions
Send them by email before the demo. A demo is a controlled environment where the vendor drives, and several of these have an honest answer that a good salesperson can talk past in real time. In writing, a vague answer stays vague.
Score each answer specific (a number, a definition, a named surface), hedged (true but unfalsifiable), or absent. You are not looking for a perfect card - you are looking for a vendor whose absent answers are ones you can live with, and who told you they were absent before you asked twice.
Four carry disproportionate weight for agency work: question 3 (stored evidence), question 7 (client isolation), question 10 (billing unit) and question 12 (export). Those decide whether the tool can carry a service line. The rest decide whether you can defend a number in a meeting.
Method: what did you actually run
Question 1. How do you define the sampling frame?
Why ask: every visibility number is a percentage of something. The denominator has four dimensions - prompts, engines, geography, window. If the vendor does not declare all four, the number is not comparable to anything, including its own value last month.
A good answer names all four, shows where each is set in the product, and volunteers the limitation: a score on your prompt set is not comparable to a score on someone else's. See AI share of voice.
A vague answer is "we track thousands of prompts across all major engines." A large denominator you cannot inspect is worse than a small one you can, because you cannot tell whether the score moved or the sample did.
Question 2. What exactly happens in one run?
Why ask: pricing, freshness and accuracy all collapse into this definition. "One prompt, one engine, once" and "one prompt fanned across five engines and averaged over three samples" are different products at the same headline price.
A good answer is arithmetic you can multiply out: prompts x engines x samples x runs per month equals records per month, before you quote a client.
A vague answer is "unlimited prompts." The constraint has only moved somewhere you cannot see it - a rate limiter, a queue, or a refresh interval that quietly makes your data three weeks old.
Question 3. Do you store the raw answer, and can I open one right now?
Why ask: this decides whether the deliverable survives a client asking "how do you know." A chart cannot answer that. The verbatim sentence the engine produced can.
A good answer is: yes - answer text, cited sources, provider metadata (model, country, whether web search was on), timestamp - and every number in the dashboard opens back to it. Then they let you click one during the demo, on a record you picked, not one they picked.
A vague answer is "we store the evidence" with no click-through. Insist on the click - the cheapest fraud test in this category.
Coverage: which surfaces, and how they are counted
Question 4. How do you count engines, and which ones are in the plan I would actually buy?
Why ask: engine counts on marketing pages are often a superset of what the entry plan includes, and vendors differ on what counts as one engine. ChatGPT with browsing and without it can be sold as two.
A good answer is a list of named surfaces at the price you would pay, plus a plain statement of what they treat as distinct and why.
A vague answer is "all major AI engines." Ask which surfaces are gated to a higher tier - usually where the number came from.
Question 5. How is the score computed, and what happens when you change the formula?
Why ask: a composite score usually blends being named in an answer, being cited as a source, position, and sometimes sentiment. When it moves you cannot tell which input moved, so you cannot tell a client what to do. Worse, vendors tune these formulas, and a retuned index rewrites your client's history without touching the underlying data.
A good answer either publishes the formula or - better - refuses to have one and reports the underlying facts separately, the case we make in mention vs citation. Either way, get a written commitment to disclose formula changes and say whether history is restated.
A vague answer is "proprietary index." That means your client's trend line is a vendor product decision.
Question 6. Can I reproduce a number myself?
Why ask: reproducibility is what separates a measurement from a claim.
A good answer is: pick a prompt, an engine and a date, and we will show you the stored answer and the arithmetic on top of it. A good vendor also volunteers that an exact re-run today will not match, because answer engines are non-deterministic - the record is reproducible, the live engine is not.
A vague answer is "our data is proprietary." Data can be. Method cannot be, if you have to defend it.
Operations: running this across a roster
Question 7. How are clients isolated from each other?
Why ask: this decides whether the tool can carry an agency, and it is the one most often answered with a tagging feature.
A good answer is a workspace per client with its own prompt set, history and access list, plus an explicit statement on whether one client's data can appear in another's export.
A vague answer is "you can tag by client." Tags are a view, not a boundary. One bad filter and you hand Client A a spreadsheet with Client B in it.
Question 8. What is your data retention policy, as a number?
Why ask: your retained deliverable is a diff against the client's own history. If raw answers roll off at ninety days, the year-over-year story you sold is gone, and you find out in month four.
A good answer is two numbers - raw answers, derived metrics - plus what happens to both on cancellation.
A vague answer is "as long as necessary." Necessary to whom.
Question 9. What does white-label actually cover, item by item?
Why ask: the word spans everything from swapping a logo on a PDF to a custom domain with a client login. Both get called white-label on a pricing page.
A good answer is an itemized list: what carries your brand, what still says the vendor's name, static export or live portal, and whether the client can log in at all. The structure to compare against is in white-label GEO reports.
A vague answer is "fully white-label" with no list. Ask for a sample export carrying your logo before you sign.
Commercials: what you are billed for
Question 10. What is the unit of billing, and what pushes me over it?
Why ask: seats, brands, prompts, runs and credits are not interchangeable, and a plan that is cheap at one client can eat the whole margin of the service line at ten.
A good answer is one unit, a stated conversion for every action that spends it, and a worked example at ten clients rather than one.
A vague answer is "contact sales" on every tier. Unpublished pricing is not automatically a red flag, but it does mean you cannot model your margin before the call.
Question 11. Do I pay for runs that fail?
Why ask: engines rate-limit, time out, and sometimes refuse to answer. Somebody absorbs that cost, and if the contract is silent it is you.
A good answer is a definition of "failed," a statement that failed runs are not billed, and a log so you can check.
A vague answer is silence. Ask the sharp version: what happens when the engine returns a refusal rather than an error.
Question 12. Can I export the full history, and in what shape?
Why ask: this is your exit cost, and what happens when a client leaves and wants their data. If the only export is a PDF, the vendor owns your client history.
A good answer is a raw-record export - API or file - with answer text, sources, timestamps and metadata, available while the account is live and for a stated window after cancellation.
A vague answer is "you can download reports." A report is a rendering; you want the records behind it.
The twelve as a checklist
| # | Question | What a bad answer sounds like |
|---|---|---|
| 1 | How is the sampling frame defined | "Thousands of prompts, all major engines" |
| 2 | What is one run | "Unlimited prompts" |
| 3 | Is the raw answer stored and openable | "We store the evidence" with no click |
| 4 | Which engines at my price | "All major AI engines" |
| 5 | How is the score computed | "Proprietary index" |
| 6 | Can I reproduce a number | "Our data is proprietary" |
| 7 | How are clients isolated | "You can tag by client" |
| 8 | Retention, as a number | "As long as necessary" |
| 9 | What does white-label cover | "Fully white-label", unitemized |
| 10 | What is the billing unit | "Contact sales" on every tier |
| 11 | Are failed runs billed | Nothing in the contract either way |
| 12 | Can I export raw history | "You can download reports" |
Our own answers, including the ones we fail
Publishing an interview script and then dodging it would be worse than not publishing one. Jincove against all twelve, as of 2026-09.
- Sampling frame. Yours, and declared. You set the prompt list, the engines and the schedule; the window is stated on every report. The resulting number is not comparable to another agency's number on a different prompt set, and we say so.
- One run. One credit equals one delivered run - one prompt on one engine. Analysis is plus two credits per answer, an evidence screenshot one, an on-demand deep scan twelve. Full rate card on pricing.
- Raw answer. Stored verbatim, with cited URLs (UTM stripped) and provider metadata. Every derived number opens back to it. That is the product, not a feature of it.
- Engines. Six named surfaces: ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode, Microsoft Copilot. Studio covers five, Agency all six. The free sample audit covers three, hand-run.
- Score. We do not have one. Two independent booleans per answer - named, and cited - and the source gap and conversion gap derived from them. No formula to retune, which is the point.
- Reproducibility. Open the record. The limit, stated honestly: re-running the same prompt tomorrow will not reproduce the same answer, because the engines are non-deterministic. What is reproducible is the stored record and the arithmetic on it.
- Isolation. A workspace per client - up to five on Studio, eight on Agency, more on Custom - with its own prompt set and history.
- Retention. We fail this one. We do not publish a retention number. Our privacy policy covers retention and deletion in principle, but there is no committed figure for raw answers, and by the standard of question 8 that is not a good enough answer. If it matters to your contract, ask us to put a number in writing.
- White-label. We fail this one too, today. Per-client reports are built to be handed over, but full white-label - your domain, your logo - is listed as coming soon on Agency and Custom. If you need it live this quarter, buy something else, and the shortlist says which.
- Billing unit. One credit, published rate card, one pooled balance across the roster. Studio $249/mo, Agency $749/mo as of 2026-09; the rate per thousand credits falls as the pool grows.
- Failed runs. Not billed. A run consumes a credit when it is delivered.
- Export. REST API access is on every published plan from Studio up, so the records leave in record shape rather than as a rendered PDF.
Two more things we do not do, which belong on a call rather than on the list: we do not generate or rewrite content - we measure the output side only, as described in input-side grounding vs output-side verification - and we do not promise to change what an AI says about anyone. Any vendor that does is selling a story.
What twelve questions cannot tell you
This script tests method, evidence and commercial shape. It does not test accuracy, because accuracy here cannot be judged from outside without a controlled bake-off across vendors on an identical prompt set, and nobody - us included - has published one. Treat any accuracy claim, ours or theirs, as unverified.
It also does not price your service line. That is downstream arithmetic, worked out in what to charge for GEO services, and it comes after the prior question of whether to sell this at all - GEO for agencies.
Run the script on us first
The cheapest test of a vendor's answers is to make them produce evidence before you pay. Request a free, human-run audit: send one client URL and an email, and we hand-run ChatGPT, Perplexity and Gemini, then reply within two business days with the exact answers, the sources, and the gaps. No card, no account. Hold the output against all twelve questions and against the pricing you would actually pay. To build the prompt set yourself first, the AI visibility checker generates one in the browser, free.
Related blogs
Related Post
Expand your knowledge with these hand-picked posts.
How to Track Brand Mentions in ChatGPT in 2026: The Agency Method
ChatGPT has no Search Console - no impressions, no analytics, no export. Here is the method agencies use to track brand mentions across ChatGPT and the other answer engines without trusting a black-box score, and why "citations" and "mentions" are two different numbers.
Gan Liu

Best AI Visibility Tools for Agencies in 2026: An Honest Shortlist
Most "best GEO tool" roundups rank on engine count. Agencies buy on a different axis - client roster, white-label, and evidence you can hand a client. Here is the honest shortlist, what each tool is actually best at, and the three questions that decide it.
Gan Liu

How to Track Brand Mentions in Perplexity in 2026: The Agency Method
Perplexity shows numbered citations and a Sources list under every answer. The agency method for reading brand mention vs citation off it - no black-box score.
Gan Liu
Start with the free audit
Send us one brand. We’ll run it through the engines and send back the answers, citations, and sources — so the first report your client sees is already backed by evidence.
Free audit: ChatGPT, Perplexity and Gemini, run by hand. Paid work covers all six engines.