Methodology: how we measure AI visibility
Last updated: October 4, 2026. The short version lives on the FAQ; this page is the full method.
1. What we measure
We ask AI assistants the questions real buyers type, about a specific business location, and record whether and how the business appears in the answers. Three things are recorded for every single sample: the full answer text, the sources the assistant cited, and the model version that produced the answer. Nothing is inferred after the fact - every figure in a report traces back to stored runs.
2. How we sample
- Questions: industry prompt packs built from real buyer language (services, pricing, comparison, urgency), localized with the city and state in the question text. We do not use precise coordinates - the city string is the location signal for all engines.
- Runs: every question is asked multiple times per engine, because assistants return a different answer on every run. Single answers are never reported as findings.
- Engines: assistants that expose a public API - which lets us sample reproducibly and label the exact model version behind every answer. The engines sampled for a given report are listed in that report; nothing is claimed for a surface we did not sample.
3. Sample sizes by product tier
| Product | Runs per question per engine | What the numbers are for |
|---|---|---|
| One-time scan | 3 (single engine) | A first read - rates with confidence intervals, treated as indicative |
| Deep audit | 5 or more (every engine sampled for that report) | Full distributions with 95% confidence intervals, competitor gap, fix list |
| Monthly report | 3 (every engine sampled for that report) | Trend monitoring only - month-over-month comparisons, never single-number conclusions |
The tiers exist on purpose: an audit buys enough samples for tight intervals; a monthly trend buys repeatability. What a monthly report will never do is call a change from a single run.
4. Why monthly re-runs beat daily checks
AI assistants generate a fresh answer every time, so daily checks mostly measure noise: a spike or dip on one day is usually sampling variation, not a real change in how the internet describes a business. Monthly re-runs with fixed questions do two things daily checks cannot. First, they make comparisons meaningful - the same questions, the same method, a month apart, with confidence intervals on both sides. Second, they make the comparison honest about uncertainty: when the intervals from two months overlap, we say no real change is visible yet, even if the raw percentage moved. When they stop overlapping, that is a shift worth acting on. This is also why "we checked today" claims from other tools deserve skepticism - ask what n was behind them.
5. How to read our numbers
- Rates, not verdicts: "mentioned in 4 of 10 runs (40%, 95% CI [19%–65%])" is the reporting unit. The interval is part of the number - a 40% rate on 10 samples is a wide range, and we say so.
- Confidence intervals: we use the Wilson score interval at 95% for every rate, including small samples where the normal approximation would mislead.
- Model versions: every sample stores the model version that produced it. When an engine updates and numbers shift, the shift is visible and attributable instead of silent.
- Citations: cited sources are recorded per run and classified (profiles, directories, review sites, news, forums, the business's own site), which is what the AI source leaderboard and competitor gap are built from.
6. Engine coverage
We sample assistants through their public APIs and record the model version behind every answer. Coverage is not a fixed promise on this page: the authoritative list for any given report is that report itself, which names every engine and model version behind its numbers. If a surface is sampled in your report, it is listed; if it is not, no rate is attributed to it. That way this page cannot drift out of step with what we actually run.
| Surface | Access path | How it is sampled |
|---|---|---|
| Gemini | Public API with Google Search grounding | Model version recorded per run; cited sources captured per run |
| ChatGPT | Public API with the web-search tool | Model version recorded per run |
| Perplexity | Public API (Sonar) | Model version recorded per run |
| Google AI Mode / AI Overviews | Not available | No public API exists; we add it when there is a compliant way to measure it. Reports say this explicitly. |
| Copilot / Grok | Planned via adapter | The engine layer is adapter-based; adding a surface is a configuration change, not a rebuild. |
7. Applied to ourselves: the public citation baseline
The same interval-first reporting applies to our own domain. Our public AI Visibility Index publishes which domains AI answers cite for local-AEO questions - and our own mention rate at baseline: 0/45 runs, Wilson 95% CI [0, 0.079]. The baseline uses its own documented method (Gemini via Google AI Studio with Google Search grounding, 15 questions, 3 independent runs each, collected October 4, 2026) and is re-run weekly; the product sampling described above is unchanged.
8. What we deliberately do not do
- We do not promise placements in AI answers - no honest method can, and our terms say so.
- We do not scrape unofficial interfaces for surfaces that have no public API.
- We do not report a single run as a finding, at any tier.
- We do not present a rate without its sample size and interval.
9. Limits you should know about
Sampling measures what assistants say when asked - it is not a measurement of rankings, traffic or revenue, and it does not capture what a specific user sees in a personalized session. Engine-side changes outside our control can shift answers between runs; that is exactly why every sample carries its model version and every report carries its sample size. Questions about the method? Contact us - we answer within 24–48 hours (business days).