How we measure AI search visibility
A visibility benchmark is a sample of answers to a defined set of questions. It can show where your brand appears, which competitors appear alongside it, and which sources an API returns. It cannot predict every answer a buyer will see.
What the hosted product measures
We run your questions three times across five grounded provider APIs: OpenAI, Anthropic, Google Gemini, Perplexity, and xAI. Grounded means web search is enabled. Keep the questions and competitors stable when comparing runs, and inspect the report’s provider and model details.
A brand mention in answer text and a cited source URL are distinct observations. Inspect the captured answer and source evidence rather than treating a single score as proof of a recommendation. A provider failure is not evidence that a brand is absent.
API answers are not consumer-app measurements
Consumer products can use different models, search tools, location, personalization, and conversation history. Our provider API benchmarks do not directly measure Google AI Overviews, Google AI Mode, or the exact ChatGPT screen a customer sees.
The July 2026 published study
The original study reports 10 software buying questions × five providers × three repeated samples = 150 answers. The source analysis reports that 149 answers returned web sources.
The published methodology names OpenAI gpt-5.4-mini, Anthropic claude-haiku-4-5, Google gemini-3.5-flash, Perplexity sonar, and xAI grok-4.3. These are the versions reported in the July article, not a statement of the current product’s model lineup.
- Brand counts use names and known aliases in answer text. Matching can miss indirect references or count ambiguous ones.
- Publisher counts deduplicate domains within each answer. G2 combines g2.com and learn.g2.com. Counts across publishers can exceed 150 because one answer can contain multiple sources.
- The published table excludes Gemini’s vertexaisearch.cloud.google.com redirect wrapper. This may underrepresent publishers behind unresolved redirects; provider comparisons need resolved URLs.
- Ten software questions are a narrow convenience sample. Three repeats per provider do not support precise population estimates or causal claims about ranking factors.
Download the published aggregates
These downloads transcribe the existing articles. The public evidence does not include raw answers, exact collection timestamps, full request configuration, or provider-level source counts. Consequently, the historical counts cannot be independently reconstructed from these files alone. We have not filled those gaps with generated data.
How to run a useful comparison
- Record the questions, competitor aliases, models, web-search settings, and run dates.
- Retain original answers, cited URLs, errors, and successful-response denominators.
- Inspect ambiguous mentions and resolve source redirects before aggregating publishers.
- Repeat the same question set after a documented change. Model and retrieval changes can also affect the results.
- Measure qualified visits and purchases separately: more AI mentions do not automatically mean more customers.
The open-source CLI exposes the implementation for inspection. Our comparison articles disclose that we build one of the products; competitor claims link to vendor documentation and are not presented as independent hands-on tests.
Inspect the output before you buy
The sample report uses illustrative data and shows the report format.
View a sample report Start tracking — $49/month