AI Visibility Benchmarks by Industry 2026
AI Visibility Benchmarks 2026: Industry-Wide Score Data
Historical Q1 peer distributions and a reproducible Q3 five-engine source benchmark, with the measurement boundaries needed to compare either one honestly.
A useful AI visibility benchmark holds the prompt set, engines, time window, and scoring version constant. Foglift's historical Q1 industry medians range from 48 to 62. The separate Q3 source benchmark found that 75% of the domains in the five engines' combined top-25 lists were exclusive to one engine.
Q3 2026 benchmark update: market baselines need five engines
Foglift sent the same 75 brand-neutral buyer questions across 25 verticals to ChatGPT, Claude, Gemini, Google AI Overview, and Perplexity. The completed Q3 panel contains 375 answers and 3,412 cited URLs across 1,510 distinct domains. This is a source-behavior benchmark, so it complements the industry score distribution below instead of replacing it.
| Q3 finding | Measured result | Decision it supports |
|---|---|---|
| Leading sources differ by engine | 69 of 92 domains in the combined top-25 lists appeared in one engine's top 25 | Measure each engine separately before assigning content or distribution work |
| Cross-engine overlap is limited | Mean pairwise Jaccard similarity was 0.094 | Do not treat one engine's citation set as a proxy for the other four |
| AI answers cite specific pages | Only 1.5% of cited URLs were homepages; 74.7% were at least two path segments deep | Track page-level citations and the exact guides, comparisons, or documentation that earned them |
Read the full Q3 2026 AI Search Citation Benchmark for the methodology, per-engine top domains, limitations, cite-as block, and downloadable aggregate CSV.
How to use an AI visibility benchmark
A benchmark is useful only when your result and the comparison cohort share the same denominator. Match the prompt set, engine list, collection window, scoring version, and brand scope before treating a difference as performance. A healthcare score collected from four engines cannot be compared directly with a SaaS score collected from five engines or with a single live checker result.
| Benchmark input | Keep constant | Why it matters |
|---|---|---|
| Questions | Prompt text, intent, language, and category | A brand can lead one buying job and disappear from another |
| Engines | The same named answer engines | The Q3 panel found little overlap between leading source sets |
| Time | Collection dates and observation cadence | Answers and citations change as retrieval systems refresh |
| Score contract | Formula, weights, exclusions, and version | Two scores on a 0 to 100 scale can measure different things |
Two benchmark contracts, kept separate
The Q1 industry table is an archived peer cohort. It covers 4,217 brands, 150 or more industry-specific prompts, and four engines: ChatGPT, Perplexity, Claude, and Google AI Overview. Its 0 to 100 composite combined citation frequency, recommendation rank, sentiment, contextual relevance, and cross-engine consistency. Use it for historical peer context inside the five published industries.
The Q3 study answers a different question: which pages and domains the five engines cite for the same frozen buyer questions. It adds Gemini, publishes the frozen panel and aggregate CSV, and does not assign the industry scores shown below. Foglift's current product score is a third contract: a versioned six-tier composite built from a closed 14-day window. Read the exact current formula at How Foglift Scores AI Visibility.
Comparison boundary
Do not compare a current workspace score or a one-result checker output numerically with the Q1 industry table. Establish a current baseline, keep the same prompts and engines, and measure change against that baseline. The historical medians remain useful for category context, not score parity.
AI Visibility Benchmarks 2026 Table
Source: Foglift Q1 2026 aggregate benchmark dataset, 4,217 brands evaluated with 150+ industry-specific prompts across ChatGPT, Perplexity, Claude, and Google AI Overviews. AI Readiness benchmark scores measure extraction readiness on the site. AI Visibility benchmark data measures actual citation, rank, sentiment, and relevance across AI engines. For the current site-readiness baseline, see Foglift's Q2 2026 AEO Readiness study.
| Industry | Sample size | Median AI Readiness Score | Median AI Visibility Score | Top-quartile threshold |
|---|---|---|---|---|
| SaaS / B2B Software | 1,148 | 68 | 62 | 84 |
| E-commerce / DTC | 739 | 51 | 48 | 73 |
| Healthcare / Health Tech | 604 | 61 | 55 | 79 |
| Agencies / Consultancies | 802 | 56 | 51 | 74 |
| Education / EdTech | 924 | 63 | 58 | 81 |
AI visibility benchmarks by industry
Below are the median, top-quartile, and bottom-quartile observations from the archived Q1 2026 cohort. The tables preserve the original four-engine measurement contract.
1. SaaS / B2B Software
SaaS and B2B software recorded the highest median AI Visibility score in the cohort at 62. The bottom-quartile value was 38 and the top-quartile threshold was 84, a 46-point spread. Use the prompt-level evidence behind your own score to find whether the gap comes from mentions, citations, recommendation rank, or engine coverage.
| Metric | Bottom 25% | Median | Top 25% |
|---|---|---|---|
| AI Visibility Score | 38 | 62 | 84 |
| ChatGPT Citation Rate | 12% | 34% | 61% |
| Perplexity Mention Rate | 8% | 28% | 53% |
| AI Overview Inclusion | 5% | 19% | 42% |
| Avg. Recommendation Rank | #7+ | #4 | #1-2 |
What to inspect next: comparison coverage, product and API documentation, independent category mentions, and differences between engines. These are diagnostic checks, not causal claims from the cohort.
2. E-commerce / DTC
E-commerce and DTC recorded the lowest median in the five-industry cohort at 48, while the top-quartile threshold reached 73. The 49-point spread matters more than the category label alone: compare shopping, product-comparison, and brand prompts separately before deciding where the gap sits.
| Metric | Bottom 25% | Median | Top 25% |
|---|---|---|---|
| AI Visibility Score | 24 | 48 | 73 |
| Product Recommendation Rate | 6% | 18% | 44% |
| Perplexity Shopping Citations | 3% | 14% | 37% |
| AI Overview Product Inclusion | 2% | 11% | 29% |
| Avg. Recommendation Rank | #8+ | #5 | #2 |
What to inspect next: buying guides, product-comparison evidence, review sources, category-page answers, and the exact pages cited for shopping prompts.
3. Healthcare / Health Tech
Healthcare and health tech recorded a median of 55 and a top-quartile threshold of 79. Health answers require careful source review because a mention, a citation, and a recommendation have different implications in a high-stakes category.
| Metric | Bottom 25% | Median | Top 25% |
|---|---|---|---|
| AI Visibility Score | 31 | 55 | 79 |
| ChatGPT Health Citation Rate | 9% | 26% | 52% |
| Trust Signal Score | 22/50 | 35/50 | 46/50 |
| AI Overview Health Inclusion | 3% | 15% | 38% |
| Author Authority Index | Low | Medium | High |
What to inspect next: named medical review, author credentials, primary clinical sources, structured entity evidence, and which pages each engine cites.
4. Agencies / Consultancies
Agencies and consultancies recorded a median of 51 and a top-quartile threshold of 74. Segment prompts by service, industry, location, and buyer stage so a broad agency average does not hide a strong or weak specialist position.
| Metric | Bottom 25% | Median | Top 25% |
|---|---|---|---|
| AI Visibility Score | 26 | 51 | 74 |
| ChatGPT Agency Citation Rate | 5% | 19% | 41% |
| Case Study Indexing Rate | 11% | 32% | 58% |
| Thought Leadership Score | 15/50 | 29/50 | 43/50 |
| Avg. Recommendation Rank | #9+ | #5 | #2 |
What to inspect next: public case studies, named expertise, client evidence, independent review profiles, and the sources engines use for service recommendations.
5. Education / EdTech
Education and EdTech recorded the second-highest median at 58, with a top-quartile threshold of 81. Compare course discovery, program evaluation, outcomes, and institutional-trust prompts separately because each intent can retrieve a different source set.
| Metric | Bottom 25% | Median | Top 25% |
|---|---|---|---|
| AI Visibility Score | 33 | 58 | 81 |
| ChatGPT Education Citation Rate | 10% | 30% | 56% |
| Content Depth Score | 20/50 | 36/50 | 47/50 |
| Course/Program Schema Adoption | 8% | 29% | 64% |
| Avg. Recommendation Rank | #7+ | #4 | #1-2 |
What to inspect next: course and program details, instructor identity, crawlable curriculum pages, outcomes evidence, and engine-specific citations.
Cross-industry findings with reproducible evidence
The following conclusions come from published Foglift research or the cited academic experiment. They do not assign causality to the historical industry cohort.
- Extraction readiness still lags SEO readiness. The industry table above uses Foglift's Q1 2026 benchmark cohort of 4,217 brands. A separate Q2 2026 AEO Readiness study analyzed 1,386 scans across 344 broader-market domains. In that Q2 population, the 311 domains with full AEO scoring had a median AI Readiness Score of 46/100 versus a median SEO score of 86/100, and 44.5% of SEO-strong domains still scored below 50 on AEO.
- Cross-engine consistency is rare. Foglift's Q3 2026 AI Search Citation Benchmark found 1,510 distinct cited domains across 375 buyer-intent answers. Of the 92 domains in the combined engine top-25 lists, 69 appeared in one engine's top 25 and none appeared in all five.
- Answer construction can change source visibility. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande (“GEO: Generative Engine Optimization,” KDD 2024) measured source-visibility gains from tactics such as cited sources, quotations, and statistics in their benchmark environment. Treat those experimental results as evidence for clear sourcing, not a guaranteed lift in production answer engines.
Industry Comparison Summary
| Industry | Median Score | Top 25% | First diagnostic |
|---|---|---|---|
| SaaS / B2B | 62 | 84 | Compare engines and buyer intents |
| Education / EdTech | 58 | 81 | Split course and institution prompts |
| Healthcare | 55 | 79 | Review cited clinical sources |
| Agencies | 51 | 74 | Segment service and industry prompts |
| E-commerce / DTC | 48 | 73 | Separate shopping and brand prompts |
Build a benchmark you can compare over time
The strongest operational benchmark is your own versioned baseline. Keep the inputs stable, then use the answer evidence to decide what to change.
- Freeze the question set. Group prompts by discovery, comparison, review, and decision intent. Record the exact text and language.
- Choose the engines. Free workspaces monitor Perplexity weekly while active. Launch adds ChatGPT, Claude, Gemini, and Google AI Overview with cadence up to daily.
- Separate the outcomes. Record brand mentions, recommendation position, sentiment, and provider-returned citations as different events. A cited page does not prove the brand was recommended.
- Keep the collection window closed. Compare equal periods and exclude provider errors from the denominator. Foglift's current score stores a closed 14-day window and the eligible prompt IDs with every snapshot.
- Inspect the source layer. Review the exact URLs each engine cites. The Q3 study found that the leading domains differ enough that one engine cannot stand in for the other four.
- Connect visibility to visits. Count recognized AI referrals by landing page separately from crawler fetches. A live fetch is useful retrieval evidence, but it does not prove a citation or a human visit.
What Foglift keeps with every current score
The current product score is designed for auditability. Each snapshot keeps the evidence needed to explain a change.
Current score evidence
- Scoring version: the formula contract used for the snapshot
- Closed window: the 14-day start and end timestamps
- Eligible evidence: successful observations, prompt counts, and exact prompt IDs
- Tier detail: Technical, Authority, Brand and Product, Review Presence, Comparison, and Share of Voice inputs
- Answer evidence: the engine, mention, position, sentiment, competitor, and citation details behind the monitored tiers
The full AI Visibility scoring methodology publishes the six weights, prompt eligibility rules, error handling, sentiment mappings, and separate competitive share-of-voice denominator.
Frequently Asked Questions About AI Visibility Benchmarks
What's a good AI Readiness benchmark score?
Use a benchmark collected with the same scoring version and page scope. Foglift's separate Q2 2026 readiness study found a 46/100 median across 311 domains with complete scoring. That technical-readiness baseline is not interchangeable with the historical industry AI Visibility scores on this page.
How does my industry compare on AI visibility benchmarks?
In the archived Q1 cohort, SaaS and B2B software had the highest median AI Visibility score at 62/100, followed by education and EdTech at 58, healthcare at 55, agencies and consultancies at 51, and e-commerce and DTC at 48. Compare those figures only inside the same Q1 measurement contract.
Where do these AI visibility benchmarks come from?
The historical industry scores come from Foglift's Q1 2026 dataset of 4,217 brands. Each brand was evaluated with at least 150 industry-specific prompts across ChatGPT, Perplexity, Claude, and Google AI Overview. The separate Q3 source benchmark uses 75 frozen buyer questions across five engines and publishes its methodology plus aggregate CSV.
How current are these AI visibility benchmarks?
This page separates two measurement contracts. The five-industry score distribution is the Q1 2026 baseline collected from January 15 through March 15. The current source-behavior evidence comes from Foglift's Q3 citation benchmark, collected from August 1 through August 10 across 375 answers from five engines. Active Free workspaces monitor Perplexity weekly. Our paid plans add four engines and faster cadence.
Can I compare these industry scores with my current Foglift score?
Use the Q1 scores only as historical peer context. The current product score uses a versioned six-tier composite and a closed 14-day monitoring window, so it is not numerically interchangeable with the archived Q1 cohort. Compare current snapshots with other snapshots that use the same scoring version, prompts, engines, and window.
Methodology Note
The industry score distribution is based on Foglift's archived Q1 2026 dataset of 4,217 brands across five industries. Each brand was evaluated using at least 150 industry-specific prompts across ChatGPT, Perplexity, Claude, and Google AI Overview. Scores reflect a 30-day rolling average collected between January 15 and March 15, 2026. The Q3 source benchmark uses a separate frozen panel of 75 buyer-intent prompts across five engines, collected from August 1 through August 10. The current product score uses a third, versioned six-tier contract. Keep all three separate when comparing results.
Sources & Further Reading
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, “GEO: Generative Engine Optimization,” KDD 2024, Princeton/IIT Delhi. Up to 40% source-visibility lift in generative-engine responses (30–40% Position-Adjusted Word Count lift; 15–30% Subjective Impression lift). Cite Sources, Quotation Addition, and Statistics Addition top three tactics; Quotation Addition strongest single intervention. arxiv.org/abs/2311.09735
- Foglift Research, “AEO Readiness Across 311 Websites: The Median Site Scores 46/100,” 2026. foglift.io
- Foglift Research, “AI Search Citation Benchmark: Q2 2026,” 2026. 75 buyer-intent prompts, five production AI search engines, 375 responses, and 1,119 distinct cited domains. foglift.io
- Foglift Research, “AI Search Citation Benchmark: Q3 2026,” 2026. 75 frozen buyer-intent prompts, five AI answer engines, 375 responses, 1,510 distinct cited domains, and a downloadable aggregate CSV. foglift.io
- Foglift, “How Foglift Scores AI Visibility,” 2026. Six published weights, a closed 14-day monitoring window, eligible prompt rules, error handling, and answer evidence. foglift.io
Build a current, repeatable baseline
Track a fixed prompt set, keep the engine and window boundaries visible, and compare like with like. Active Free workspaces monitor Perplexity weekly. Launch adds all five engines and cadence up to daily.
Fundamentals: Learn about GEO (Generative Engine Optimization) and AEO (Answer Engine Optimization) (the two frameworks for optimizing your content for AI search engines).
Related reading
What Is an AI Visibility Score?
Understanding the metric that measures how visible your brand is in AI search.
AI Search Trends 2026
10 predictions every marketer needs to know about AI-powered search.
Enterprise AI Monitoring
How large brands track AI search visibility at scale.
GEO Strategy Framework
A complete framework for building your generative engine optimization strategy.