Skip to main content
← Back to Blog

AI Search Strategy

How to A/B Test AI Search Visibility Changes

A controlled experiment protocol for teams that need to know whether a content change moved mention rate, citation rate, or share of voice. Last updated July 16, 2026.

AI search testing is a repeated measurement problem. Preserve one prompt-engine panel, change one public variable, keep a matched control, and compare changes rather than isolated scores.

Why a Classic Website A/B Test Is the Wrong Mental Model

A conversion test can send half of human visitors to variant A and half to variant B. An AI answer engine does not join that visitor split. It retrieves whichever public version its crawler and index can access, then generates an answer from a changing source set. Cookie-based variants are especially weak for this job because Googlebot generally does not support cookies.

The useful design is a staggered, controlled before-and-after experiment. Measure a fixed panel, publish one change to the treatment pages, leave matched control pages unchanged, and run the same panel again. The treatment effect is the change in the treatment group minus the change in the control group. That difference-in-differences step removes some category-wide movement, such as an engine update that raises citation rates for both groups.

Run-to-run variation still matters. NIST's 2026 draft practices for language-model evaluation state that models are generally sampled nondeterministically and that repeated attempts can reduce uncertainty while revealing how much uncertainty comes from model sampling. A 2026 AI visibility study found changing responses and sources from identical queries across Perplexity, OpenAI search, and Gemini. Its design used three consumer-product topics, daily sampling over nine days, and a separate ten-minute sampling regime. A single before screenshot and one after screenshot cannot support a causal claim.

Write the Experiment Card Before Editing the Page

Precommitting the outcome, panel, and decision rule keeps the analysis honest. It prevents a weak mention-rate result from being reframed as a successful sentiment test after the data arrives. Use this compact plan:

FieldExample
HypothesisAdding a sourced comparison table will increase first-party citation rate for shortlist prompts.
Primary outcomeFirst-party citation rate per prompt and engine. Mention rate and competitor share are secondary.
TreatmentOne comparison table added to the selected pages. No title, schema, or internal-link changes.
ControlMatched pages with similar topic, baseline visibility, age, and authority that remain unchanged.
PanelThe same buyer prompts, engines, locales, and evaluation rules before and after release.
Window14 days before and 14 days after indexability is confirmed, extended if planned runs are incomplete.
Decision ruleShip broadly only when the estimated lift clears the predeclared minimum useful effect and the interval excludes a material loss.

Foglift's five-engine source-divergence study shows why engine is part of the panel. Across 62 complete five-engine prompt sets, brand-mention agreement exceeded 90% for most engine pairs while citation-domain overlap remained much lower. An aggregate lift can conceal a ChatGPT gain and a Perplexity loss. Preserve both the engine-level estimates and the aggregate.

Choose One Variable You Can Defend

The strongest tests connect one edit to one mechanism. Retrieval tests change whether a page is likely to enter the candidate set. Extraction tests change whether a selected page presents a clear, supported answer. Bundling five edits may improve the page, but it does not tell you which mechanism worked.

CandidateTestable hypothesisGuardrail
Title and main headingSharper query alignment improves retrieval for one defined intent.Keep the promise accurate and preserve the canonical URL.
Entity statementA concise company, category, audience, and use-case statement reduces ambiguity.Use the same approved statement across the page and structured data.
Sourced comparison tableExtractable facts improve citation or absorption for shortlist prompts.Date every price and capability fact; change no other content block.
FAQ blockDirect answers improve extraction for recurring buyer questions.Measure the visible answers. FAQPage markup does not guarantee AI selection.
Citation formattingPrimary-source links placed beside claims make evidence easier to inspect and reuse.Hold the underlying claim and source constant when testing placement alone.
Structured dataSupported markup may clarify a visible entity or article fact.Google requires markup to match visible copy and documents no special AI schema.

The original GEO paper, accepted at KDD 2024, reported visibility gains of up to 40% inside its benchmark and found that effects varied by domain. That study established that content interventions can change visibility under a controlled generative-engine setup. It did not establish that every edit transfers to organic, longitudinal AI search. Treat each candidate as a hypothesis for your own pages and prompts.

Google's current AI-feature guidance adds an important boundary for schema experiments. AI Overviews and AI Mode require no special AI markup, and structured data should match visible text. A schema-only test can measure whether supported markup clarifies a page, but a null result does not mean the visible answer block failed.

Build a Stable Prompt and Control Panel

Start with 20 to 30 independent buyer prompts per segment when budget permits. Include discovery, shortlist, comparison, and problem prompts only if you will analyze those intents separately. Repeating one prompt 30 times helps estimate sampling variance for that prompt. It does not create 30 independent buyer questions.

Match each treatment page with an untreated page that has a similar topic, baseline citation or mention rate, age, internal-link depth, and authority. If page-level matching is impossible, use a staggered rollout across a homogeneous cluster: update half now and half after the first measurement window. Exclude pages touched by another release, digital PR campaign, migration, or major link acquisition during the test.

Lock the exact prompt text, engine, locale, account state, and grading rubric. Record the answer, brand mention, cited URLs, competitor mentions, retrieval failure, and timestamp for every run. NIST recommends saving complete evaluation logs with the system version and grouping logs intended for comparison. Screenshots are supporting evidence; structured row-level logs are the analysis dataset.

Use 14 Days as a Minimum Window, Then Check the Denominator

Collect 14 days of baseline observations. Publish the treatment, verify that the changed page is server-rendered, indexable, canonical, and available to the relevant search crawler, then begin the post period after the new version can be observed. Collect the same weekdays for at least 14 days. Extend the window if indexing is delayed or the planned prompt-engine runs are incomplete.

Fourteen days supplies a practical cadence floor. Statistical power still depends on baseline rate, the minimum lift worth acting on, prompt-to-prompt variation, and cost. Define the minimum useful effect first. A team that would only act on a 10-point mention-rate lift needs a different sample from a team trying to detect two points.

Report the number of unique prompts, engines, repeated runs, eligible responses, failures, and cited pages beside every rate. Foglift's July 16 dogfood snapshot measured 28% visibility across 1,000 responses, 21 prompts, and five engines, with 49 cited pages. The 28% becomes interpretable because the denominator and panel are visible. It remains a monitoring snapshot until a treatment, control, and predeclared comparison turn it into an experiment.

Analyze Prompt-Level Changes and Their Uncertainty

For a practical analysis, compute the outcome rate for every prompt in each period, average within engine, and calculate the treatment-minus-control change. Resample whole prompts, rather than individual repeated runs, to build a bootstrap interval. This keeps all repeated observations for one prompt together and respects the prompt as the independent cluster.

Larger programs can use a generalized linear mixed model with treatment, post-period, and treatment-by-post terms, plus random effects for prompt and engine. NIST's 2026 AI evaluation report tested this family of models across 22 frontier models and three benchmarks, showing how it can quantify uncertainty and separate item difficulty from system effects. The model is useful when the panel is unbalanced. A simple prompt-cluster bootstrap is easier to audit and is often sufficient for a content team.

Read the result in this order

  1. Did the primary metric move in the intended direction?
  2. Does the uncertainty interval exclude a material loss?
  3. Does the estimated lift clear the minimum effect worth shipping?
  4. Is the effect consistent across engines and buyer-intent segments?
  5. Did control pages, source domains, failures, or indexing state change at the same time?

A narrow positive interval supports rollout to the next matched batch. A wide interval means the test is inconclusive. A split result by engine supports an engine-specific follow-up. None of those outcomes proves that the same edit will work on every topic.

Keep the Test Safe for Search

Avoid duplicate public variants when a sequential or page-cluster test will answer the question. If separate URLs are necessary, Google's search-testing guidance recommends canonical tags pointing to the original and temporary 302 redirects for redirected experiments. Do not serve a special content version to crawlers. End the experiment once the decision criterion is met and remove temporary variants.

Log every content commit, deployment, indexing request, recrawl observation, model-access change, and external campaign during the window. If one of those events changes the treatment or control asymmetrically, pause the decision and rerun the affected period. A clean inconclusive test is more valuable than a precise number built on contaminated conditions.

A Reusable 28-Day Protocol

  1. Write one hypothesis, one primary metric, one minimum useful effect, and one decision rule.
  2. Select treatment pages, matched controls, and at least 20 to 30 buyer prompts per analyzed segment when budget permits.
  3. Collect 14 days of fixed prompt-engine observations and preserve complete answer and citation logs.
  4. Publish one content change. Confirm crawl access, canonical state, server-rendered text, and indexability.
  5. Collect the matching 14-day post period, extending it until planned observations are complete.
  6. Estimate treatment-minus-control movement and an uncertainty interval with prompts kept as clusters.
  7. Roll out, revert, or run a narrower follow-up according to the predeclared rule.

Use the GEO strategy framework to choose the next experiment, the AI search KPI guide to define the metric, and Foglift's dashboard as the repeated measurement layer. The discipline comes from the protocol: fixed panels, complete evidence, controlled changes, and qualified conclusions.

Sources

Frequently Asked Questions

Can you really A/B test AI search visibility?

Yes, with a controlled before-and-after design. Keep the prompt, engine, locale, model access path, and scoring rule fixed; change one public content variable; retain matched untreated pages or prompts; and compare the treatment change with the control change. A visitor-level split alone does not control which public version an AI engine retrieves.

How long should an AI search visibility test run?

Use 14 days before and 14 days after the changed page is confirmed indexable as a practical minimum, then extend the window when runs are sparse or the page has not been recrawled. The calendar alone is not the stopping rule. Stop when the planned prompt-engine panel has enough repeated observations to estimate the effect with a useful uncertainty interval.

How many prompts do I need for an AI visibility experiment?

Choose the sample from the smallest change you need to detect, your baseline mention rate, and the number of repeated runs you can afford. A practical starting panel is at least 20 to 30 independent buyer prompts per segment, measured repeatedly on each target engine. Treat prompts as the independent clusters when computing uncertainty; repeated runs of one prompt do not create many independent prompts.

What should I change in an AI search experiment?

Change one retrieval or extraction variable at a time: the title and main heading, a concise entity statement, a sourced comparison table, an FAQ block, citation formatting, or structured data that matches visible copy. Keep the URL, surrounding page, internal-link plan, and other tested variables stable during the measurement window.

What is the best metric for an AI search A/B test?

Choose one primary metric before the test. Mention rate works for recommendation visibility, first-party citation rate works for source selection, and share of voice works for competitive prompts. Report engine-level results and uncertainty alongside the aggregate because a blended rate can hide opposite movement across engines.

Fundamentals: Learn about GEO (Generative Engine Optimization) and AEO (Answer Engine Optimization) (the two frameworks for optimizing your content for AI search engines).

Related reading

Free tool

Run a free Technical Audit for your AI Readiness Score

Audit any URL in 30 seconds. See scores for SEO, AI Readiness, performance, security, and accessibility.

Free Technical Audit

No signup required. Results in 30 seconds.