AI SEO Testing · Evidence → Action

What should brands
test in AI search now?

An evidence library across four buckets: official guidance, large-scale studies, controlled experiments, and technical labs. Each is mapped to a test plan your team can scope, defend, and measure.

Hank Azarian Consulting · Strategic SEO, AI search, and growth testing.

usage grows · referrals stay flat
Sample / demo data Brief 01

Gen AI: Growth in Usage & Referrals (US Web, Jan 25 – Jan 26)

1.5B1B.5B0 Jan 25 Jan 26 Referrals →

Bars: visits · Line: referrals · Source: Similarweb, March 2026 · Illustrative adaptation

AI Tools vs Search Engines across the Purchase Journey

Discovery Research Narrowing Evaluation Purchase

Purple: AI tools · Orange: search engines · Source: Similarweb, March 2026 · Illustrative adaptation

Which of these findings should our team actually test?

The Problem

AI SEO research is everywhere. A working evidence system is not.

Teams are reading studies from platforms, tool vendors, publishers, and independent researchers, alongside official guidance from Google and OpenAI. The findings rarely arrive separated by evidence type, scored for confidence, or mapped to a test you can actually run.

"We separate official guidance from vendor studies, controlled experiments, risk analyses, and technical labs, then map each to a test you can execute."

The System

Two capabilities,
one working method.

01 Evidence Library

Standardized Learning Cards across four evidence buckets.

Every published finding is normalized into a Learning Card with source, evidence type, trust role, finding, business implication, test idea, confidence, primary metric, risks, and an activation path.

  • Citation Mechanics: how AI engines pick sources
  • Controlled Tests: paired experiments with clean setups
  • Technical Guardrails: eligibility before experimentation
  • Risk Signals + Technical Labs: competitive risk and edge tests
02 Guided Test Planning

Each card maps to a real test plan template.

Click Customize this test on any card. A guided refinement flow uses the card's recommended template, your goal, timeframe, and resources to generate a structured, source-cited test plan.

  • Pick a card from the library
  • Confirm goal, timeframe, resources, and audience
  • Get hypothesis, setup, measurement, success, and watchouts
  • Move from research to action, with source attribution

Evidence Library · Launch Set

Four buckets. Eight evidence cards. One testing method.

The goal is not to chase every AI SEO trend. It is to separate guidance, studies, experiments, risk analyses, and technical labs, then identify which findings are worth testing in your business.

Citation Mechanics

How AI engines pick, cluster, and re-cite sources.

2 evidence cards
Launch Large-scale study

How ChatGPT Sources the Web

Citation mechanics & co-citation behavior

ChatGPT web citations concentrate early in conversations, use multiple sources per cited turn, and often cluster sources by topic or domain category.

Test idea Build a first-turn prompt set, identify citation neighbors, then test whether new or refreshed assets increase brand inclusion or co-citation share.
Primary metricShare of cited answers across first-turn prompts
Evidence notes
Business implication. Brands should optimize for the first research question and understand which sources they appear beside, not just whether they are cited once.

Secondary metrics. Co-citation frequency · Citation position · Citation neighbor overlap · Brand mention sentiment

Best fit

Enterprise B2BSaaSPublishersPR and brand teams

Watchouts & limits

  • Source is a tool vendor
  • Dataset is specific to U.S. English ChatGPT conversations
  • Citation behavior can change by model or product update

Why this confidence: Large dataset and clear citation mechanics, but platform behavior may shift as ChatGPT search changes.

Launch Large-scale study

AI Overview Citations vs. Organic Rankings

SERP overlap & query fan-out evidence

AI Overview citations don't simply mirror the top organic results. Ahrefs' updated study found a lower top-10 overlap than its earlier work, suggesting more fan-out influence.

Test idea For each priority query, compare organic top-10 pages against AI Overview citations, then build a fan-out map of missing subtopics and source types.
Primary metricAI Overview citation inclusion
Evidence notes
Business implication. Ranking for the exact query is not enough. Teams need to test coverage across related subqueries, comparison intents, and supporting pages.

Secondary metrics. Organic rank overlap · Fan-out topic coverage · Supporting-page diversity · Citation source type

Best fit

Enterprise SEOPublishersSaaSEcommerce

Watchouts & limits

  • Google changes AIO systems frequently
  • Citation visibility can vary by personalization, location, and query class

Why this confidence: Large-scale SERP and citation dataset, but Google AI Overview behavior changes quickly.

Controlled Tests

Paired experiments isolating one variable at a time.

4 evidence cards
Launch Controlled experiment

HTML vs. Markdown for AI Visibility

Clean format test (anti-shortcut)

In Otterly's 14-day test, paired Markdown pages received no AI crawler visits or citations, while the HTML versions did receive AI crawler activity and citations.

Test idea Run a controlled format test on a small set of pages before investing in Markdown mirrors, llms-full mirrors, or alternate machine-readable copies.
Primary metricAI crawler visits by URL format
Evidence notes
Business implication. Markdown mirrors should not be treated as a default AI SEO shortcut for SaaS, service, or content sites.

Secondary metrics. Citation pickup · Log-file AI bot requests · Prompt-level answer inclusion

Best fit

SaaSService businessesContent-led B2BTechnical SEO teams

Watchouts & limits

  • Single-site experiment
  • Markdown may behave differently for developer docs or GitHub-native content
  • Short 14-day window

Why this confidence: Clean controlled setup with same content and discovery path, but limited to one brand/site context.

Launch Controlled experiment

Schema Quality and AI Overview Visibility

Structured-data quality test

A small controlled test found the well-implemented schema page was the only one to appear in an AI Overview.

Test idea Select matched page sets and test strong schema improvements against weak or missing schema, measuring AI Overview inclusion and traditional organic movement.
Primary metricAI Overview inclusion
Evidence notes
Business implication. Promising but not conclusive. Schema quality (not just presence) may matter. Replicate across matched page sets before scaling.

Secondary metrics. Indexation · Organic rank movement · Rich-result eligibility · Crawl frequency

Best fit

Technical SEOPublishersSaaSEcommerce

Watchouts & limits

  • Small sample
  • Cannot fully isolate schema from indexation effects
  • Needs replication across more pages

Why this confidence: Very clean design, but only three pages, promising rather than conclusive.

Expansion Controlled experiment

Evidence Density vs. Generic Claims

Page-level credibility treatment test

The GEO benchmark found that Cite Sources, Quotation Addition, and Statistics Addition improved generative engine visibility. AutoGEO later surfaced similar preference rules around source citation, factual accuracy, and specific evidence. The practical hypothesis: pages that support claims with named sources, concrete data, and attributed quotes may be easier to cite than pages that rely on generic assertions.

Test idea Select 10 to 20 comparable pages. Create evidence-rich rewrites for the treatment group by adding named citations, statistics, and attributed quotes. Keep a matched control group unchanged, then measure citation inclusion across ChatGPT, Google AI Overviews, and Perplexity over 60 days.
Primary metricCitation inclusion rate
Evidence notes
Business implication. Content teams should test whether stronger evidence density raises AI citation inclusion more than publishing additional generic content. The useful question is not whether the page sounds authoritative, but whether its claims are specific, attributable, and easy to verify.

Secondary metrics. Citation position · Answer influence score · Citation quality · AI referral signals

Best fit

PublishersSaaSEnterprise B2BService businesses

Watchouts & limits

  • Evidence must be accurate and attributable. Do not invent, inflate, or pad citations.
  • Adding sources should make the page more useful for readers, not just more citation-shaped.
  • Citation patterns vary by engine, query class, and source type.

Why this confidence: Consistent benchmark signal across GEO and AutoGEO, but not direct revenue proof. Visibility lift still needs to be tied to traffic, assisted conversions, or lead quality in production.

Expansion Controlled experiment

Answer-First Structure vs. Delayed Intros

Extractability and opening-structure test

AutoGEO surfaced preference rules around conclusion-first, comprehensive, and in-depth content. The practical hypothesis: pages that answer the core question in the opening block may be easier for generative engines to extract, paraphrase, and cite than pages that begin with delayed narrative intros.

Test idea Select 10 to 20 FAQ, glossary, explainer, or best-practice pages. Rewrite the opening block so each page answers the core query directly, then support that answer with structured explanation and evidence. Keep a matched control group unchanged, then measure answer-paraphrase overlap over 60 days.
Primary metricAnswer-paraphrase overlap
Evidence notes
Business implication. Teams should test whether leading with the answer improves AI answer influence before committing to full-scale rewrites. Extractability may depend more on opening structure than on total page length.

Secondary metrics. Citation inclusion · Citation share · Answer position · Scroll depth

Best fit

PublishersSaaSEnterprise B2BService businesses

Watchouts & limits

  • Do not turn pages into shallow summaries. Preserve depth, evidence, and supporting context.
  • Answer-first structure may not suit narrative pages, case studies, or opinion-led content.
  • Effects may vary by query intent, engine, and page type.

Why this confidence: Strong directional signal from AutoGEO preference-rule extraction, but it still needs page-type validation. The effect may vary by content format, query intent, and engine.

Technical Guardrails

Official guidance and eligibility validation. Necessary, not sufficient.

2 evidence cards
Launch Official guidance

AI Search Eligibility & Crawler Access Baseline

Platform guardrail & crawl-access baseline

Google says SEO fundamentals remain relevant for AI features and that AI Overviews / AI Mode may use query fan-out. OpenAI documents separate crawlers for search, user-triggered browsing, and model training.

Test idea Create an AI access baseline: Google eligibility, snippet controls, robots rules, OpenAI bot access, server-log verification, and page-level crawl status.
Primary metricCrawler & indexability readiness
Evidence notes
Business implication. Eligibility before experimentation. Verify indexability, snippets, robots controls, and AI crawler access before running speculative AI SEO tests.

Secondary metrics. Robots allow/deny status · Snippet eligibility · OAI-SearchBot requests · GPTBot requests

Best fit

Enterprise SEOTechnical SEOPublishersRegulated industries

Watchouts & limits

  • Official docs do not reveal full ranking systems
  • Crawler access does not guarantee AI citation visibility

Why this confidence: Official platform documentation. Defines eligibility and access controls, not ranking causality.

Launch Correlation study

Does llms.txt Improve AI Citations?

Anti-hype technical validation

SE Ranking found low adoption of llms.txt and no observed correlation between the file and AI citation frequency.

Test idea Implement llms.txt as a low-cost technical baseline, then monitor AI bot requests, citation frequency, and whether any crawlers actually fetch the file.
Primary metricAI citation frequency before/after implementation
Evidence notes
Business implication. llms.txt can be a low-effort baseline, but it should not displace higher-confidence AI visibility work. Measure before scaling.

Secondary metrics. AI crawler requests to /llms.txt · Bot access status · Indexed priority pages · Server log evidence

Best fit

Technical SEO teamsEnterprise sitesPublishersDeveloper-led content teams

Watchouts & limits

  • May become more useful later
  • Domain-level analysis may miss page-level effects
  • Major platforms do not appear to rely on it today

Why this confidence: Large domain-level dataset, but correlation studies cannot prove non-effect across every platform or future implementation.

Risk Signals + Technical Labs

Competitive risk analyses and edge / infrastructure experiments.

2 evidence cards
Launch Risk analysis

Self-Promotional Listicles in AI Search

Comparison-page & reputation-risk evidence

Self-promotional listicles still appear in AI citations, especially in software/review contexts, but Peec does not recommend the tactic because of reputational and algorithmic risk.

Test idea Audit AI answers for category, comparison, and “best X” prompts to measure whether answers cite neutral third parties, owned comparison pages, or biased competitor assets.
Primary metricCitation share across comparison prompts
Evidence notes
Business implication. Brands need to understand how comparison content appears in AI answers without blindly copying manipulative SEO tactics.

Secondary metrics. Self-promotional source rate · Competitor mentions · Sentiment · Third-party source mix

Best fit

SaaSReview-heavy categoriesMarketplace brandsReputation teams

Watchouts & limits

  • Software/review vertical bias
  • Does not mean self-promotional listicles are a safe strategy
  • Reputational risk may outweigh short-term visibility

Why this confidence: Large citation dataset with fixed prompts over time, but vertical-specific to software and review-style queries.

Technical lab Technical lab

Structural Metadata Through HTTP Headers

Crawler discovery & structural metadata experiment

A site owner can expose internal links and headings through HTTP response headers, enabling fast crawl/audit workflows without parsing full HTML.

Test idea Deploy structural headers on a staging directory, enforce byte budgets, monitor 5xx risk, validate crawl/audit utility, then compare AI crawler behavior before any production rollout.
Primary metricStructural metadata extraction success rate
Evidence notes
Business implication. Promising for owned-site technical auditing and machine-readable structure, but should not be presented as a proven AI citation lever.

Secondary metrics. Header size · 5xx rate · AI crawler requests · Internal-link graph coverage

Best fit

Technical SEOLarge programmatic sitesPublisher engineering teamsEdge/CDN teams

Watchouts & limits

  • Header-size limits can break pages
  • Requires dev/SRE/security involvement
  • Not a proven ranking or citation factor

Why this confidence: Medium for crawl/audit utility, low for AI visibility impact, not a proven AI citation lever.

All cards link to source material. Some sources are studies, some are official guidance, some are experiments, and some are technical labs. Confidence ratings reflect study design, dataset size, and replication risk.

Interactive · Test Customization

Customize a test plan in four questions.

Pick any Learning Card, confirm your context, and generate a plan tied to that study's recommended template: source, hypothesis, setup, measurement, success, and watchouts.

Customizing test from How ChatGPT Sources the Web Tool vendor · Profound Large-scale study Confidence: Medium-high · 4/5
  1. 01 Goal
  2. 02 Timeframe
  3. 03 Resources
  4. 04 Audience
What is your primary AI search goal?

This sets the test objective and success metric. The card's default goal is suggested.

What timeframe are you working with?

Determines measurement windows and target sample size. The card's suggested timeframe is highlighted.

Which resources are available?

Plans are scoped to what you can staff. Card-suggested resources are highlighted; pick all that apply.

What type of site or team are you optimizing?

Shapes the recommended setup. The card's best-fit audiences are highlighted.

Who it's for

Built for teams making disciplined AI search decisions.

  • Enterprise SEO teams
  • Content-led SaaS companies
  • Publishers
  • Ecommerce brands
  • Agencies evaluating AI visibility
  • Leadership teams prioritizing AI search investment
  • Analytics teams measuring assisted AI influence

Get Started

Stop collecting AI SEO studies.
Start running better tests.

Turn the evidence library into a prioritized testing roadmap your team can scope, defend, and measure.

Start with one evidence-backed test plan. Use it to decide what to test, how to measure it, and what not to overclaim.