Launch Controlled experiment
HTML vs. Markdown for AI Visibility
Clean format test (anti-shortcut)
In Otterly's 14-day test, paired Markdown pages received no AI crawler visits or citations, while the HTML versions did receive AI crawler activity and citations.
Test idea Run a controlled format test on a small set of pages before investing in Markdown mirrors, llms-full mirrors, or alternate machine-readable copies.
Confidence High · 5/5
Primary metricAI crawler visits by URL format
Evidence notes
Business implication. Markdown mirrors should not be treated as a default AI SEO shortcut for SaaS, service, or content sites.
Secondary metrics. Citation pickup · Log-file AI bot requests · Prompt-level answer inclusion
Best fit
SaaSService businessesContent-led B2BTechnical SEO teams
Watchouts & limits
- Single-site experiment
- Markdown may behave differently for developer docs or GitHub-native content
- Short 14-day window
Why this confidence: Clean controlled setup with same content and discovery path, but limited to one brand/site context.
Launch Controlled experiment
Schema Quality and AI Overview Visibility
Structured-data quality test
A small controlled test found the well-implemented schema page was the only one to appear in an AI Overview.
Test idea Select matched page sets and test strong schema improvements against weak or missing schema, measuring AI Overview inclusion and traditional organic movement.
Confidence Medium-low · 2/5
Primary metricAI Overview inclusion
Evidence notes
Business implication. Promising but not conclusive. Schema quality (not just presence) may matter. Replicate across matched page sets before scaling.
Secondary metrics. Indexation · Organic rank movement · Rich-result eligibility · Crawl frequency
Best fit
Technical SEOPublishersSaaSEcommerce
Watchouts & limits
- Small sample
- Cannot fully isolate schema from indexation effects
- Needs replication across more pages
Why this confidence: Very clean design, but only three pages, promising rather than conclusive.
Expansion Controlled experiment
Evidence Density vs. Generic Claims
Page-level credibility treatment test
The GEO benchmark found that Cite Sources, Quotation Addition, and Statistics Addition improved generative engine visibility. AutoGEO later surfaced similar preference rules around source citation, factual accuracy, and specific evidence. The practical hypothesis: pages that support claims with named sources, concrete data, and attributed quotes may be easier to cite than pages that rely on generic assertions.
Test idea Select 10 to 20 comparable pages. Create evidence-rich rewrites for the treatment group by adding named citations, statistics, and attributed quotes. Keep a matched control group unchanged, then measure citation inclusion across ChatGPT, Google AI Overviews, and Perplexity over 60 days.
Confidence Medium · 3/5
Primary metricCitation inclusion rate
Evidence notes
Business implication. Content teams should test whether stronger evidence density raises AI citation inclusion more than publishing additional generic content. The useful question is not whether the page sounds authoritative, but whether its claims are specific, attributable, and easy to verify.
Secondary metrics. Citation position · Answer influence score · Citation quality · AI referral signals
Best fit
PublishersSaaSEnterprise B2BService businesses
Watchouts & limits
- Evidence must be accurate and attributable. Do not invent, inflate, or pad citations.
- Adding sources should make the page more useful for readers, not just more citation-shaped.
- Citation patterns vary by engine, query class, and source type.
Why this confidence: Consistent benchmark signal across GEO and AutoGEO, but not direct revenue proof. Visibility lift still needs to be tied to traffic, assisted conversions, or lead quality in production.
Expansion Controlled experiment
Answer-First Structure vs. Delayed Intros
Extractability and opening-structure test
AutoGEO surfaced preference rules around conclusion-first, comprehensive, and in-depth content. The practical hypothesis: pages that answer the core question in the opening block may be easier for generative engines to extract, paraphrase, and cite than pages that begin with delayed narrative intros.
Test idea Select 10 to 20 FAQ, glossary, explainer, or best-practice pages. Rewrite the opening block so each page answers the core query directly, then support that answer with structured explanation and evidence. Keep a matched control group unchanged, then measure answer-paraphrase overlap over 60 days.
Confidence Medium · 3/5
Primary metricAnswer-paraphrase overlap
Evidence notes
Business implication. Teams should test whether leading with the answer improves AI answer influence before committing to full-scale rewrites. Extractability may depend more on opening structure than on total page length.
Secondary metrics. Citation inclusion · Citation share · Answer position · Scroll depth
Best fit
PublishersSaaSEnterprise B2BService businesses
Watchouts & limits
- Do not turn pages into shallow summaries. Preserve depth, evidence, and supporting context.
- Answer-first structure may not suit narrative pages, case studies, or opinion-led content.
- Effects may vary by query intent, engine, and page type.
Why this confidence: Strong directional signal from AutoGEO preference-rule extraction, but it still needs page-type validation. The effect may vary by content format, query intent, and engine.