A GEO Playbook From Analyzing 70 Million AI Search Citations

A GEO Playbook From Analyzing 70 Million AI Search Citations
Sponsored

Generative Engine Optimization has no shortage of tactical advice. The harder problem is separating recommendations supported by large-scale observation from practices that sound plausible because they fit a convenient story about how language models work. Pepper has entered that debate with a playbook drawn from what it says is an analysis of 70 million AI Search citations and experiments conducted across hundreds of clients.

The May 1 edition of Pepper’s AI Native newsletter summarizes seven lessons presented by GEO research lead Kishan Panpalia at the company’s Index ’26 summit in San Francisco. The recommendations range from entity optimization and topical-cluster coverage to third-party distribution, page structure and platform prioritization.

The scale makes the analysis worth examining, but the methodological boundary matters. Pepper’s article does not publish a reproducible 70-million-citation dataset, sampling methodology, prompt distribution, confidence intervals or controlled experimental protocol for every claim. The findings should therefore be read as proprietary observational evidence and practitioner guidance rather than seven universal laws of AI search.

GEO is becoming an organizational strategy rather than another SEO checklist

Pepper’s first argument is structural: GEO should be treated less like an isolated channel and more like product marketing for machines. AI answers reconstruct brands from a wider information environment than a company’s own website, so SEO, content, PR, social and customer evidence can all contribute to what a system understands.

That framing matches an increasingly visible feature of generative discovery. A model recommending software, hotels or professional services can draw on official pages, reviews, directories, editorial lists, community discussions and other third-party material. Optimizing only the corporate domain leaves much of that evidence environment outside the GEO program.

This does not mean every organization needs a separate GEO department. It means ownership becomes cross-functional. Technical accessibility may belong to SEO, extractable evidence to editorial teams, reputation signals to PR and product facts to product marketing.

Entities matter, but “AI thinks in entities, not keywords” is too absolute

Pepper’s second lesson is that the unit of optimization is shifting from keywords toward entities: identifiable brands, people, products, categories, concepts and the relationships among them. The practical recommendation is useful. Content should make clear what a thing is, what category it belongs to and how it relates to other concepts.

The underlying explanation should be treated more carefully. Modern AI search systems do not simply replace keyword matching with a single knowledge graph. Retrieval can involve lexical search, embeddings, ranking systems, entity signals, query fan-out and model reasoning in different combinations.

The actionable part survives that nuance. Ambiguous entity references make machine interpretation harder. Consistent naming, explicit product-category relationships, factual descriptions and clear connections between organizations, products and use cases give retrieval and generation systems less semantic ambiguity to resolve.

Topical coverage may matter more than owning one hero query

One of Pepper’s most interesting examples compares a hypothetical brand ranking first for one major project-management query with another brand appearing between positions four and seven across five related queries. Using reciprocal rank fusion, Pepper’s example gives the broadly visible brand a score of 0.0767 versus 0.0164 for the single-query leader—nearly a fivefold difference.

The lesson is not that Google rankings mechanically determine AI citations. Rather, a brand that is consistently visible across a cluster of related information needs creates more opportunities to enter the evidence environment surrounding a topic.

This aligns with query fan-out behavior documented in modern AI search. One visible user question can produce multiple underlying retrieval queries, meaning a page optimized only for one exact phrase may compete across a much wider semantic neighborhood.

NetContentSEO has examined this directly in Google AI Mode’s increasingly visible query-fan-out interface. As AI systems decompose a request into related subproblems, topical breadth becomes an eligibility strategy rather than simply an old pillar-and-cluster SEO tactic with a new label.

Pepper says source preferences differ sharply by AI platform

The fourth lesson is that AI engines should not be treated as one homogeneous citation market. Pepper reports that its research found different source patterns across platforms. In its presentation, ChatGPT was described as drawing 41% of the relevant signal from authoritative list mentions, 16% from reviews and 14% from customer examples, while Claude was said to draw 68% from traditional directories.

Those percentages are potentially useful but should not be generalized beyond Pepper’s dataset without more information about the prompts, categories, geography, time period and classification methodology. Source mixes can change substantially depending on what users ask and which product mode is being tested.

The broader conclusion is better supported: citation strategies are platform-dependent. NetContentSEO’s coverage of Ahrefs data comparing citation mixes across major AI engines found exactly that fragmentation, with different domains dominating different products.

Pepper recommends starting with ChatGPT and Google—but that is a prioritization choice

Pepper argues that marketers unable to optimize everywhere should begin with ChatGPT and Google because source overlap can provide spillover coverage into other systems. This is a pragmatic resource-allocation recommendation rather than a universal ranking rule.

The right priority will depend on the audience. A B2B software company whose buyers heavily use Claude could rationally allocate more attention there. An ecommerce brand may care more about Google’s increasingly agentic shopping environment. A publisher may prioritize whichever assistants produce measurable referrals or citations in its category.

The important practice is to measure engines separately. Combining all AI citations into one visibility score can hide the fact that a brand is strong in one discovery ecosystem and nearly absent in another.

The “ugly content” claim contains a useful idea wrapped in an overstatement

Pepper’s fifth recommendation is deliberately provocative: LLMs prefer “ugly” content to polished prose. The suggested pattern includes a TL;DR, definitions, explicit entity relationships, tables, infographics, video and visible FAQs, with roughly one core fact per block.

The useful principle is extractability, not ugliness. Retrieval and generation systems often need passages that can stand on their own as evidence. A paragraph containing a clear definition, number, comparison or procedure is easier to reuse than several paragraphs of atmospheric introduction that delay the factual payload.

But there is no reason to conclude that good human writing and machine readability are inherently opposed. Clear editorial prose can remain engaging while also making claims explicit, keeping referents unambiguous and organizing evidence into coherent sections.

Recent research covered by NetContentSEO makes the distinction particularly important. In 252,000 paired GEO tests across six models, topical relevance, source position, explicit price information and recency emerged as stronger citation factors than cosmetic formatting. Structure can help extraction, but formatting alone is not a substitute for useful evidence.

One fact per block is better understood as evidence modularity

Pepper’s “one core fact per block” advice deserves a more precise interpretation. AI retrieval systems frequently work at passage or chunk level. If a passage mixes several unrelated claims, the relevant evidence can become harder to isolate. If each section has a coherent informational purpose, retrieval systems have a cleaner unit to match against a question.

This does not require writing one sentence per paragraph or producing robotic pages. A professionally edited paragraph can develop one coherent factual idea across several sentences. The objective is semantic concentration, not visual fragmentation.

For publishers, a useful test is whether a retrieved passage still makes sense when removed from the rest of the page. If it relies on vague pronouns, unexplained context or an argument established many paragraphs earlier, its standalone retrieval value may be lower.

PR matters because AI visibility extends beyond owned domains

Pepper’s sixth recommendation reframes PR as a GEO growth function. Panpalia says he has observed press-release distribution and recurring third-party mentions being cited heavily by language models, leading Pepper to recommend a more consistent monthly distribution cadence rather than one or two large annual campaigns.

The broad idea has strong strategic logic. AI systems frequently encounter brands through third-party sources, and those sources can establish category membership, comparisons, product claims and reputation context that the brand’s own website cannot independently validate.

But the recommendation should not be interpreted as permission to flood press-release networks with low-value material. Search engines have long treated manipulative link schemes and mass-distributed promotional content cautiously, and AI systems can change source preferences. The sustainable objective is credible, fact-rich third-party evidence rather than citation volume for its own sake.

Third-party visibility can matter even when it does not send the click

One of the most consequential differences between GEO and conventional SEO is that an external page can improve brand visibility without sending a conventional referral. An AI system may retrieve a comparison, review or directory page, extract information about a company and mention the company in its answer.

That creates an indirect visibility path: third-party source → AI retrieval → brand recommendation. The user may never visit the third-party source or the brand’s website.

NetContentSEO has seen this pattern in commercial experiments where third-party pages generated most observed brand mentions while citation volume and referral traffic diverged. That makes distribution and reputation part of the retrieval problem, not merely a link-building exercise.

Flattening accordion FAQs is plausible advice, but not a universal crawler law

Pepper’s seventh recommendation is to remove accordion-style FAQs and display the content openly because hidden sections allegedly create extra work for AI crawlers. This is the playbook’s easiest recommendation to overgeneralize.

Many modern crawlers and rendering systems can process content that is present in the DOM even when it is visually collapsed. Whether an accordion causes a retrieval problem depends on how it is implemented, whether the content exists in the initial HTML, whether JavaScript is required to fetch it and which crawler or agent is accessing the page.

The safer rule is not “all accordions are bad.” It is to ensure important evidence is present in crawlable, rendered content and does not depend on a user interaction that the relevant retrieval system cannot perform.

The 70-million-citation headline needs methodological restraint

Large numbers create an impression of certainty, but dataset size is only one dimension of evidence quality. Pepper says it has analyzed 70 million AI Search citations, yet the newsletter does not disclose enough methodological detail to calculate how representative those citations are of the wider AI-search population.

Seventy million citations from a narrow set of commercial prompts would answer a different question from 70 million citations sampled across consumer, informational, local, medical, technical and transactional queries. Similarly, a dataset dominated by one geography or one model configuration could produce source distributions that shift elsewhere.

The playbook is therefore most useful as a hypothesis generator. Its seven recommendations provide concrete things to test against a brand’s own prompt set rather than constants that should be implemented blindly.

Controlled GEO research sometimes contradicts practitioner heuristics

This distinction matters because recent controlled studies have shown that apparently sensible GEO edits can have unexpected effects. NetContentSEO covered SAGEO Arena, where rewriting a page to make it more attractive to a generator could simultaneously reduce its chance of surviving retrieval.

That result exposes the danger of optimizing only for the final citation step. A page must first be crawled or otherwise available, retrieved for the relevant query, survive ranking or reranking and enter the model context. Only then can the generator decide whether to cite or absorb it.

A tactic that improves one stage can damage another. That is why Pepper’s closing emphasis on experimentation may be more durable than any individual tactic in the list.

Speed of experimentation may be the strongest lesson in the playbook

Panpalia’s final framing is that the difference between stronger and weaker GEO programs is not necessarily the software they buy but how quickly they can run experiments. That is difficult to disagree with in a discovery environment whose models, retrieval systems, source preferences and interfaces change frequently.

A practical GEO program should therefore treat recommendations as testable hypotheses. Select a strategically important page cluster, define the prompt set, record baseline citations and brand mentions, make a controlled change where possible and observe whether visibility changes across engines over time.

The measurement should also extend beyond citations. A page can be retrieved without being cited, cited without materially shaping the answer and cited repeatedly without producing valuable traffic. Citation frequency is one observable layer of a much larger system.

The real playbook is evidence, entities, distribution and testing

Pepper’s seven lessons can be condensed into a broader operating model. Make entities and relationships explicit. Build topical coverage across the cluster of questions surrounding a category rather than chasing one hero keyword. Produce passages with concentrated, reusable evidence. Ensure the brand exists in credible third-party environments. Measure AI engines separately. Keep important information technically accessible. Then test whether those changes actually alter retrieval, citations and recommendations.

None of those principles requires sacrificing human readability or turning every page into a database dump. The strongest GEO content can serve both audiences: people receive clear, useful editorial information, while machines encounter unambiguous entities and extractable evidence.

Pepper’s 70-million-citation analysis is valuable because it turns a rapidly evolving practice into specific hypotheses marketers can test. Its limitations are equally instructive. AI Search does not yet offer a stable optimization formula, and proprietary citation datasets cannot substitute for controlled measurement on the prompts that matter to a particular business.

The most defensible GEO playbook is therefore not a fixed list of tricks. It is an evidence-driven process for discovering which sources, passages and brand signals survive the full retrieval-to-answer pipeline—and updating that process as the engines change.

0%