Headings and lists may change how an AI answer engine distributes citation credit after sources have already been retrieved, but a new controlled experiment does not show that formatting reliably gets a source cited in the first place. That distinction is the central finding of CITECHOICE, a September 2026 preprint that tries to separate causal effects from the correlations that dominate much of today's GEO advice.
The CITECHOICE preprint, by Sriram Selvam and Anneswa Ghosh, analyzes 113 pairs of documents that supported the same pre-specified fact within authentic agentic-search transcripts. When one source was presented in a structured form using headings and lists, it received an average of 0.50 additional citation markers per answer. But the more important question — whether that source was cited at least once — increased by only 4.5 percentage points, with a confidence interval that crossed zero. The authors therefore treat the source-admission result as inconclusive.
A Search Engine Journal analysis highlights the same practical warning: structure appears capable of concentrating citation credit, but the experiment does not establish a dependable tactic for getting an otherwise uncited page into an AI answer. The paper is also a preprint and has not been peer-reviewed.
CITECHOICE tries to test citation allocation rather than observe it
Many AI-search studies begin by collecting answers, identifying which pages were cited and then comparing characteristics of cited and uncited pages. That can reveal correlations, but it cannot establish that a particular characteristic caused the citation. A highly relevant source might rank higher, contain schema, use clearer headings and receive more citations simultaneously, without any one of those characteristics independently causing the outcome.
CITECHOICE was designed to reduce that problem. The researchers collected transcripts from a GPT-5.4 search agent using Exa as its search provider. From 129 answered everyday queries, they identified document pairs that appeared in the same search call and independently supported the same fact. The selection process did not use the documents' rank or eventual citation outcome, and a blinded human review later confirmed 103 of the 113 pairs as genuine matches.
The researchers then replayed the frozen conversations while manipulating two properties: the relative order of the matched documents and whether the target document appeared as prose or as a structured rewrite. Only the final answer was regenerated. This design allows the study to ask a narrower causal question than a normal citation survey: when the retrieval context is held fixed, does changing how a source is presented alter the credit the model assigns to it?
Structured pages received more citation markers
The clearest result was citation count. Structured rendering increased the target document's citation count by 0.50 citations per answer on average, with a 95% confidence interval from +0.20 to +0.84. The Holm-adjusted p-value was .033, making this the study's strongest statistically supported presentation effect.
That does not mean the answers suddenly contained more citations overall. According to the paper, total citation volume did not increase, and the competing matched source did not lose a comparable amount of credit. The authors characterize the result as citation credit becoming more concentrated on the structured version.
This matters because “more citations” can describe two very different outcomes. A source can begin appearing in answers where it was previously absent, or a source that was already being used can receive additional citation markers across multiple claims. CITECHOICE provides stronger evidence for the second mechanism than the first.
The +4.5-point source-admission result was not conclusive
The study's pre-specified primary outcome was binary citation incidence: did the target document receive at least one citation in the answer? Structured text increased that probability by 4.5 percentage points, but the 95% confidence interval ranged from -1.4 to +10.4 points and the reported p-value was .168.
Because that interval includes zero, the experiment cannot rule out the possibility that the observed increase came from sampling variation. The researchers also note that the study was powered to reliably detect effects of roughly 8.5 percentage points or larger, meaning a smaller genuine effect could exist without this experiment being large enough to establish it.
For GEO practitioners, this is the result that should prevent an easy headline such as “use lists to get cited by AI.” The data do not support that claim. Structure may influence citation allocation after a source is already inside the retrieved context, but reliable source admission remains unproven.
The experiment does not isolate formatting alone
There is another important limitation. The structured and prose conditions were generated as separate rewrites rather than being guaranteed word-for-word identical content with only HTML or visual formatting changed. Grok 4.3 produced nearly all of the rewrites, with GPT-5.4 used as a fallback for one pair, and the researchers used a separate fidelity check to ensure that the underlying facts remained equivalent.
That means the observed effect cannot be attributed purely to heading tags, bullet syntax or layout. The structured version may have changed phrasing, segmentation or emphasis in ways that made individual claims easier for the answer model to associate with the document.
The authors explicitly avoid claiming a pure formatting mechanism. That caution is important for publishers because changing an article from paragraphs to bullets is not necessarily equivalent to the experimental treatment used here.
A stricter formatting test did not produce a stable advantage
The researchers also attempted a more constrained comparison in which wording remained the same while sentences were placed into list rows. Across all 113 pairs, that manipulation initially appeared to increase citation incidence. But when the experiment was repeated on a subset, the direction of the effect reversed.
That instability reinforces the paper's central warning. Even when a presentation choice looks favorable in one run, the result may not survive repetition. The researchers describe their findings as an attribution-sensitivity warning rather than an optimization recipe.
Fifteen percent of citation decisions changed on rerun
CITECHOICE also measured something GEO experiments often ignore: the noise floor of the model itself. The researchers reran 120 responses with the same inputs and found that the binary decision about whether the target source was cited changed in 15% of cases.
In other words, approximately one in seven source-level citation decisions flipped even though the underlying input was unchanged. The aggregate citation-count effect remained comparatively stable, but the authors estimate that decoding randomness accounted for about 45% of the variance in a single generation's family-level effect.
This has immediate implications for AI visibility testing. A single prompt run is weak evidence that an optimization “won” or “lost” a citation. If the same frozen context can produce different citation choices on repeated generations, before-and-after screenshots can easily overstate causal effects.
Source position also looked more powerful in observational data than under intervention
The paper's order experiment offers a useful parallel. In the original search results, documents appearing in the first position were cited 85.1% of the time, compared with 42.8% for documents in fifth position — a raw gap of 42.3 percentage points.
It would be tempting to conclude that moving a source upward causes a similarly enormous citation increase. The controlled replay produced a much smaller result. Moving the same page higher within its matched pair increased citation incidence by 7.9 percentage points in the main replay, but that result did not remain statistically significant after adjustment for multiple tests. A held-out order-only test on 56 pairs estimated an effect of exactly 0.0 percentage points, with a 95% confidence interval from -5.4 to +5.4.
The explanation is familiar from traditional search research: position and quality are confounded in observational data. Search providers generally place more relevant documents higher, so the raw position gap contains both ranking and document-quality effects. CITECHOICE shows why citation correlations should not automatically be converted into optimization instructions.
This was a frozen post-retrieval experiment
Perhaps the biggest practical limitation is that the study did not modify live web pages and then ask whether an AI search system crawled, retrieved and selected them differently. The researchers replayed saved transcripts in which the documents were already present. That deliberately isolates citation allocation, but it removes several earlier stages of the answer-engine pipeline.
A real publisher has to pass through crawling or acquisition, indexing or corpus inclusion, retrieval, ranking or source selection, answer generation and finally citation allocation. CITECHOICE mainly studies the last part of that sequence while holding retrieval context fixed.
Consequently, the study cannot tell us whether restructuring a live article makes it more likely to be retrieved by an answer engine. It also cannot establish that a structured page will outperform prose across different search providers, models or production systems.
The useful GEO lesson is narrower — and better
The absence of a simple optimization trick does not make the result unimportant. It gives content teams a more precise model of what structure might be doing. Headings, lists and clear segmentation can make information easier to parse and associate with individual claims. In this experiment, that changed how citation credit was distributed once equivalent sources were already available to the model.
That is different from saying structure creates retrieval eligibility or guarantees selection. A beautifully structured page that never enters the relevant retrieval set cannot receive citation credit. Likewise, a page can be retrieved and still lose the final attribution decision to another source supporting the same fact.
For publishers, headings and lists therefore remain sensible editorial tools because they improve navigation, readability and information architecture for human readers. CITECHOICE adds evidence that presentation can also affect machine attribution in a controlled setting. What it does not provide is a reason to restructure every page solely to chase AI citations.
AI citation tests need repetitions, not screenshots
The 15% rerun flip rate may ultimately be one of the study's most useful findings for practitioners. GEO measurement often treats a citation as a deterministic ranking result: the brand appeared today, therefore something worked; it disappeared tomorrow, therefore something broke.
Generative systems do not behave that cleanly. Citation measurement should therefore include repeated runs, consistent prompts, controlled conditions and reporting of variability. Teams should distinguish citation count from citation incidence and, where possible, separate retrieval from final attribution.
Those distinctions make experimentation harder, but they also make the conclusions more credible. A strategy that produces a small apparent gain in one answer should not be promoted as an AI-search best practice until the effect survives repetition and can be separated from model randomness and other variables.
Structure can move credit without reliably opening the door
CITECHOICE does not show that headings and lists are irrelevant. Its strongest result says the opposite: presentation changed the number of citation markers assigned to an equivalent source within a frozen retrieval context. That is evidence of causal sensitivity in the attribution stage.
But the experiment also demonstrates why that finding needs careful language. The structured treatment did not conclusively increase the probability that the target source was cited at all, the stricter formatting test was unstable on repetition, and 15% of binary citation decisions changed when identical inputs were rerun.
The practical conclusion is that structure can redistribute AI citation credit, but it is not a reliable ticket into the answer. For GEO, that is a more useful result than another formatting checklist: it separates a measurable attribution effect from the much larger and still unresolved problem of getting a source retrieved, selected and cited consistently in the first place.