A team at Sprinklr just ran the largest controlled test of AI citation ranking factors anyone has published: 252,000 trials across six large language models, isolating one variable at a time until only four things came out with odds ratios above 10,000x. Everything else you've been told to fix on a comparison page — structure, tone, social proof — barely moved the needle. Here's what changes when you rewrite a page in the order the data actually supports, walked through step by step on an illustrative example.
A 252,000-Trial Study Just Reordered the AI Citation Ranking Factors List
Rahul Vishwakarma, Shushant Kumar, and Ratnesh Jamidar, researchers at Sprinklr, published the study at the 49th International ACM SIGIR Conference this July under the title "What Gets Cited: Competitive GEO in AI Answer Engines." The design is unusually clean for this kind of research. Instead of scraping citations across the open web and guessing at causes, they built a controlled two-document retrieval-augmented generation testbed: exactly two candidate sources go into the model's context, the two differ in precisely one factor, and the researchers record which one the model cites first. Brand names were anonymized. Source order was counterbalanced. That removes the two biggest confounds in most GEO research — brand familiarity and position bias — and leaves you with something close to a real experiment.
The scale backs it up. 1,440 base scenarios, each run against three query paraphrases, produced 4,320 scenario-query pairs. Each pair ran five independent times across six models: Gemini-2.5-Flash, GPT-5-Nano, GPT-5-Mini, GPT-5.2, Claude-3.5-Sonnet, and Kimi-K2-Thinking. That's 252,000 trials testing 18 factors grouped into six categories — content match, completeness, trustworthiness, readability, competitive standing, and freshness. Run one factor at a time, over and over, across six different model architectures, and the noise cancels out. What's left is a hierarchy, and it doesn't match the priority order most GEO checklists ship with.
The Four Gatekeeper Factors That Decide Citation Odds Before Anything Else Matters
Four of the eighteen factors produced effects so large the paper reports them as odds ratios exceeding 10,000x, and all six models agreed on all four regardless of how selective each model otherwise was. Kimi-K2-Thinking treated 83% of the eighteen factors as statistically significant. Gemini-2.5-Flash recognized only 33%, the most selective of the six. But every model, from the most responsive to the most conservative, flagged the same four gatekeepers. When six differently trained systems agree on almost nothing else but agree unanimously on four factors, that's not a stylistic preference. That's a floor.
Notice what isn't on this list. Nothing about paragraph structure, nothing about bullet points versus prose, nothing about how many customer logos sit on the page. The four gatekeepers are binary, checkable facts about a page: does it match the topic, does it state a price, does it carry a current date, does it present itself as the first choice or the backup. None of them require creative writing. All four are near-mechanical fixes once you know to look for them, which is exactly why they're worth walking through on a real page.
Rewriting One Comparison Page, Gatekeeper by Gatekeeper
Here's an illustrative example — a composite, not an audited client page — to make the priority order concrete. Say you run marketing for a mid-market observability platform, and you maintain a page titled "[Your Product] vs. Datadog: Which Is Right for Growing Teams?" It's the kind of comparison page every SaaS company builds and few maintain. Before touching a word of the prose, run it through the four gatekeepers in order.
Step one, topic match. Pull the actual query patterns the page is supposed to answer — "observability pricing comparison," "Datadog alternative for a 50-person engineering team," "log aggregation cost comparison" — and check whether the page's content actually addresses them or has drifted into generic platform marketing. A page titled as a comparison that spends most of its copy on your product's roadmap and only a fraction of it actually comparing anything is a topic mismatch waiting to lose a citation it should have won by default.
Step two, price. If the page doesn't state a number — even a starting price, even a range — fix that before anything else on this list. "Flexible pricing to fit your team" is not a price. "Custom quotes based on usage" is not a price. The study's finding is blunt: when a query implies cost matters, the source with a stated number beats the source without one, and it isn't close. Put an actual figure on the page, with whatever caveats you need, and the gatekeeper is cleared.
Step three, timestamp. Comparison pages rot fast — pricing tiers change, competitors ship features, the entire premise of the page can go stale in a quarter. Check the visible date. If it says last year, or has no date at all, update it and mean it: refresh the pricing, refresh the feature comparison, refresh the section that name-checks a competitor's plan you haven't verified in months. A visible, accurate, current date is worth more to citation odds than a rewrite of the page's opening paragraph.
Step four, list position. Read the page as if you'd never seen your own product before, and ask what your language implies about rank. "A strong alternative to Datadog for teams on a budget" concedes second place before the reader finishes the sentence. Reframe around what the page is actually for — the comparison, not a concession — and drop any construction that reads as "we're the backup option."
| GATEKEEPER FACTOR | WHAT IT COSTS YOU | THE FIX |
|---|---|---|
| Topic Mismatch | Citation odds collapse even on an otherwise strong page | Rewrite to answer the actual query intent, not adjacent marketing copy |
| Price Not Mentioned | Loses to any source stating a number, vague ranges included | State an actual price or starting tier, not "contact us" |
| Recent vs. Old Timestamp | A visibly stale page loses to a visibly fresh one on identical content | Update the visible date and the content behind it, together |
| Lower List Position | Self-framing as the secondary option gets treated as the secondary option | Remove "alternative to" and "runner-up" framing; lead with the comparison |
“Get any one of these four wrong and the study's odds ratios say it doesn't matter what else you did — structure, tone, evidence, comparisons. The gatekeeper failure dominates.”
A priority order, not a linear percentage scale. Gatekeeper effects were measured in odds ratios above 10,000x; secondary effects ranged roughly 1.6x–750x depending on the model; the bottom tier showed weak or inconsistent effects.
Then, and Only Then, Fix the Secondary AI Citation Ranking Factors
Seven more factors showed real effects, just far smaller and far more dependent on which model you're optimizing for. Missing Specifications, Hedged Language, Keyword Gap, Claims Without Evidence, Internal Contradictions, No Comparisons, and Less Comprehensive Coverage all moved citation odds, with odds ratios reported roughly in the 1.6x to 750x range depending on the model. That's a real range, worth closing. It just isn't in the same universe as the four gatekeepers, and it shouldn't be the first thing you touch.
Back on the observability comparison page: once the four gatekeepers are cleared, work through the secondary list. Missing specifications means naming actual limits — log retention windows, ingestion rate caps, seat minimums — instead of leaving a feature row blank. Hedged language means replacing "may help reduce alert fatigue" with a claim you're willing to state plainly and support. Claims without evidence means every performance or cost claim on the page needs something backing it: a benchmark, a customer number, a documented default. Internal contradictions means checking that the pricing figure you just added in step two matches every other place price is mentioned on the page — comparison tables are where these quietly drift apart. No comparisons and less comprehensive coverage both mean the same thing in practice: don't leave categories out of the comparison table because the answer is unflattering. A visibly incomplete comparison reads as a hidden one.
Model sensitivity varies more here than it does with the gatekeepers. Kimi-K2-Thinking treated 83% of all eighteen factors as significant, the most responsive model in the study, picking up on secondary and even some weak-effect factors that other models ignored. GPT-5.2 and its siblings sat in the middle. Claude-3.5-Sonnet flagged about half the factors as significant. Gemini-2.5-Flash was the most selective, responding to only a third. If you're optimizing primarily for citation in one engine, that variance matters. If you're optimizing across all of them, the secondary factors are still worth fixing, just not worth fixing before the gatekeepers.
| FACTOR | TIER | WHAT THE STUDY FOUND |
|---|---|---|
| Missing Specifications | Secondary | Real, model-dependent effect on citation odds |
| Hedged Language | Secondary | Real, model-dependent effect on citation odds |
| Keyword Gap | Secondary | Real, model-dependent effect on citation odds |
| Claims Without Evidence | Secondary | Real, model-dependent effect on citation odds |
| Internal Contradictions | Secondary | Real, model-dependent effect on citation odds |
| No Comparisons | Secondary | Real, model-dependent effect on citation odds |
| Less Comprehensive Coverage | Secondary | Real, model-dependent effect on citation odds |
| Content Structure | Weak / inconsistent | Effect did not hold consistently across models |
| Scattered Information | Weak / inconsistent | Effect did not hold consistently across models |
| Overly Promotional Tone | Weak / inconsistent | Effect did not hold consistently across models |
| Weaker Value Proposition | Weak / inconsistent | Effect did not hold consistently across models |
| Weaker Social Proof | Weak / inconsistent | Effect did not hold consistently across models |
| No vs. Old Timestamp | Weak / inconsistent | Effect did not hold consistently across models |
| Recent vs. No Timestamp | Weak / inconsistent | Effect did not hold consistently across models |
What Barely Moved the Needle in 252,000 Trials
The seven weak-effect factors are the part of this study worth sitting with, because they include several things GEO consultants — this agency included, in earlier pieces — have told clients to prioritize. Content structure. Scattered information. Overly promotional tone. Weaker value proposition. Weaker social proof. No-timestamp-versus-old-timestamp and recent-versus-no-timestamp comparisons. Across 252,000 trials and six models, none of these produced a consistent, reliable effect on which source got cited.
That doesn't mean structure and social proof are worthless. It means they aren't gatekeepers, and treating them as the starting point of a GEO rewrite gets the sequence backwards. A beautifully structured page with scannable headers and a wall of logos still loses the citation if it's missing a price, running a stale date, or muddling the topic. The study's contribution isn't "structure doesn't matter." It's that structure work delivers its value only after the four binary, checkable basics are already handled. Spend the first hour of a comparison-page rewrite on structure and you may be polishing a page that was never going to get cited regardless.
The Checklist, Reordered
Apply this to any comparison, alternative, or versus page you maintain, not just the illustrative one above. Run the four gatekeepers first: confirm the page's actual content matches the queries you want it to answer, confirm it states a real price, confirm the visible date is current and accurate, confirm nothing on the page frames your product as the secondary option. Only after all four are clean should you move to specifications, evidence, contradictions, and comparison completeness. Structure, tone, and social proof come last — worth doing, not worth doing first.
Pull up your highest-traffic comparison page today and run the four-question audit before you touch anything else: does it match the query, does it state a price, is the date current, does the framing concede second place. If it fails any one of those, fix that before you reach for a content-structure pass. For a deeper look at how comparison pages win citations once the basics are in place, or how to refresh old content for AI citations without a full rebuild, both go further into execution. If you're auditing effort and structure signals separately, see the breakdown of what content effort means to Google's ranking systems. For teams running this across a SaaS comparison library, our generative engine optimization work and tech and SaaS engagements start with exactly this four-factor audit before any rewrite begins.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.