Something Inc.Schedule a free consultation
TECHNICAL SEO

Answer Engine Optimization: Retrieval Mechanics or Operating Discipline?

Two credible researchers published opposing emphases on answer engine optimization in the same July — one says citation odds are set by training-data mechanics, the other says they're set by how well you run the work. Both are right, at different layers.

TTTyler TruffiManaging Partner · AUG 2, 2026 · 9 MIN READ

Answer engine optimization has split into two camps that sound like they're arguing about the same thing and aren't. One camp says your odds of getting cited by an AI engine are mostly set by mechanics you can't move — what got crawled, what got licensed, what made it into training data. The other camp says your odds are mostly set by discipline you build and keep running — structuring content, publishing, measuring, revising, on a loop. Two credible researchers staked out those positions within days of each other in July 2026. Neither is wrong. They're describing different layers of the same system, and conflating those layers is the actual mistake most practitioners make.

TL;DR · 60 SECONDSDan Petrovic of DEJAN argues (July 21, 2026) that AI citation patterns are downstream of training-data composition, crawl access, and licensing — mechanics a site has limited direct control over. Josh Blyskal of Profound argues (July 26, 2026) that answer engine optimization only works as a continuously run operating loop, not a one-time technical project. Something Inc.'s verdict: machine access is a floor you fix once; operating discipline is the compounding lever you run indefinitely. Treat the two as sequential, not competing.

What the answer engine optimization debate is actually about

In the same week of late July 2026, two people who spend their working lives studying how AI engines choose sources published pieces that read, on the surface, like they contradict each other. One is a retrieval and technical-SEO researcher explaining why a widely repeated claim about AI's source preferences is wrong. The other is a GEO strategist proposing a repeatable operating model for running answer engine optimization inside a marketing team. Put side by side, they look like a disagreement about what actually wins AI citations.

It isn't really a disagreement. It's two people looking at different layers of the same funnel and each describing the layer they specialize in as if it were the whole system. Retrieval researchers spend their time on what gets into the candidate pool an engine draws from in the first place. Strategists spend their time on what happens once your content is already inside that pool and competing for the citation slot. Both layers matter. Most teams doing answer engine optimization right now only work on one of them, usually without realizing the other exists.

The retrieval-mechanics case: what Dan Petrovic's Reddit research shows

Dan Petrovic, founder of the Australian SEO and GEO research consultancy DEJAN, published a post on dejan.ai/blog on July 21, 2026 titled "No, AI doesn't prefer Reddit. Search does." His argument runs against a popular assumption in the GEO industry: that large language models have some inherent preference for Reddit as a source. Petrovic's research points the other direction. AI training-data pipelines systematically underrepresent Reddit relative to its actual size and influence online, largely because of licensing restrictions and crawl limits on Reddit's own content. The outsized presence of Reddit in AI-generated answers, he argues, is mostly a downstream effect of Google's heavy reliance on Reddit within its own search results, with Google-rooted AI surfaces inheriting that bias rather than models themselves favoring the platform.

The broader point behind that finding, and behind adjacent DEJAN retrieval research, is the one worth sitting with. A meaningful share of whether a source gets cited at all is fixed upstream of anything a marketing team does on a given day. It's set by training-data composition, crawl access, and licensing arrangements a site has limited ability to influence directly. That's a different problem than writing better content. It's a machine-access problem, and it's the layer our own research on how AI engines decide what to cite has found teams consistently underweight — they'll rewrite a page five times before they check whether the engine's crawler can even reach it.

THE RETRIEVAL-MECHANICS READCitation odds aren't purely a content-quality contest. A real portion of the outcome is decided before your page is even evaluated, by whether the engine's crawlers, licensing deals, and training corpus gave your domain a fair shot at being in the pool at all.

The operating-discipline case: Josh Blyskal's four-stage loop

Josh Blyskal, an AI strategist at the GEO analytics company Profound, published research on joshblyskal.com on July 26, 2026 introducing what he calls SAGE for AEO: a four-stage operating loop for running answer engine optimization. The specifics of the framework aren't the interesting part. The premise underneath it is. Blyskal treats answer engine optimization the way a marketing team treats a reporting cycle — something you structure, act on, measure, and revise continuously — rather than a technical project you scope, ship, and close out.

Blyskal's framing, in plain terms: an operating model for answer engine optimization doesn't get installed once and left running in the background. It gets run, on a cadence, by a team that keeps showing up to it.

That's a pointed correction to how most companies currently budget for this work. Schema gets added once. A content audit happens once. A vendor tool gets switched on and left to report a number nobody acts on. Blyskal's argument, and the reason it matters more than the specific stage names in his loop, is that none of those one-time moves compound. Structural and content decisions do, because among the pages an engine has already decided are eligible to be cited, the ongoing work of making one page more extractable than a competitor's is exactly the kind of decision our piece on why schema markup doesn't get you cited walks through — the technical box-checking isn't nothing, but it isn't the lever people market it as either.

Something Inc.'s verdict on answer engine optimization: two layers, not one fight

Both researchers are right, at different layers, and treating their positions as competing theories is where practitioners go wrong. Technical and training-data accessibility is the floor. It's necessary, and once you've done the machine-access work, it's largely fixed — this is Petrovic's layer, and it's where a lot of teams should be spending their first real answer engine optimization budget instead of skipping straight to content. Operating discipline is where the ongoing, compounding leverage actually lives, because most of what decides whether your specific page gets chosen among the technically-accessible candidates is structural and content work you keep doing — this is Blyskal's layer, and it's the one that never finishes.

LAYERWHOSE RESEARCH MAPS TO ITWHAT YOU ACTUALLY DO ABOUT IT
Machine access — crawl, licensing, training-data presencePetrovic / DEJAN, retrieval-mechanics campAudit bot access, robots directives, and feed availability once. Fix it. Stop treating it as a lever you pull every quarter.
Structural & content operating disciplineBlyskal / Profound, operating-discipline campRun a continuous loop: structure content for extraction, publish, measure what got cited, revise. Indefinitely, not once.
Source-trust logic, engine by engineNeither camp alone — this is engine-specificTrack which engines actually cite you and why, since each one runs its own idiosyncratic trust model rather than a shared standard.
The floor
What the retrieval-mechanics camp controlsCrawl access, robots.txt and bot policy, licensing arrangements, feed availability, and the structured data that makes a page machine-readable. Fix these once. They don't reward daily attention the way content does.
The compounding lever
What the operating-discipline camp controlsContent structure, answer-first formatting, freshness, depth of comparison content, and the cadence of measuring and revising based on what actually gets cited. This is where a team keeps winning or keeps losing.
Machine access (fixed once, then largely stable)30%
Operating discipline (run continuously, compounds over time)70%

Something Inc.'s working model of where AI citation outcomes get decided — our internal framework, not a measured external study

That split also explains why platforms disagree so much about who's winning. Each engine runs its own retrieval and source-trust logic on top of whatever training data it inherited, which is exactly what we found across engines in why AI engines disagree on sources — the same domain can be a trusted source in one engine's answers and functionally invisible in another's, for reasons that have nothing to do with the quality of a given page and everything to do with how that engine assembled its candidate pool.

Running answer engine optimization as sequential work, not competing theories

The practical move is to stop asking which camp is correct and start sequencing the work the way both camps' research actually implies. Fix machine access first, because it's a one-time project with a clear finish line — crawl permissions, licensing where it applies, feed hygiene, structured data that removes ambiguity for a parser. Then run an operating loop indefinitely, because that's where the compounding advantage sits and where most competitors quietly stop investing after the initial project wraps.

HOW A QUESTION BECOMES A CITATION
Fix machine accesscrawl, licensing, feeds — once
Structure the contentanswer-first, extractable
Publish & instrumenttrack grounding, citation, mention separately
Evolve the looprevise based on what actually got cited

The instrumentation step matters more than most teams give it credit for. Citation, grounding, and mention are three separate outcomes that get flattened into one vague "AI visibility" number by most vendor dashboards, and our guide on measuring GEO grounding, citation, and mention as separate metrics walks through why collapsing them hides exactly the signal you need to know whether the operating loop is working or just spinning.

DO THIS NEXTRun the machine-access audit once — crawl permissions, licensing exposure, feed and structured-data hygiene — and close it out. Then stand up a real operating loop: structure, publish, measure citation and mention separately, revise, repeat. If you need both the audit and the ongoing loop built and run, that's the full scope of our generative engine optimization service. Don't keep re-running the audit instead of running the loop. That's the mistake both camps are, in their own way, warning you about.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.