This piece was prompted by a LinkedIn post from Ivana Drakulevska, who rounded up three genuinely significant developments in SEO and AI search from the past week and proposed a nine-stage framework for diagnosing AI visibility. Credit for surfacing these three stories together, and for the pipeline model discussed below, belongs to her post — we verified each claim independently and expanded on the context, methodology and practical application below.

Three things happened in AI search over the past two weeks that, taken individually, look like routine industry news. Taken together, they describe a real shift in how content actually gets found, cited and acted on — and why the standard SEO advice being handed out right now (add FAQs, add schema, use bullet points) increasingly misses where visibility is actually breaking down.

ChatGPT is now, legally, a search engine in the European Union

On August 31, 2026, the European Commission formally designated ChatGPT a Very Large Online Search Engine (VLOSE) under the Digital Services Act, placing it alongside Reddit and Roblox, which were designated Very Large Online Platforms in the same announcement. The number behind the designation is specific: services cross this threshold once they report at least 45 million average monthly users in the EU, and ChatGPT came in well past it, at roughly 159.1 million average monthly EU users.

The mechanism matters more than the headline. The designation isn’t based on ChatGPT being an AI chatbot in some general sense — it’s specifically tied to ChatGPT’s live web-search function, the capability that lets it retrieve and cite current information from the open web rather than answering purely from training data. That’s a capability-based test, not a category-based one, and regulators now have a template to apply the identical test to Gemini, Claude and Perplexity as each of those products’ web-search features scale to similar usage levels in the EU.

The practical stakes are real: OpenAI has until the end of December 2026 to comply with VLOSE obligations covering minor-safety protections, algorithmic transparency and independent risk audits, with non-compliance risking fines of up to 6% of global annual turnover — the same regulatory regime that produced a €120 million fine against X in December 2025 sets the enforcement baseline here. For anyone whose traffic or citations increasingly come through ChatGPT rather than traditional search, this is the first time a regulator has formally agreed, in a binding legal sense, that the comparison to “search engine” isn’t just a loose industry metaphor.

A 24-hour bug just proved that citations and retrieval are two separate things

Google shipped Gemini 3.8 Flash inside AI Mode on September 2, 2026. Within days, reports surfaced that AI Mode responses — specifically for paid Google AI Pro and Ultra subscribers who had manually selected the 3.8 Flash model — were coming back without citations or source links for a meaningful share of queries, particularly top-of-funnel searches. Robby Stein, Google Search’s VP of Product, confirmed on September 4 that the behavior wasn’t intentional and that a fix was coming; citations had largely reappeared within about 24 hours of the first reports.

The bug itself was narrow and short-lived. What it exposed is not: AI Mode was still retrieving and synthesizing content from the sites it was drawing on during the bug window — it simply wasn’t showing where that content came from. That’s a distinction that matters enormously for anyone trying to measure their own AI visibility. A site’s content can be actively influencing an AI answer — shaping the synthesis, contributing to the response a user actually reads and acts on — while showing zero visible citations, zero referral traffic, and zero evidence in any citation-tracking tool that it was used at all. Retrieval and citation are two different pipeline stages, and this bug is the cleanest public proof yet that they can fail, or succeed, independently of each other.

SearchGPT and Google are pulling from almost entirely different sources

The third data point is the one most directly useful for anyone doing this work day to day. Independent research comparing SearchGPT’s citations against Google’s results found only about 24-25% overlap between SearchGPT’s top-100 cited domains and both Google’s organic search results and Google’s AI Overviews — meaning roughly three out of every four sources SearchGPT cites don’t appear in the equivalent Google results at all. A separate controlled study found SearchGPT’s citation patterns align far more closely with Bing (87% match) than with Google (56% match), which tracks with what’s publicly known about the underlying retrieval infrastructure several AI products lean on.

The source-type pattern is just as telling as the overlap number. SearchGPT leans heavily on reference and news sources — Wikipedia shows up in roughly half of its citations, with Britannica, Reuters and AP News also well represented. Google’s organic results and AI Overviews, by contrast, pull much more heavily from user-generated platforms: Reddit alone appears in 42% of organic results and 7% of AI Overviews, with Quora and Facebook also showing up at meaningfully higher rates than in SearchGPT’s citations. In practice, this means “ranking well in Google” and “getting cited by ChatGPT’s search feature” are measuring exposure to genuinely different retrieval systems with different source preferences, not two views of the same underlying visibility. Repeated identical prompts against the same AI system can even produce materially different citation sets run to run, adding a layer of measurement noise that traditional rank tracking never had to deal with.

How Talmyn verified this, and why that process matters here specifically

Given how much of this piece rests on numbers — user thresholds, fine percentages, overlap rates — it’s worth being explicit about how each claim was checked before it went into this article, rather than asking readers to simply trust a LinkedIn summary or, for that matter, trust us.

For the EU designation, we cross-checked the announcement against the European Commission’s own Digital Services Act newsroom coverage and independent reporting from outlets that specialize in EU tech regulation, confirming the August 31, 2026 date, the 45-million-user threshold, ChatGPT’s reported 159.1 million EU average monthly users, the December 2026 compliance deadline, and the 6%-of-global-turnover fine ceiling, along with the €120 million X precedent cited as the enforcement baseline. For the Gemini 3.8 Flash citation bug, we confirmed it against reporting that directly quoted Google Search VP Robby Stein’s own public acknowledgment that the behavior wasn’t intentional, rather than relying on unconfirmed third-party speculation about what caused it — a real, sourced admission from the company involved is a meaningfully stronger form of confirmation than an SEO commentator’s interpretation of unusual AI Mode output. For the SearchGPT/Google overlap figures, we checked the headline 24-25% overlap number against more than one independent study rather than a single source, since citation-overlap research is exactly the kind of number that varies by methodology (query set, sample size, time window) and is easy to overstate if only one study’s figure is repeated as settled fact.

That last point is worth sitting with: this entire article is, in a sense, a demonstration of its own argument. A LinkedIn post is not a primary source, however accurate it turns out to be — it’s a pointer to real developments that still need independent verification before they’re safe to build a strategy on. The same discipline applies to AI citations themselves. A model citing a source doesn’t make that source correct, complete, or even representative of a consensus; it means the model’s retrieval and selection process favored that source for that specific query, at that specific moment, under whatever ranking logic that system currently applies. Verifying a claim and being cited by an AI system are not the same act, and conflating them is one of the quieter ways bad information compounds in an AI-mediated information environment.

The framework worth borrowing: nine stages, not one metric

Drakulevska’s proposed response to all three of these developments is a nine-stage pipeline for thinking about AI visibility, rather than treating “did we get cited” as a single pass/fail metric: Accessibility, Discovery, Retrieval, Selection, Synthesis, Entity Representation, Citation, Recommendation, and User Action. The framing that matters most here is the core claim underneath the model: citation is an outcome, not a strategy. A site can fail at any one of those nine stages and produce the exact same visible symptom — no citation, no AI-driven traffic — for completely different underlying reasons. A page an AI system can’t technically access (Accessibility) looks identical from the outside to a page it can access and index, but never selects as authoritative enough to cite (Selection), or a page it does draw from but simply doesn’t attribute due to a bug like the one Gemini 3.8 Flash just had (Citation).

That distinction is exactly why generic tactical advice — add more FAQ schema, use more bullet points, write a clearer meta description — increasingly doesn’t move the needle on its own. Those tactics mostly target the Accessibility and Discovery stages, which for most established, technically sound sites were never actually the problem. If the real bottleneck for a given site sits further down the pipeline — at Selection, where an AI system has access to the content but doesn’t judge it authoritative enough to draw from, or at Entity Representation, where the site or its authors aren’t well-established as a recognizable entity in the systems these models actually reference — no amount of schema markup fixes that. Diagnosing which stage is actually failing, rather than defaulting to the same handful of surface-level fixes regardless of the actual cause, is the more useful discipline these three stories point toward.

Where the nine-stage model is genuinely useful, and where it’s harder than it looks

It’s worth stress-testing the framework rather than just adopting it wholesale, because it has a real, practical limitation: several of its nine stages aren’t independently observable from outside the AI system doing the retrieving. A site owner can reasonably check Accessibility (can a crawler reach and render the page at all) and, with effort, Discovery (is the page appearing in the systems that feed these models’ retrieval indexes). Retrieval, Selection and Synthesis, by contrast, happen entirely inside a closed model’s inference process — there’s no dashboard that shows a site “failed at Selection” the way Google Search Console shows a page has low impressions. In practice, most teams will only ever be able to infer which stage is failing indirectly, by testing the same prompts repeatedly across different AI systems and reasoning backward from the pattern of results — which is genuinely useful, but slower and less precise than the clean nine-stage diagram makes it look.

The other honest caveat: because different AI systems draw from different retrieval infrastructure and different source-type preferences — the SearchGPT-versus-Google overlap data above makes that concrete — a site can be doing everything right at every one of the nine stages for one AI system and still be invisible to another, for reasons that have nothing to do with the site’s own quality. A reference-heavy, Wikipedia-adjacent authority signal that helps with SearchGPT’s citation patterns won’t necessarily help with a system that leans more on user-generated platforms like Reddit and Quora. There is not, at least not yet, a single unified way to be “visible to AI search” — there are several separate systems with different preferences, and the nine-stage model has to be run separately, and honestly, against each one a site actually cares about.

What this actually means for agencies and site owners, stage by stage

Turning the framework from a diagnostic model into something actionable means treating each stage as its own, separately testable question, rather than a single “are we AI-visible” checkbox.

Accessibility and Discovery are the two stages every technical SEO audit already covers, and they’re still worth confirming rather than assuming: can the AI system’s crawler actually reach the page (check robots.txt and any bot-blocking rules against the specific crawler user-agents these AI products use, not just Googlebot), and is the page actually appearing in whatever index or retrieval layer feeds that system. If a site is confident these two stages are solid, that’s exactly the signal to stop spending further effort here and look further down the pipeline.

Retrieval and Selection are best tested empirically rather than theoretically: run the same set of real, representative prompts a target customer might actually type, across each AI system that matters to the business, repeated more than once given the measurement variability noted above, and track which competitors or sources get cited instead. If a site is consistently retrieved-but-not-selected against a specific competitor, the gap is almost always about depth, specificity or demonstrated expertise on that exact query — not technical accessibility.

Entity Representation is the stage most agencies currently underinvest in relative to its apparent importance: does the site, its authors and its organization exist as a recognizable, well-sourced entity across the reference material these models actually draw on (structured data, a real Wikipedia or Wikidata presence where warranted, consistent author identity across a body of published work, genuine third-party citations of the site elsewhere) rather than existing only as a collection of individually optimized pages with no larger identity tying them together.

Citation, Recommendation and User Action are the stages closest to the eventual business outcome, and also the hardest to fully control, since they depend on decisions happening inside the model at inference time — which is exactly why the Gemini 3.8 Flash bug is such a useful cautionary example. A drop in visible citations for a site with no fault of its own is a real, recurring possibility, not a hypothetical, and it argues for tracking actual referral behavior and brand mentions in AI answers over time rather than reacting to any single day’s citation count as if it were a stable, permanent signal.

The throughline across all of this: the sites and agencies that adapt fastest here won’t be the ones chasing the newest single tactic, but the ones who build a real, repeated process for testing where their own visibility actually breaks down, system by system, and keep re-running that test as these products keep changing underneath them — because as the Gemini bug and the EU designation both show in different ways, the systems themselves are still actively being built, regulated and fixed in real time.