Forking this conversation, because I don't think we have a shared framing
of what we are competing for in a "Google" Zero landscape dominated by
ChatBot/AI search style RAG citations (i.e.
https://en.wikipedia.org/wiki/Retrieval-augmented_generation).

For the last 9 months, I've been examining how civil society content should
be showing up in AI search, and there is a missing perspective in our "we
will build it and they will come" approach to Wikipedia. I don't think we
can wait for the big tech companies or European regulatory bodies to adopt
a different idea of how RAG should work.   Here is my take from what I have
been engaging with in the AIO/AEO/SEO space:

*Wikipedia doesn't have content that is SEO/AEO optimized*

Part of the problem, even if we did have an MCP server is that the
models (at least in my tracking), are pushing many citations away from
"factual" websites, towards "authoritative" websites. This authoritative
content includes:

   - Expert original, analysis that makes strong claims based on facts
   (i.e. blog posts by authoritative companies or recently published
   ScienceDirect articles)
   - Content that has been updated recently, with the biggest "hot takes"
   (i.e. I have monitored a couple of prompt pools where citations shift to
   newer content after 2-3 months)
   - Content that helps users make a decision between different choices
   (i.e. review websites, etc)


This is following Google's longer-term push towards "human centered and
useful" content (sometimes called  E-E-A-T an abbreviation of  experience,
expertise, authoritativeness, and trustworthiness, in SEO world).
https://developers.google.com/search/docs/fundamentals/creating-helpful-content


To win in an AI optimization battle -- its less about Wikipedia doing well
in the keyword search indexes that led to our content being visible (which
is why we have a reputation as a "fact checking" website) and more about
"winning" in the criteria for what makes a good RAG citation --- and our
content format, is the exact opposite of the EAAT criteria:

   - Wikipedia is not authoriative, but rather points to other authorities
   - We ground our content in anonymity instead of named experts or
   instutional process/opinoin
   - We rarely do original analysis instead summarizing the experience and
   expertise of others,
   - Alot of our content is out of date, and self-aware of its gaps (i.e.
   maintenance tags), so also is likely to be undermining its own
   trustworthiness


*All the data points to us being used, but without an official roundup we
are all talking in the dark about different assumed reputation losses*

RAG unlike Google Search Indexing, seems to be using Wikipedia for a
fraction of a fraction of responses, favoring these other kinds of
sources:

   - Only 5% of AI overviews have Wikipedia in them:
   https://ahrefs.com/blog/most-cited-domains-ai-overviews/
   - I have access to SERanking's corpus of prompt monitoring across 5
   models (ChatGPT, Perplexity, Google Models and they suggest that in May
   ~16% of prompts included Wikipdia, and in their most recent month (June),
   ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial
   AIO data -- so could have gaps.
   - Studies from earlier in the year put Wikipedia at about ~13% of
   CHATGPT citations (
   
https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over-25-of-chatgpt-citations-in-the-us-new-5w-research-finds--wsj-nyt-and-bloomberg-do-not-appear-in-the-top-20-302768339.html
   but chatgpt on average includes >20 sources in a response, compared to
   googles 5-10 and doesn't expose it in the interface very well)
   - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend to
   represent a greater % of content (in the SERanking data pool nearly 30% of
   responses had a Youtube Video cited for instance)
   - Domain specific citation pools have pretty significant differences in
   "which" sources are being called, with Wikipedia doing well on some prompt
   pools: https://generativepulse.ai/report/




*RAG/AI search optimization focuses more on intent than keywords, and we
aren't very effective at serving intent, and we don't know where our
optimization options are*
What we need is an understanding of "which actual user reader behavior are
we seeking to serve?". In the past we were extremely lazy, because keyword
search always delivered Wikipedia as "a first". Now we need our content to
be more optimized for the kind of user curiosity driving their use of a
chatbot/search tool:

   - What percentage of prompts or AI searches are informational vs opinion
   forming? Are we even a competitor for grounding opinion based questions or
   only the informational ones?
   - How many of the interactions are two or three steps down a chain of
   more "specific" interactions with the chatbot and thus no longer need
   "general knowledge" information from Wikipedia, but rather the kinds of
   stuff that we rely on our citations to provide ?
   - How much are the AI companies optimizing for "sales" or "addiction"
   rather than for leading users to reliable content? (I was tracking a series
   of informational topics about food that (on ChatGPT and Google), kept
   wanting me to continue the conversation by *inviting me to go to local
   hamburger restraunts)*. Do we even have a reasonable chance to be in
   those searches?
   - How much is geolocation forcing more and more responses into "local"
   sources rather than "global" websites? In one dataset I tracked, in Global
   South countries citations were overwhelmingly to Facebook and Instagram
   despite more authoritative academic, news and Wikipedia-type sites in the
   same searches from the UK.


*We may need to radically change the "readable signals" on our content
pages, meaning changing the Manual of Style, Editing Practices, and AI
enabled enrichment.*

If we are trying to market Wikipedia's content into AI interfaces, we also
can't do what most AI optimization/marketing agencies would suggest:
writing listicle/FAQ type content that closely matches the user-queries
that folks are giving the IA models (i.e. analysis like:
https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/
).

We would then have to experiment with other content types, that _no longer
look like the encyclopedia_. Or we would need to be reconfiguring the
Encyclopedic content to expose enrichments to paragraphs or sections within
the encyclopedia  that pretty radically change editorial assupmtions and
our Manual of Style (i.e. instead of simple 1-2 word section headings, like
"History" we may need intent-focused headings like "What is the history of
[x topic]?).

If we want to compete in the shifting AI search landscape -- we would need
a lot more data from the Foundation on where we are succeeding or not, and
then consider *_radically different_ *ways of exposing our content in terms
of treating RAG systems as a user that needs correct paths to Wikipedia
pages.

However, this doesn't necessarily need to change the *human reader
experience *, but would need to be about configuring the content (beyond an
MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which
I haven't seen addressed in any WMF publications or community
conversations. Without a firm theory of "What kind of consumer is an AI
search agent/RAG index?" and "How does our content need to serve that AI
audience?" the editing community won't be able to adjust its editing
practices or weigh in on feature recommendations that make our content "AI
useful".

As I have written elsewhere, I think there is a inherent audience for
editing/using the Wikis organically:
https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed
-- but its a different question than competing "with other information
sources" for AI as an audience.





On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l <
[email protected]> wrote:

>
>
> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l <
> [email protected]> wrote:
>
>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l
>> <[email protected]> wrote:
>>
>> > Yah a search engine that actually gives real references that supports
>> the statements in question would be amazing.
>>
>> Almost like .. a Knowledge Engine. ;-)
>>
>
> Is the WMF building an MCP server to connect Wikipedia and Wikidata
> directly to Gemini, Claude, and ChatGPT? This is a more lightweight,
> backdoor way to leverage the audience of those platforms but present
> structured outputs to AI chats for users based on Wikimedia knowledge. If
> we did so, we could present citations within the returned responses to
> users, and those platforms make it transparent to the user when they are
> calling a particular tool.
>
> Steven Walling
>
> Sadly, the only realistic path I see there would be through
>> acquisition, and even if that was financially feasible, you'd begin by
>> inheriting a lot of corporate practices that aren't really consistent
>> with Wikimedia values.
>>
>> But perhaps there is a middle ground where Wikimedia seeks to define
>> more clearly the terms of engagement that it wants with search engines
>> (clear attribution, clear and correct references, calls-to-edit,
>> etc.), and then finds and recognizes search partners who implement
>> those. To Luis' point, that need not be done by WMF.
>>
>> Warmly,
>>
>> Erik
>> _______________________________________________
>> Wikimedia-l mailing list -- [email protected], guidelines
>> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
>> https://meta.wikimedia.org/wiki/Wikimedia-l
>> Public archives at
>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/HNDPDZBBDMILF7WGMUBVJVIAZYQ7OXOS/
>> To unsubscribe send an email to [email protected]
>
> _______________________________________________
> Wikimedia-l mailing list -- [email protected], guidelines
> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
> https://meta.wikimedia.org/wiki/Wikimedia-l
> Public archives at
> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/IUDAKJRN5EQT5CCWEEYQXKDSVXCDBXR6/
> To unsubscribe send an email to [email protected]
_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: 
https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and 
https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at 
https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/YODBKTCLA2XJC24Y23I2SUN3PUYA4QXE/
To unsubscribe send an email to [email protected]

Reply via email to