We sure do know what a good AI strategy looks like for Wikipedia. No AI on
Wikipedia.

We have not succeeded by being FaceGramTwitTube. We have succeeded by not
being like them.

So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any
reason, period. Wikipedia is written by people for people.

Todd

On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l <
[email protected]> wrote:

> What has motivated me to spend time writing Wikipedia over the years is
> writing for humans. The fact that the content is openly licensed and the
> work is supported by an NGO is also key.
>
> Personally i do not feel any motivation to write primarily for trillion
> dollar machines surrounded by venture capital folks hoping to make a
> killing. If the machines want to adapt to human facing content sure.
>
> The approaches you mention is how Healthline succeeded, they basically
> have dozens of articles covering the same topic just addressing it from a
> slightly different question.
>
> J
>
>
> Sent from Gmail Mobile
>
> On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l <
> [email protected]> wrote:
>
>> Forking this conversation, because I don't think we have a shared framing
>> of what we are competing for in a "Google" Zero landscape dominated by
>> ChatBot/AI search style RAG citations (i.e.
>> https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
>>
>> For the last 9 months, I've been examining how civil society content
>> should be showing up in AI search, and there is a missing perspective in
>> our "we will build it and they will come" approach to Wikipedia. I don't
>> think we can wait for the big tech companies or European regulatory bodies
>> to adopt a different idea of how RAG should work.   Here is my take from
>> what I have been engaging with in the AIO/AEO/SEO space:
>>
>> *Wikipedia doesn't have content that is SEO/AEO optimized*
>>
>> Part of the problem, even if we did have an MCP server is that the
>> models (at least in my tracking), are pushing many citations away from
>> "factual" websites, towards "authoritative" websites. This authoritative
>> content includes:
>>
>>    - Expert original, analysis that makes strong claims based on facts
>>    (i.e. blog posts by authoritative companies or recently published
>>    ScienceDirect articles)
>>    - Content that has been updated recently, with the biggest "hot
>>    takes" (i.e. I have monitored a couple of prompt pools where citations
>>    shift to newer content after 2-3 months)
>>    - Content that helps users make a decision between different choices
>>    (i.e. review websites, etc)
>>
>>
>> This is following Google's longer-term push towards "human centered and
>> useful" content (sometimes called  E-E-A-T an abbreviation of  experience,
>> expertise, authoritativeness, and trustworthiness, in SEO world).
>> https://developers.google.com/search/docs/fundamentals/creating-helpful-content
>>
>>
>> To win in an AI optimization battle -- its less about Wikipedia doing
>> well in the keyword search indexes that led to our content being visible
>> (which is why we have a reputation as a "fact checking" website) and more
>> about "winning" in the criteria for what makes a good RAG citation --- and
>> our content format, is the exact opposite of the EAAT criteria:
>>
>>    - Wikipedia is not authoriative, but rather points to other
>>    authorities
>>    - We ground our content in anonymity instead of named experts or
>>    instutional process/opinoin
>>    - We rarely do original analysis instead summarizing the experience
>>    and expertise of others,
>>    - Alot of our content is out of date, and self-aware of its gaps
>>    (i.e. maintenance tags), so also is likely to be undermining its own
>>    trustworthiness
>>
>>
>> *All the data points to us being used, but without an official roundup we
>> are all talking in the dark about different assumed reputation losses*
>>
>> RAG unlike Google Search Indexing, seems to be using Wikipedia for a
>> fraction of a fraction of responses, favoring these other kinds of
>> sources:
>>
>>    - Only 5% of AI overviews have Wikipedia in them:
>>    https://ahrefs.com/blog/most-cited-domains-ai-overviews/
>>    - I have access to SERanking's corpus of prompt monitoring across 5
>>    models (ChatGPT, Perplexity, Google Models and they suggest that in May
>>    ~16% of prompts included Wikipdia, and in their most recent month (June),
>>    ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial
>>    AIO data -- so could have gaps.
>>    - Studies from earlier in the year put Wikipedia at about ~13% of
>>    CHATGPT citations (
>>    
>> https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over-25-of-chatgpt-citations-in-the-us-new-5w-research-finds--wsj-nyt-and-bloomberg-do-not-appear-in-the-top-20-302768339.html
>>    but chatgpt on average includes >20 sources in a response, compared to
>>    googles 5-10 and doesn't expose it in the interface very well)
>>    - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend
>>    to represent a greater % of content (in the SERanking data pool nearly 30%
>>    of responses had a Youtube Video cited for instance)
>>    - Domain specific citation pools have pretty significant differences
>>    in "which" sources are being called, with Wikipedia doing well on some
>>    prompt pools: https://generativepulse.ai/report/
>>
>>
>>
>>
>> *RAG/AI search optimization focuses more on intent than keywords, and we
>> aren't very effective at serving intent, and we don't know where our
>> optimization options are*
>> What we need is an understanding of "which actual user reader behavior
>> are we seeking to serve?". In the past we were extremely lazy, because
>> keyword search always delivered Wikipedia as "a first". Now we need our
>> content to be more optimized for the kind of user curiosity driving their
>> use of a chatbot/search tool:
>>
>>    - What percentage of prompts or AI searches are informational vs
>>    opinion forming? Are we even a competitor for grounding opinion based
>>    questions or only the informational ones?
>>    - How many of the interactions are two or three steps down a chain of
>>    more "specific" interactions with the chatbot and thus no longer need
>>    "general knowledge" information from Wikipedia, but rather the kinds of
>>    stuff that we rely on our citations to provide ?
>>    - How much are the AI companies optimizing for "sales" or "addiction"
>>    rather than for leading users to reliable content? (I was tracking a 
>> series
>>    of informational topics about food that (on ChatGPT and Google), kept
>>    wanting me to continue the conversation by *inviting me to go to
>>    local hamburger restraunts)*. Do we even have a reasonable chance to
>>    be in those searches?
>>    - How much is geolocation forcing more and more responses into
>>    "local" sources rather than "global" websites? In one dataset I tracked, 
>> in
>>    Global South countries citations were overwhelmingly to Facebook and
>>    Instagram despite more authoritative academic, news and Wikipedia-type
>>    sites in the same searches from the UK.
>>
>>
>> *We may need to radically change the "readable signals" on our content
>> pages, meaning changing the Manual of Style, Editing Practices, and AI
>> enabled enrichment.*
>>
>> If we are trying to market Wikipedia's content into AI interfaces, we
>> also can't do what most AI optimization/marketing agencies would suggest:
>> writing listicle/FAQ type content that closely matches the user-queries
>> that folks are giving the IA models (i.e. analysis like:
>> https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/
>> ).
>>
>> We would then have to experiment with other content types, that _no
>> longer look like the encyclopedia_. Or we would need to be reconfiguring
>> the Encyclopedic content to expose enrichments to paragraphs or sections
>> within the encyclopedia  that pretty radically change editorial assupmtions
>> and our Manual of Style (i.e. instead of simple 1-2 word section headings,
>> like "History" we may need intent-focused headings like "What is the
>> history of [x topic]?).
>>
>> If we want to compete in the shifting AI search landscape -- we would
>> need a lot more data from the Foundation on where we are succeeding or not,
>> and then consider *_radically different_ *ways of exposing our content
>> in terms of treating RAG systems as a user that needs correct paths to
>> Wikipedia pages.
>>
>> However, this doesn't necessarily need to change the *human reader
>> experience *, but would need to be about configuring the content (beyond
>> an MCP server or Enterpise APIs) *for an AI audience/consumer experience
>> -- *which I haven't seen addressed in any WMF publications or community
>> conversations. Without a firm theory of "What kind of consumer is an AI
>> search agent/RAG index?" and "How does our content need to serve that AI
>> audience?" the editing community won't be able to adjust its editing
>> practices or weigh in on feature recommendations that make our content "AI
>> useful".
>>
>> As I have written elsewhere, I think there is a inherent audience for
>> editing/using the Wikis organically:
>> https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed
>> -- but its a different question than competing "with other information
>> sources" for AI as an audience.
>>
>>
>>
>>
>>
>> On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l <
>> [email protected]> wrote:
>>
>>>
>>>
>>> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l <
>>> [email protected]> wrote:
>>>
>>>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l
>>>> <[email protected]> wrote:
>>>>
>>>> > Yah a search engine that actually gives real references that supports
>>>> the statements in question would be amazing.
>>>>
>>>> Almost like .. a Knowledge Engine. ;-)
>>>>
>>>
>>> Is the WMF building an MCP server to connect Wikipedia and Wikidata
>>> directly to Gemini, Claude, and ChatGPT? This is a more lightweight,
>>> backdoor way to leverage the audience of those platforms but present
>>> structured outputs to AI chats for users based on Wikimedia knowledge. If
>>> we did so, we could present citations within the returned responses to
>>> users, and those platforms make it transparent to the user when they are
>>> calling a particular tool.
>>>
>>> Steven Walling
>>>
>>> Sadly, the only realistic path I see there would be through
>>>> acquisition, and even if that was financially feasible, you'd begin by
>>>> inheriting a lot of corporate practices that aren't really consistent
>>>> with Wikimedia values.
>>>>
>>>> But perhaps there is a middle ground where Wikimedia seeks to define
>>>> more clearly the terms of engagement that it wants with search engines
>>>> (clear attribution, clear and correct references, calls-to-edit,
>>>> etc.), and then finds and recognizes search partners who implement
>>>> those. To Luis' point, that need not be done by WMF.
>>>>
>>>> Warmly,
>>>>
>>>> Erik
>>>> _______________________________________________
>>>> Wikimedia-l mailing list -- [email protected],
>>>> guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines
>>>> and https://meta.wikimedia.org/wiki/Wikimedia-l
>>>> Public archives at
>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/HNDPDZBBDMILF7WGMUBVJVIAZYQ7OXOS/
>>>> To unsubscribe send an email to [email protected]
>>>
>>> _______________________________________________
>>> Wikimedia-l mailing list -- [email protected], guidelines
>>> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
>>> https://meta.wikimedia.org/wiki/Wikimedia-l
>>> Public archives at
>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/IUDAKJRN5EQT5CCWEEYQXKDSVXCDBXR6/
>>> To unsubscribe send an email to [email protected]
>>
>> _______________________________________________
>> Wikimedia-l mailing list -- [email protected], guidelines
>> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
>> https://meta.wikimedia.org/wiki/Wikimedia-l
>> Public archives at
>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/YODBKTCLA2XJC24Y23I2SUN3PUYA4QXE/
>> To unsubscribe send an email to [email protected]
>
> _______________________________________________
> Wikimedia-l mailing list -- [email protected], guidelines
> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
> https://meta.wikimedia.org/wiki/Wikimedia-l
> Public archives at
> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/5MPMZGZ2MXLLHER3NTNUS6KHIMAXIIBD/
> To unsubscribe send an email to [email protected]
_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: 
https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and 
https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at 
https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/QAGHRXXTGSXKVFOALBRQHP5G27NUIFDD/
To unsubscribe send an email to [email protected]

Reply via email to