Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < [email protected]> wrote: > We sure do know what a good AI strategy looks like for Wikipedia. No AI on > Wikipedia. > > We have not succeeded by being FaceGramTwitTube. We have succeeded by not > being like them. > > So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any > reason, period. Wikipedia is written by people for people. > > Todd > > On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < > [email protected]> wrote: > >> What has motivated me to spend time writing Wikipedia over the years is >> writing for humans. The fact that the content is openly licensed and the >> work is supported by an NGO is also key. >> >> Personally i do not feel any motivation to write primarily for trillion >> dollar machines surrounded by venture capital folks hoping to make a >> killing. If the machines want to adapt to human facing content sure. >> >> The approaches you mention is how Healthline succeeded, they basically >> have dozens of articles covering the same topic just addressing it from a >> slightly different question. >> >> J >> >> >> Sent from Gmail Mobile >> >> On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < >> [email protected]> wrote: >> >>> Forking this conversation, because I don't think we have a shared >>> framing of what we are competing for in a "Google" Zero landscape dominated >>> by ChatBot/AI search style RAG citations (i.e. >>> https://en.wikipedia.org/wiki/Retrieval-augmented_generation). >>> >>> For the last 9 months, I've been examining how civil society content >>> should be showing up in AI search, and there is a missing perspective in >>> our "we will build it and they will come" approach to Wikipedia. I don't >>> think we can wait for the big tech companies or European regulatory bodies >>> to adopt a different idea of how RAG should work. Here is my take from >>> what I have been engaging with in the AIO/AEO/SEO space: >>> >>> *Wikipedia doesn't have content that is SEO/AEO optimized* >>> >>> Part of the problem, even if we did have an MCP server is that the >>> models (at least in my tracking), are pushing many citations away from >>> "factual" websites, towards "authoritative" websites. This authoritative >>> content includes: >>> >>> - Expert original, analysis that makes strong claims based on facts >>> (i.e. blog posts by authoritative companies or recently published >>> ScienceDirect articles) >>> - Content that has been updated recently, with the biggest "hot >>> takes" (i.e. I have monitored a couple of prompt pools where citations >>> shift to newer content after 2-3 months) >>> - Content that helps users make a decision between different choices >>> (i.e. review websites, etc) >>> >>> >>> This is following Google's longer-term push towards "human centered and >>> useful" content (sometimes called E-E-A-T an abbreviation of experience, >>> expertise, authoritativeness, and trustworthiness, in SEO world). >>> https://developers.google.com/search/docs/fundamentals/creating-helpful-content >>> >>> >>> To win in an AI optimization battle -- its less about Wikipedia doing >>> well in the keyword search indexes that led to our content being visible >>> (which is why we have a reputation as a "fact checking" website) and more >>> about "winning" in the criteria for what makes a good RAG citation --- and >>> our content format, is the exact opposite of the EAAT criteria: >>> >>> - Wikipedia is not authoriative, but rather points to other >>> authorities >>> - We ground our content in anonymity instead of named experts or >>> instutional process/opinoin >>> - We rarely do original analysis instead summarizing the experience >>> and expertise of others, >>> - Alot of our content is out of date, and self-aware of its gaps >>> (i.e. maintenance tags), so also is likely to be undermining its own >>> trustworthiness >>> >>> >>> *All the data points to us being used, but without an official roundup >>> we are all talking in the dark about different assumed reputation losses* >>> >>> RAG unlike Google Search Indexing, seems to be using Wikipedia for a >>> fraction of a fraction of responses, favoring these other kinds of >>> sources: >>> >>> - Only 5% of AI overviews have Wikipedia in them: >>> https://ahrefs.com/blog/most-cited-domains-ai-overviews/ >>> - I have access to SERanking's corpus of prompt monitoring across 5 >>> models (ChatGPT, Perplexity, Google Models and they suggest that in May >>> ~16% of prompts included Wikipdia, and in their most recent month (June), >>> ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial >>> AIO data -- so could have gaps. >>> - Studies from earlier in the year put Wikipedia at about ~13% of >>> CHATGPT citations ( >>> >>> https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over-25-of-chatgpt-citations-in-the-us-new-5w-research-finds--wsj-nyt-and-bloomberg-do-not-appear-in-the-top-20-302768339.html >>> but chatgpt on average includes >20 sources in a response, compared to >>> googles 5-10 and doesn't expose it in the interface very well) >>> - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend >>> to represent a greater % of content (in the SERanking data pool nearly >>> 30% >>> of responses had a Youtube Video cited for instance) >>> - Domain specific citation pools have pretty significant differences >>> in "which" sources are being called, with Wikipedia doing well on some >>> prompt pools: https://generativepulse.ai/report/ >>> >>> >>> >>> >>> *RAG/AI search optimization focuses more on intent than keywords, and we >>> aren't very effective at serving intent, and we don't know where our >>> optimization options are* >>> What we need is an understanding of "which actual user reader behavior >>> are we seeking to serve?". In the past we were extremely lazy, because >>> keyword search always delivered Wikipedia as "a first". Now we need our >>> content to be more optimized for the kind of user curiosity driving their >>> use of a chatbot/search tool: >>> >>> - What percentage of prompts or AI searches are informational vs >>> opinion forming? Are we even a competitor for grounding opinion based >>> questions or only the informational ones? >>> - How many of the interactions are two or three steps down a chain >>> of more "specific" interactions with the chatbot and thus no longer need >>> "general knowledge" information from Wikipedia, but rather the kinds of >>> stuff that we rely on our citations to provide ? >>> - How much are the AI companies optimizing for "sales" or >>> "addiction" rather than for leading users to reliable content? (I was >>> tracking a series of informational topics about food that (on ChatGPT and >>> Google), kept wanting me to continue the conversation by *inviting >>> me to go to local hamburger restraunts)*. Do we even have a >>> reasonable chance to be in those searches? >>> - How much is geolocation forcing more and more responses into >>> "local" sources rather than "global" websites? In one dataset I tracked, >>> in >>> Global South countries citations were overwhelmingly to Facebook and >>> Instagram despite more authoritative academic, news and Wikipedia-type >>> sites in the same searches from the UK. >>> >>> >>> *We may need to radically change the "readable signals" on our content >>> pages, meaning changing the Manual of Style, Editing Practices, and AI >>> enabled enrichment.* >>> >>> If we are trying to market Wikipedia's content into AI interfaces, we >>> also can't do what most AI optimization/marketing agencies would suggest: >>> writing listicle/FAQ type content that closely matches the user-queries >>> that folks are giving the IA models (i.e. analysis like: >>> https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ >>> ). >>> >>> We would then have to experiment with other content types, that _no >>> longer look like the encyclopedia_. Or we would need to be reconfiguring >>> the Encyclopedic content to expose enrichments to paragraphs or sections >>> within the encyclopedia that pretty radically change editorial assupmtions >>> and our Manual of Style (i.e. instead of simple 1-2 word section headings, >>> like "History" we may need intent-focused headings like "What is the >>> history of [x topic]?). >>> >>> If we want to compete in the shifting AI search landscape -- we would >>> need a lot more data from the Foundation on where we are succeeding or not, >>> and then consider *_radically different_ *ways of exposing our content >>> in terms of treating RAG systems as a user that needs correct paths to >>> Wikipedia pages. >>> >>> However, this doesn't necessarily need to change the *human reader >>> experience *, but would need to be about configuring the content >>> (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer >>> experience -- *which I haven't seen addressed in any WMF publications >>> or community conversations. Without a firm theory of "What kind of consumer >>> is an AI search agent/RAG index?" and "How does our content need to serve >>> that AI audience?" the editing community won't be able to adjust its >>> editing practices or weigh in on feature recommendations that make our >>> content "AI useful". >>> >>> As I have written elsewhere, I think there is a inherent audience for >>> editing/using the Wikis organically: >>> https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed >>> -- but its a different question than competing "with other information >>> sources" for AI as an audience. >>> >>> >>> >>> >>> >>> On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < >>> [email protected]> wrote: >>> >>>> >>>> >>>> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < >>>> [email protected]> wrote: >>>> >>>>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l >>>>> <[email protected]> wrote: >>>>> >>>>> > Yah a search engine that actually gives real references that >>>>> supports the statements in question would be amazing. >>>>> >>>>> Almost like .. a Knowledge Engine. ;-) >>>>> >>>> >>>> Is the WMF building an MCP server to connect Wikipedia and Wikidata >>>> directly to Gemini, Claude, and ChatGPT? This is a more lightweight, >>>> backdoor way to leverage the audience of those platforms but present >>>> structured outputs to AI chats for users based on Wikimedia knowledge. If >>>> we did so, we could present citations within the returned responses to >>>> users, and those platforms make it transparent to the user when they are >>>> calling a particular tool. >>>> >>>> Steven Walling >>>> >>>> Sadly, the only realistic path I see there would be through >>>>> acquisition, and even if that was financially feasible, you'd begin by >>>>> inheriting a lot of corporate practices that aren't really consistent >>>>> with Wikimedia values. >>>>> >>>>> But perhaps there is a middle ground where Wikimedia seeks to define >>>>> more clearly the terms of engagement that it wants with search engines >>>>> (clear attribution, clear and correct references, calls-to-edit, >>>>> etc.), and then finds and recognizes search partners who implement >>>>> those. To Luis' point, that need not be done by WMF. >>>>> >>>>> Warmly, >>>>> >>>>> Erik >>>>> _______________________________________________ >>>>> Wikimedia-l mailing list -- [email protected], >>>>> guidelines at: >>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>> Public archives at >>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/HNDPDZBBDMILF7WGMUBVJVIAZYQ7OXOS/ >>>>> To unsubscribe send an email to [email protected] >>>> >>>> _______________________________________________ >>>> Wikimedia-l mailing list -- [email protected], >>>> guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines >>>> and https://meta.wikimedia.org/wiki/Wikimedia-l >>>> Public archives at >>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/IUDAKJRN5EQT5CCWEEYQXKDSVXCDBXR6/ >>>> To unsubscribe send an email to [email protected] >>> >>> _______________________________________________ >>> Wikimedia-l mailing list -- [email protected], guidelines >>> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/YODBKTCLA2XJC24Y23I2SUN3PUYA4QXE/ >>> To unsubscribe send an email to [email protected] >> >> _______________________________________________ >> Wikimedia-l mailing list -- [email protected], guidelines >> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/5MPMZGZ2MXLLHER3NTNUS6KHIMAXIIBD/ >> To unsubscribe send an email to [email protected] > > _______________________________________________ > Wikimedia-l mailing list -- [email protected], guidelines > at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/QAGHRXXTGSXKVFOALBRQHP5G27NUIFDD/ > To unsubscribe send an email to [email protected]
_______________________________________________ Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/6RKSJMWMVUHUUVV7WGKPOCOHCN4NBQK4/ To unsubscribe send an email to [email protected]
