Hey Todd, James and Sage (referring to his recent post about adding video to interfaces)
Your comments highlight exatly where we actually very much are in concensus: Wikipedia is first and foremost curated knowledge by humans for humans (this is the core of my essay here: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed) I am not arguing that we should have AI writing, or that AI is the primary audience for our content, but rather whether we are entering an era of Zero Click Google, and if the big tech companies see through their vision (i.e. implementaiton of sloppy, adhoc commercially controlled interfaces on top of the open internet i.e.: https://blog.google/products-and-platforms/products/search/search-io-2026/ ) we need to think about this as a distribution channel not an invisible enemy. We cannot afford to miss this modality change in the same way we missed other big changes on the internet. such as social media and the TikTokified enshiftification towards endless scrolling, because we had a guaranteed distribution channel: Google. Some more thoughts (and realizing that this should be an op-ed/blogpost somewhere). *Interfaces are cheap, informed curators are expensive* Agentic AI, Coding tools and LLMs are making the cost of new interfaces *extremely cheap, *so cheap that I am commissioning a complex knowledge repository for a fraction of the cost and time it would take otherwise. Interface projects like WikiProject Med's offline medical Wikipedia App, which used to take several months of highly specialized software development, can now be spun up in a long weekend with a Claude Code Max subscription. Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing. What is not cheap for all of the other platforms (and is often paid for by ads), but which we have in abundance, is motivated humans who can continue curating the knowledge (and more than 25 years of experiments on facilitating knowledge equity focused gap filling). The questions we need to address are: - Do the curators understand the future of distribution to other humans we need to be building for across multiple future internet scenerios? - Do the curation practices serve diverse forms of access (languages, geographies, topics) from that distribution? - Do the curators understand demand and how curation choices affect distribution to that demand? - Can we recruit the next generation of curators who don't "assume" that the pageview metric is the reason we contribute? *We need to see how distribution to humans is changing * Our metrics infrastructure overemphasizes our historical abundance of pageviews—80% of which were driven by Google-- without seeing our distribution to these other platforms. We need to *see the distribution* in order to make any decisions about our curatorial practices. . Once we understand the distribution, we will also see that AI tools building interfaces need more than just access (i.e. Enterprise API or an open license and scraper access), they require knowledge formatting and organization practices (i.e. markdown files, vectorized search interfaces, AI skills, content chunking that makes the RAG search step easier, etc) that necessarily require the content curators to change some of their workflows and practices (emphasizing the authority of original authors in citations, metadata on kind of "human questions" a section describes, etc). Waiting for a handful of people at the Foundation to figure out which of these content organization tactics are important is just not feasible—we, as the curators, need to *as the curators* imagine this future and implement it in our content updates (especially when WMF's payroll relies on the nostalgic pageview-to-fundraising business model). *We need a Wikimedia specific strategy for the future, not copy our peers* Most websites are dividing the "content curation" from the representation layer (modern headless CMS's https://en.wikipedia.org/wiki/Headless_content_management_system are increasingly the go to for other publishing websites). Demanding that the Foundation (or Wikimedia projects) maintain our interface as the source of reader interactions is likely not sustainable, or consistent with the way in which knowledge curation now works on the internet. Our peers in textual content curation implemented radically different strategies 3-5 years ago that build on this seperation of curation and consumption: - New York Times doubled down on a "captive in an App" strategy because it was a pillar of their approach—a strategy many news organizations adopted as well. This approach is highly inappropriate for the Wikimedia model; we have abundant research showing that our users seek "public service utility" content from us, not timely or trustworthy content. - Britanica has shifted towards an education-market-first model that allows them to build interfaces appropriate to educational needs and garuntee a pipeline of funding. - Reddit optimized for AI tools answering constructive user question -- this also is not our goal, we are curators not "authorities" for answers. - Companies like Healthline optimized for SEO optimization, which gravitates for "generic assumptions of public search" instead of high quality verfiable content (I have found misinformation on healthline multiple times). My question is: What is the Wikimedia specific business model that allows our curated content to reach the humans we want to reach in this new distribution environment dictated by AI interfaces like RAG and slop-Apps ? On Sat, Jul 11, 2026 at 12:04 PM James Heilman via Wikimedia-l < [email protected]> wrote: > The blue "Return to article" button works just fine. But yes thanks for > pointing out that the back button within the Google chrome browser does not > work. Will work on fixing that. > > J > > On Sat, Jul 11, 2026 at 4:54 PM Todd Allen via Wikimedia-l < > [email protected]> wrote: > >> Are you talking about that godawful thing that I just clicked on in the >> "wheat" article, that the back button doesn't work on? >> >> I'm pulling that out. Sorry, but "back button works" is a basic thing of >> Web functionality. I should not need to click a "return to article" button >> to get back where I was. >> >> Todd >> >> On Sat, Jul 11, 2026 at 8:32 AM James Heilman via Wikimedia-l < >> [email protected]> wrote: >> >>> English Wikipedia is a conservative organization / project as are many >>> other large versions of Wikipedia. Many stakeholders need to be convinced / >>> brought on board to make even relatively minor changes. Innovating within >>> smaller versions of Wikipedia or in other projects outside Wikipedia is >>> much easier. And successes can occasionally be brought into the larger >>> Wikipedias such as we did with Our World in Data interactive graphs... >>> >>> https://en.wikipedia.org/wiki/Wheat#Production_and_consumption >>> >>> We have now added 100s of these in various languages. And they are >>> getting thousands of plays a day. >>> >>> James >>> >>> On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < >>> [email protected]> wrote: >>> >>>> Wikimedia-I used to be a premium listserve of high minded ideas about >>>> how to move the Foundation into the future. The past few months it has >>>> drifted into a lot of whining about AI with no real effort to accomplish >>>> anything or even plan to do so. >>>> >>>> - Charles >>>> >>>> On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < >>>> [email protected]> wrote: >>>> >>>>> We sure do know what a good AI strategy looks like for Wikipedia. No >>>>> AI on Wikipedia. >>>>> >>>>> We have not succeeded by being FaceGramTwitTube. We have succeeded by >>>>> not being like them. >>>>> >>>>> So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for >>>>> any reason, period. Wikipedia is written by people for people. >>>>> >>>>> Todd >>>>> >>>>> On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < >>>>> [email protected]> wrote: >>>>> >>>>>> What has motivated me to spend time writing Wikipedia over the years >>>>>> is writing for humans. The fact that the content is openly licensed and >>>>>> the >>>>>> work is supported by an NGO is also key. >>>>>> >>>>>> Personally i do not feel any motivation to write primarily for >>>>>> trillion dollar machines surrounded by venture capital folks hoping to >>>>>> make >>>>>> a killing. If the machines want to adapt to human facing content sure. >>>>>> >>>>>> The approaches you mention is how Healthline succeeded, they >>>>>> basically have dozens of articles covering the same topic just addressing >>>>>> it from a slightly different question. >>>>>> >>>>>> J >>>>>> >>>>>> >>>>>> Sent from Gmail Mobile >>>>>> >>>>>> On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < >>>>>> [email protected]> wrote: >>>>>> >>>>>>> Forking this conversation, because I don't think we have a shared >>>>>>> framing of what we are competing for in a "Google" Zero landscape >>>>>>> dominated >>>>>>> by ChatBot/AI search style RAG citations (i.e. >>>>>>> https://en.wikipedia.org/wiki/Retrieval-augmented_generation). >>>>>>> >>>>>>> For the last 9 months, I've been examining how civil society content >>>>>>> should be showing up in AI search, and there is a missing perspective in >>>>>>> our "we will build it and they will come" approach to Wikipedia. I don't >>>>>>> think we can wait for the big tech companies or European regulatory >>>>>>> bodies >>>>>>> to adopt a different idea of how RAG should work. Here is my take from >>>>>>> what I have been engaging with in the AIO/AEO/SEO space: >>>>>>> >>>>>>> *Wikipedia doesn't have content that is SEO/AEO optimized* >>>>>>> >>>>>>> Part of the problem, even if we did have an MCP server is that the >>>>>>> models (at least in my tracking), are pushing many citations away from >>>>>>> "factual" websites, towards "authoritative" websites. This authoritative >>>>>>> content includes: >>>>>>> >>>>>>> - Expert original, analysis that makes strong claims based on >>>>>>> facts (i.e. blog posts by authoritative companies or recently >>>>>>> published >>>>>>> ScienceDirect articles) >>>>>>> - Content that has been updated recently, with the biggest "hot >>>>>>> takes" (i.e. I have monitored a couple of prompt pools where >>>>>>> citations >>>>>>> shift to newer content after 2-3 months) >>>>>>> - Content that helps users make a decision between different >>>>>>> choices (i.e. review websites, etc) >>>>>>> >>>>>>> >>>>>>> This is following Google's longer-term push towards "human centered >>>>>>> and useful" content (sometimes called E-E-A-T an abbreviation of >>>>>>> experience, expertise, authoritativeness, and trustworthiness, in SEO >>>>>>> world). >>>>>>> https://developers.google.com/search/docs/fundamentals/creating-helpful-content >>>>>>> >>>>>>> >>>>>>> To win in an AI optimization battle -- its less about Wikipedia >>>>>>> doing well in the keyword search indexes that led to our content being >>>>>>> visible (which is why we have a reputation as a "fact checking" website) >>>>>>> and more about "winning" in the criteria for what makes a good RAG >>>>>>> citation >>>>>>> --- and our content format, is the exact opposite of the EAAT criteria: >>>>>>> >>>>>>> - Wikipedia is not authoriative, but rather points to other >>>>>>> authorities >>>>>>> - We ground our content in anonymity instead of named experts or >>>>>>> instutional process/opinoin >>>>>>> - We rarely do original analysis instead summarizing the >>>>>>> experience and expertise of others, >>>>>>> - Alot of our content is out of date, and self-aware of its gaps >>>>>>> (i.e. maintenance tags), so also is likely to be undermining its own >>>>>>> trustworthiness >>>>>>> >>>>>>> >>>>>>> *All the data points to us being used, but without an official >>>>>>> roundup we are all talking in the dark about different assumed >>>>>>> reputation >>>>>>> losses* >>>>>>> >>>>>>> RAG unlike Google Search Indexing, seems to be using Wikipedia for a >>>>>>> fraction of a fraction of responses, favoring these other kinds of >>>>>>> sources: >>>>>>> >>>>>>> - Only 5% of AI overviews have Wikipedia in them: >>>>>>> https://ahrefs.com/blog/most-cited-domains-ai-overviews/ >>>>>>> - I have access to SERanking's corpus of prompt monitoring >>>>>>> across 5 models (ChatGPT, Perplexity, Google Models and they suggest >>>>>>> that >>>>>>> in May ~16% of prompts included Wikipdia, and in their most recent >>>>>>> month >>>>>>> (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 >>>>>>> in >>>>>>> commercial AIO data -- so could have gaps. >>>>>>> - Studies from earlier in the year put Wikipedia at about ~13% >>>>>>> of CHATGPT citations ( >>>>>>> >>>>>>> https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over-25-of-chatgpt-citations-in-the-us-new-5w-research-finds--wsj-nyt-and-bloomberg-do-not-appear-in-the-top-20-302768339.html >>>>>>> but chatgpt on average includes >20 sources in a response, compared >>>>>>> to >>>>>>> googles 5-10 and doesn't expose it in the interface very well) >>>>>>> - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn >>>>>>> tend to represent a greater % of content (in the SERanking data pool >>>>>>> nearly >>>>>>> 30% of responses had a Youtube Video cited for instance) >>>>>>> - Domain specific citation pools have pretty significant >>>>>>> differences in "which" sources are being called, with Wikipedia >>>>>>> doing well >>>>>>> on some prompt pools: https://generativepulse.ai/report/ >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> *RAG/AI search optimization focuses more on intent than keywords, >>>>>>> and we aren't very effective at serving intent, and we don't know where >>>>>>> our >>>>>>> optimization options are* >>>>>>> What we need is an understanding of "which actual user reader >>>>>>> behavior are we seeking to serve?". In the past we were extremely lazy, >>>>>>> because keyword search always delivered Wikipedia as "a first". Now we >>>>>>> need >>>>>>> our content to be more optimized for the kind of user curiosity driving >>>>>>> their use of a chatbot/search tool: >>>>>>> >>>>>>> - What percentage of prompts or AI searches are informational vs >>>>>>> opinion forming? Are we even a competitor for grounding opinion based >>>>>>> questions or only the informational ones? >>>>>>> - How many of the interactions are two or three steps down a >>>>>>> chain of more "specific" interactions with the chatbot and thus no >>>>>>> longer >>>>>>> need "general knowledge" information from Wikipedia, but rather the >>>>>>> kinds >>>>>>> of stuff that we rely on our citations to provide ? >>>>>>> - How much are the AI companies optimizing for "sales" or >>>>>>> "addiction" rather than for leading users to reliable content? (I was >>>>>>> tracking a series of informational topics about food that (on >>>>>>> ChatGPT and >>>>>>> Google), kept wanting me to continue the conversation by *inviting >>>>>>> me to go to local hamburger restraunts)*. Do we even have a >>>>>>> reasonable chance to be in those searches? >>>>>>> - How much is geolocation forcing more and more responses into >>>>>>> "local" sources rather than "global" websites? In one dataset I >>>>>>> tracked, in >>>>>>> Global South countries citations were overwhelmingly to Facebook and >>>>>>> Instagram despite more authoritative academic, news and >>>>>>> Wikipedia-type >>>>>>> sites in the same searches from the UK. >>>>>>> >>>>>>> >>>>>>> *We may need to radically change the "readable signals" on our >>>>>>> content pages, meaning changing the Manual of Style, Editing Practices, >>>>>>> and >>>>>>> AI enabled enrichment.* >>>>>>> >>>>>>> If we are trying to market Wikipedia's content into AI interfaces, >>>>>>> we also can't do what most AI optimization/marketing agencies would >>>>>>> suggest: writing listicle/FAQ type content that closely matches the >>>>>>> user-queries that folks are giving the IA models (i.e. analysis like: >>>>>>> https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ >>>>>>> ). >>>>>>> >>>>>>> We would then have to experiment with other content types, that _no >>>>>>> longer look like the encyclopedia_. Or we would need to be reconfiguring >>>>>>> the Encyclopedic content to expose enrichments to paragraphs or sections >>>>>>> within the encyclopedia that pretty radically change editorial >>>>>>> assupmtions >>>>>>> and our Manual of Style (i.e. instead of simple 1-2 word section >>>>>>> headings, >>>>>>> like "History" we may need intent-focused headings like "What is the >>>>>>> history of [x topic]?). >>>>>>> >>>>>>> If we want to compete in the shifting AI search landscape -- we >>>>>>> would need a lot more data from the Foundation on where we are >>>>>>> succeeding >>>>>>> or not, and then consider *_radically different_ *ways of exposing >>>>>>> our content in terms of treating RAG systems as a user that needs >>>>>>> correct >>>>>>> paths to Wikipedia pages. >>>>>>> >>>>>>> However, this doesn't necessarily need to change the *human reader >>>>>>> experience *, but would need to be about configuring the content >>>>>>> (beyond an MCP server or Enterpise APIs) *for an AI >>>>>>> audience/consumer experience -- *which I haven't seen addressed in >>>>>>> any WMF publications or community conversations. Without a firm theory >>>>>>> of >>>>>>> "What kind of consumer is an AI search agent/RAG index?" and "How does >>>>>>> our >>>>>>> content need to serve that AI audience?" the editing community won't be >>>>>>> able to adjust its editing practices or weigh in on feature >>>>>>> recommendations >>>>>>> that make our content "AI useful". >>>>>>> >>>>>>> As I have written elsewhere, I think there is a inherent audience >>>>>>> for editing/using the Wikis organically: >>>>>>> https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed >>>>>>> -- but its a different question than competing "with other information >>>>>>> sources" for AI as an audience. >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>> On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < >>>>>>> [email protected]> wrote: >>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < >>>>>>>> [email protected]> wrote: >>>>>>>> >>>>>>>>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l >>>>>>>>> <[email protected]> wrote: >>>>>>>>> >>>>>>>>> > Yah a search engine that actually gives real references that >>>>>>>>> supports the statements in question would be amazing. >>>>>>>>> >>>>>>>>> Almost like .. a Knowledge Engine. ;-) >>>>>>>>> >>>>>>>> >>>>>>>> Is the WMF building an MCP server to connect Wikipedia and Wikidata >>>>>>>> directly to Gemini, Claude, and ChatGPT? This is a more lightweight, >>>>>>>> backdoor way to leverage the audience of those platforms but present >>>>>>>> structured outputs to AI chats for users based on Wikimedia knowledge. >>>>>>>> If >>>>>>>> we did so, we could present citations within the returned responses to >>>>>>>> users, and those platforms make it transparent to the user when they >>>>>>>> are >>>>>>>> calling a particular tool. >>>>>>>> >>>>>>>> Steven Walling >>>>>>>> >>>>>>>> Sadly, the only realistic path I see there would be through >>>>>>>>> acquisition, and even if that was financially feasible, you'd >>>>>>>>> begin by >>>>>>>>> inheriting a lot of corporate practices that aren't really >>>>>>>>> consistent >>>>>>>>> with Wikimedia values. >>>>>>>>> >>>>>>>>> But perhaps there is a middle ground where Wikimedia seeks to >>>>>>>>> define >>>>>>>>> more clearly the terms of engagement that it wants with search >>>>>>>>> engines >>>>>>>>> (clear attribution, clear and correct references, calls-to-edit, >>>>>>>>> etc.), and then finds and recognizes search partners who implement >>>>>>>>> those. To Luis' point, that need not be done by WMF. >>>>>>>>> >>>>>>>>> Warmly, >>>>>>>>> >>>>>>>>> Erik >>>>>>>>> _______________________________________________ >>>>>>>>> Wikimedia-l mailing list -- [email protected], >>>>>>>>> guidelines at: >>>>>>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>>>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>>>>>> Public archives at >>>>>>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/HNDPDZBBDMILF7WGMUBVJVIAZYQ7OXOS/ >>>>>>>>> To unsubscribe send an email to >>>>>>>>> [email protected] >>>>>>>> >>>>>>>> _______________________________________________ >>>>>>>> Wikimedia-l mailing list -- [email protected], >>>>>>>> guidelines at: >>>>>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>>>>> Public archives at >>>>>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/IUDAKJRN5EQT5CCWEEYQXKDSVXCDBXR6/ >>>>>>>> To unsubscribe send an email to >>>>>>>> [email protected] >>>>>>> >>>>>>> _______________________________________________ >>>>>>> Wikimedia-l mailing list -- [email protected], >>>>>>> guidelines at: >>>>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>>>> Public archives at >>>>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/YODBKTCLA2XJC24Y23I2SUN3PUYA4QXE/ >>>>>>> To unsubscribe send an email to >>>>>>> [email protected] >>>>>> >>>>>> _______________________________________________ >>>>>> Wikimedia-l mailing list -- [email protected], >>>>>> guidelines at: >>>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>>> Public archives at >>>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/5MPMZGZ2MXLLHER3NTNUS6KHIMAXIIBD/ >>>>>> To unsubscribe send an email to [email protected] >>>>> >>>>> _______________________________________________ >>>>> Wikimedia-l mailing list -- [email protected], >>>>> guidelines at: >>>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>>> Public archives at >>>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/QAGHRXXTGSXKVFOALBRQHP5G27NUIFDD/ >>>>> To unsubscribe send an email to [email protected] >>>> >>>> _______________________________________________ >>>> Wikimedia-l mailing list -- [email protected], >>>> guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines >>>> and https://meta.wikimedia.org/wiki/Wikimedia-l >>>> Public archives at >>>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/6RKSJMWMVUHUUVV7WGKPOCOHCN4NBQK4/ >>>> To unsubscribe send an email to [email protected] >>> >>> >>> >>> -- >>> James Heilman >>> MD, CCFP-EM, Wikipedian >>> _______________________________________________ >>> Wikimedia-l mailing list -- [email protected], guidelines >>> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/FRGCURSOK3KPMT3UWZFOR6Z4TMZFE7KF/ >>> To unsubscribe send an email to [email protected] >> >> _______________________________________________ >> Wikimedia-l mailing list -- [email protected], guidelines >> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/XILTKQ2VLVFZKTY5BEYP55426JVAU5VY/ >> To unsubscribe send an email to [email protected] > > > > -- > James Heilman > MD, CCFP-EM, Wikipedian > _______________________________________________ > Wikimedia-l mailing list -- [email protected], guidelines > at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/FUYVW2DYO4VKDTCOBIRAICU5KKIHXI3U/ > To unsubscribe send an email to [email protected]
_______________________________________________ Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/6M3FINLF4MVXB4O5YZWMV4ICBUTHQHXK/ To unsubscribe send an email to [email protected]
