[TL;DR: lovely things exist; we can build more; keen to see some of you next week.]
While preparing for the Wikimania preconference on Wiki AI <https://meta.wikimedia.org/wiki/Artificial_intelligence/2026_Wiki_AI> (ht to Alex Ostrovsky; come join us!) it struck me how completely we've stopped talking about AI's steady evolution as practical knowledge infrastructure. Recent discussions have focused on AI as a challenge for the open web, as norms around knowledge-seeking and scraping change. But we have hardly taken any time at all to savor and apply improvements that could advance our mission, since 2023 (the last time we updated our sweet in-house machine learning model cards <https://meta.wikimedia.org/wiki/Machine_learning_models>). Let us return to building <https://www.youtube.com/watch?v=_hk6KLD-0tg&t=309s> wiki-quality AI, accessible, transparent, and accountable; neutral, proportionate, and impactful. We need a family of wiki-principled language models <https://diff.wikimedia.org/2026/07/11/wikimedia-principled-llms/> (to borrow Jan Ainali's phrase): with curated open data, open source code, and open weights. Available to readers and contributors, hosted in our own green data centers, and integrated into Wikimedia projects (as ORES was), so details of model use can be easily linked from relevant edit summaries.* We can get much of what we want with better data sourcing, translation, and post-training. These are some of the more accessible and inexpensive parts of the model-building stack. And we are not alone in wanting this: each of the last <https://en.wikipedia.org/wiki/AI_Action_Summit_2025> two <https://en.wikipedia.org/wiki/India_AI_Impact_Summit_2026> AI Action Summits has centered the need to build AI in the public interest, including advancing human knowledge and protecting & expanding the commons. In that spirit, I invite you all to channel some of the energy around AI debates and forecasts towards a few evergreen goals: 1) Write about your vision for the future you want. We need better stories about the future. Publish it somewhere and link to it from m:AI#Talks_and_Readings <https://meta.wikimedia.org/wiki/Artificial_intelligence#Talks_and_Readings>. Submit it to a <https://futurevisionxprize.com/> contest <https://protopianprize.com/>. I linked Jan's blog post above as an example... describe things you want even if you don't yet know how to build them: a benchmark, a tool, a coalition, a campus. Our communities do well with the charismatic megafauna of collective visions: we are willing to try things that don't scale, and are always prepared <https://en.wikiquote.org/wiki/Umberto_Eco> to rewrite our environments. 2) Consider that someone knocking on your library window a billion times a week asking for an article or an entire back-catalog is not a burden, it's a relationship. We have a responsibility toward AI systems that rely on Wikipedia, and also an opportunity to shape how they learn from and interact with the commons. In an ideal world (where CDN issues are solved), how should we best serve agents and other AI tools who turn up looking for grounding or guidance or bulk materials? How should we influence the default skills and settings of agents to advance the commons? 3) There are global non-profit networks working towards transparent green public AI <https://en.wikipedia.org/wiki/Colorless_green_ideas_sleep_furiously>, including Humanity AI, Current AI, and Public AI.** Let's find ways to inform and benefit from aligned efforts, and present a unified front in proposing changes to current systems and designing new ones.*** Wikimedians should define the components we need within our own ecosystem, from benchmarks and datasets to curation norms and tools. Then we can explore post-training processes for neutral, proportional, epistemically humble langauge models. All watched over by wikis of loving grace, SJ * Even if, on some projects, model use is one step removed from mainspace edits (as with the AIlogbot <https://en.wikipedia.org/wiki/User:Fermiboson/AIlog>). ** These include the collaborative *AIPotluck <https://www.aipotluck.org/> *project, an effort to build a fully transparent AI stack (whose current biggest gap <https://www.aipotluck.org/map> is open data); and the *Public AI network <http://publicai.network>*, a coalition of model builders and advocates that DWeb <https://blog.archive.org/2022/02/15/the-decentralized-web-an-introduction/> friends and I started a few years ago, working towards public options at every layer of the stack. *** I wrote about things we might design first, for the Next25 discussion Christophe organized (on meta <https://meta.wikimedia.org/wiki/User:Sj/Design_chats/AI>). Briefly: [*] We should give editors access to the most useful <https://www.tomshardware.com/software/linux/linus-torvalds-rebukes-anti-ai-stances-in-the-linux-kernel-code-review-process-says-linux-is-not-one-of-those-anti-ai-projects-creator-embraces-ai-as-just-a-tool-and-clearly-a-useful-one> open tools, including AI: including for classification, language support, and citation checking, [⁑] We should improve existing public models to produce a family of wiki AI models that serve our characteristic needs [⁂] We need shared playgrounds for experimenting with AI systems 🧸 [✣] We need a bigger wiki to experiment with new forms of automation... including via AI and agentic systems. This requires new tools for high-volume review and is not suitable for testing on production wikis. [*] As bandwidth becomes an issue again: We should integrate better with repositories like Zenodo, Github, 🤗, and Openverse. Then we don't have to replicate their storage infrastructure just to have convenient and editable metadata. == Références == 0. Wikimania preconference day on Wiki AI (still time to sign up today): https://meta.wikimedia.org/wiki/Artificial_intelligence/2026_Wiki_AI 1. Machine learning model cards: https://meta.wikimedia.org/wiki/Machine_learning_models 2. Building Language Models for Wikimedia (Bob West, 2024) - https://www.youtube.com/watch?v=_hk6KLD-0tg&t=309s 3. Wikimedia-principled LLMs - https://diff.wikimedia.org/2026/07/11/wikimedia-principled-llms/ 4. AI Action Summit 2025 - https://en.wikipedia.org/wiki/AI_Action_Summit_2025 5. AI Impact Summit 2026 - https://en.wikipedia.org/wiki/India_AI_Impact_Summit_2026 6. AI Talks and Readings (meta) - https://meta.wikimedia.org/wiki/Artificial_intelligence#Talks_and_Readings 7a. https://futurevisionxprize.com/ 7b. Protopian Fiction Prize for Public AI - https://protopianprize.com/ 8. "the cultivated person's first duty is to be always prepared to rewrite the encyclopaedia" – Eco https://en.wikiquote.org/wiki/Umberto_Eco 9. "transparent green ideas sleep furiously" (Bohr) - https://en.wikipedia.org/wiki/Colorless_green_ideas_sleep_furiously 10. AI Potluck (Current AI foundation) - https://www.aipotluck.org/ 11. A gap map for open source AI - https://www.aipotluck.org/map 12. http://publicai.network/ 13. https://blog.archive.org/2022/02/15/the-decentralized-web-an-introduction/ 14. Building Wiki AI - https://meta.wikimedia.org/wiki/User:Sj/Design_chats/AI 15. https://www.phoronix.com/news/Linux-Is-Not-Anti-AI 16. https://en.wikipedia.org/wiki/All_Watched_Over_by_Machines_of_Loving_Grace_(TV_series) -- Samuel Klein @metasj w:user:sj
_______________________________________________ Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/W6RUCFY4RBIVHF73OEIVGS223TPTHAMX/ To unsubscribe send an email to [email protected]
