On 8/10/26 7:32 PM, Theodore Tso wrote:
On Mon, Aug 10, 2026 at 05:35:33AM -0500, Pirate Praveen wrote:
1. Public AI is AI as public infrastructure like highways, water, or
electricity. An example of Public AI is https://publicai.co...

Actually, publicai.cois not only an "example" but the *name* of the
open-source project hosted at https://publicai.co.

As such, your trying to use Public AI has a term of art, when its
pre-existing use is as the name of an open-source project which you
are trying to use as an example is likely to cause... confusion.


My idea was to base this definition on how Public AI actually works to a more easily understood general criteria. I have now made that clear in v3.0 of the draft,

"So to evaluate an AI web service, we will use the definition of Open Source AI by OSI, but some requirements may be relaxed."

...For AI to be accessible to people, it has to be
available easily to people and not for profit collaborations between
multiple organizations and public funding is essential to keep it
sustainable.

3. At this point, we don't want to insist on training data being freely
available, similar to how we don't insist on Free Hardware Designs.

So you *could* use the Open Source AI definition from the OSI
institute, but it's not an overlap with what you are going for, since
(a) the OSAID doesn't require that the LLM be publically funded, which
you stated in your point (1) --- but since it's not clear whether is
the key part of your definition of "Public AI", you might not consider
this to be a problem, and (b) the OSAID has a much stricter set of
requirements that would disallow Open Weight models that would
otherwise be allowable.

I have clarified that point in v3.0. Public funding may be desirable, but not a requirement.

"For AI to be accessible to people, it has to be available easily to people, like Wikipedia or Lets Encrypt. Not for profit collaborations between multiple organizations and public funding is essential to keep it sustainable, though public funding is not a requirement. How we are going to fund such Public AI will evolve in the future."

So for example, some of the most commonly used Open Weight models such
as Qwen 3 would run afoul of the OSAID requirements.  And there are
very few models indeed aside from the Apertus that would meet your
publicly funded requirement.  An example of a model that would meet
the OSAID requirement might be the T5 model from the 2019 Google
Paper, "Exploring the Limits of Transfer Learning with a Unified
Text-to-Text Transforme", but which wasn't publically funded.  It also
isn't actively being developed anymore, since it was an LLM that was
created just to explore some research ideas.  But I'm not sure it
really matteres if an academic LLM was funded by a company or by a
grant from a government agency.

2. Public AI must use Free Software models.

If we strip out your publically funded model, then you could just
simply use the word "Open Weights", although that's not exactly
formally specified as a definition either.  So for example, for
Debian, do we care that the LLaMa has a restriction on commercial use?
(After all, it's not at all applicable to how Debian would use the
LLM.)  What about the restriction in the LLaMa license which forbids
using the model outputs to train competing LLM's? etc.

I have now used Open Weights as well.

In contrast, we might have the Gemma 4 model, which is released under
the Apache license, which would allow its use to create derivitive
LLM's, but (oh, horrors!) it was funded and released to the world by
Google.  (Again, I'm not sure why having an LLM funded by government
agency makes it any difference from an LLM funded by a company.  Ask
someone who used to receive an NSF grant how capricious and subject to
change government grant funding can be.)

I have clarified this as well.

This is why having a formal specification and definition of what is
allowed is important, and why I don't think your proposal is quite
ready for prime time.  If it were passed, the amount of confusion and
debate caused by the fact that the definition is not crisp and precise
would be harmful to the Debian project, IMHO.

Does "Open Source AI definition with requirement for availability of training data relaxed" removes ambiguities?

I have also mentioned about preserving copyleft - though I'm not fully sure if we need that as how copyright law evolves is not really in our control and I could remove that requirement if people feel that should be removed as well.

Best regards,

                                                - Ted

Attachment: OpenPGP_0x8F53E0193B294B75.asc
Description: OpenPGP public key

Attachment: OpenPGP_signature.asc
Description: OpenPGP digital signature

Reply via email to