| Micru added a comment. |
There is some criticism about using a fixed set of codes for the representation of languages and variants:
https://en.wikipedia.org/wiki/ISO_639-3#Criticism
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.824.7083&rep=rep1&type=pdf
The ISO standard that was supposed to encode variants, ISO 639-6:2009, was withdrawn in 2014 with no planned replacement.
However, there is the need to administer the range of options to avoid partisan edits. The current model used by the monolingual datatype works well in that case because it omits language variants. Wiktionaries sometimes mention the language variant on the pronunciation or on the definition. For instance in en-wikt tomaca only lists two pronunciations, however ca-wikt tomaca lists two pronunciations plus the dialects where the word is used (català central, and valencià).
A possible solution could be to have a mandatory field for language from a fixed list administered by the community, and an optional field for variants encoded as items. Constrain checks could be performed by bots to ensure that the selected variants are in the subclass tree of a given language.
Another option is to represent the language variant as a statement or qualifier. While more flexible, it adds complexity.
Cc: Nikki, Lydia_Pintscher, WMDE-leszek, thiemowmde, Denny, Micru, Aklapper, Lexicographical data, daniel, D3r1ck01, MuhammadShuaib, Izno, Psychoslave, Wikidata-bugs, aude, Gryllida, Shizhao, Arrbee, Mbch331, Jay8g
_______________________________________________ Wikidata-bugs mailing list [email protected] https://lists.wikimedia.org/mailman/listinfo/wikidata-bugs
