On Mon, Jul 24, 2017 at 5:22 PM, Stuart A. Yeates <[email protected]> wrote:
> Sorry it's taken me so long to get back to this.
> https://pdfs.semanticscholar.org/dea9/142b39bdc2c3738e0f9cb7c6d117750ef2f7.pdf
> and https://meta.wikimedia.org/wiki/Beyond_categories are good places to
> start on the issues with cats on en.wiki.

very helpful. Thanks!

Leila

> cheers
> stuart
>
> --
> ...let us be heard from red core to black sky
>
> On 12 July 2017 at 02:53, Leila Zia <[email protected]> wrote:
>
>> Hi Stuart,
>>
>> On Mon, Jul 10, 2017 at 6:45 PM, Stuart A. Yeates <[email protected]>
>> wrote:
>> > The category system on en.wiki is not an IS-A system and there have been
>> > several discussions about making it it based on mathematical principals
>> > which have come to nothing because the consensus of editors is against
>> it.
>> > The best way to think about categories is as a locally-faceted related
>> > links system.
>>
>> It would be great if you can share a link to one or more of those
>> conversations, if it's not too hard to find them. This is a
>> conversation that comes up often and I'd like to educate myself with
>> this background. (and to confirm: on our end the goal is not to change
>> the category system on enwiki, but to make it machine understandable
>> for specific applications.)
>>
>> > Having said that, Category:Wikipedia maintenance is an important root
>> > probably useful for separating  the wheat from the chaff. Most of these
>> are
>> > also hidden categories. I'm not sure whether this flag appears in the
>> SQL,
>> > but see
>> > https://en.wikipedia.org/wiki/Wikipedia:Categorization#Hiding_categories
>>
>> Looking into these. thanks!
>>
>> Best,
>> Leila
>>
>> > cheers
>> > stuart
>> >
>> > --
>> > ...let us be heard from red core to black sky
>> >
>> > On 11 July 2017 at 13:20, Leila Zia <[email protected]> wrote:
>> >
>> >> Hi all,
>> >>
>> >> [If you are not interested in discussions related to the category system
>> >> (on English Wikipedia)
>> >> , you can stop here. :)]
>> >>
>> >> We have run into a problem that some of you may have thought about or
>> >> addressed before. We are trying to clean up the category system on
>> English
>> >> Wikipedia by turning the category structure to an IS-A hierarchy. (The
>> >> output of this work can be useful for the research on template
>> >> recommendation [1], for example, but the use-cases won't stop there).
>> One
>> >> issue that we are facing is the following:
>> >>
>> >> We are currently
>> >> using
>> >>  SQL dumps to extract categories associated with every article on
>> English
>> >> Wikipedia (main namespace). [2]
>> >> Using this approach, we get 5 categories associated with Flow cytometry
>> >> bioinformatics article [3]:
>> >>
>> >> Flow_cytometry
>> >> Bioinformatics
>> >>
>> >> Wikipedia_articles_published_in_peer-reviewed_literature
>> >> Wikipedia_articles_published_in_PLOS_Computational_Biology
>> >> CS1_maint:_Multiple_names:_authors_list
>> >>
>> >> The problem is that only the first two categories are the ones we are
>> >> interested in. We have one cleaning step through which we only keep
>> >> categories that belong to category Article and that step removes the
>> last
>> >> category above, but the other two Wikipedia_... remain there. We need to
>> >> somehow prune the data and clean it from those two categories.
>> >>
>> >> One way we could do the above would be to parse wikitext instead of the
>> SQL
>> >> dumps and focus on extracting categories marked by pattern
>> [[Category:XX]],
>> >> but in that case, we would lose a good category such as
>> >> Guided_missiles_of_Norway
>> >> because that's generated by a template.
>> >>
>> >> Any ideas on how we can start with a "cleaner" dataset of categories
>> >> related to the topic of the articles as opposed to maintenance related
>> or
>> >> other types of categories?
>> >>
>> >> Thanks,
>> >> Leila
>> >>
>> >> [1] https://meta.wikimedia.org/wiki/Research:Expanding_Wikipedia
>> >> _stubs_across_languages
>> >>
>> >> [2] The exact code we use is
>> >>
>> >> SELECT p.page_id id, p.page_title title, cl.cl_to category
>> >> FROM categorylinks cl
>> >> JOIN page p
>> >> on cl.cl_from = p.page_id
>> >> where cl_type = 'page'
>> >> and page_namespace = 0
>> >> and page_is_redirect = 0
>> >>
>> >> and the edges of the category graph are extracted with
>> >>
>> >> *SELECT p.page_title category, cl.cl_to parent *
>> >> *FROM categorylinks cl *
>> >> *JOIN page p *
>> >> *ON p.page_id = cl.cl_from *
>> >> *where p.page_namespace = 14*
>> >>
>> >>
>> >> [3] https://en.wikipedia.org/wiki/Flow_cytometry_bioinformatics
>> >> _______________________________________________
>> >> Wiki-research-l mailing list
>> >> [email protected]
>> >> https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
>> >>
>> > _______________________________________________
>> > Wiki-research-l mailing list
>> > [email protected]
>> > https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
>>
>> _______________________________________________
>> Wiki-research-l mailing list
>> [email protected]
>> https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
>>
> _______________________________________________
> Wiki-research-l mailing list
> [email protected]
> https://lists.wikimedia.org/mailman/listinfo/wiki-research-l

_______________________________________________
Wiki-research-l mailing list
[email protected]
https://lists.wikimedia.org/mailman/listinfo/wiki-research-l

Reply via email to