[ 
https://issues.apache.org/jira/browse/LUCENE-8462?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16589157#comment-16589157
 ] 

Ryadh Dahimene commented on LUCENE-8462:
----------------------------------------

I've just pushed a new `TestSnowballVocabData.zip` file that includes a lighter 
dataset (1165 words) generated from this input file:
[https://github.com/ibnmalik/golden-corpus-arabic/blob/develop/core/words.txt]

The source corpus "Golden Arabic Corpus" is licensed under a BSD-2 license: 
[https://github.com/ibnmalik/golden-corpus-arabic/blob/develop/LICENSE]

I've also included the original BSD-2 LICENSE file in the zip archive under the 
`arabic/` folder.

 

 

 

 

> New Arabic snowball stemmer
> ---------------------------
>
>                 Key: LUCENE-8462
>                 URL: https://issues.apache.org/jira/browse/LUCENE-8462
>             Project: Lucene - Core
>          Issue Type: Improvement
>            Reporter: Ryadh Dahimene
>            Priority: Trivial
>              Labels: Arabic, snowball, stemmer
>          Time Spent: 10m
>  Remaining Estimate: 0h
>
> Added a new Arabic snowball stemmer based on 
> [https://github.com/snowballstem/snowball/blob/master/algorithms/arabic.sbl]
> As well an Arabic test dataset in `TestSnowballVocabData.zip` from the 
> snowball-data available here 
> [https://github.com/snowballstem/snowball-data/tree/master/arabic]
> Link to the corresponding Github PR:
> [https://github.com/apache/lucene-solr/pull/439]
>  
>  



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to