[ 
https://issues.apache.org/jira/browse/LUCENE-9929?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17322005#comment-17322005
 ] 

Adrien Grand commented on LUCENE-9929:
--------------------------------------

+1 to avoid introducing options. I'd rather like to keep a no-option 
ScandinavianXXXFilter or split it into filters per language if we want to have 
different defaults for Norwegian, Swedish and Danish, and ask users who have 
more specific needs to write their own token filter or use something like 
MappingCharFilter instead?

> Make ScandinavianNormalizationFilter configurable wrt foldings
> --------------------------------------------------------------
>
>                 Key: LUCENE-9929
>                 URL: https://issues.apache.org/jira/browse/LUCENE-9929
>             Project: Lucene - Core
>          Issue Type: Improvement
>          Components: modules/analysis
>            Reporter: Jan Høydahl
>            Assignee: Jan Høydahl
>            Priority: Major
>          Time Spent: 40m
>  Remaining Estimate: 0h
>
> The ScandinavianNormalizationFilter applies foldings for aa, ao, ae, oe and 
> oo. But all those five do not make sense for both Norwegian, Swedish and 
> Danish. Implement an optional configuration option where users can select 
> which of them to apply. I.e. for Norwegian, a user would then configure (in 
> Solr):
> {code:java}
> <filter class="solr.ScandinavianNormalizationFilterFactory 
> foldings="ae,oe,aa"/>
> {code}
> This would activate foldings for ae->æ, oe->ø, aa->å, but not oo->o and ao->a.
> The default will be to activate all five as before, so it will be backward 
> compatible.



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to