Hi Andrea,
Since the facets are already modified in order to have
lower case values, followed by a separator, followed by their
real value, ex.:
"aboriginal
rights\n|||\naboriginal rights",1,
"abricot\n|||\nabricot",3,
"access to
care\n|||\naccess to care",1,
"access to
healthcare\n|||\naccess to healthcare",1,
"accessibilité aux
soins\n|||\naccessibilité aux soins",1,
"accountability\n|||\naccountability",1,
"accès aux soins de
santé\n|||\naccès aux soins de santé",1,
"actes de
colloque\n|||\nactes de colloque",1,
"activité
motrice\n|||\nActivité motrice",1,
We have altered the first form (that is before the
separator), in order to also have the accented characters
replaced by their unaccented counterpart.
We have added the following dependency in
[SOURCES]/dspace-api/pom.xml:
<dependency>
<groupId>commons-lang</groupId>
<artifactId>commons-lang3</artifactId>
<version>3.4</version>
</dependency>
And we have changed all occurrences of
OrderFormat
within
the [SOURCES]/dspace-api/src/main/java/org/dspace/browse/SolrBrowseCreateDAO.java
for
org.apache.commons.lang3.StringUtils.stripAccents(OrderFormat
(...)).
And within
[SOURCES]/dspace-api/src/main/java/org/dspace/discovery/SolrServiceImpl.java
where values are set to lower case (a few occurrences), we
have added the stripAccents, which will give something similar
to:
org.apache.commons.lang3.StringUtils.stripAccents(value.toLowerCase())
org.apache.commons.lang3.StringUtils.stripAccents(facetValue.toLowerCase())
org.apache.commons.lang3.StringUtils.stripAccents(indexValue.toLowerCase())
etc....
We did reindex and now we have:
"aboriginal
rights\n|||\naboriginal rights",1,
"abricot\n|||\nabricot",3,
"acces aux soins de
sante\n|||\naccès aux soins de santé",1,
"access to
care\n|||\naccess to care",1,
"access to
healthcare\n|||\naccess to healthcare",1,
"accessibilite aux
soins\n|||\naccessibilité aux soins",1,
"accountability\n|||\naccountability",1,
"actes de
colloque\n|||\nactes de colloque",1,
"activite
motrice\n|||\nActivité motrice",1,
|
Economic, social and cultural rights [1]
Economical Europe [1]
Économie [1]
Économie et droit de l'homme [1]
Economie sanction [1]
Economy [1]
ecosystem approach [1]
Ecrit [1]
Écrit [2]
écrit [1]
(...)
|
|
|
|
Is that what you are looking for? I can send you
our modified files if you wish.
regards,
Marie-Hélène Vézina
Université de Montréal
|
|
|
|
|
|
|
Le mercredi 24 février 2016 23:18:49 UTC-5, Andrea Schweer
a écrit :
Hi
all,
has anyone had any luck getting the subject browse/facets to
ignore
diacritics when sorting the subject terms? Right now (5.x,
XMLUI,
discovery), we get a sort order like this:
manager
Manuka Leaf Oil
Manukau
Manukau Harbour
marginal groups
marginal value theorem
Maze procedure
Mānuka Honey
in browse and
manager
Manuka Leaf Oil
Manukau
Manukau Harbour
marginal groups
marginal value theorem
Maze procedure
mythology
Mānuka Honey
in the facet listing.
I've been asked whether it's possible to change the sort order
such that
"Mānuka" gets sorted between "manager" and "Manukau" -- that
is, as if
there was no diacritic. The subject terms should remain listed
with
their actual values though (it's not an option to simply strip
out the
diacritics during indexing).
I have vague memories that we managed to tweak this behaviour
by
configuring the collation locale in PostgreSQL, back in the
database-backed browse. Obviously that doesn't work with
Discovery.
I tried a few changes to the Solr schema for the discovery
core, but
none appeared to do what I want. Based on the UnicodeCollation
page in
the Solr wiki [1], it looks like configuring collation for a
sort field
might be the way to go; it isn't possible to set up collation
for the
subject field/s because then the stored value changes to a
binary
representation of the metadata value that includes the
collation key. I
imagine this would work for title browse, where the results
(=DSpace
items) are sorted by one of the configured sort fields.
However, the
subject browse listings (where the results are metadata
values)
ultimately come from a facet query, and facet.sort only takes
"count" or
"index" as its values [2] -- it doesn't look like it's
possible to add
collation into the mix there.
Does anyone have any thoughts on this?
cheers,
Andrea
[1]: https://wiki.apache.org/solr/UnicodeCollation
[2]: https://wiki.apache.org/solr/SimpleFacetParameters#facet.sort
--
Dr Andrea Schweer
Lead Software Developer, ITS Information Systems
The University of Waikato, Hamilton, New Zealand
+64-7-837 9120
--