Hi Tomas,
Did you mean filtering out \p{Z} (Separator) and \p{C} (Other)?
http://perldoc.perl.org/perlunicode.html#Unicode-Character-Properties
I think such filtering would be problematic. Here are a few examples:
. The RTL marker U+200F is in the C category - this marker is
essential for the processing of bidirectional languages like Arabic or
Hebrew; bonus xkcd joke: http://xkcd.com/1137/
. Private use characters that might be used by some people are in
the C category
. The space character U+0020 is in the Z category
(a good tool to look up properties is http://rishida.net/scripts/uniview -
kudos to Richard Ishida to make this available)
So I think even if the removal of some character categories will make work
for some languages easier, it might break the use for some other languages.
In my opinion the training and decoding tool chain in Moses should not make
any assumptions about characters other than it processes tokens that are
separated by whitespace (separators \p{Z}). If you get too much into which
character classes are processed in which way, things can get really
complicated with different Unicode versions, libraries.
Of course any Moses user is free to add pre- and post-processing to
remove/escape characters outside of the tool chain. Frankly I think the
tokenizer is already doing too much - some people might want to process the
characters &, <, > in their raw form.
Thanks
Achim
From: [email protected] [mailto:[email protected]]
On Behalf Of Tomas Hudik
Sent: Friday, March 08, 2013 6:13 AM
To: Barry Haddow
Cc: [email protected]
Subject: Re: [Moses-support] training process and special unicode characters
Thanks Barry,
I didn't realize there are perl scripts before Moses binary yesterday.
And after a short discussion with Tom Hoar it seems it would really be
better to filter out all non-text chars (unicode regex: \p{z}, \p{c}). Or,
if some char is important/common in particular dataset re-tag it. Would be
difficult to trace all changes in processing of Unicode by various perl
versions, boost library and potentially some other components involved in
moses workflow.
Cheers, t.
From: Barry Haddow [mailto:[email protected]]
Sent: Friday, March 08, 2013 11:30 AM
To: Tomas Hudik
Cc: [email protected]
Subject: Re: [Moses-support] training process and special unicode characters
Hi Tomas
There's some support in the tokeniser for removing special characters, but
it may not deal with all the unicode control characters. In general, it's
better to remove such characters before you pass the data to Moses as they
can cause problems with the training pipeline. Since a lot of the text
processing is in perl, it really depends on how perl handles them,
cheers - Barry
On 07/03/13 10:53, Tomas Hudik wrote:
Hi,
How is moses treating special Unicode control or white space characters?
e.g. http://www.fileformat.info/info/unicode/char/2028/index.htm, or
http://en.wikipedia.org/wiki/Unicode_control_characters
Is it excluded or they become part of phrase table?
Not sure if this question wouldn't be better put somewhere in boost library
forum .
Thanks, Tomas
_______________________________________________
Moses-support mailing list
[email protected]
http://mailman.mit.edu/mailman/listinfo/moses-support
_______________________________________________
Moses-support mailing list
[email protected]
http://mailman.mit.edu/mailman/listinfo/moses-support