On Tuesday, October 12 2004 23:20, Ilya Konstantinov wrote:
> Shai Berger wrote:
> >The situation where the question comes up is a single document containing
> > both Yiddish and Hebrew. Nobody would switch layouts.
>
> Since Hebrew is a subset of Yiddish, so you can keep typing Hebrew with
> an Yiddish keyboard?
>
And the Yiddish alphabet is close enough to a subset of Hebrew (the 
exceptions, as far as I know, are connected ligatures, not new characters).

> And yet, knowing the IME language the character was input with does help
> in cases when the document uses two languages, neither of which is a
> subset of the other -- e.g. Hebrew and English. In this case, the user
> is forced to give us the 'IME language switched' hint.
>
But then these languages have completely disjoint sets of characters; deciding 
between them is trivial. The only case where IME would be really helpful is 
where the languages have enough common characters that some words in each can 
be written in the common subset, and yet enough distinct characters to 
justify separate layouts. There are such cases, e.g. Latin-1 vs. Latin-2 
(Cyrillic) languages. But relying on IME is not very robust even in these 
cases.

> And anyway, even if Hebrew is a subset of Yiddish, we surely know it's
> an RTL language (either it's Hebrew or Yiddish -- both are RTL) so we
> need to set RTL base direction for this consequetive set of Hebrew
> characters.
>
But Unicode will already tell you that.

>
> >>We have language info anyway, so we might as well use it.
> >>In the information exported to other programs, via clipboard or export
> >>formats, we should add the direction characters to replicate any
> >>directional formatting the language info implies to the OOo rendered.
> >
> >I say again, we should strive to get rid of language info.
>
> We cannot! Our language tools (speller, hyphenation) cannot tell a latin
> text's language (English or Dutch) unless its marked. They haven't grown
> artificial intelligence yet to look at the entire text and guess what
> language its in; and even if we had the AI, the user should've still had
> the option to override it.

Ok. We need language info. Still, IME language is a poor source for this info. 
Further, language info applies to words at a minimum, not to characters, and 
this should be reflected in the structure of information we keep. Using IME 
for gaining language info is a passable solution, but storing it as character 
traits seems broken.
==============================================================
To unsubscribe, send mail to [EMAIL PROTECTED]
with "unsubscribe hebrew" in the message body.

Reply via email to