Reply to myself: yes:

ENV['LANG'] = 'en_US.utf8'

Did the job.

Thanks!

On 7/19/06, Julio Cesar Ody <[EMAIL PROTECTED]> wrote:
> Does it take anything other than simply:
>
> $KCODE = 'u'
>
> right in the beginning of the script?
>
> I have that in place already.
>
> (it's CGI we're talking about)
>
> On 7/19/06, David Balmain <[EMAIL PROTECTED]> wrote:
> > On 7/19/06, Julio Cesar Ody <[EMAIL PROTECTED]> wrote:
> > > Just sharing my experience and asking another question.
> > >
> > > I tried the analyzer suggested here:
> > > http://www.ruby-forum.com/topic/72086#101764. It works fine if you
> > > specify the search field you want to use (anyway, it seems that's how
> > > it's suppose to work).
> > >
> > > # CODE
> > > analyzer = 
> > > Ferret::Analysis::PerFieldAnalyzer.new(Ferret::Analysis::StandardAnalyzer.new)
> > > analyzer["chinese"] = Ferret::Analysis::RegExpAnalyzer.new(/./, false)
> > >
> > > index = Index::Index.new(:path => '/var/index', :analyzer => analyzer,
> > > :default_field => "*")
> > >
> > > ...
> > >
> > > index.search_each("chinese: #{val}") do |doc, score|  #val is a chinese 
> > > char
> > >  puts "#{doc} - #{score}"
> > > end
> > > # END CODE
> > >
> > > This works OK. However, if you try searching like this:
> > >
> > > # CODE
> > > index.search_each(val) do |doc, score|  #val is a chinese char
> > >  puts "#{doc} - #{score}"
> > > end
> > > # END CODE
> > >
> > > I get in my lighttpd error log:
> > >
> > > /var/www/localhost/htdocs/cgi-bin/search_chinese.ruby:19:in
> > > `search_each': : Error occured at <analysis.c>:701 (StandardError)
> > > Error: exception 2 not handled: Error decoding input string. Check
> > > that you have the locale set correctly
> > >         from /var/www/localhost/htdocs/cgi-bin/search_chinese.ruby:19
> > >
> > > Which MAKES SENSE, since the docs I created before are created like this:
> > >
> > > doc = { "author" => "englishchars", "title" => "more regular chars",
> > > "chinese" => "新闻"}
> > > index << doc
> > >
> > > and I think search_each is going through all the fields (since I
> > > explicitly said it should when I issued :default_field => "*" up
> > > there), finding english chars, and trying to match them against the
> > > chinese ones I supplied as a search query.
> >
> > Actually, it's not because of there is a comparison between Chinese
> > and English characters. That shouldn't cause an error. The error is
> > being thrown because val can't be decoded using the StandardAnalyzer.
> > Again, you need to check that val is correctly encoded and you have
> > your locale set correctly.The only times tokenizing happens are when
> > you add documents to the index and when you run a query through the
> > query parser. Apart from that, all operations on strings are done at
> > the byte level. I hope that makes sense.
> >
> > > So alright, I can use the suggested analyzer. But my question is: is
> > > there a way to use an analyzer that would work with both character
> > > types (english, and asian) simply by not returning matches them as
> > > opposed to giving me an error?
> > >
> > > Thanks a ton for any help.
> >
> > The answer to this question is that it already should work correctly.
> > Just make sure the locale is set correctly when the search method is
> > called and that whatever you pass as a query to the search method is
> > correctly encoded according to the locale.
> >
> > Cheers,
> > Dave
> > _______________________________________________
> > Ferret-talk mailing list
> > [email protected]
> > http://rubyforge.org/mailman/listinfo/ferret-talk
>
>
> --
> Julio C. Ody
> http://rootshell.be/~julioody
>


-- 
Julio C. Ody
http://rootshell.be/~julioody
_______________________________________________
Ferret-talk mailing list
[email protected]
http://rubyforge.org/mailman/listinfo/ferret-talk

Reply via email to