Does it take anything other than simply:

$KCODE = 'u'

right in the beginning of the script?

I have that in place already.

(it's CGI we're talking about)

On 7/19/06, David Balmain <[EMAIL PROTECTED]> wrote:
> On 7/19/06, Julio Cesar Ody <[EMAIL PROTECTED]> wrote:
> > Just sharing my experience and asking another question.
> >
> > I tried the analyzer suggested here:
> > http://www.ruby-forum.com/topic/72086#101764. It works fine if you
> > specify the search field you want to use (anyway, it seems that's how
> > it's suppose to work).
> >
> > # CODE
> > analyzer = 
> > Ferret::Analysis::PerFieldAnalyzer.new(Ferret::Analysis::StandardAnalyzer.new)
> > analyzer["chinese"] = Ferret::Analysis::RegExpAnalyzer.new(/./, false)
> >
> > index = Index::Index.new(:path => '/var/index', :analyzer => analyzer,
> > :default_field => "*")
> >
> > ...
> >
> > index.search_each("chinese: #{val}") do |doc, score|  #val is a chinese char
> >  puts "#{doc} - #{score}"
> > end
> > # END CODE
> >
> > This works OK. However, if you try searching like this:
> >
> > # CODE
> > index.search_each(val) do |doc, score|  #val is a chinese char
> >  puts "#{doc} - #{score}"
> > end
> > # END CODE
> >
> > I get in my lighttpd error log:
> >
> > /var/www/localhost/htdocs/cgi-bin/search_chinese.ruby:19:in
> > `search_each': : Error occured at <analysis.c>:701 (StandardError)
> > Error: exception 2 not handled: Error decoding input string. Check
> > that you have the locale set correctly
> >         from /var/www/localhost/htdocs/cgi-bin/search_chinese.ruby:19
> >
> > Which MAKES SENSE, since the docs I created before are created like this:
> >
> > doc = { "author" => "englishchars", "title" => "more regular chars",
> > "chinese" => "新闻"}
> > index << doc
> >
> > and I think search_each is going through all the fields (since I
> > explicitly said it should when I issued :default_field => "*" up
> > there), finding english chars, and trying to match them against the
> > chinese ones I supplied as a search query.
>
> Actually, it's not because of there is a comparison between Chinese
> and English characters. That shouldn't cause an error. The error is
> being thrown because val can't be decoded using the StandardAnalyzer.
> Again, you need to check that val is correctly encoded and you have
> your locale set correctly.The only times tokenizing happens are when
> you add documents to the index and when you run a query through the
> query parser. Apart from that, all operations on strings are done at
> the byte level. I hope that makes sense.
>
> > So alright, I can use the suggested analyzer. But my question is: is
> > there a way to use an analyzer that would work with both character
> > types (english, and asian) simply by not returning matches them as
> > opposed to giving me an error?
> >
> > Thanks a ton for any help.
>
> The answer to this question is that it already should work correctly.
> Just make sure the locale is set correctly when the search method is
> called and that whatever you pass as a query to the search method is
> correctly encoded according to the locale.
>
> Cheers,
> Dave
> _______________________________________________
> Ferret-talk mailing list
> [email protected]
> http://rubyforge.org/mailman/listinfo/ferret-talk


-- 
Julio C. Ody
http://rootshell.be/~julioody
_______________________________________________
Ferret-talk mailing list
[email protected]
http://rubyforge.org/mailman/listinfo/ferret-talk

Reply via email to