> I'm not sure what you're actually looking for.

What I'm looking for is a standard library (as in glibc) function that 
identifies UTF-8 characters and their length. That or consice and 
precise rules for UTF-8 that would allow me to implement it myself. I've 
never manipulated UTF-8 so I have no idea where to look (not that you 
should go out of your way to find that for me).
I could of course use a regular expression but there should be an easier 
and neatier way to do that.

> BTW, I can't tell for sure if it's a bug or not. The empty pattern
> simply seems to work at byte level, not at character level, like the
> operator \C. So, it made PCRE engine stopping further process after
> its first trigger as it made the (UTF-8) following byte \x82 orphaned
> by removing the first byte \xC3.

Exactly. It's the plugin that makes the empty pattern "work at a byte 
level" (more precisely: skip a byte and declare it to be unmatched each 
time there's a zero-sized match). This is someting I did because it made 
sense and because Luciano said other regex implementations did the same 
thing... but I didn't think about UTF-8 or CRLF back then obviously.
I'd say it's a bug because, if you're offering UTF-8 services, you 
should handle those characters properly and not as multiple ASCII 
characters.

Of course I could invalidate zero-char matches instead like Sheri 
suggests (I don't need the library to support it) but, last I heard, 
this is not how such cases are usually handled. I don't know if this is 
spelled out in some reference document somewhere or not.

Reply via email to