> I'm not sure what you're actually looking for. What I'm looking for is a standard library (as in glibc) function that identifies UTF-8 characters and their length. That or consice and precise rules for UTF-8 that would allow me to implement it myself. I've never manipulated UTF-8 so I have no idea where to look (not that you should go out of your way to find that for me). I could of course use a regular expression but there should be an easier and neatier way to do that.
> BTW, I can't tell for sure if it's a bug or not. The empty pattern > simply seems to work at byte level, not at character level, like the > operator \C. So, it made PCRE engine stopping further process after > its first trigger as it made the (UTF-8) following byte \x82 orphaned > by removing the first byte \xC3. Exactly. It's the plugin that makes the empty pattern "work at a byte level" (more precisely: skip a byte and declare it to be unmatched each time there's a zero-sized match). This is someting I did because it made sense and because Luciano said other regex implementations did the same thing... but I didn't think about UTF-8 or CRLF back then obviously. I'd say it's a bug because, if you're offering UTF-8 services, you should handle those characters properly and not as multiple ASCII characters. Of course I could invalidate zero-char matches instead like Sheri suggests (I don't need the library to support it) but, last I heard, this is not how such cases are usually handled. I don't know if this is spelled out in some reference document somewhere or not.
