swzoh wrote:
> So, if a first byte or following byte(s) is missing or orphaned, then
> PCRE engine would probably detect it and produce an error. Thus, I
> reckon there would be no need to worry about it in UTF-8 case.

Sorry I sent you the wrong pseudo-expression... it's more specific cases 
that are problematic actually: stuff like ([^x]+|)
It seems it simply stops matching in this case so no error is raised by 
the plugin. Oops!
I should look into it really but I think I already know what the problem 
is so...

Does someone know the best way to reliably identify (and skip) the whole 
UTF-8 character in such cases?

To reproduce the bug, replace your line:
> local szRes=regex.umg(zText;;+
> ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ")
with this:
regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ")

> My previous example didn't require other plugins except the file
> plugin.

You used a plugin to call MS functions. But yeah, my PP is shamefully old.

Reply via email to