swzoh wrote:
>> Does someone know the best way to reliably identify (and skip) the
>> whole UTF-8 character in such cases?
> 
> I think you can use, like the one below:
> [^\x01-\x{10FFFF}]

Is a regex really the best way to identify a UTF-8 character and 
determine its size?
I don't know what \x{} does so I can't say if it's realiable but there 
has to be a better way... I think I'd rather study the UTF-8 rules and 
implement them myself! Surely there's a standard library function that 
can take care of this.

>> To reproduce the bug, replace your line:
>>> local szRes=regex.umg(zText;;+
>>> ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ")
>> with this:
>> regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ")
> 
> Is this what you had in mind?
> 
> local szRes=regex.umg(zText;;+
> ,"([^"++esc(?"\xC3\x82",?"\")++?"]+|[^\x01-\x{10FFFF}])",?"\0 ")

No, what I had in mind is exactly what I wrote.

What you wrote doesn't seem to trigger any bug... do you think it does?

Reply via email to