--- In [email protected], "swzoh" <[EMAIL PROTECTED]> wrote: > > --- In [email protected], Julien Pierrehumbert <julp@> wrote: > > > > Does someone know the best way to reliably identify (and skip) the > > whole UTF-8 character in such cases? > > I think you can use, like the one below: > [^\x01-\x{10FFFF}]
What happens if you add {2,} to the end of the pattern, to require at
least two characters (bytes?) ?
Regards,
Sheri
