--- In [email protected], "swzoh" <[EMAIL PROTECTED]> wrote:
>
> --- In [email protected], Julien Pierrehumbert <julp@> wrote:
> >
> > Does someone know the best way to reliably identify (and skip) the
> > whole UTF-8 character in such cases?
> 
> I think you can use, like the one below:
> [^\x01-\x{10FFFF}]

What happens if you add {2,} to the end of the pattern, to require at
least two characters (bytes?) ?

Regards,
Sheri

Reply via email to