--- In [email protected], "Sheri" <[EMAIL PROTECTED]> wrote:
>
> --- In [email protected], "swzoh" <seanzoh@> wrote:
> >
> > --- In [email protected], Julien Pierrehumbert <julp@> wrote:
> > >
> > > Does someone know the best way to reliably identify (and skip) the
> > > whole UTF-8 character in such cases?
> >
> > I think you can use, like the one below:
> > [^\x01-\x{10FFFF}]
>
> What happens if you add {2,} to the end of the pattern, to require at
> least two characters (bytes?) ?
They are already multi-byte characters. Adding {2} to the above will
pick up two characters, not two bytes in UTF-8 mode.
You should differentiate between byte and character which will be
rather awkward at first for the single-byte (ANSI) character users.
Sean