--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]> 
wrote:
> 
> 
> To rebuild PCRE with UTF8 support, I had to read a bit
> of the doc (not much mind you) and I spotted some
> interesting options... I thought the stuff regarding
> newlines in particular would be of interest for Sheri
> who asked be to build 6.7 especially for that. I didn't
> hear from you BTW: did my build of 6.7 help at all?

The advantage to having pcre recognize CRLFs as newlines would be in
the usefulness of dot and dollar shortcut patterns. When not using the
"(?s)" option, dot should represent any character except a line ending
character. But using latest beta, dot matches carriage returns and
dollar (which should match at line ends) matches between carriage
returns and linefeeds. Which lineending is used might also affect the
use of ^ pattern (ok right now because it matches at the start of
string or after a linefeed character).

I know how to make patterns that don't use ^, dot, or dollar patterns,
but pattern making would be easier (especially for less experienced
regex users) if they worked properly.

The pcre documentation says "at build time it is conventional to use
the standard for your operating system," which for Windows would be
CRLF. But it can be changed by the caller (the plugin), so if not
compiled with CRLF, the call needs a way to specify CRLF. I guess the
call needs a way to specify whichever lineending the user is dealing
with (e.g., we sometimes read unix text on a windows platform).

> 
> If anyone would want to try a build with some of those
> options set and you can't compile PCRE yourself, let me
> know (preferably off-list as this is borderline off-
> topic). None of those options were used in my builds (I
> only used --enable-utf8 in the last one).
> 

I'd be happy to check one built with CRLF.

Regards,
Sheri

Reply via email to