--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]> wrote: > > > To rebuild PCRE with UTF8 support, I had to read a bit > of the doc (not much mind you) and I spotted some > interesting options... I thought the stuff regarding > newlines in particular would be of interest for Sheri > who asked be to build 6.7 especially for that. I didn't > hear from you BTW: did my build of 6.7 help at all?
The advantage to having pcre recognize CRLFs as newlines would be in the usefulness of dot and dollar shortcut patterns. When not using the "(?s)" option, dot should represent any character except a line ending character. But using latest beta, dot matches carriage returns and dollar (which should match at line ends) matches between carriage returns and linefeeds. Which lineending is used might also affect the use of ^ pattern (ok right now because it matches at the start of string or after a linefeed character). I know how to make patterns that don't use ^, dot, or dollar patterns, but pattern making would be easier (especially for less experienced regex users) if they worked properly. The pcre documentation says "at build time it is conventional to use the standard for your operating system," which for Windows would be CRLF. But it can be changed by the caller (the plugin), so if not compiled with CRLF, the call needs a way to specify CRLF. I guess the call needs a way to specify whichever lineending the user is dealing with (e.g., we sometimes read unix text on a windows platform). > > If anyone would want to try a build with some of those > options set and you can't compile PCRE yourself, let me > know (preferably off-list as this is borderline off- > topic). None of those options were used in my builds (I > only used --enable-utf8 in the last one). > I'd be happy to check one built with CRLF. Regards, Sheri
