To rebuild PCRE with UTF8 support, I had to read a bit of the doc (not much mind you) and I spotted some interesting options... I thought the stuff regarding newlines in particular would be of interest for Sheri who asked be to build 6.7 especially for that. I didn't hear from you BTW: did my build of 6.7 help at all? Anyway, here's a few lines from PCRE's README:
> . If, in addition to support for UTF-8 character strings, you want to > include > support for the \P, \p, and \X sequences that recognize Unicode > character > properties, you must add --enable-unicode-properties to the "configure" > command. This adds about 30K to the size of the library (in the form > of a > property table); only the basic two-letter properties such as Lu are > supported. > > . You can build PCRE to recognize either CR or LF or the sequence CRLF as > indicating the end of a line. Whatever you specify at build time is the > default; the caller of PCRE can change the selection at run time. > The default > newline indicator is a single LF character (the Unix standard). You can > specify the default newline indicator by adding --newline-is-cr or > --newline-is-lf or --newline-is-crlf to the "configure" command, > respectively. > > . When called via the POSIX interface, PCRE uses malloc() to get > additional > storage for processing capturing parentheses if there are more than > 10 of > them. You can increase this threshold by setting, for example, > > --with-posix-malloc-threshold=20 > > on the "configure" command. > > . PCRE has a counter that can be set to limit the amount of resources > it uses. > If the limit is exceeded during a match, the match fails. The > default is ten > million. You can change the default by setting, for example, > > --with-match-limit=500000 > > on the "configure" command. This is just the default; individual > calls to > pcre_exec() can supply their own value. There is discussion on the > pcreapi > man page. > > . There is a separate counter that limits the depth of recursive > function calls > during a matching process. This also has a default of ten million, > which is > essentially "unlimited". You can change the default by setting, for > example, > > --with-match-limit-recursion=500000 > > Recursive function calls use up the runtime stack; running out of > stack can > cause programs to crash in strange ways. There is a discussion about > stack > sizes in the pcrestack man page. > > . The default maximum compiled pattern size is around 64K. You can > increase > this by adding --with-link-size=3 to the "configure" command. You can > increase it even more by setting --with-link-size=4, but this is > unlikely > ever to be necessary. If you build PCRE with an increased link size, > test 2 > (and 5 if you are using UTF-8) will fail. Part of the output of > these tests > is a representation of the compiled pattern, and this changes with > the link > size. If anyone would want to try a build with some of those options set and you can't compile PCRE yourself, let me know (preferably off-list as this is borderline off-topic). None of those options were used in my builds (I only used --enable-utf8 in the last one).
