To rebuild PCRE with UTF8 support, I had to read a bit of the doc (not 
much mind you) and I spotted some interesting options... I thought the 
stuff regarding newlines in particular would be of interest for Sheri 
who asked be to build 6.7 especially for that. I didn't hear from you 
BTW: did my build of 6.7 help at all? Anyway, here's a few lines from 
PCRE's README:

> . If, in addition to support for UTF-8 character strings, you want to 
> include
>   support for the \P, \p, and \X sequences that recognize Unicode 
> character
>   properties, you must add --enable-unicode-properties to the "configure"
>   command. This adds about 30K to the size of the library (in the form 
> of a
>   property table); only the basic two-letter properties such as Lu are
>   supported.
>
> . You can build PCRE to recognize either CR or LF or the sequence CRLF as
>   indicating the end of a line. Whatever you specify at build time is the
>   default; the caller of PCRE can change the selection at run time. 
> The default
>   newline indicator is a single LF character (the Unix standard). You can
>   specify the default newline indicator by adding --newline-is-cr or
>   --newline-is-lf or --newline-is-crlf to the "configure" command,
>   respectively.
>
> . When called via the POSIX interface, PCRE uses malloc() to get 
> additional
>   storage for processing capturing parentheses if there are more than 
> 10 of
>   them. You can increase this threshold by setting, for example,
>
>   --with-posix-malloc-threshold=20
>
>   on the "configure" command.
>
> . PCRE has a counter that can be set to limit the amount of resources 
> it uses.
>   If the limit is exceeded during a match, the match fails. The 
> default is ten
>   million. You can change the default by setting, for example,
>
>   --with-match-limit=500000
>
>   on the "configure" command. This is just the default; individual 
> calls to
>   pcre_exec() can supply their own value. There is discussion on the 
> pcreapi
>   man page.
>
> . There is a separate counter that limits the depth of recursive 
> function calls
>   during a matching process. This also has a default of ten million, 
> which is
>   essentially "unlimited". You can change the default by setting, for 
> example,
>
>   --with-match-limit-recursion=500000
>
>   Recursive function calls use up the runtime stack; running out of 
> stack can
>   cause programs to crash in strange ways. There is a discussion about 
> stack
>   sizes in the pcrestack man page.
>
> . The default maximum compiled pattern size is around 64K. You can 
> increase
>   this by adding --with-link-size=3 to the "configure" command. You can
>   increase it even more by setting --with-link-size=4, but this is 
> unlikely
>   ever to be necessary. If you build PCRE with an increased link size, 
> test 2
>   (and 5 if you are using UTF-8) will fail. Part of the output of 
> these tests
>   is a representation of the compiled pattern, and this changes with 
> the link
>   size.

If anyone would want to try a build with some of those options set and 
you can't compile PCRE yourself, let me know (preferably off-list as 
this is borderline off-topic). None of those options were used in my 
builds (I only used --enable-utf8 in the last one).

Reply via email to