swzoh wrote:
> So, not wholly skipping CRLF is a bug of PCRE, not regex plugin.

Here's the reply of PCRE's author on this issue (the short version is
that wholly skipping would break other expressions):

There are horrible problems with the meaning of "." when the newline
sequence is more than one character long. I think Perl bypasses all of
this by converting the line ending into a single character on input and
re-converting on output. There is some discussion of this in

http://unicode.org/unicode/reports/tr18/#Line_Boundaries

One thing it says is this:

   Note: For some implementations, there may be a performance impact in
   recognizing CRLF as a single entity, such as with an arbitrary pattern
   character ("."). To account for that, an implementation may satisfy
   R1.6 if there is a mechanism available for converting the sequence
   CRLF to a single line boundary character before regex processing.

I think Perl is banking on that, but PCRE cannot do such a conversion.

When "." is not matching newline, what it actually means is that dot
won't match if the current position is at the start of a newline
sequence. This does not affect the external logic for a non-anchored
regex, which moves on by one character if the match has failed.

In practical terms, at the outer level, it doesn't know that the match
failed because dot failed at the start. For a general regex, I don't
think it would be easy to implement this.

Also, I don't think you can just automatically skip both CR and LF when
moving on (when CRLF is a line ending) because that would break the
expression "\nfoo".

I am in the final stages of preparing PCRE 7.0 for release. I will take
a look at the code to see what the possibilities are for doing something
about this, but at the moment I am not sure that I will be able to.

If I can't change the code, I will try to make it clear in the
documentation exactly what is implemented.

Regards,
Philip

-- Philip Hazel, University of Cambridge Computing Service.

Reply via email to