I wonder if he could just simply make . the equivalent of [^\r\n] when
CRLF is in effect.

In other words not a carriage return or a line feed.

So, the CRLF remains two characters, but dot skips them both because
one is a carriage return and one is a line feed. Dot would also skip
isolated carriage returns and line feeds but such are abnormal in
files using CRLF. The external logic could stay as it is, but the dot
would have a different meaning when the CRLF option is in effect. So
searching for \n would find line feeds, \r would find carriage returns
and \r\n would find an entire line break. \nfoo would find 4
characters and \r\nfoo would find 5 (same as non CRLF line ending
choices).

But ^ would need to continue to match where after \r\n with
multiline+CRLF options and $ would need to continue to match before
\r\n with multiline+CRLF

I don't know how prevalent isolated carriage returns are in unix text
files, but it probably wouldn't hurt much if dot meant [^\r\n] there
also. I suppose it would be better not to change dot's meaning (for
non-CRLF options) and cause unwanted surprises for established user
patterns.

Would you forward this suggestion to him Julien?

Regards,
Sheri


--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]> 
wrote:
>
> swzoh wrote:
> > So, not wholly skipping CRLF is a bug of PCRE, not regex plugin.
> 
> 
> Here's the reply of PCRE's author on this issue (the
> short version is that wholly skipping would break other
> expressions):
> 
> There are horrible problems with the meaning of "."
> when the newline sequence is more than one character
> long. I think Perl bypasses all of this by converting
> the line ending into a single character on input and
> re-converting on output. There is some discussion of
> this in
> 
> http://unicode.org/unicode/reports/tr18/
> #Line_Boundaries
> 
> One thing it says is this:
> 
> Note: For some implementations, there may be a
> performance impact in recognizing CRLF as a single
> entity, such as with an arbitrary pattern character
> ("."). To account for that, an implementation may
> satisfy R1.6 if there is a mechanism available for
> converting the sequence CRLF to a single line boundary
> character before regex processing.
> 
> I think Perl is banking on that, but PCRE cannot do
> such a conversion.
> 
> When "." is not matching newline, what it actually
> means is that dot won't match if the current position
> is at the start of a newline sequence. This does not
> affect the external logic for a non-anchored regex,
> which moves on by one character if the match has
> failed.
> 
> In practical terms, at the outer level, it doesn't know
> that the match failed because dot failed at the start.
> For a general regex, I don't think it would be easy to
> implement this.
> 
> Also, I don't think you can just automatically skip
> both CR and LF when moving on (when CRLF is a line
> ending) because that would break the expression
> "\nfoo".
> 
> I am in the final stages of preparing PCRE 7.0 for
> release. I will take a look at the code to see what the
> possibilities are for doing something about this, but
> at the moment I am not sure that I will be able to.
> 
> If I can't change the code, I will try to make it clear
> in the documentation exactly what is implemented.

> 
> Regards,
> Philip
> 
> -- Philip Hazel, University of Cambridge Computing Service.
>


Reply via email to