Hi sgp,

I agree with your assessment of the situation, just trying to help
with some regex specifics if I can.

--- In [email protected], sgp <[EMAIL PROTECTED]> wrote:

> (*)In my findings, the "less compatible" aspect of the new
> plugin is due to it being compiled with line ending = CRLF by
> default. This is conceptually more correct on Windows, but
> practically it broke some of my older regular expressions that,
> quite sloppily, used

That's false, it is not compiled with CRLF by default :)

The choices when compiling are:

CRLF, LF, CR, ANY and ANYCRLF

default can be only one of the above (old one was LF; there used to be
fewer options)

CRLF was introduced with PCRE 6.7.

ANY was introduced with PCRE 7.0, and in addition to supporting all of
CRLF, LF and CR, it recognizes certain Unicode line endings which are
not compatible with Windows/Ansi (for example, ANY thinks a horizontal
ellipsis into a linebreak).

(Julien made available some 6.7 and 7.0 libraries compiled with CRLF
and ANY, but those were never part of any actual releases).

ANYCRLF was introduced with PCRE 7.1, and that's the default we've
been using. It supports all of CRLF, LF and CR and none of the unicode
line endings. It depends on your input data. If the actual linebreaks
ARE CRLFs it uses CRLFs. In the long run, that is really much better
than needing to dodge carriage returns when they precedes linefeeds.
But if your input uses LFs without carriage returns it uses LF. Ditto
for CR. The setting establishes what constitutes a linebreak. If a
single file had a combination of CRLFs, and isolated CRs and LFs it
would recognize all of them as linebreaks.

(?m)pattern.$ to

> match and end of line \n character. With the new plugin I had to
> change the above to

Your pattern above was not previously matching \n character. It was
matching up to a \n. In other words it matched pattern++"\r" (followed
by \n that was not part of the match)

I suppose if you had no carriage return it would match where there
were two \n in a row, e.g., 

pattern++"\n" (followed by another "\n" that was not part of the match)

> (?m)pattern[\n|\r|\r\n|\n\r]$

What you wrote in this second pattern looks like it is intended to be
a regular expression so the part in the square brackets is a character
class. A character class is not a series of characters, it matches one
character that is any one of the included characters. So after
"pattern", it is matching a linefeed or carriage return or a vertical
bar! The dollar sign doesn't match any characters, it matches empty
space at the end of a line just before a line break. So this does
indeed end up splitting the carriage return and linefeed just as you
used to. I don't know why you want to do that? If you wanted to
include the linebreak in the match, putting:

\r?\n

no square brackets, no dollar sign, e.g,

(?m)pattern\r?\n

would do it.

Regards,
Sheri

Reply via email to