> Personally, I'm rather against using CRLF as newline, at least with
> the current implementation in PCRE.
I'm also rather against it even if implemented properly... but I also
believe in freedom all that good stuff so I was planning to provide both
builds and let the user pick and choose (assuming the CRLF build it
turns out to be useful of course).
> dot only excludes CRLF, not CR, LF, or even LFCR. So, the following
> can happen. dot stops before CRLF, but doesn't skip the whole CRLF as
> these are two characters. Looks like it's happening indeed:
>
> win.debug(regex.mg(esc("0r0n",0)++"test",?".+",?"\0|"))
>
Not really... what you're describing is this case actually:
regex.mg("test"++esc("0r0n",0),?".+",?"\0|")
In any case there is indeed a problem. Given that we're using the first
PCRE version which has this feature, bugs are to be excepted. I'll
report it unless someone wants to do it or can show us that this is
actually not a PCRE bug.
Even if this was fixed in PCRE, the problem will not be solved in all
cases because the plugin's *g services also treat CRLF as two characters
and there's nothing PCRE can do about that.
Does this happen with UTF-8 as well BTW?