OK, I think this should apply and hopefully Julien could fix it. It
would appear the plugin is sending a shortened string for testing
multiple matches when it should be sending startoffsets?

Regards,
Sheri

(what follows comes from pcreapi doc)

The string to be matched by pcre_exec()

The subject string is passed to pcre_exec() as a pointer in subject, a
length in length, and a starting byte offset in startoffset. In UTF-8
mode, the byte offset must point to the start of a UTF-8 character.
Unlike the pattern string, the subject may contain binary zero bytes.
When the starting offset is zero, the search for a match starts at the
beginning of the subject, and this is by far the most common case.

A non-zero starting offset is useful when searching for another match
in the same subject by calling pcre_exec() again after a previous
success. Setting startoffset differs from just passing over a
shortened string and setting PCRE_NOTBOL in the case of a pattern that
begins with any kind of lookbehind. For example, consider the pattern

  \Biss\B

which finds occurrences of "iss" in the middle of words. (\B matches
only if the current position in the subject is not a word boundary.)
When applied to the string "Mississipi" the first call to pcre_exec()
finds the first occurrence. If pcre_exec() is called again with just
the remainder of the subject, namely "issipi", it does not match,
because \B is always false at the start of the subject, which is
deemed to be a word boundary. However, if pcre_exec() is passed the
entire string again, but with startoffset set to 4, it finds the
second occurrence of "iss" because it is able to look behind the
starting point to discover that it is preceded by a letter.

If a non-zero starting offset is passed when the pattern is anchored,
one attempt to match at the given offset is made. This can only
succeed if the pattern does not require the match to be at the start
of the subject.


--- In [email protected], "swzoh" <[EMAIL PROTECTED]> wrote:
>
> --- In [email protected], "Sheri" <sherip99@> wrote:
> >
> > --- In pow [EMAIL PROTECTED], Julien Pierrehumbert <julp@> wrote:
> 
> > > That's right... and unfortunately using other regex
> > > features as fancy lookbehinds isn't likely to help (in
> > > most cases anyway). The basic problem is that the regex
> > > library won't be aware of what has been consumed by
> > > previous matches. There's no easy way around that AFAIK.
> > 
> > I just need to see an example that demonstrates what doesn't work.
> 
> There was an example of this in the past:
> 
> http://tech.groups.yahoo.com/group/power-pro/message/27613
> 
> Sean
>


Reply via email to