Martin, I made them an epic with linked issues. Do we need them to be
sub-issues?

https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues

But glad we're both fans of epics.

On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]> wrote:

> I’m also +1 for this direction. We should aim for this to be achieved for
> the 3.0.0 (GA) release in November (see roadmap in Jira).
> Think a separate epic with sub-issues might be helpful.
>
> Thanks,
> Martin
>
>
> > Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>:
> >
> > I'm in favor of removing regex usage where possible for the reasons
> > you gave, and because regex usage is often a source of CVEs.
> >
> > Thanks,
> > Jeff
> >
> > On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]>
> wrote:
> >>
> >> Hey everyone,
> >>
> >> While reviewing the code I had noticed a few places where we are still
> >> using regex unnecessarily.  These improvements aren't a major rush, and
> I
> >> have one epic tracking all of the necessary changes.  These improvements
> >> are all heavily tested. Only the first one is stacked because it
> improves
> >> StringUtil, which the others depend on. That means no major rush merging
> >> them and none are blockers for a 3.0 release (but would be a great
> story to
> >> tell
> >>
> >> So OPENNLP-1928 is the only one stacked; the rest will hang off of main
> and
> >> merge cleanly.
> >>
> >> After the first ticket is done, I'll create a speed test to show any
> >> performance/memory improvement from the fix.  As we've seen from the
> other
> >> tickets where we reduced RegEx, using cursors and string pointers over
> >> regex gives us far less memory churn and better speed.  I also have an
> >> easier time understanding it as I never really "get" a lot of regex.
> >>
> >> Feel free to chime in, make changes, test, or review...
> >>
> >> Here's the list:
> >> Ticket Scope PR
> >> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928> part
> 1:
> >> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace,
> >> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit
> >> apache/opennlp#1275
> >> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930> part
> 2:
> >> Arvores Deitadas markup parsing apache/opennlp#1276
> >> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931> part
> 3:
> >> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277
> >> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932> part
> 4:
> >> wildcard matching in the model resolver apache/opennlp#1278
> >> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933> part
> 5:
> >> per-call String regex splits and replacements apache/opennlp#1279
> >> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug:
> >> BasicContextGenerator splits on its separator as a regular expression
> >> apache/opennlp#1280
> >> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934> part
> 6:
> >> tokenizer alphanumeric pattern evaluated as a character set
> >> apache/opennlp#1281
> >> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935> part
> 7:
> >> checkstyle guard and the exempt list in checkstyle-suppressions.xml
> >> apache/opennlp#1282
>
>

Reply via email to