I’m also +1 for this direction. We should aim for this to be achieved for the 3.0.0 (GA) release in November (see roadmap in Jira). Think a separate epic with sub-issues might be helpful.
Thanks, Martin > Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>: > > I'm in favor of removing regex usage where possible for the reasons > you gave, and because regex usage is often a source of CVEs. > > Thanks, > Jeff > > On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]> wrote: >> >> Hey everyone, >> >> While reviewing the code I had noticed a few places where we are still >> using regex unnecessarily. These improvements aren't a major rush, and I >> have one epic tracking all of the necessary changes. These improvements >> are all heavily tested. Only the first one is stacked because it improves >> StringUtil, which the others depend on. That means no major rush merging >> them and none are blockers for a 3.0 release (but would be a great story to >> tell >> >> So OPENNLP-1928 is the only one stacked; the rest will hang off of main and >> merge cleanly. >> >> After the first ticket is done, I'll create a speed test to show any >> performance/memory improvement from the fix. As we've seen from the other >> tickets where we reduced RegEx, using cursors and string pointers over >> regex gives us far less memory churn and better speed. I also have an >> easier time understanding it as I never really "get" a lot of regex. >> >> Feel free to chime in, make changes, test, or review... >> >> Here's the list: >> Ticket Scope PR >> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928> part 1: >> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace, >> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit >> apache/opennlp#1275 >> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930> part 2: >> Arvores Deitadas markup parsing apache/opennlp#1276 >> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931> part 3: >> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277 >> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932> part 4: >> wildcard matching in the model resolver apache/opennlp#1278 >> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933> part 5: >> per-call String regex splits and replacements apache/opennlp#1279 >> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug: >> BasicContextGenerator splits on its separator as a regular expression >> apache/opennlp#1280 >> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934> part 6: >> tokenizer alphanumeric pattern evaluated as a character set >> apache/opennlp#1281 >> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935> part 7: >> checkstyle guard and the exempt list in checkstyle-suppressions.xml >> apache/opennlp#1282
