I’m also +1 for this direction. We should aim for this to be achieved for the 
3.0.0 (GA) release in November (see roadmap in Jira).
Think a separate epic with sub-issues might be helpful.

Thanks,
Martin


> Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>:
> 
> I'm in favor of removing regex usage where possible for the reasons
> you gave, and because regex usage is often a source of CVEs.
> 
> Thanks,
> Jeff
> 
> On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]> wrote:
>> 
>> Hey everyone,
>> 
>> While reviewing the code I had noticed a few places where we are still
>> using regex unnecessarily.  These improvements aren't a major rush, and I
>> have one epic tracking all of the necessary changes.  These improvements
>> are all heavily tested. Only the first one is stacked because it improves
>> StringUtil, which the others depend on. That means no major rush merging
>> them and none are blockers for a 3.0 release (but would be a great story to
>> tell
>> 
>> So OPENNLP-1928 is the only one stacked; the rest will hang off of main and
>> merge cleanly.
>> 
>> After the first ticket is done, I'll create a speed test to show any
>> performance/memory improvement from the fix.  As we've seen from the other
>> tickets where we reduced RegEx, using cursors and string pointers over
>> regex gives us far less memory churn and better speed.  I also have an
>> easier time understanding it as I never really "get" a lot of regex.
>> 
>> Feel free to chime in, make changes, test, or review...
>> 
>> Here's the list:
>> Ticket Scope PR
>> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928> part 1:
>> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace,
>> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit
>> apache/opennlp#1275
>> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930> part 2:
>> Arvores Deitadas markup parsing apache/opennlp#1276
>> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931> part 3:
>> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277
>> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932> part 4:
>> wildcard matching in the model resolver apache/opennlp#1278
>> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933> part 5:
>> per-call String regex splits and replacements apache/opennlp#1279
>> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug:
>> BasicContextGenerator splits on its separator as a regular expression
>> apache/opennlp#1280
>> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934> part 6:
>> tokenizer alphanumeric pattern evaluated as a character set
>> apache/opennlp#1281
>> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935> part 7:
>> checkstyle guard and the exempt list in checkstyle-suppressions.xml
>> apache/opennlp#1282

Reply via email to