Thanks for understanding.  I'll minimize the shape and move the
functionality to a branch in addons, just as our canary in the coal mine.

https://issues.apache.org/jira/browse/OPENNLP-1924 - Epic for ADDONS

My next commits are in
https://github.com/apache/opennlp/tree/OPENNLP-1885-sentencepiece -
stripping down the sentencepiece for the API shape and then putting the
code in addons tomake org.apache.opennlp.the `org.apache.opennlp.addons`
package and the OPENNLP-1885-subword-addon which will come from a private
repository and enter the current addons repository via a formal PR.

https://github.com/apache/opennlp-addons/tree/OPENNLP-1224-canary-addon -
the addons revamp

Does that make sense?

On Wed, Sep 2, 2026 at 3:30 PM Richard Zowalla <[email protected]> wrote:

> Sounds good to me.
>
> On 2026/09/02 19:21:27 Kristian Rickert wrote:
> > It does - but not only embeddings -
> >
> > OpenNLP needs SentencePiece to run SentencePiece-backed neural annotators
> > over sentences - like our own onnx features - while preserving exact
> > offsets into the original text.  This is the main reason I wrote it.  All
> > analysis, even in Python packages in this area does not handle this well
> > (if at all), making it hard to create truly accurate highlighters without
> > it.  However, that analysis and offset tracking belongs in a core API
> > layer.  So it's not adding a dependency; it's just an interface that
> tracks
> > the offsets for it and perhaps some associated metadata at that level.
> >
> > I'll take a second look at SentencePiece and carve out the API/core
> stuff;
> > the rest goes into addon.
> >
> > That's the reasoning: embeddings and SentencePiece are already central to
> > much NLP work, so the interface can exist there. As new technologies
> emerge
> > to fill this gap, we can pivot (like how the gRPC server uses OpenVINO,
> > TEI, or Model2Vec, for example).
> >
> > Thoughts?
> >
> >
> > On Tue, Sep 1, 2026, 9:02 AM Richard Zowalla <[email protected]> wrote:
> >
> > > Hss todo with embeddings, no?
> > > That would contract in Martin mail because he wrote "PRs around
> > > "embeddings" should for now live and be maintained in opennlp-addons.“.
> > >
> > > Gruß
> > > Richard
> > >
> > > > Am 01.09.2026 um 14:33 schrieb Kristian Rickert <[email protected]
> >:
> > > >
> > > > Quick clarification:
> > > >
> > > > In the core/addons thread, Martin wrote that OPENNLP-1885,
> OPENNLP-1893,
> > > > OPENNLP-1909, and possibly OPENNLP-1894 should fit in core, and that
> he
> > > > wanted to put effort into them so M6 could mature.
> > > >
> > > > The M6 proposal includes OPENNLP-1893, OPENNLP-1909, and
> OPENNLP-1894,
> > > but
> > > > not OPENNLP-1885. Since the two messages were  posted about a minute
> > > apart,
> > > > I suspect they may have crossed?
> > > >
> > > >
> > > >
> > > >
> > > >
> > > > On Tue, Sep 1, 2026 at 3:12 AM Richard Zowalla <[email protected]>
> wrote:
> > > >
> > > >> Subject: [DISCUSS] 3.0.0-M6 release
> > > >>
> > > >> Hi all,
> > > >>
> > > >> M5 went out on 2026-07-21, so we're about six weeks out and main has
> > > >> accumulated ~30 commits since. I'd like to cut 3.0.0-M6 rather than
> let
> > > >> that pile grow further and we have a target around 08/18.
> > > >>
> > > >> What's already on main since M5:
> > > >>
> > > >> - OPENNLP-1888 Document annotation container with typed,
> offset-anchored
> > > >>  layers (#1182). This is the biggest one and the one several other
> PRs
> > > >>  build on, so getting it into users' hands early matters.
> > > >> - OPENNLP-1903 BeamSearch: chain nodes instead of per-candidate
> Sequence
> > > >>  copies (#1205)
> > > >> - OPENNLP-1906 sentence detector abbreviation veto was quadratic in
> > > >>  document length (#1208)
> > > >> - OPENNLP-1904 / -1905 default-locale case folding breaking
> Morfologik
> > > >>  lookup and lemmatizer models (#1206, #1207)
> > > >> - OPENNLP-1922 ReDoS / stack exhaustion in the
> RegexNameFinderFactory
> > > >>  EMAIL and URL patterns (#1250)
> > > >> - OPENNLP-1899 hardened SymSpell deserialization (#1197)
> > > >> - OPENNLP-1902 cached models verified against their SHA-512 checksum
> > > >>  (#1239)
> > > >> - OPENNLP-1900 documented and published security model (#1198)
> > > >> - OPENNLP-1913 / -1915 Maven Wrapper in the source release, mvnw.cmd
> > > >>  line endings (#1229, #1231)
> > > >>
> > > >> That's already a reasonable milestone on its own. On top of that I'd
> > > >> propose we try to land these four before we cut:
> > > >>
> > > >> 1. OPENNLP-1921 Paragraph normalizer for Gutenberg-style and
> Markdown
> > > >>   text (#1249, tusharsaini18899). Approved already, small, and it
> slots
> > > >>   into the normalizer chain we cleaned up in OPENNLP-1916.
> > > >>
> > > >> 2. OPENNLP-1893 Hunspell affix dictionaries for stemming (#1190,
> > > >>   krickert). Classified as core in the scope thread without
> objection:
> > > >>   fits the existing abstractions, no new dependencies, no bundled
> data.
> > > >>   Currently "changes requested" - if that round gets addressed I'd
> like
> > > >>   it in M6.  (#1191, krickert).
> > > >>
> > > >> 3. OPENNLP-1909 Verified installer for user-supplied third-party
> > > >>   resources (#1211, krickert). Interface in core,
> > > >>   concrete downloaders in the modules they need it - M6 is a good
> place
> > > >> to
> > > >>   try that    split out in practice rather than on the list.
> > > >>
> > > >> 4. OPENNLP-1894 Dictionary-based tokenization for Japanese, Korean
> and
> > > >>   Chinese (#1191, krickert); open up for Asian language feels
> natural in
> > > >> a globalized world.
> > > >>
> > > >> To avoid a misunderstanding: putting something
> > > >> into a milestone is not the same as deciding it belongs in core for
> > > >> 3.0.0 final. Milestones are explicitly not API-frozen - that is
> what the
> > > >> M is for. If CJK tokenization or the installer ships in M6 and we
> later
> > > >> agree it belongs in addons, moving it before the final release is
> > > >> exactly the kind of change a milestone series is supposed to
> absorb. I'd
> > > >> rather we get these in front of users and collect feedback than
> hold a
> > > >> release while the scope thread converges.
> > > >>
> > > >> So M6 does not preempt
> > > >>
> > > >>  https://lists.apache.org/thread/vc8mnvfp6rmr5qb2s9y6hw0zpcxm6t68
> > > >>
> > > >> and I'd like that thread to keep running in parallel. The open
> questions
> > > >> there - the review bar for addons, the compatibility contract, who
> > > >> drives discoverability - are release-cadence questions for addons,
> not
> > > >> or a core milestone. Same for the addons cleanup and getting
> > > >> it to Maven Central: worth doing, independent of M6.
> > > >>
> > > >> If anything on the list above isn't ready, it drops from M6 and goes
> > > >> into M7 (or GA???). I'd rather cut on a date than on a feature list.
> > > >>
> > > >> Objections, additions, or things I've missed that should be in M6?
> > > >>
> > > >> Gruß
> > > >> Richard
> > >
> > >
> >
>

Reply via email to