[
https://issues.apache.org/jira/browse/OPENNLP-1833?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18087996#comment-18087996
]
ASF GitHub Bot commented on OPENNLP-1833:
-----------------------------------------
krickert opened a new pull request, #493:
URL: https://github.com/apache/opennlp-sandbox/pull/493
See https://issues.apache.org/jira/browse/OPENNLP-1833
Draft for early feedback on the gRPC expansion. This branch restructures the
opennlp-grpc sandbox module around a v1 document-centric API and builds it
out
into a working analysis server.
## API / contract
- AnalyzeDocument RPC over a v1 proto contract (opennlp_document,
opennlp_pipeline, opennlp_service)
- Explicit offset encoding selection (UTF-8 byte / UTF-16 / code point) with
server-side mapping
- AnalysisProfile / AnalysisOptions with strict validation: unsupported
steps, backends and
models are rejected with meaningful gRPC status codes instead of being
silently ignored
- Model bundle catalog (ListModelBundles) reflecting actually loaded
components
## Embeddings
- EmbeddingProvider abstraction with ONNX Runtime CPU and CUDA
implementations behind a
strict factory (model.embedder.backend=onnx|cuda)
- Models declared per id in server config, loaded eagerly at startup,
dimension read from
ONNX session metadata
- Standard single-segment BERT encoding (mask=1, types=0), OOV-to-UNK
mapping, truncation
at 512 wordpieces, deterministic native resource management
- GPU build via -Dgpu swaps onnxruntime for onnxruntime_gpu so the CPU and
CUDA runtimes
never coexist on the classpath
## Chunking (RAG-style segmentation)
- sentence, token-window and semantic algorithms via chunk_embed_configs and
PIPELINE_STEP_CHUNK
- Per-chunk embeddings with multiple models per group, plus group statistics
- Semantic chunking on consecutive-sentence cosine similarity with
percentile or fixed
thresholds and min/max chunk size constraints
## Testing
- 41 unit tests covering providers, factory, chunkers, offset mapping and
analyzer policies,
green on both CPU and GPU build flavors
- End-to-end verified against the public all-MiniLM-L6-v2 ONNX export:
server embeddings are
bit-identical to the corrected SentenceVectorsDL output (see OPENNLP-1836)
## Known follow-ups
- Classpath model discovery (opennlp-models-*) breaks inside the shaded jar
because each
model jar ships a root-level model.properties; the shaded server currently
needs explicit
model.sentence_detector.path / model.tokenizer.path config
- OpenVINO and remote/composite providers are future work behind the same
provider interface
> Add document-centric gRPC API — evolve opennlp-sandbox POC with canonical
> OpenNlpDocument and AnalyzeDocument RPC
> -----------------------------------------------------------------------------------------------------------------
>
> Key: OPENNLP-1833
> URL: https://issues.apache.org/jira/browse/OPENNLP-1833
> Project: OpenNLP
> Issue Type: New Feature
> Components: gRPC binding
> Affects Versions: 3.0.0-M3
> Reporter: Kristian Rickert
> Assignee: Kristian Rickert
> Priority: Major
>
> h3. Problem
> Apache OpenNLP is primarily an *in-process Java library* (API, CLI, UIMA).
> The README notes embedding in distributed pipelines (Flink, NiFi, Spark), but
> there is *no standard wire contract* for cross-language clients or remote
> inference.
> A proof-of-concept exists in the sandbox:
> * *Repository:*
> [https://github.com/apache/opennlp-sandbox/tree/main/opennlp-grpc]
> * *Current scope:* Three separate gRPC services
> ({{{}SentenceDetectorService{}}}, {{{}TokenizerTaggerService{}}},
> {{{}PosTaggerService{}}}) with string-based requests and {{model_hash}} per
> call
> * *Gap:* No unified *document* message, no pipeline orchestration, no
> NER/chunking/embeddings, and clients must chain multiple RPCs
> Main OpenNLP ({{{}apache/opennlp{}}}) has *no gRPC modules* on {{{}main{}}}.
> OpenNLP 3.0 brings thread-safe {{*ME}} classes (JDK 21+), which makes a
> long-lived gRPC server practical. The {{opennlp-dl}} / {{opennlp-dl-gpu}}
> modules already support ONNX inference (including sentence embeddings via
> {{{}SentenceVectorsDL{}}}).
> h3. Proposal
> Evolve the sandbox POC into ASF-native modules (target: main repo after
> consensus):
> ||Module||Purpose||
> |{{opennlp-grpc-api}}|Protocol Buffers + generated stubs (Java first;
> descriptors for other languages)|
> |{{opennlp-grpc-server}}|gRPC server, model bundle registry, pipeline
> orchestration|
> |{{opennlp-grpc-examples}}|Sample clients (e.g. Python)|
> *Core API change:* Introduce a canonical *{{OpenNlpDocument}}* message (1:1
> text document in, enriched document out) and a primary *{{AnalyzeDocument}}*
> RPC that runs a configurable NLP pipeline server-side—similar in spirit to
> the existing UIMA {{OpenNlpTextAnalyzer}} composite, but as a
> language-neutral contract.
> *Package naming (proposed):* {{org.apache.opennlp.grpc.v1}}
> h3. Non-goals (v1 RFC)
> * Binary/PDF document parsing (Tika, etc.) — callers supply {{raw_text}}
> * Training, evaluation, or model-update RPCs
> * Embedding {{.bin}} model bytes in request messages (models remain
> server-side)
> * Authentication / multi-tenancy in the core API (deployment concern: mTLS,
> reverse proxy)
> * Coreference (documented in manual but not implemented in current codebase)
> h3. Compatibility
> * *Additive* Maven modules; no breaking changes to {{opennlp-api}} /
> {{opennlp-runtime}}
> * Sandbox granular services may be deprecated or moved to
> {{opennlp.legacy.v1}} after migration
> h3. Phased delivery (high level)
> ||Phase||Scope||
> |*0*|This JIRA + community RFC (this ticket)|
> |*1*|Design document + full {{.proto}} definitions (server coded along the
> way for demo but open for changes - phase 2 attempts to lock in the proto)|
> |*2+*|Implementation: orchestrator, server, tests, graduation from sandbox to
> main repo|
> |*Later*|GPU embeddings ({{{}opennlp-dl-gpu{}}}), optional OpenVINO/DJL
> inference backends|
> h3. Design highlights
> # *Three proto layers (NLP-only):* domain types ({{{}OpenNlpDocument{}}}),
> pipeline config ({{{}AnalysisProfile{}}}), service
> ({{{}OpenNlpAnalysisService{}}})
> # *Offset contract:* All exported spans use *character offsets in the
> original {{raw_text}}* ({{{}CHAR_DOCUMENT{}}}), half-open {{[start, end)}}
> matching {{opennlp.tools.util.Span}}
> # *Model bundles:* Replace per-RPC {{model_hash}} with {{ModelBundleRef}} +
> server-defined profiles (reuse sandbox model discovery patterns)
> # *Thread safety:* Leverage OpenNLP 3.0 thread-safe {{*ME}} instances cached
> per model bundle
> h3. Sample protobuf (illustrative — full spec in design doc)
> The following is a *sketch* for discussion; field numbers and optional
> messages may change during RFC.
> {code:java|title=opennlp.proto}
> syntax = "proto3";
> package org.apache.opennlp.grpc.v1;
> option java_package = "org.apache.opennlp.grpc.v1";
> option java_multiple_files = true;
> // --- Layer 1: Document ---
> message OpenNlpDocument {
> string doc_id = 1;
> string raw_text = 2;
> optional string detected_language = 3;
> optional float language_confidence = 4;
> repeated AnnotatedSentence sentences = 5;
> map<string, string> metadata = 6;
> }
> message AnnotatedSentence {
> CharSpan sentence_span = 1;
> repeated Token tokens = 2;
> repeated NamedEntity entities = 3;
> }
> message Token {
> string text = 1;
> CharSpan char_span = 2;
> optional string pos_tag = 3;
> }
> message NamedEntity {
> CharSpan char_span = 1;
> string entity_type = 2;
> optional double prob = 3;
> }
> message CharSpan {
> int32 start = 1;
> int32 end = 2;
> CoordinateSpace space = 3;
> optional string type = 4;
> optional double prob = 5;
> }
> enum CoordinateSpace {
> COORDINATE_SPACE_UNSPECIFIED = 0;
> CHAR_DOCUMENT = 1;
> }
> // --- Layer 2: Pipeline ---
> enum PipelineStep {
> PIPELINE_STEP_UNSPECIFIED = 0;
> LANGUAGE_DETECT = 1;
> SENTENCE_DETECT = 2;
> TOKENIZE = 3;
> POS_TAG = 4;
> NER = 5;
> }
> message AnalysisProfile {
> string profile_id = 1;
> repeated PipelineStep steps = 2;
> ModelBundleRef model_bundle = 3;
> }
> message ModelBundleRef {
> string bundle_id = 1;
> }
> message AnalysisOptions {
> bool include_probabilities = 1;
> bool clear_adaptive_data = 2;
> }
> // --- Layer 3: Service ---
> service OpenNlpAnalysisService {
> rpc AnalyzeDocument(AnalyzeDocumentRequest) returns
> (AnalyzeDocumentResponse);
> rpc GetServiceInfo(GetServiceInfoRequest) returns (GetServiceInfoResponse);
> }
> message AnalyzeDocumentRequest {
> OpenNlpDocument document = 1;
> AnalysisProfile profile = 2;
> AnalysisOptions options = 3;
> }
> message AnalyzeDocumentResponse {
> OpenNlpDocument document = 1;
> repeated ProcessingDiagnostic diagnostics = 2;
> }
> message ProcessingDiagnostic {
> PipelineStep step = 1;
> string message = 2;
> DiagnosticSeverity severity = 3;
> }
> enum DiagnosticSeverity {
> DIAGNOSTIC_SEVERITY_UNSPECIFIED = 0;
> INFO = 1;
> WARNING = 2;
> ERROR = 3;
> }
> message GetServiceInfoRequest {}
> message GetServiceInfoResponse {
> string opennlp_version = 1;
> string api_version = 2;
> repeated string available_profile_ids = 3;
> }
> {code}
> h3. Comparison: sandbox vs proposed
> ||Aspect||Sandbox POC||Proposed||
> |Services|3 (sent / token / POS)|1 primary ({{{}OpenNlpAnalysisService{}}})|
> |I/O|Strings + {{StringList}}|{{OpenNlpDocument}}|
> |Models|{{model_hash}} per RPC|{{ModelBundleRef}} + profiles|
> |Pipeline|Client-side chaining|Server-side {{AnalysisProfile}}|
> |Package|{{package opennlp}}|{{org.apache.opennlp.grpc.v1}}|
> h3. References
> * Sandbox POC:
> [https://github.com/apache/opennlp-sandbox/tree/main/opennlp-grpc]
> * Current sandbox proto:
> [https://github.com/apache/opennlp-sandbox/blob/main/opennlp-grpc/opennlp-grpc-api/opennlp.proto]
> * UIMA composite pipeline:
> {{opennlp-extensions/opennlp-uima/descriptors/OpenNlpTextAnalyzer.xml}}
> * ONNX / GPU: {{{}opennlp-dl{}}}, {{{}opennlp-dl-gpu{}}},
> {{SentenceVectorsDL}}
> * Full design document (companion): {{docs/rfc/opennlp-grpc-design.md}} in
> contributor branch or attachment
> h3. Questions for the community
> # Should v1 expose *only* {{{}AnalyzeDocument{}}}, or retain sandbox
> granular RPCs under a legacy package?
> # Target release: *3.0.x* (additive) vs {*}3.1{*}?
> # Proto tooling: Maven {{protobuf-maven-plugin}} only, is a gradle build OK?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)