[ 
https://issues.apache.org/jira/browse/OPENNLP-1833?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18087996#comment-18087996
 ] 

ASF GitHub Bot commented on OPENNLP-1833:
-----------------------------------------

krickert opened a new pull request, #493:
URL: https://github.com/apache/opennlp-sandbox/pull/493

   See https://issues.apache.org/jira/browse/OPENNLP-1833
   
   Draft for early feedback on the gRPC expansion. This branch restructures the
   opennlp-grpc sandbox module around a v1 document-centric API and builds it 
out
   into a working analysis server.
   
   ## API / contract
   - AnalyzeDocument RPC over a v1 proto contract (opennlp_document, 
opennlp_pipeline, opennlp_service)
   - Explicit offset encoding selection (UTF-8 byte / UTF-16 / code point) with 
server-side mapping
   - AnalysisProfile / AnalysisOptions with strict validation: unsupported 
steps, backends and
     models are rejected with meaningful gRPC status codes instead of being 
silently ignored
   - Model bundle catalog (ListModelBundles) reflecting actually loaded 
components
   
   ## Embeddings
   - EmbeddingProvider abstraction with ONNX Runtime CPU and CUDA 
implementations behind a
     strict factory (model.embedder.backend=onnx|cuda)
   - Models declared per id in server config, loaded eagerly at startup, 
dimension read from
     ONNX session metadata
   - Standard single-segment BERT encoding (mask=1, types=0), OOV-to-UNK 
mapping, truncation
     at 512 wordpieces, deterministic native resource management
   - GPU build via -Dgpu swaps onnxruntime for onnxruntime_gpu so the CPU and 
CUDA runtimes
     never coexist on the classpath
   
   ## Chunking (RAG-style segmentation)
   - sentence, token-window and semantic algorithms via chunk_embed_configs and 
PIPELINE_STEP_CHUNK
   - Per-chunk embeddings with multiple models per group, plus group statistics
   - Semantic chunking on consecutive-sentence cosine similarity with 
percentile or fixed
     thresholds and min/max chunk size constraints
   
   ## Testing
   - 41 unit tests covering providers, factory, chunkers, offset mapping and 
analyzer policies,
     green on both CPU and GPU build flavors
   - End-to-end verified against the public all-MiniLM-L6-v2 ONNX export: 
server embeddings are
     bit-identical to the corrected SentenceVectorsDL output (see OPENNLP-1836)
   
   ## Known follow-ups
   - Classpath model discovery (opennlp-models-*) breaks inside the shaded jar 
because each
     model jar ships a root-level model.properties; the shaded server currently 
needs explicit
     model.sentence_detector.path / model.tokenizer.path config
   - OpenVINO and remote/composite providers are future work behind the same 
provider interface




> Add document-centric gRPC API — evolve opennlp-sandbox POC with canonical 
> OpenNlpDocument and AnalyzeDocument RPC
> -----------------------------------------------------------------------------------------------------------------
>
>                 Key: OPENNLP-1833
>                 URL: https://issues.apache.org/jira/browse/OPENNLP-1833
>             Project: OpenNLP
>          Issue Type: New Feature
>          Components: gRPC binding
>    Affects Versions: 3.0.0-M3
>            Reporter: Kristian Rickert
>            Assignee: Kristian Rickert
>            Priority: Major
>
> h3. Problem
> Apache OpenNLP is primarily an *in-process Java library* (API, CLI, UIMA). 
> The README notes embedding in distributed pipelines (Flink, NiFi, Spark), but 
> there is *no standard wire contract* for cross-language clients or remote 
> inference.
> A proof-of-concept exists in the sandbox:
>  * *Repository:* 
> [https://github.com/apache/opennlp-sandbox/tree/main/opennlp-grpc]
>  * *Current scope:* Three separate gRPC services 
> ({{{}SentenceDetectorService{}}}, {{{}TokenizerTaggerService{}}}, 
> {{{}PosTaggerService{}}}) with string-based requests and {{model_hash}} per 
> call
>  * *Gap:* No unified *document* message, no pipeline orchestration, no 
> NER/chunking/embeddings, and clients must chain multiple RPCs
> Main OpenNLP ({{{}apache/opennlp{}}}) has *no gRPC modules* on {{{}main{}}}. 
> OpenNLP 3.0 brings thread-safe {{*ME}} classes (JDK 21+), which makes a 
> long-lived gRPC server practical. The {{opennlp-dl}} / {{opennlp-dl-gpu}} 
> modules already support ONNX inference (including sentence embeddings via 
> {{{}SentenceVectorsDL{}}}).
> h3. Proposal
> Evolve the sandbox POC into ASF-native modules (target: main repo after 
> consensus):
> ||Module||Purpose||
> |{{opennlp-grpc-api}}|Protocol Buffers + generated stubs (Java first; 
> descriptors for other languages)|
> |{{opennlp-grpc-server}}|gRPC server, model bundle registry, pipeline 
> orchestration|
> |{{opennlp-grpc-examples}}|Sample clients (e.g. Python)|
> *Core API change:* Introduce a canonical *{{OpenNlpDocument}}* message (1:1 
> text document in, enriched document out) and a primary *{{AnalyzeDocument}}* 
> RPC that runs a configurable NLP pipeline server-side—similar in spirit to 
> the existing UIMA {{OpenNlpTextAnalyzer}} composite, but as a 
> language-neutral contract.
> *Package naming (proposed):* {{org.apache.opennlp.grpc.v1}}
> h3. Non-goals (v1 RFC)
>  * Binary/PDF document parsing (Tika, etc.) — callers supply {{raw_text}}
>  * Training, evaluation, or model-update RPCs
>  * Embedding {{.bin}} model bytes in request messages (models remain 
> server-side)
>  * Authentication / multi-tenancy in the core API (deployment concern: mTLS, 
> reverse proxy)
>  * Coreference (documented in manual but not implemented in current codebase)
> h3. Compatibility
>  * *Additive* Maven modules; no breaking changes to {{opennlp-api}} / 
> {{opennlp-runtime}}
>  * Sandbox granular services may be deprecated or moved to 
> {{opennlp.legacy.v1}} after migration
> h3. Phased delivery (high level)
> ||Phase||Scope||
> |*0*|This JIRA + community RFC (this ticket)|
> |*1*|Design document + full {{.proto}} definitions (server coded along the 
> way for demo but open for changes - phase 2 attempts to lock in the proto)|
> |*2+*|Implementation: orchestrator, server, tests, graduation from sandbox to 
> main repo|
> |*Later*|GPU embeddings ({{{}opennlp-dl-gpu{}}}), optional OpenVINO/DJL 
> inference backends|
> h3. Design highlights
>  # *Three proto layers (NLP-only):* domain types ({{{}OpenNlpDocument{}}}), 
> pipeline config ({{{}AnalysisProfile{}}}), service 
> ({{{}OpenNlpAnalysisService{}}})
>  # *Offset contract:* All exported spans use *character offsets in the 
> original {{raw_text}}* ({{{}CHAR_DOCUMENT{}}}), half-open {{[start, end)}} 
> matching {{opennlp.tools.util.Span}}
>  # *Model bundles:* Replace per-RPC {{model_hash}} with {{ModelBundleRef}} + 
> server-defined profiles (reuse sandbox model discovery patterns)
>  # *Thread safety:* Leverage OpenNLP 3.0 thread-safe {{*ME}} instances cached 
> per model bundle
> h3. Sample protobuf (illustrative — full spec in design doc)
> The following is a *sketch* for discussion; field numbers and optional 
> messages may change during RFC.
> {code:java|title=opennlp.proto}
> syntax = "proto3";
> package org.apache.opennlp.grpc.v1;
> option java_package = "org.apache.opennlp.grpc.v1";
> option java_multiple_files = true;
> // --- Layer 1: Document ---
> message OpenNlpDocument {
>   string doc_id = 1;
>   string raw_text = 2;
>   optional string detected_language = 3;
>   optional float language_confidence = 4;
>   repeated AnnotatedSentence sentences = 5;
>   map<string, string> metadata = 6;
> }
> message AnnotatedSentence {
>   CharSpan sentence_span = 1;
>   repeated Token tokens = 2;
>   repeated NamedEntity entities = 3;
> }
> message Token {
>   string text = 1;
>   CharSpan char_span = 2;
>   optional string pos_tag = 3;
> }
> message NamedEntity {
>   CharSpan char_span = 1;
>   string entity_type = 2;
>   optional double prob = 3;
> }
> message CharSpan {
>   int32 start = 1;
>   int32 end = 2;
>   CoordinateSpace space = 3;
>   optional string type = 4;
>   optional double prob = 5;
> }
> enum CoordinateSpace {
>   COORDINATE_SPACE_UNSPECIFIED = 0;
>   CHAR_DOCUMENT = 1;
> }
> // --- Layer 2: Pipeline ---
> enum PipelineStep {
>   PIPELINE_STEP_UNSPECIFIED = 0;
>   LANGUAGE_DETECT = 1;
>   SENTENCE_DETECT = 2;
>   TOKENIZE = 3;
>   POS_TAG = 4;
>   NER = 5;
> }
> message AnalysisProfile {
>   string profile_id = 1;
>   repeated PipelineStep steps = 2;
>   ModelBundleRef model_bundle = 3;
> }
> message ModelBundleRef {
>   string bundle_id = 1;
> }
> message AnalysisOptions {
>   bool include_probabilities = 1;
>   bool clear_adaptive_data = 2;
> }
> // --- Layer 3: Service ---
> service OpenNlpAnalysisService {
>   rpc AnalyzeDocument(AnalyzeDocumentRequest) returns 
> (AnalyzeDocumentResponse);
>   rpc GetServiceInfo(GetServiceInfoRequest) returns (GetServiceInfoResponse);
> }
> message AnalyzeDocumentRequest {
>   OpenNlpDocument document = 1;
>   AnalysisProfile profile = 2;
>   AnalysisOptions options = 3;
> }
> message AnalyzeDocumentResponse {
>   OpenNlpDocument document = 1;
>   repeated ProcessingDiagnostic diagnostics = 2;
> }
> message ProcessingDiagnostic {
>   PipelineStep step = 1;
>   string message = 2;
>   DiagnosticSeverity severity = 3;
> }
> enum DiagnosticSeverity {
>   DIAGNOSTIC_SEVERITY_UNSPECIFIED = 0;
>   INFO = 1;
>   WARNING = 2;
>   ERROR = 3;
> }
> message GetServiceInfoRequest {}
> message GetServiceInfoResponse {
>   string opennlp_version = 1;
>   string api_version = 2;
>   repeated string available_profile_ids = 3;
> }
> {code}
> h3. Comparison: sandbox vs proposed
> ||Aspect||Sandbox POC||Proposed||
> |Services|3 (sent / token / POS)|1 primary ({{{}OpenNlpAnalysisService{}}})|
> |I/O|Strings + {{StringList}}|{{OpenNlpDocument}}|
> |Models|{{model_hash}} per RPC|{{ModelBundleRef}} + profiles|
> |Pipeline|Client-side chaining|Server-side {{AnalysisProfile}}|
> |Package|{{package opennlp}}|{{org.apache.opennlp.grpc.v1}}|
> h3. References
>  * Sandbox POC: 
> [https://github.com/apache/opennlp-sandbox/tree/main/opennlp-grpc]
>  * Current sandbox proto: 
> [https://github.com/apache/opennlp-sandbox/blob/main/opennlp-grpc/opennlp-grpc-api/opennlp.proto]
>  * UIMA composite pipeline: 
> {{opennlp-extensions/opennlp-uima/descriptors/OpenNlpTextAnalyzer.xml}}
>  * ONNX / GPU: {{{}opennlp-dl{}}}, {{{}opennlp-dl-gpu{}}}, 
> {{SentenceVectorsDL}}
>  * Full design document (companion): {{docs/rfc/opennlp-grpc-design.md}} in 
> contributor branch or attachment
> h3. Questions for the community
>  # Should v1 expose *only* {{{}AnalyzeDocument{}}}, or retain sandbox 
> granular RPCs under a legacy package?
>  # Target release: *3.0.x* (additive) vs {*}3.1{*}?
>  # Proto tooling: Maven {{protobuf-maven-plugin}} only, is a gradle build OK?



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to