Source: tesseract Version: 5.5.0-1.1 Severity: important Tags: security upstream X-Debbugs-Cc: [email protected], Debian Security Team <[email protected]>
Hi, The following vulnerabilities were published for tesseract. CVE-2026-88047[0]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, Classify::ReadNormProtos in src/classify/normmatch.cpp | parses the NORMPROTO component of a .traineddata file and uses | std::istream::operator>>(char*) to extract a whitespace-delimited | token into a fixed 61-byte stack buffer without setting a stream | width. The 100-byte line buffer can carry a token of up to 99 | characters, so a token longer than 60 characters writes up to 39 | attacker-controlled bytes past the buffer during TessBaseAPI::Init | of the legacy engine, causing stack corruption, denial of service, | and potentially control-flow hijacking on affected standard-library | implementations. Builds using Apple's libc++ C++20 bounded array | overload are incidentally protected, while typical libstdc++ builds | remain affected. No fixed release is available as of this review. CVE-2026-88048[1]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, FullyConnected::DeSerialize in src/lstm/fullyconnected.cpp | does not validate the deserialized layer scalars ni_ and no_ against | the weight-matrix dimensions. During FullyConnected::Forward, | MatrixDotVector in src/lstm/weightmatrix.cpp writes w.dim1() results | into temp_line, which is sized from no_, and reads w.dim2() minus | one inputs from curr_input, which is sized from ni_. A crafted | .traineddata NT_SOFTMAX layer can therefore use inconsistent | dimensions to cause a heap out-of-bounds write and read on the | default LSTM engine, resulting in heap corruption, a crash, | information disclosure, or potentially controlled corruption. No | fixed release is available as of this review. CVE-2026-88049[2]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, prior .traineddata hardening added bounds checks to | NetworkIO::CopyTimeStepGeneral and NetworkIO::Randomize in | src/lstm/networkio.cpp but left NetworkIO::WriteTimeStepPart and | NetworkIO::AddTimeStepPart unchecked. In LSTM::Forward in | src/lstm/lstm.cpp, source_ is sized from the independently | deserialized na_ field while the WriteTimeStepPart count is ns_, | which comes from the CI gate WeightMatrix dim1() value. A crafted | NT_LSTM layer can make ns_ much larger than na_, causing a heap out- | of-bounds write during the first recognition step on the default | LSTM engine and resulting in heap corruption, a crash, or | potentially controlled corruption. No fixed release is available as | of this review. CVE-2026-88050[3]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, RecodedCharID::DeSerialize in src/ccutil/unicharcompress.h | validates length_ but accepts negative code_ values from a crafted | .traineddata recoder component. UnicharCompress::ComputeCodeRange in | src/ccutil/unicharcompress.cpp can consequently produce code_range_ | equal to zero, after which SetupDecoder indexes is_valid_start_ with | the negative code on a size-zero vector. The resulting out-of-bounds | bit write uses a large wrapped index and reliably causes a wild- | address crash or allocation failure on the default LSTM engine. No | fixed release is available as of this review. CVE-2026-88051[4]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, the callback form of GenericVector::read in | src/ccutil/genericvector.h reads the independent int32 fields | reserved and size_used_ from a .traineddata model without a cap or | an invariant check. reserve(reserved) allocates the backing array, | but the callback loop writes size_used_ elements. A crafted | TESSDATA_INTTEMP component with version_id 4 or later can therefore | set reserved to a small value and size_used_ to a large value when | fontinfo_table_.read(fp, read_info) is called from | src/classify/intproto.cpp, causing a heap out-of-bounds write of | FontInfo structures, heap corruption, a crash, or potentially | controlled corruption. No fixed release is available as of this | review. CVE-2026-88052[5]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, UNICHARSET::load_via_fgets in src/ccutil/unicharset.cpp | trusts the declared unichar count as a loop bound and uses id as an | unchecked index into the unichars vector. | unichar_insert_backwards_compatible can leave the vector unchanged | for an empty, duplicate, or already-encodable representation, | causing id to become larger than unichars.size(). Subsequent set_* | calls and the write to unichars[id].properties.enabled then write | UNICHAR_PROPERTIES beyond the vector during initialization in both | the default LSTM and legacy engines, causing heap corruption, a | crash, or potentially controlled corruption. No fixed release is | available as of this review. CVE-2026-88053[6]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, Classify::ReadIntTemplates in src/classify/intproto.cpp | reads NumClassPruners, NumClasses, and NumProtoSets from the | TESSDATA_INTTEMP component of a crafted .traineddata file and uses | those values as loop bounds without validating them against | MAX_NUM_CLASS_PRUNERS, MAX_NUM_CLASSES, and MAX_NUM_PROTO_SETS. The | loops store heap pointers into fixed-capacity ClassPruners and | ProtoSets arrays in INT_TEMPLATES_STRUCT and INT_CLASS_STRUCT, so an | oversized count causes heap out-of-bounds pointer writes during | legacy-classifier initialization before OCR begins, resulting in | heap corruption, a crash, or potentially controlled corruption. No | fixed release is available as of this review. CVE-2026-88054[7]: | Tesseract is an open source OCR engine. In version 5.5.3 and | earlier, Plumbing::DeSerialize in src/lstm/plumbing.cpp rejects | excessively large network stacks but accepts a zero-length stack for | NT_SERIES, NT_PARALLEL, or NT_REVERSED layers in a crafted | .traineddata model. During LSTMRecognizer initialization in | src/lstm/lstmrecognizer.cpp, CacheXScaleFactor(XScaleFactor()) | reaches Series::CacheXScaleFactor in src/lstm/series.cpp, which | dereferences stack_[0] on the empty vector and invokes a virtual | method through an invalid Network pointer. This causes a | deterministic crash and denial of service at model load. No fixed | release is available as of this review. If you fix the vulnerabilities please also make sure to include the CVE (Common Vulnerabilities & Exposures) ids in your changelog entry. For further information see: [0] https://security-tracker.debian.org/tracker/CVE-2026-88047 https://www.cve.org/CVERecord?id=CVE-2026-88047 [1] https://security-tracker.debian.org/tracker/CVE-2026-88048 https://www.cve.org/CVERecord?id=CVE-2026-88048 [2] https://security-tracker.debian.org/tracker/CVE-2026-88049 https://www.cve.org/CVERecord?id=CVE-2026-88049 [3] https://security-tracker.debian.org/tracker/CVE-2026-88050 https://www.cve.org/CVERecord?id=CVE-2026-88050 [4] https://security-tracker.debian.org/tracker/CVE-2026-88051 https://www.cve.org/CVERecord?id=CVE-2026-88051 [5] https://security-tracker.debian.org/tracker/CVE-2026-88052 https://www.cve.org/CVERecord?id=CVE-2026-88052 [6] https://security-tracker.debian.org/tracker/CVE-2026-88053 https://www.cve.org/CVERecord?id=CVE-2026-88053 [7] https://security-tracker.debian.org/tracker/CVE-2026-88054 https://www.cve.org/CVERecord?id=CVE-2026-88054 Regards, Salvatore

