Source: tesseract
Version: 5.5.0-1.1
Severity: important
Tags: security upstream
X-Debbugs-Cc: [email protected], Debian Security Team <[email protected]>

Hi,

The following vulnerabilities were published for tesseract.

CVE-2026-88047[0]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, Classify::ReadNormProtos in src/classify/normmatch.cpp
| parses the NORMPROTO component of a .traineddata file and uses
| std::istream::operator>>(char*) to extract a whitespace-delimited
| token into a fixed 61-byte stack buffer without setting a stream
| width. The 100-byte line buffer can carry a token of up to 99
| characters, so a token longer than 60 characters writes up to 39
| attacker-controlled bytes past the buffer during TessBaseAPI::Init
| of the legacy engine, causing stack corruption, denial of service,
| and potentially control-flow hijacking on affected standard-library
| implementations. Builds using Apple's libc++ C++20 bounded array
| overload are incidentally protected, while typical libstdc++ builds
| remain affected. No fixed release is available as of this review.


CVE-2026-88048[1]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, FullyConnected::DeSerialize in src/lstm/fullyconnected.cpp
| does not validate the deserialized layer scalars ni_ and no_ against
| the weight-matrix dimensions. During FullyConnected::Forward,
| MatrixDotVector in src/lstm/weightmatrix.cpp writes w.dim1() results
| into temp_line, which is sized from no_, and reads w.dim2() minus
| one inputs from curr_input, which is sized from ni_. A crafted
| .traineddata NT_SOFTMAX layer can therefore use inconsistent
| dimensions to cause a heap out-of-bounds write and read on the
| default LSTM engine, resulting in heap corruption, a crash,
| information disclosure, or potentially controlled corruption. No
| fixed release is available as of this review.


CVE-2026-88049[2]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, prior .traineddata hardening added bounds checks to
| NetworkIO::CopyTimeStepGeneral and NetworkIO::Randomize in
| src/lstm/networkio.cpp but left NetworkIO::WriteTimeStepPart and
| NetworkIO::AddTimeStepPart unchecked. In LSTM::Forward in
| src/lstm/lstm.cpp, source_ is sized from the independently
| deserialized na_ field while the WriteTimeStepPart count is ns_,
| which comes from the CI gate WeightMatrix dim1() value. A crafted
| NT_LSTM layer can make ns_ much larger than na_, causing a heap out-
| of-bounds write during the first recognition step on the default
| LSTM engine and resulting in heap corruption, a crash, or
| potentially controlled corruption. No fixed release is available as
| of this review.


CVE-2026-88050[3]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, RecodedCharID::DeSerialize in src/ccutil/unicharcompress.h
| validates length_ but accepts negative code_ values from a crafted
| .traineddata recoder component. UnicharCompress::ComputeCodeRange in
| src/ccutil/unicharcompress.cpp can consequently produce code_range_
| equal to zero, after which SetupDecoder indexes is_valid_start_ with
| the negative code on a size-zero vector. The resulting out-of-bounds
| bit write uses a large wrapped index and reliably causes a wild-
| address crash or allocation failure on the default LSTM engine. No
| fixed release is available as of this review.


CVE-2026-88051[4]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, the callback form of GenericVector::read in
| src/ccutil/genericvector.h reads the independent int32 fields
| reserved and size_used_ from a .traineddata model without a cap or
| an invariant check. reserve(reserved) allocates the backing array,
| but the callback loop writes size_used_ elements. A crafted
| TESSDATA_INTTEMP component with version_id 4 or later can therefore
| set reserved to a small value and size_used_ to a large value when
| fontinfo_table_.read(fp, read_info) is called from
| src/classify/intproto.cpp, causing a heap out-of-bounds write of
| FontInfo structures, heap corruption, a crash, or potentially
| controlled corruption. No fixed release is available as of this
| review.


CVE-2026-88052[5]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, UNICHARSET::load_via_fgets in src/ccutil/unicharset.cpp
| trusts the declared unichar count as a loop bound and uses id as an
| unchecked index into the unichars vector.
| unichar_insert_backwards_compatible can leave the vector unchanged
| for an empty, duplicate, or already-encodable representation,
| causing id to become larger than unichars.size(). Subsequent set_*
| calls and the write to unichars[id].properties.enabled then write
| UNICHAR_PROPERTIES beyond the vector during initialization in both
| the default LSTM and legacy engines, causing heap corruption, a
| crash, or potentially controlled corruption. No fixed release is
| available as of this review.


CVE-2026-88053[6]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, Classify::ReadIntTemplates in src/classify/intproto.cpp
| reads NumClassPruners, NumClasses, and NumProtoSets from the
| TESSDATA_INTTEMP component of a crafted .traineddata file and uses
| those values as loop bounds without validating them against
| MAX_NUM_CLASS_PRUNERS, MAX_NUM_CLASSES, and MAX_NUM_PROTO_SETS. The
| loops store heap pointers into fixed-capacity ClassPruners and
| ProtoSets arrays in INT_TEMPLATES_STRUCT and INT_CLASS_STRUCT, so an
| oversized count causes heap out-of-bounds pointer writes during
| legacy-classifier initialization before OCR begins, resulting in
| heap corruption, a crash, or potentially controlled corruption. No
| fixed release is available as of this review.


CVE-2026-88054[7]:
| Tesseract is an open source OCR engine. In version 5.5.3 and
| earlier, Plumbing::DeSerialize in src/lstm/plumbing.cpp rejects
| excessively large network stacks but accepts a zero-length stack for
| NT_SERIES, NT_PARALLEL, or NT_REVERSED layers in a crafted
| .traineddata model. During LSTMRecognizer initialization in
| src/lstm/lstmrecognizer.cpp, CacheXScaleFactor(XScaleFactor())
| reaches Series::CacheXScaleFactor in src/lstm/series.cpp, which
| dereferences stack_[0] on the empty vector and invokes a virtual
| method through an invalid Network pointer. This causes a
| deterministic crash and denial of service at model load. No fixed
| release is available as of this review.


If you fix the vulnerabilities please also make sure to include the
CVE (Common Vulnerabilities & Exposures) ids in your changelog entry.

For further information see:

[0] https://security-tracker.debian.org/tracker/CVE-2026-88047
    https://www.cve.org/CVERecord?id=CVE-2026-88047
[1] https://security-tracker.debian.org/tracker/CVE-2026-88048
    https://www.cve.org/CVERecord?id=CVE-2026-88048
[2] https://security-tracker.debian.org/tracker/CVE-2026-88049
    https://www.cve.org/CVERecord?id=CVE-2026-88049
[3] https://security-tracker.debian.org/tracker/CVE-2026-88050
    https://www.cve.org/CVERecord?id=CVE-2026-88050
[4] https://security-tracker.debian.org/tracker/CVE-2026-88051
    https://www.cve.org/CVERecord?id=CVE-2026-88051
[5] https://security-tracker.debian.org/tracker/CVE-2026-88052
    https://www.cve.org/CVERecord?id=CVE-2026-88052
[6] https://security-tracker.debian.org/tracker/CVE-2026-88053
    https://www.cve.org/CVERecord?id=CVE-2026-88053
[7] https://security-tracker.debian.org/tracker/CVE-2026-88054
    https://www.cve.org/CVERecord?id=CVE-2026-88054

Regards,
Salvatore

Reply via email to