https://bz.apache.org/SpamAssassin/show_bug.cgi?id=8429

            Bug ID: 8429
           Summary: Text-based handlers see raw bytes even when
                    normalize_charset is on
           Product: Spamassassin
           Version: 4.0.3
          Hardware: PC
                OS: Mac OS X
            Status: NEW
          Severity: normal
          Priority: P2
         Component: Libraries
          Assignee: [email protected]
          Reporter: [email protected]
  Target Milestone: Undefined

A UTF-16 .js file reached the JavaScript handler NUL-interleaved, because the
JavaScript handler called decode() instead of decode_and_normalize(), as the
HTML handler already did. The same applies to the other text-based handlers
(SVG & ICS).

_normalize() had to be expanded to support an undefined charset, because
handlers can return parts with no charset (e.g. a .js file inside a RAR
archive). decode_and_normalize() now passes an undefined charset through
instead of substituting us-ascii (the us-ascii default is kept for
normalize_charset 0).

Changes:

- Handlers that process text (JavaScript, SVG, ICS) now call
decode_and_normalize() and encode the result back to UTF-8 bytes for rules,
URIs and child parts. Binary handlers (PDF, Image, Archive) still call
decode(). SVG sets utf8_mode when given bytes, as the HTML handler does, so
entities decode to UTF-8.

- _normalize() accepts an undefined charset. For undeclared or declared-UTF-16
data it calls detect_utf16() once. Undeclared data that isn't UTF-16 goes to
the existing fallbacks.

- detect_utf16() now decides both whether data is UTF-16 and its byte order,
from the first 1024 bytes: a BOM, or NULs in at least a quarter of the code
units at one byte position and few at the other. Otherwise it returns undef.
Previously it assumed its input was UTF-16 and reported UTF-16BE even for plain
ASCII.

- A part declared as UTF-16 (charset=UTF-16, UTF-16LE or UTF-16BE) with no BOM
or NUL pattern is decoded using the declared byte order, or as big-endian for
plain UTF-16 (RFC 2781). Previously a declared LE/BE was ignored.

- Mail::SpamAssassin::HTML now passes <script> text and non-base64 data: URI
content to child parts as UTF-8 bytes, not Perl characters.

Behavior change:

- Parts with no declared charset that are UTF-16 (BOM or NUL pattern) are now
decoded. Undeclared UTF-16 with almost no ASCII and no BOM (e.g. CJK) is still
not detected.

- A declared UTF-16 part with no BOM or NUL pattern is decoded using the
declared byte order, or as big-endian. Previously the byte order was guessed
heuristically, or the part fell through to the fallbacks.

- With normalize_charset 0, nothing changes: handlers still see the part's
original bytes.

Tests: t/node_utf16.t (detection, stray NULs, the 1024-byte window, declared
UTF-16 charsets), t/handler_javascript_utf16.t, t/handler_decode_normalize.t.
The patch includes their fixtures: t/data/nice/handler_archive_utf16_bom,
handler_archive_utf16_nobom, handler_javascript_utf16_attach,
handler_javascript_html_utf8, handler_svg_utf8, handler_svg_cp1251,
handler_svg_entity and handler_ics_cp1251.

-- 
You are receiving this mail because:
You are the assignee for the bug.

Reply via email to