https://bz.apache.org/SpamAssassin/show_bug.cgi?id=8429

Kent Oyer <[email protected]> changed:

           What    |Removed                     |Added
----------------------------------------------------------------------------
   Attachment #6099|0                           |1
        is obsolete|                            |

--- Comment #5 from Kent Oyer <[email protected]> ---
Created attachment 6100
  --> https://bz.apache.org/SpamAssassin/attachment.cgi?id=6100&action=edit
Patch with lenient transcoding

Thanks for the upvotes. There is one minor change I need to make. Malicious
JavaScript is frequently obfuscated by splitting the code into chunks and
concatenating them back together. With UTF-16 source code, sometimes the chunks
are split in the middle of a surrogate pair. That leaves a lone surrogate on
both ends. If we transcode with FB_CROAK then it fails and falls back to
Windows-1251 and the whole thing is corrupted. 

The solution is to do lenient transcoding in cases where we're fairly certain
the data is UTF-16 (BOM or NUL byte pattern). That replaces the lone surrogates
with the Unicode replacement character (U+FFFD) but keeps most of the code
intact. 

If we're not sure about the charset, then still use FB_CROAK. 

It's a small change, plus another test case. I'll go ahead and commit, but if
it's problem let me know.

--- a/lib/Mail/SpamAssassin/Message/Node.pm
+++ b/lib/Mail/SpamAssassin/Message/Node.pm
@@ -650,13 +650,20 @@ sub _normalize {
     # Failing that, a UTF-16 label is trusted: its byte order if it names one,
     # else big-endian (RFC 2781).  Undeclared text that is not UTF-16 is left
     # to the detection fallbacks below.
+    #
+    # With a BOM or the NUL pattern the data is known to be UTF-16, so decode
+    # leniently: a lone surrogate (legal in a JavaScript string, e.g. a pair
+    # split across two string literals) becomes U+FFFD instead of failing the
+    # whole part.  A label alone is weaker evidence, so that decode stays
strict.
     my $decoder = detect_utf16( $_[0] );
+    my $check = Encode::FB_DEFAULT;
     if (!defined $decoder && $charset_declared ne '') {
       $decoder = Encode::find_encoding(
                    $charset_declared =~ /LE\z/i ? 'UTF-16LE' : 'UTF-16BE');
+      $check = Encode::FB_CROAK;
     }
     if (defined $decoder) {
-      if (eval { $rv = $decoder->decode($_[0], Encode::FB_CROAK |
Encode::LEAVE_SRC); defined $rv }) {
+      if (eval { $rv = $decoder->decode($_[0], $check | Encode::LEAVE_SRC);
defined $rv }) {
         dbg("message: decoded as charset %s, declared %s",
           $decoder->name, $charset_declared);
         utf8::encode($rv) if !$return_decoded;

-- 
You are receiving this mail because:
You are the assignee for the bug.

Reply via email to