On Mon, Oct 5, 2026, at 15:15, Stephen J. Turnbull wrote:
> Bron Gondwana via Mailman-Developers writes:
> 
> > This is my first time posting to this list.  I'm one of the authors
> > on the DKIM2 spec, and I'm keen to have mailman support DKIM2 even
> > in the early drafts.
> 
> Thank you for contacting us.  Though it really would have been nice to
> have been invited to participate before you had enough standard to
> pretend coding implementations was useful.  We really got blind-sided
> by DMARC, especially the part where they repurposed it from anti-
> phishing to anti-spam.  I mean, I personally was sleeping 2-3 hours a
> night for several days level of blind-sided, and I doubt I was the
> only one.

Yes, I'm sorry I should have pulled the Mailman community in earlier.  I didn't 
want to drag you through the really early stages of "we're not even sure which 
headers need to be signed" pre-IETF discussions, and of course since the work 
went to the IETF, everyone is welcome!  You don't need an invite.

> > The proposed changes to Mailman in order to support DKIM2 are
> > focused very much on the last one, the recipes.  Since you have to
> > describe the changes you are making, it is sensible to make the
> > minimum possible changes!
> 
> That's true.  Of course, the absolute minimum difference description
> is "the bytes are all different, trust us, it's OK".  (That's ARC, of
> course.)  My guess is you're not down with that. ;-)

There is a way to do that with DKIM2 (for the body, you can't do valid DKIM2 at 
all without describing all your header field changes), but it's strongly 
discouraged.  The whole point of DKIM2 is to avoid having to trust anyone, 
because you can verify what they did.

> >  *Preserve the original Content-Transfer-Encoding when decorating*
> > 
> >     When a header or footer is added to a single-part text message,
> >     keep the body's existing encoding instead of letting the email
> >     library choose a new one and silently re-encode every line:
> 
> I assume you've already added code to control CTE in email flattening
> in your Mailman fork.  Does this (and other such changes) require
> monkey-patching the Python email package?  

Actually no.  It's purely within decorate.py.  The simpler method would be to 
always use MIME wrapping for DKIM2-signed incoming messages, but there's a 
reason why mailing lists don't just do that with every message anyway.

> If so, you should talk to
> them, too, as DKIM2 is going to piss off a lot of people who massage
> email before passing it on, and it would be nice if ALL THE THINGS had
> email packages that support DKIM2.  Last I looked R David Murray was
> the guy for Python.  Barry Warsaw was a major contributor and would
> know who to talk to.
> 
> In fact I wouldn't be surprised if most of what you describe shouldn't
> be pushed upstream to the Python email package.  Or maybe splitting it
> into a PyPI package would be most efficient, the cPython folks can
> incorporate it at their leisure.

The existing wrapper code in decorate.py explicitly takes any existing CTE and 
rewrites the part as unicode.  I GUESS we could have the email package have a 
mode which took that and back-converted it identically, but that seems very 
inefficient at best.

We have a similar problem in calendaring and contacts, where you can take an 
existing icalendar or vcard file and either perform the minimal surgical change 
on it, or convert the whole thing to an intermediate datastructure, manipulate 
that, and serialise the datastructure back out to a normalised card.  The 
second one is easier in a lot of aways, but creates much larger diffs and tends 
to lose details which the intermediate structure doesn't represent faithfully.

> >     - 7bit/8bit: concatenate the decoded text and store the result
> >       with the same CTE, upgrading 7bit to 8bit only if the result
> >       contains high bytes.
> 
> Unclear.  I guess you are claiming that the pre-concatenation text
> remains byte-for-byte identical, so the changes description just needs
> to know the length of prepended material and of the original, and
> maybe the length of appended material.  (Does it need to say anything
> about the CTE change?)

Yes, the changes description just needs to know the number of lines to copy 
from the output data to recreate the input data; and (if any lines are changed) 
it has to include the input version of that line.  It's basically an "undo" 
diff.

If the CTE is changed, then that's a header diff, and the description of that 
will be included in the header recipe.

> >     - quoted-printable: concatenate at the raw QP level.  Only the
> >       header and footer text is freshly QP-encoded; the original
> >       body lines remain byte-identical, including unnecessarily
> >       quoted characters (=48 for 'H') and non-standard soft line
> >       break positions.
> 
> Do we even keep that undecoded text around?  I seem to recall we
> don't.

Worst case (without adding ANY support to Mailman at all) my wrapper milters 
capture the message on the way in, and on the way out.  You can diff ANY two 
sets of data of course, so you can calculate the diff - but the more that 
Mailman re-encodes and fiddles with things, the bigger the diff.  A lot of this 
pre-cursor work is just reducing the size of the diff to the smallest 
manipulation required.

>   I'm not sure I like the implications of multiplying the
> storage requirements for large messages that way, especially for
> resource-constrained instances (eg, living in a Linode Nanode). 

Do you mean storage requirements for the archive, or storage requirements 
during the message transit?

> And I
> know of at least two sites with lists whose content frequently
> includes large (GBs) attachments.  (Yes, I have seen multimegabyte
> files QP-encoded, although my correspondents are too polite to send me
> GBs of attachments so I can't testify to that level of aggravation.)

That shouldn't cost any more if those parts aren't edited - the lines will just 
be copied verbatim.

> Which reminds me: are senders required to describe their MIME
> structure in sufficient detail so that we can delete prohibited parts
> and translate undesired subtypes (specifically, text/html) to
> preferred subtypes?  I would guess not since the draft I read, like
> DKIM v1, treats the body as a binary blob.

Strictly, DKIM2 treats the body like an array of lines.  We did think about 
doing MIME-structure based recipes. There are some real advantages to 
understanding the structure better, but we concluded that the added complexity 
didn't earn its keep.

But this leads to another point.

> Obviously, it is at best highly undesirable to include those deleted
> parts verbatim in a change description.

Deleting parts is a sin.  A list which strips attachments should probably just 
not participate in DKIM2; or should re-originate the message (strip any 
existing DKIM2, and just start with a hashes-only `Message-Instance: m=1` on 
egress).

Otherwise yes, you need to describe the entire attachment.  At that point your 
choice is either to encode it in the recipe (which will make the header ungodly 
long) or hide it in the MIME postamble.  Either way, the wire size will still 
include the entire size of the attachment, negating most of the benefits of 
stripping it in the first place.

> In any case where a message violates site or list policy but without
> DKIM2 we could massage it into conformance, we will have an option to
> reject it.

Indeed, you always have that possibility!  Or to forward it without adding 
DKIM2 -- at least while the world still allows that.

> >     - base64: decode, concatenate, and re-encode at the line width
> >       the original used.  Base64 is deterministic, so every
> >       complete line before the original's last line is unchanged;
> >       only that last line (its padding disappears) and the appended
> >       footer lines differ.  The message stays single-part instead
> >       of being MIME-wrapped into multipart/mixed.
> 
> >     Any other CTE, or a failure in one of these paths, falls
> >     through to MIME wrapping as before.  The decorate.rst doctest
> >     for mixed-charset messages is updated for the QP-preserving
> >     behaviour.
> 
> To me, this all looks overengineered.  

Yes, it is a little.  It's done that way for a reason, to make sure we minimise 
the amount we change the lines on the wire.  Certainly just going straight to 
MIME wrapping for everything works (it was my first pass); and I'd be very 
happy for that to be an answer that everyone was happy with.  You probably know 
better than me how many lists just MIME-wrap everything.  If that's the future, 
then I'm all in favour.

> How about just check for a
> DKIM2 header?  If so memoize the body verbatim.  If the body was
> compound, you need to do the regular processing, and perhaps reject on
> the basis of "our rules say mutate but DKIM2 can't deal, so please
> reformulate your message to satisfy policy or turn off DKIM2 and we'll
> do it".  Finally, MIME-wrap the original body, and compute the change
> description.  In the cases of text/plain 7bit or 8bit bodies, include
> a note in the MIME prelude that "this message body was MIME-wrapped
> because the sender uses DKIM2, you need a MIME-capable MUA to view
> this as the sender intended."

Works for me.

> If no DKIM2 header, just do our thing.
> 
> > NOTE: this is largely Claude's work,
> 
> Thanks for mentioning that.  It doesn't change our review process or
> standards.
> 
> > with me guiding the output and reviewing the shape, but not the raw
> > code itself.  I'm prepared to spend more time on human review and
> > making this high quality if the project is willing to take my work
> > on it.
> 
> I trust you on the promised human review (but no promises of
> acceptance until it passes our review).
> 
> We're always happy to have people contribute high-quality code that
> improves deliverability or standard conformance, and is in Mailman
> style (nomenclature, module structure, etc.)  We reserve the right to
> require it come wrapped in an option so site or list administrators
> can decide for themselves if it makes their lives better.

Excellent.  I'm not _as_ experienced with Python as I am with Perl or C, but 
I've written bits of it (including a chunk of work with offlineimap back in the 
day).

Right now my branches have both an option to turn on dkim2 support entirely, 
and a per-list switch to control which lists it's enabled for.  I agree that 
this is the best approach; and I'd strongly agree with defaulting it to being 
turned off until DKIM2 is fully standardised.

Bron.

--
  Bron Gondwana, CEO, Fastmail Pty Ltd / Fastmail US LLC
  [email protected]
_______________________________________________
Mailman-Developers mailing list -- [email protected]
To unsubscribe send an email to [email protected]
https://mail.python.org/mailman3/lists/mailman-developers.python.org/
Mailman FAQ: https://wiki.list.org/x/AgA3

Security Policy: https://wiki.list.org/x/QIA9

Reply via email to