On Fri, Aug 14, 2026 at 04:42:32PM +0200, Patrice Dumas wrote:
> The second is about easing the implementation.  We convert to UTF-8 upon
> reading and do everything in this encoding, in particular determining
> line width for cutting lines, or centering or flushing right.  In Info
> output, however, we also need to track the bytes count in the output
> encoding to fill the Info tag tables that allows to find anchors and
> nodes in Info file.  Therefore, we do the output in two phases, first
> conversion in UTF-8 and then encoding to determine bytes counts.  In
> some cases it is not easy to do it in two phases, we need to convert
> back already encoded strings to UTF-8 to do the counting.  Outputting
> in UTF-8 only would remove
> * the need to do two phases
> * the need to convert back

Thinking more about the issue, I think that first it is a good thing to
keep the two phases, as one phase for conversion of characters as text,
with possibly dome reorganizations of the text and additions of leading
spaces or end of lines for example, followed by a phase considering the
text as bytes and doing the bytes counting and adding binary quoting
seems to be a good design independently of whether encoding takes place.

I think that it would be good to revisit the design of the code that
needs to add spaces and move text around (@center, @flushright and
@multitable), such that
* when formatting the inside, character text segments and point anchors
  are produced and possibly moved around.
* at the top level, go through the final segments and do the bytes
  counting (and possibly encoding).

I think that it would be a much cleaner design.

Any comment?

-- 
Pat

Reply via email to