Date: Sun, 20 Feb 2000 05:32:55 -0300
From: hammer <[EMAIL PROTECTED]>
Subject: Re: Character conversion

Lars Nordstrom asked some pertinent questions about ReRead and
character conversion:

> Can this extended functionality convert any given, existing file for the 
> purpose of sending it as an e-mail?
> 
Yes, as long it's a text file (8-bit, "extended ASCII"). The only
obnoxious char.s in the source file would be ^Z <g>.

> Will it write to an e-mail in the charset chosen?
>
Yes.  It copies (=appends) the text to any file: either as a whole in
its original form (and char-set) or - marked parts - from the
*displayed* format: the latter only remaps chars.  So it's to chose.

> Will it write to an e-mail in the charset chosen?
>
ReRead doesn't write. It allows to use any editor of your choice.
You write with *your* editor which uses *your* charset.

(Though if you do not want to send it like that, you could indeed pass
it through ReRead to remap the whole set and send the output copy as
mailout.)

> Will it indicate in the header what charset is used?
> 
No.
ReRead is basically an offline *reader* - it has the "reply" function
for creating a reply-mail item (file) which can be used as such with
Nettamer or Netmail Pro. But then it's one of these (and I have a
PMail version on the workbench) which, as mailer clients, do the header
formatting (and it would be easy to add a configuration tag which
could add a header line).  Nettamer for instance would set a header line
for 8-bit chars if one uses the /8 switch.  Default is "US-ASCII", i.e.
IBM/CodePage 437.

> Will the resulting e-mail be able to pass unharmed and unhindered through 
> the transport mechanisms of the internet?
> 
Ha! - good question: *theoretically* (and since about 1991) all net
transport through TCP/IP should be "8-bit transparent" (ITU standard).
For SMTP (mail transport), default is "US-ASCII" too (so this wouldn't
be needed to be "declared" in a header.)

This is true in *practice* too: all 8-bit chars *do* arrive in their
original form at destination.
(The first 32 chars in the *7*-bit range of the hitherto usual
256-charsets are another cattle of fish -which is why there is still the
"binary" distinction for zipped files, for instance).

The problem is entirely with the "ethnocentricity" of applications, and
their rigidity (or outright sloppiness: evidently German language M$ware
does *not* recognize "undeclared" US-ASCII - which I use - as the
standard default but *interpretes* the 8-bit chars in its weird German
way; and as so many Germans don't know of the existence of other
languages and charsets they complain.)

Sure the traditional limit of 256 units (255 in fact, as the "null"
char is unchangeable) for sets of characters is insufficient - and the
solution of UCS/UNICODE is long since on its way, though it will take
another millenium to arrive - but the means for easy and comfortable
choices to remap had been there all the time, and good programmers'
editors do offer simple and instant character remapping.

> As stated in my original post, QP is understood by all mailreaders I know 
> of. And easily converted.

a.) this is a lead-footed workaround (look only at the havoc of
line-breaks; UCS will do a better job with the chars but wouldn't
improve the originally bad, M$void "text" formatting), and

b.) it does *not* solve the underlying problem: QP translates the German
Win$-definition of the left-opening single quotation mark or " ' " from
what it is in *that* char-set definition, namely position 145 of the
256-char set or chr$(145) or 91hex, which gets "=91" in QP; this "=91"
arrives duely at the destination and is back-translated there from QP
into what the *local* position is ... "garble" (in case of the Danish
or Norwegian languages, the ligated "ae", namely indeed "their"
chr$[145] if the US-ASCII char-set is used there, for instance), and
this then should again be converted to chr$(39) to make sense.

But despite of all header declarations, this last step is very often
not, or not correctly, done by many applications.  I see masses of the
resulting garbage on other people's screens around here (Brussels is in
the privileged position to get drowned by the German/Scandinavian/Roman/
Anglo-saxon character salad all at once, not to speak of Greek, and now
even the Slavic languages arrive here).

As the situation (without UCS implementation) is at present, there's
even a "philosophical" aspect involved: to fiddle or not to fiddle
with the "original" char sequence received. Many mailer applications
do not save them in their original state but in an already - more or
less - "converted", if not saladded state. (If you re-send that to the
origninal e-ddress you'd get some complaints from there...) ReRead
doesn't change anything in the original, but allows to "export" the
character-remapped displayed interpretation to whatever you want; and
then do what you want with such *copy*.

All that's needed for this usually is a comparably small remapping table
(for Umlaute in German/Swedish text, accented letters in French) which
is executed fast, and can be expanded/changed ad lib. And a comparably
small effort and addition in coding to use it in ReRead.
QED.

> If the answer is "yes" to these questions I will withdraw the package from
> my website and start using and promoting RR.

"No" and no, for heaven's sake.
And this was not the point of the exercise either.
Rather that I think that we should reflect about a condition which
makes us look for solutions in the range of hundreds of KB instead
ofcoping with the trivial relativity of a 256-byte charset.

BTW, I did a full matrix of 8-bit chars (and some others in the 7-bit
ASCII range) based on the UCS/UNICODE definition for some seven most
common Western European codepages (incl. US-ASCII and Win and Mac, and
a number of "national" printer codes; it's part of the registration
"bonus" for ReRead, as there isn't much difference else to the free
available set). It's some 18 KB and I would consider it a shamefull
bloat to integrate all that into the working configuration of a simple
text reader.  But it's fairly easy to build all sorts of really needed,
short remapping tables with it.  (In my working setup for ReRead I get
along with almost all European mail/text sources with about three dozen
entries to the list in the ini-file; I didn't succeed to measure a time
difference between reading-in the original as-is and a re-mapped
display, as this is below of what I can catch with an ordinary
stopwatch.) Though writing - and correcting - that matrix took not days
but months, and this is outrageous: how many users had wasted their time
to find out, and even to correct manually, only bits and pieces of those
inconsistencies ?!

// Heimo Claasen   //   <[EMAIL PROTECTED]>   //   Brussels 2000-02-19
HomePage of ReRead - and much to read ==> http://www.inti.be/hammer

All answers qeustioned here.



********************************************************
To unsubscribe from this list,
send a message to [EMAIL PROTECTED] with the single word
                     Unsubscribe
as the subject.
You MUST use the same address with which you subscribed!
********************************************************

Reply via email to