On 2014-01-21 07:45, Henri Sivonen wrote : > (In reply to André Pirard from comment #17) >> I think that the first thing for Character Encoding Autodetect to be less >> confusing is ti say what it does. >> Assuming that it means that any indication of a character set is ignored ans >> that it is guessed by the contents... > It means: ... > > How would you make the menu "say" this? Language to [auto]detect encoding for Language to suit [auto-detection] ... something like that
The key hint is to understand that it's a language. I'm the reporter of this bug and you made me discover this explanation after 5 years. I, and everybody according to the bug title, had looked for it all over the place in vain. Pity there are no HTTP links on system menus. Graphic things (No doc. Any questions?) badly need it. Please note that "Universal" is not a language and that I still do not understand what that means. My guess was that it meant utf-8 but what would that mean? it's not a language either. Also, (Off) could be "no encoding auto-detection" to make it very clear what we're about. Note: I did not report this problem at all. See Description and read below. Alexander Sachs changed the subject to mean a problem of his. Strange doings. I opened another bug to be able say what I meant. I was accused of saying things that did not happen (but that some 6 other persons met). I was even accused of tweaking the encoding identification by forcing the encoding of the preceding page in a test. As if the encoding of one page influenced the encoding of the next one. I finally shuddered and turned away to something else. >> Also, picking the character code from the HTTP request is an error because >> the contents of the page MUST specify the encoding, it knows better than an >> Apache server > Indeed, Ruby's Postulate generally holds. Unfortunately, HTTP disagreed and > it's too late to change that, because it would break pages that currently > work due to Ruby's Postulate not being true for them. > http://www.intertwingly.net/slides/2004/devcon/69.html > > And besides, all browser now agree on the precedence of HTTP over > <meta>, so it's not worthwhile to break interoperability. That is wrong. MIME was intended to describe the single character set of a file that does not provide it. HTML self-describes it and can contain many character sets that MIME is unable to describe. It's like saying "he speaks English" of someone who says "Je parle français ik spreek vlaams и я говорю по русский" >> and the browser won't update the page when it's written to a >> file. > Firefox is supposed to if you choose the "complete" option in Save As... Right and you made me notice it. But why is it correct with a "complete" page and surprisingly incorrect with a "HTML" one? In fact, I met so many character handling bugs in my life that I no longer care reporting anything. Like that craze of removing http:// from Firefox URL bar. This caused a tons of bugs and I still have a stock of 12 or so. Why the hell do that when it was going so well, everyone in the street knew what http:// was and started asking why one removed it? >> The only case where character encoding mangling is necessary is when, for >> example, displaying a text file of which the character set is specified >> nowhere > Or when displaying an HTML file whose character encoding is specified > nowhere. :-( Sorry to say that if a HTML file contains no specification it *must* be ISO8859-1. That default has been decided one day and must be respected to remain compatible with existing pages. I was perfectly astounded by the W3C validator which stated that it was using UTF-8 by default. My bug report was that Firefox displayed the wrong character set, and it was probably only when there was no specified character code. I'm not sure that bug is corrected. I see much less such errors, but also less pages without a charset specification. -- You received this bug notification because you are a member of Desktop Packages, which is subscribed to firefox in Ubuntu. https://bugs.launchpad.net/bugs/206884 Title: fuzzy/confusing firefox View -> Character encoding menu semantics Status in The Mozilla Firefox Browser: Fix Released Status in The Great Mass of Obsolete Junk: Won't Fix Status in “firefox” package in Ubuntu: Triaged Bug description: Please note that I am not the reporter of this bug any longer. Alexander Sack is. He changed the title to his own understanding. I personnally understand View -> Character encoding perfectly. What I say is that FF does not always display ISO8859-1 by default. André. I have seen this since long with both Firefox 2.x and 3.0b3. Rem: Please note that I don't say that Firefox always uses the wrong encoding. Please read my followup to see how to reproduce the problem. I display, for example, http://atilf.atilf.fr/tlf.htm Its header is <HEAD> <TITLE> Le Trésor de la Langue Française Informatisé </TITLE> <link rel="stylesheet" type="text/css" href="atilf.css"> </HEAD> Hence, its encoding should be ISO8859-1 by default as it has always been. As the uploaded attachment shows Firefox displays it using UTF-8. In Edit|Preferences|Content|Font & Colors|Default font|Advanced|Character Encoding there's an option named "Default character encoding" documented as follows The character encoding selected here will be used to display pages that do not specify which encoding to use. What's the use of this setting if the default must ALWAYS be ISO8859-1? Otherwise said, what would be the definition of a changing default? It can only cause people to _produce_ the error I describe. Hence, produce confusion. I saw people say that the wrong behavior I describe is caused by a wrong setting. There should obviously be no user setting for a necessary default. How could the heck a user know what default to set in his browser before being able to read a page if the only place it can be said is in that page he could only read by setting the correct default ;-) And this option was left to ISO8859-1 in my browser, of course. Search www.w3.org/TR/html401/charset.html for "default" and you will learn that a HTML document character code that should obviously be specified within the document is designed to be specified in the HTTP header (without saying BTW how it is specified when FTP is used) with ISO8859-1 as the default. Note that this blunder attributed to HTTP servers accused of not being able to detect the character code of files they store or of being misconfigured has been circumvented by introducing a META directive able to provide -- from the HTML document itself -- HTTP header data and hence the character code. But note that this is done without concluding that ISO8859-1 is the default code of META too, and hence of the document, without regard to the following question. Question : how the heck could a HTML "user agent" that ignores the default character set work any better than my posting this if you and I didn't know that we have to use ASCII? Answer : no better than the page display I show in my attachment. And finally, note that if the reliability of the expected result of a standard lies in this phrase : "By combining these mechanisms, an author can greatly improve the chances that, when the user retrieves a resource, the user agent will recognize the character encoding." the conclusion is : "OK, OK, that was only my bad luck again, it's a random game, bug dismissed, Firefox within said specs, I have to try again". Or should we try to see why Firefox didn't display ISO8859-1? I've see browsers do that for years. To manage notifications about this bug go to: https://bugs.launchpad.net/firefox/+bug/206884/+subscriptions -- Mailing list: https://launchpad.net/~desktop-packages Post to : [email protected] Unsubscribe : https://launchpad.net/~desktop-packages More help : https://help.launchpad.net/ListHelp

