On 8/08/26 4:33 pm, Veek M wrote:
it's the raw utf-8 encoded as bytes but since it is a unicode string why
doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make
it '\\x'

Because you decoded it as ascii with backslashreplace. It's replacing each byte that's outside the ascii range with four characters: a backslash, an 'x', and two hex digits. The backslashes are being doubled when you print the string and its repr() gets computed.

Since the file is actually utf-8 and not ascii, that's the appropriate way to decode it:

  fh = open('/tmp/x', 'rt', encoding = 'utf-8')

Then your ffff should come through as a single character in the string and print as '\uffff'.

--
Greg



--
https://mail.python.org/mailman3//lists/python-list.python.org

Reply via email to