The Question :
314 people think this question is useful
Why is the below item failing? Why does it succeed with “latin-1” codec?
o = "a test of \xe9 char" #I want this to remain a string as this is what I am receiving
v = o.decode("utf-8")
Which results in:
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
line 16, in decode
return codecs.utf_8_decode(input, errors, True) UnicodeDecodeError:
'utf8' codec can't decode byte 0xe9 in position 10: invalid continuation byte
The Question Comments :
The Answer 1
284 people think this answer is useful
In binary, 0xE9 looks like
1110 1001. If you read about UTF-8 on Wikipedia, you’ll see that such a byte must be followed by two of the form
10xx xxxx. So, for example:
But that’s just the mechanical cause of the exception. In this case, you have a string that is almost certainly encoded in latin 1. You can see how UTF-8 and latin 1 look different:
(Note, I’m using a mix of Python 2 and 3 representation here. The input is valid in any version of Python, but your Python interpreter is unlikely to actually show both unicode and byte strings in this way.)
The Answer 2
306 people think this answer is useful
I had the same error when I tried to open a CSV file by
The solution was change the encoding to
pd.read_csv('ml-100k/u.item', sep='|', names=m_cols , encoding='latin-1')
The Answer 3
66 people think this answer is useful
It is invalid UTF-8. That character is the e-acute character in ISO-Latin1, which is why it succeeds with that codeset.
If you don’t know the codeset you’re receiving strings in, you’re in a bit of trouble. It would be best if a single codeset (hopefully UTF-8) would be chosen for your protocol/application and then you’d just reject ones that didn’t decode.
If you can’t do that, you’ll need heuristics.
The Answer 4
46 people think this answer is useful
Because UTF-8 is multibyte and there is no char corresponding to your combination of
\xe9 plus following space.
Why should it succeed in both utf-8 and latin-1?
Here how the same sentence should be in utf-8:
'a test of \xc3\xa9 char'
The Answer 5
16 people think this answer is useful
If this error arises when manipulating a file that was just opened, check to see if you opened it in
The Answer 6
12 people think this answer is useful
Use this, If it shows the error of UTF-8
The Answer 7
10 people think this answer is useful
utf-8 code error usually comes when the range of numeric values exceeding 0 to 127.
the reason to raise this exception is:
1)If the code point is < 128, each byte is the same as the value of the code point.
2)If the code point is 128 or greater, the Unicode string can’t be represented in this encoding. (Python raises a UnicodeEncodeError exception in this case.)
In order to to overcome this we have a set of encodings, the most widely used is “Latin-1, also known as ISO-8859-1”
So ISO-8859-1 Unicode points 0–255 are identical to the Latin-1 values, so converting to this encoding simply requires converting code points to byte values; if a code point larger than 255 is encountered, the string can’t be encoded into Latin-1
when this exception occurs when you are trying to load a data set ,try using this format
Add encoding technique at the end of the syntax which then accepts to load the data set.
The Answer 8
7 people think this answer is useful
This happened to me also, while i was reading text containing Hebrew from a
file -> save as and I saved this file as a
The Answer 9
4 people think this answer is useful
Well this type of error comes when u are taking input a particular file or data in pandas such as :-
Then the error is displaying like this :-
UnicodeDecodeError: ‘utf-8’ codec can’t decode byte 0xf4 in position 1: invalid continuation byte
So to avoid this type of error can be removed by adding an argument
The Answer 10
-1 people think this answer is useful
In this case, I tried to execute a .py which active a path/file.sql.
My solution was to modify the codification of the file.sql to “UTF-8 without BOM” and it works!
You can do it with Notepad++.
i will leave a part of my code.
con=psycopg2.connect(host = sys.argv,
port = sys.argv,dbname = sys.argv,user = sys.argv, password = sys.argv)
cursor = con.cursor()
sqlfile = open(path, ‘r’)