Scraping

Character Encodings in Scraped Data: Detecting Charsets and Fixing Mojibake

We fetched 193 top homepages in six countries. A common Python default garbled the text of 37 of them. How to decode scraped pages the way browsers do.

Chris Collins

Chris Collins

October 1, 2026 · 9 min read

Mojibake is the garbled text you get when bytes are decoded with the wrong character set: “Café” instead of “Café”, or a Japanese title turned into a row of stray Latin letters. In a browser it is rare, because browsers follow a precise, standardised procedure to work out a page’s encoding. In scraping pipelines it is common, because the HTTP library usually does something simpler.

We measured how often that matters on real sites, and the answer was more often than most teams assume. This guide covers what we found, why it happens, and decoding code that follows the same order browsers do.

Key takeaways

  • On 1 October 2026 we fetched the homepages of the top 50 domains in six country domains (Japan, Korea, China, Russia, Taiwan and Germany). 193 returned a page.
  • Nine of the 193 were not UTF-8: Shift_JIS in Japan, EUC-KR in Korea, windows-1251 in Russia and ISO-8859-1 in Germany.
  • The bigger problem was not legacy encodings but a library default. Python’s requests decoded 37 of the 193 pages, about 19%, as ISO-8859-1, garbling every non-ASCII character, because the server did not name a charset in its header.
  • Legacy labels do not mean what their names say. Browsers decode “ISO-8859-1” as windows-1252 and “Shift_JIS” as the Windows variant, and strict decoders get some characters wrong.
  • Decode in the browser’s order: byte order mark, HTTP header, meta tag, then validity and detection.

What we measured

We took the top 50 domains ending in .jp, .kr, .cn, .ru, .tw and .de from the Tranco list of popular domains, fetched each homepage once, and kept the 193 that returned an HTML page with status 200. For each page we recorded how it declared its encoding, what it actually was, and what the requests library’s .text produced.

Country domainPagesNot UTF-8Garbled by requests’ .text
Japan (.jp)364 (Shift_JIS)8
Korea (.kr)312 (EUC-KR)4
China (.cn)33010
Russia (.ru)252 (windows-1251)4
Taiwan (.tw)2906
Germany (.de)391 (ISO-8859-1)5
Total193937

UTF-8 has clearly won: 184 of 193 pages used it. But the right-hand column shows that winning the encoding war has not ended the problem. Most of the garbled pages were UTF-8. They declared it in a <meta> tag rather than the HTTP header, and the library never looked there.

Why the library got it wrong

When a response says Content-Type: text/html with no charset, requests sets the encoding to ISO-8859-1, following the old HTTP/1.1 default for text types. It does not read the page’s <meta charset> tag. Every byte above 127 is then decoded as a Latin-1 character, so UTF-8 text in any language comes out garbled:

OriginalBytes decoded as Latin-1
CaféCafé
価格 (Japanese, “price”)ä¾¡æ ¼
JRA日本中央競馬会 (a Shift_JIS page)JRA followed by control characters and stray Latin letters

Nothing raises an error. The text is still a valid string, it still parses as HTML, and selectors still match. The damage shows up later, as unsearchable product names, broken deduplication and garbage in a model’s training data, which makes it a classic case of a successful request returning the wrong content.

Legacy labels do not mean what they say

The second trap is subtler. When a page does declare a legacy encoding, browsers do not use the strict standard of that name. The WHATWG Encoding Standard, which browsers implement, maps several labels to supersets:

Declared labelWhat browsers actually decode
iso-8859-1, latin1, asciiwindows-1252
gb2312, gbkgb18030
shift_jisShift_JIS with the Windows extensions (code page 932)
euc-krEUC-KR with the Windows extensions (code page 949)
big5Big5 with the Hong Kong supplementary characters

Decoding with the strict codec of the same name produces errors or, worse, different characters. On one Japanese entertainment site that declared Shift_JIS, Python’s strict shift_jis codec turned every fullwidth tilde ”~” (U+FF5E) into a wave dash ”〜” (U+301C), as many times as the page used it. The Encoding Standard’s index maps that byte pair to the fullwidth tilde, which is what visitors see. The two look almost identical and compare as different strings, so a price range like “1,000~2,000” silently stops matching.

The other mappings fail more loudly. A page labelled gb2312 that uses a character outside that standard, or an euc-kr page with one of the extended Hangul syllables, raises a decoding error under Python’s strict codecs. And a page labelled ISO-8859-1 that uses the euro sign gets an invisible control character instead, because the euro exists only in windows-1252.

Decode the way browsers do

The HTML Standard defines the order: a byte order mark wins, then the encoding named by the transport layer (the HTTP header), then a <meta> declaration found by scanning the start of the document, which authors are required to place within the first 1024 bytes. Following the same order, with the browser’s label mappings and a detection fallback, gives this:

import codecs
import re

from charset_normalizer import from_bytes

# Legacy labels mean what browsers decode them as (WHATWG Encoding Standard), not their strict namesakes.
BROWSER_DECODER = {
    "iso-8859-1": "cp1252", "latin1": "cp1252", "ascii": "cp1252", "us-ascii": "cp1252",
    "gb2312": "gb18030", "gbk": "gb18030", "x-gbk": "gb18030",
    "shift_jis": "cp932", "sjis": "cp932", "x-sjis": "cp932", "windows-31j": "cp932",
    "euc-kr": "cp949", "ks_c_5601-1987": "cp949",
    "big5": "big5hkscs",
}
HEADER_CHARSET = re.compile(r"charset\s*=\s*[\"']?([^;\"'\s]+)", re.I)
META_CHARSET = re.compile(rb"""<meta[^>]+charset\s*=\s*["']?\s*([A-Za-z0-9_.:-]+)""", re.I)


def python_codec(label):
    """Map a declared charset label to a Python codec name, or None if it is unknown."""
    label = label.strip().lower()
    label = BROWSER_DECODER.get(label, label)
    try:
        return codecs.lookup(label).name
    except LookupError:
        return None


def decode_html(body, content_type=""):
    """Decode HTML bytes in the browser's order: BOM, HTTP header, <meta>, then UTF-8, then detection."""
    if body.startswith(codecs.BOM_UTF8):
        return body[3:].decode("utf-8", "replace"), "utf-8 (BOM)"
    header = HEADER_CHARSET.search(content_type or "")
    meta = META_CHARSET.search(body[:1024])  # the HTML spec requires the declaration there
    labels = (("header", header and header.group(1)), ("meta", meta and meta.group(1).decode("ascii", "replace")))
    for source, label in labels:
        codec = label and python_codec(label)
        if codec:
            return body.decode(codec, "replace"), f"{codec} ({source})"
    try:
        return body.decode("utf-8"), "utf-8 (valid)"
    except UnicodeDecodeError:
        guess = from_bytes(body).best()
        codec = guess.encoding if guess else "cp1252"
        return body.decode(codec, "replace"), f"{codec} (detected)"

Pass it the raw bytes (response.content in requests) and the Content-Type header, never response.text. It returns the text and a note of how the encoding was chosen, which is worth storing alongside each page so that a bad decision can be traced later.

Run over all 193 pages, it decoded 192 without a single replacement character. The remaining page declared UTF-8 and contained two invalid byte sequences, which a browser shows as replacement characters too. It also produced the correct text for the 37 pages the library’s default had garbled, and for the 9 pages in legacy encodings. Tested on constructed edge cases, it decodes a gb2312-labelled page containing a GBK-only character, an euc-kr page with an extended Hangul syllable, an ISO-8859-1 page with a euro sign and an undeclared Shift_JIS page, all without errors.

What else turned up

  • Declarations in the wrong place. In our first pass, 17 pages put their <meta charset> after the first 1024 bytes, against the HTML specification. Thirteen also named the charset in the header, and the other four were valid UTF-8, so nothing broke, but a decoder that relied on the meta tag alone would have missed it.
  • Conflicting declarations. One Taiwanese site sent UTF-8 in the header and Big5 in its meta tag. The header wins, as it does in browsers, and in this case it was right.
  • Byte order marks. Four pages started with a UTF-8 byte order mark. Some with no charset in the header were still decoded as Latin-1 by the library, and even where the header was correct, the library left the invisible mark at the start of the text.

Practical rules

  • Decode bytes yourself. Keep the raw response, and if you archive raw responses, you can re-decode them later if your rules change.
  • Never assume Latin-1. Treat a missing header charset as “look further”, not as ISO-8859-1.
  • Store everything as UTF-8. Decode once at the edge of the pipeline, record which encoding was used, and keep UTF-8 from then on.
  • Watch for replacement characters. A sudden rise in U+FFFD in a field is a cheap, reliable alarm, and belongs next to the checks described in schema drift in scraped data.
  • Normalise after decoding. Correct characters can still be written in more than one way, such as fullwidth digits or different tildes. Normalise before comparing or deduplicating, as covered in normalising prices, numbers and dates across locales.

The bottom line

Nearly the whole web is UTF-8 now, and pipelines still garble it, because the commonest failure is not an exotic encoding but a library default that ignores the page’s own declaration. On our sample, that default garbled about one page in five, and none of them raised an error.

Decode from bytes, in the order browsers use, with the labels mapped the way browsers map them, and record which rule decided. That turns an invisible data-quality problem into a solved one.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started