Unicode
In short: An international standard that assigns a unique number (code point) to every character from practically every writing system in the world.
In more detail: Solves the problem that older character sets (e.g. ASCII) only covered a few characters and were unusable for other languages/symbols (emoji, Cyrillic or Asian characters). How these code points are stored as bytes is governed by encodings like UTF-8 or UTF-16 — Unicode itself only defines the mapping character ↔ number, not the byte format.
In Depth
ASCII (from the 1960s) only encoded 128 characters in 7 bits — enough for the English keyboard, but not for umlauts, accents, or non-Latin scripts at all. Every country/manufacturer then developed its own, incompatible extensions (e.g. ISO-8859-1 for Western Europe), which led to broken characters (“mojibake”, e.g. “ä” instead of “ä”) when exchanging data between systems with different encodings. Unicode solves this by defining ONE single, globally valid number (a code point, written as U+00E4 for “ä”) per character — independent of language, platform, or program.
UTF-8 has become the dominant encoding on the web, because it’s backward-compatible with ASCII: the first 128 characters take up exactly 1 byte (identical to ASCII), further characters (umlauts, emoji, Chinese characters) take 2 to 4 bytes depending on complexity. This makes UTF-8 storage-efficient for predominantly English text while remaining fully compatible with every Unicode character in the world:
"ä".encode("utf-8") # b'\xc3\xa4' - 2 bytes for one character
"A".encode("utf-8") # b'A' - 1 byte, identical to ASCIIA common practical pitfall: the number of Unicode characters in a text doesn’t necessarily correspond to the number of bytes it’s stored in — anyone who miscalculates text lengths or storage space by confusing characters with bytes produces subtle bugs with database column lengths or file uploads containing special characters.
See also: String