EMZETT.
Login

Bytes

In short: A group of 8 bits — the basic unit in which storage sizes and file sizes are given (KB, MB, GB, TB).

In more detail: A byte can represent 256 different values (0–255), for example a single ASCII character. For larger units there are two competing systems: decimal (1 KB = 1,000 bytes, used by hard drive manufacturers) and binary (1 KiB = 1,024 bytes, technically exact, still often shown as “KB” by operating systems) — the cause of many “missing” gigabytes on new storage devices.

In Depth

Why exactly 8 bits?

The byte was historically established as the smallest sensibly addressable storage unit, because 8 bits offer just enough combinations (256) to uniquely represent, for example, a single character in the ASCII character set (letters, digits, punctuation, control characters). Earlier computer systems did experiment with other byte sizes (6-bit or 7-bit bytes occurred historically), but 8 bits became established from the 1960s/70s as a practical compromise between storage efficiency and representability, and has since become a practically universal standard. Modern character encodings such as Unicode (in the UTF-8 variant) use several bytes in a row for many international characters (e.g. umlauts or emoji need 2-4 bytes), but remain byte-based backward-compatible with ASCII for the classic English character set — plain English text looks byte-identical in ASCII and UTF-8.

Decimal vs. binary: KB vs. KiB

The confusion between decimal and binary prefixes has a concrete technical cause: memory chips are internally organised in powers of two (2^10 = 1024 is the natural “kilo” unit for binary systems, since memory addresses are binary-coded), while hard drive manufacturers traditionally calculate with decimal powers of ten (1000 = 1 KB), because that produces bigger, more marketing-friendly numbers on the packaging — a “1 TB” hard drive calculated decimally contains less storage than an identically named size calculated binarily. Since the 1990s there have officially been separate, IEC-standardised units for this — KiB/MiB/GiB/TiB (binary, exact steps of 1024) versus KB/MB/GB/TB (decimal, exact steps of 1000). In practice, however, both terms are still often mixed up cheerfully to this day (operating systems often display “GB” but actually mean “GiB”), which is why a hard drive sold as “1 terabyte” usually only shows just under 931 GiB in the operating system — not fraud, simply two different counting conventions colliding.

Typical orders of magnitude

For reference: a single text character corresponds to about 1 byte, an average text message to a few hundred bytes, a high-resolution photo to several megabytes, a feature film in good quality to several gigabytes, and the entire data holdings of large data centres are measured in petabytes or exabytes (each 1000 times the previous unit).

Memory alignment and addressing

In computer architecture, the byte is also the smallest individually addressable unit in main memory — every memory address points to exactly one byte, even though a processor often internally reads several bytes at once (e.g. 8 bytes at a time on a 64-bit architecture). Data types that occupy more than one byte (e.g. a 4-byte integer) are, for performance reasons, usually placed at addresses divisible by their own size (“memory alignment”) — an unaligned access can, on some architectures, even cause a program error, or at least be noticeably slower, since the processor then has to perform two separate memory accesses instead of a single one.

Endianness: byte order

When a value occupies several bytes, the additional question arises of in which order these are stored in memory — known as “endianness”. On little-endian systems (most modern PC processors), the least significant byte comes first; on big-endian (e.g. defined as the standard in many network protocols), the most significant byte comes first. This isn’t a purely academic detail: if raw data is exchanged between systems with different endianness (for example when directly reading a binary file or network payload) without accounting for byte order, seemingly arbitrarily wrong numeric values result — a classic, hard-to-find programming bug in low-level network or file-format code.

See also: Bits, Unicode