Integrity
In short: The property that data hasn’t been changed (unnoticed) since it was created or last changed in an authorised way.
In more detail: Integrity is usually checked technically via hash values: if the hash of the received data matches the expected hash, the data wasn’t changed. Combined with a signature, authenticity can also be proven (not just “unchanged”, but also “really from the stated sender”). Integrity is one of the three classic IT security protection goals alongside confidentiality and availability.
In Depth
A simple example shows how integrity checks work in practice — for instance when downloading software:
1. The provider publishes the file + its SHA-256 hash separately (e.g. on the website)
2. The user downloads the file
3. The user calculates the hash of the downloaded file locally:
$ sha256sum program.zip
4. Comparison: does the calculated hash match the published one?
Yes -> the file arrived unchanged
No -> the file was damaged OR manipulated (e.g. replaced by malware)
A pure hash check alone, however, only protects against UNINTENTIONAL changes (transmission errors) and against manipulation by someone who can change the file but NOT also manipulate the published hash. An active attacker who controls both the file and the stated hash (e.g. with a compromised download server) can simply swap both at the same time. That’s why a message authentication code (MAC) or a digital signature is usually used for real security: these additionally link the integrity check to a secret key that only the legitimate sender has — an attacker can then change the file, but can’t create a valid new MAC/valid new signature for it.
In network protocols such as TLS, integrity is checked automatically and continuously for every single data packet — not just once for a download, but throughout the entire connection.
Integrity in databases and distributed systems
Integrity isn’t just a networking concept, but also central to databases and distributed systems: a database system has to ensure that simultaneous, competing write accesses don’t lead to inconsistent states (e.g. two simultaneous transfers that “overwrite” each other and make money disappear) — mechanisms such as ACID transactions guarantee exactly this kind of integrity at the database level, independently of cryptographic methods. Distributed systems (e.g. blockchains), on the other hand, often use chained hash structures: each new data block contains the hash of the previous block, so that subsequently manipulating any earlier block would inevitably change all subsequent hashes and thus make it immediately recognisable.
Integrity protection for backups
A practical, often overlooked use case: backups should be checked for integrity regularly, not just when they’re created — a backup that has been damaged unnoticed over months (e.g. by a faulty storage area or creeping data corruption) is worthless in an emergency. Well-run backup strategies therefore store an integrity hash with every backup and check it regularly and automatically, instead of only relying on it working when a restore is needed.
Merkle trees for efficient integrity checks of large amounts of data
With very large amounts of data (e.g. millions of files in a software repository), it would be inefficient to recalculate and transmit the hash of the ENTIRE data set on every small change. Merkle trees (also called hash trees) solve this elegantly: data is split into small blocks, each block is hashed individually, neighbouring hashes are hashed again in pairs, until in the end a single “root hash” represents the integrity of the whole structure. If even a tiny part of the data changes, the root hash changes, but only the affected subtrees have to be retransmitted/checked — Git uses exactly this structure internally to guarantee integrity efficiently across huge commit histories, as do BitTorrent and many blockchain systems.
See also: Hashing, Authenticity