EMZETT.
Login

Strings

In short: A data type for text — a sequence of characters, treated as a reference type in most languages, even though it often feels like a simple value in everyday use.

In more detail: In many modern languages, strings are immutable — every apparent “change” (e.g. converting to uppercase) actually creates a completely new string internally, instead of changing the existing one. This makes strings safe for concurrent access by several threads, but can become inefficient for very many consecutive changes (concatenations).

In Depth

text = "hello"
uppercase = text.upper()
print(text)              # still "hello" - the original stays unchanged!
print(uppercase)  # "HELLO" - a NEW string was created

The immutability of strings is a deliberate design decision, not a technical limitation: because a string never changes after its creation, several parts of a program (or several threads) can safely share the same string, with no risk of one accidentally affecting the other — there’s simply no way for someone to change the content “behind the scenes” while someone else is currently reading it.

A string internally consists of a sequence of individual characters, which can be addressed via an index (usually starting at 0):

text = "Emzett"
text[0]     # "E"
text[-1]    # "t" (last character, possible in many languages with a negative index)
len(text)   # 6 - the length of the string

Important for international text: not every “visible character” necessarily corresponds to exactly one storage unit — some characters (emojis, certain accent combinations) are internally composed of several so-called code units, which can lead to surprising results for length calculations and index access with such characters in some languages. For most everyday Latin text this doesn’t matter, but it becomes relevant as soon as emojis or certain non-Latin writing systems are involved.

See also: String Concatenation, Special Characters, Non-primitive Types