A String Is a Sequence — and It Cannot Be Changed
A string is Python's type for text: an ordered sequence of characters. You can write one with single quotes, double quotes or triple quotes, and Python treats all three identically once the line has been read. The only reason to prefer one is convenience — choose the quote character that does not appear inside your text and you never have to escape anything. Triple quotes differ in one useful way: they may span several lines, so the line breaks you type become real newlines in the value.
The property that shapes everything else in this lesson is that strings are immutable. Once a string object exists, its characters can never be modified. Assigning to a position — name[0] = "R" — raises TypeError: 'str' object does not support item assignment. Every method that looks like it edits text actually builds a brand-new string and hands it back, leaving the original exactly as it was.
That single fact explains the most common string bug beginners write. Calling text.strip() on a line by itself does nothing you can observe: the cleaned-up string is created and immediately thrown away. You have to catch the result with text = text.strip(). Whenever a string method seems to have had no effect, this is almost always the reason.
Immutability is a deliberate design choice rather than an oversight. Because a string can never change underneath you, it is safe to use as a dictionary key, safe to hand to a function without worrying that the function will rewrite it, and safe to share between two variables. The price is paid in exactly one situation — repeatedly growing a string inside a loop — and the last section of this lesson shows the standard way around that.
# Three ways to write the same text
a = 'Namaste'
b = "Namaste"
c = """Namaste"""
# Pick the quote that is not inside your text
print("It's raining") # no escaping needed
print('She said "hello"') # no escaping needed
# Triple quotes keep real line breaks
address = """Flat 3B
Green Park
New Delhi"""
print(address)
# Strings are immutable
name = "ravi"
# name[0] = "R" # TypeError: 'str' object does not support item assignment
name = "R" + name[1:] # build a new string instead
print(name) # Ravi
# The classic no-op
text = " spaces "
text.strip() # result created, then thrown away
print(repr(text)) # ' spaces ' unchanged
text = text.strip() # catch the result
print(repr(text)) # 'spaces' repr()shows a value the way you would type it, with quotes and escape characters visible. It is the fastest way to confirm whether a string really does have that trailing space you suspect.
Indexing and Slicing
Characters are numbered from zero, so text[0] is the first one. Negative numbers count backwards from the end: text[-1] is the last character and text[-2] the one before it. That saves writing text[len(text) - 1], which works but is noise. Asking for a position that does not exist raises IndexError.
Slicing takes a range of characters with text[start:stop]. The rule to fix in your memory now — because it applies to lists, tuples and range() as well — is that start is included and stop is excluded. "Python"[0:2] is "Py", not "Pyt". It looks arbitrary until you see the payoff: the length of a slice is simply stop - start, and text[:n] followed by text[n:] always rebuilds the original with no overlap and no gap.
Either end may be left out and Python fills in the obvious default: the start of the string, or the end of it. A third number is the step. text[::2] takes every second character, and text[::-1] walks backwards, which is the idiomatic way to reverse a string. Unlike indexing, slicing never raises an error for an out-of-range number — it simply gives you as much as exists, which is why text[:100] is a safe way to truncate a value of unknown length.
text = "Python"
print(text[0]) # P
print(text[5]) # n
print(text[-1]) # n last character
print(text[-2]) # o
# print(text[6]) # IndexError: string index out of range
print(text[0:2]) # Py stop is EXCLUDED
print(text[2:5]) # tho
print(text[:3]) # Pyt from the beginning
print(text[3:]) # hon to the end
print(text[:]) # Python a full copy
print(text[::2]) # Pto every second character
print(text[::-1]) # nohtyP reversed
# Slices never go out of range
print(text[2:100]) # thon
print(repr(text[50:60])) # '' — empty string, no error
print(len(text)) # 6 text[i]— one character; counting starts at 0text[-1]— the last character;text[-2]the second lasttext[start:stop]— start included, stop excludedtext[start:stop:step]— take everystep-th charactertext[::-1]— the whole string, backwardslen(text)— how many characters it holds
The Methods You Will Actually Use
String methods are called on a string with a dot, and each returns something new. A small handful covers most day-to-day work: trimming whitespace, changing case, replacing a substring, testing how a string begins or ends, and finding where something occurs.
The commonest real use is normalising input before comparing it. Someone typing a roll number or an email address will add a stray space or use different capitalisation, and "CS21B1042 " == "cs21b1042" is False. Strip the whitespace and lower the case once, at the point where the value enters your program, and every comparison after that is honest.
For "does this appear anywhere", use the in operator rather than find(). find() returns the index of the first match, or -1 when there is none, and that -1 is a trap. Writing if text.find("a"): looks like a membership test but is wrong twice over: it is true when the answer is -1 and the substring is absent, and false when the match sits at index 0. Compare explicitly against 0, or just use in. index() does the same job as find() but raises ValueError instead of returning -1, which is the better choice when a missing value genuinely is a bug.
raw = " CS21B1042 "
print(repr(raw.strip())) # 'CS21B1042' both ends
print(repr(raw.rstrip())) # ' CS21B1042' right end only
print(raw.strip().lower()) # cs21b1042
title = "python for placements"
print(title.upper()) # PYTHON FOR PLACEMENTS
print(title.replace("python", "Python"))
email = "asha@college.edu.in"
print(email.endswith(".in")) # True
print(email.startswith("asha")) # True
print(email.count(".")) # 2
# Finding things
print("for" in title) # True <- use this for membership
print(title.find("for")) # 7
print(title.find("java")) # -1
# print(title.index("java")) # ValueError: substring not found
# The -1 trap: "python" is at index 0, and 0 is falsy
if title.find("python"):
print("found") # never runs
if title.find("python") >= 0:
print("found") # correct .strip(),.lstrip(),.rstrip()— remove whitespace (or given characters) from the ends.lower(),.upper()— case conversion, mostly for comparison.replace(old, new)— replaces every occurrence, not just the first.startswith(s),.endswith(s)— clearer than slicing for testing a prefix or suffix.find(s)returns -1 when absent;.index(s)raises ValueError instead.count(s)— how many non-overlapping times it appears.isdigit(),.isalpha()— cheap validation before converting to a number
strip(".com")does not remove the text.com. The argument is a set of characters, so it strips any of.,c,oormfrom both ends until it meets something else. To remove a known prefix or suffix useremoveprefix()andremovesuffix(), available from Python 3.9..title()capitalises after every non-letter, so"it's fine".title()gives"It'S Fine". For people's names it is usually wrong; store them the way the person typed them.
split() and join(): The Pair Behind Most Data Cleaning
split() turns one string into a list of strings, and it behaves differently depending on whether you give it a separator. With an argument, it cuts at exactly that separator and keeps every empty piece it produces. With no argument at all, it splits on any run of whitespace and quietly drops the empties — which is what you almost always want for text a human typed.
join() is the reverse, and its shape surprises everyone at first: you call it on the separator, not on the list. ", ".join(names) reads backwards for about a week and then stops bothering you. It is designed that way because the thing being joined can be any iterable of strings — a list, a tuple, a set, or a generator expression — while the separator is always a string, so the separator is the natural place to hang the method.
The one hard rule is that join() accepts strings only. Handing it a list of integers raises TypeError: sequence item 0: expected str instance, int found. Convert first, usually with a generator expression such as ", ".join(str(m) for m in marks). This pair — split the line, clean each field, join it back — is the backbone of nearly every small data-cleaning script you will write.
line = "Asha, 87, 92, 78"
parts = line.split(",")
print(parts) # ['Asha', ' 87', ' 92', ' 78'] note the leading spaces
parts = [p.strip() for p in line.split(",")]
print(parts) # ['Asha', '87', '92', '78']
# No argument: split on ANY run of whitespace and drop the empties
messy = " too many spaces "
print(messy.split()) # ['too', 'many', 'spaces']
print(messy.split(" ")) # ['', '', '', 'too', '', '', '', 'many', ...] — many empties
# maxsplit: stop after n cuts
print("key=value=extra".split("=", 1)) # ['key', 'value=extra']
# join is called on the SEPARATOR
names = ["Asha", "Ravi", "Meera"]
print(", ".join(names)) # Asha, Ravi, Meera
print("\n".join(names)) # one name per line
print("".join(names)) # AshaRaviMeera
# join only accepts strings
marks = [87, 92, 78]
# print(", ".join(marks)) # TypeError: expected str instance, int found
print(", ".join(str(m) for m in marks)) # 87, 92, 78 - Reading a real CSV file? Use the
csvmodule from the standard library rather thansplit(","). A field that legitimately contains a comma is written inside quotes, and splitting on the comma tears it in half.
f-strings: Building Text Properly
An f-string is a normal string with the letter f in front of it. Anything you put inside curly braces is treated as a Python expression: it is evaluated where it stands and its result is dropped into the text. That is not limited to variable names — arithmetic, function calls, indexing and method calls all work, so f"Average: {sum(scores) / len(scores)}" is a single readable line.
Before f-strings existed you glued text together with +. That is still legal, but it fails the moment a piece is not a string: "Total: " + 257 raises a TypeError, so you end up sprinkling str() calls everywhere. An f-string converts automatically and, more importantly, reads in the same order as the sentence it produces, so you can check it against the intended output at a glance.
After a colon inside the braces you can add a format specifier, which controls how the value is displayed without changing the value itself. :.2f fixes the number of decimal places, :, inserts thousands separators, :.1% shows a fraction as a percentage, and a width with an alignment character lines columns up — which is how you print a readable table without any extra library.
name, marks, total = "Asha", 87, 100
print(f"{name} scored {marks} out of {total}")
print(f"That is {marks / total * 100:.1f}%") # That is 87.0%
# Any expression works inside the braces
scores = [87, 92, 78]
print(f"Best: {max(scores)}, average: {sum(scores) / len(scores):.2f}")
print(f"First subject: {scores[0]}, name in caps: {name.upper()}")
# Format specifiers come after a colon
price = 1234567.891
print(f"{price:.2f}") # 1234567.89
print(f"{price:,.2f}") # 1,234,567.89
print(f"{0.876:.1%}") # 87.6%
# Width and alignment give you tables for free
rows = [("Asha", 87), ("Ravi", 92), ("Meera", 78)]
for n, m in rows:
print(f"{n:<10}{m:>5}")
# Debugging form: prints the expression text AND its value (Python 3.8+)
total_marks = 257
print(f"{total_marks=}") # total_marks=257
# A literal curly brace is written twice
print(f"{{not a placeholder}}") # {not a placeholder} f"...{expr}..."— any expression, evaluated where it stands{x:.2f}— a fixed number of decimal places{x:,}— thousands separators{x:<10},{x:>10},{x:^10}— left, right or centred in a field ten wide{x:.1%}— show a fraction as a percentage{x=}— print the expression text along with its value, for quick debugging
{value:.2f}rounds for display only; the number itself is untouched. If you need the rounded value for further arithmetic, useround(value, 2).
Escape Sequences, Raw Strings and the Windows Path Trap
Inside a normal string a backslash does not mean "backslash". It starts an escape sequence, a two-character code for something you cannot type directly. \n is a newline, \t is a tab, \\ is one literal backslash, and \" is a quote that does not end the string. This is why the same text can be written several ways and why print() and repr() disagree about what a value looks like.
The trap follows immediately: Windows file paths are full of backslashes. Write "C:\temp\notes.txt" and Python sees a tab and a newline, not folder names, and your path silently becomes nonsense. Prefixing the string with r makes it a raw string, in which backslashes stay exactly as typed. Forward slashes also work on Windows, and for anything beyond a one-off script the pathlib module builds correct paths on every operating system.
Raw strings matter for one more thing you will meet soon: regular expressions. Patterns for the re module use backslashes heavily — \d for a digit, \s for whitespace — and without the r prefix you would have to double every one of them. Writing regex patterns as raw strings is not a style preference, it is the convention.
print("Line 1\nLine 2") # \n is a newline
print("Name\tMarks") # \t is a tab
print("A backslash: \\") # \\ is one literal backslash
print("She said \"hi\"") # escaped quotes
# The Windows path trap
path = "C:\temp\notes.txt"
print(path) # C: <tab> emp <newline> otes.txt -- not a path at all
path = r"C:\temp\notes.txt" # raw string: backslashes stay literal
print(path) # C:\temp\notes.txt
path = "C:/temp/notes.txt" # forward slashes also work on Windows
print(path)
# Raw strings are the convention for regex patterns
import re
print(re.findall(r"\d+", "Order 42 shipped on 7 May")) # ['42', '7'] - A raw string still cannot end with a single backslash, because the backslash escapes the closing quote even in raw mode. If you need a trailing separator, add it another way — or use
pathlib.Path, which handles separators for you. - Python warns about escape sequences it does not recognise, such as
\dinside a plain string. If you meant the two characters, mark the string raw.
Three String Bugs Worth Recognising
The first is is where == was meant. == asks whether two strings hold the same characters; is asks whether two names point at the same object in memory. For short strings written as literals, CPython usually reuses a single object, so is appears to work — and that is exactly what makes it dangerous. Build the same text at runtime, by joining or slicing or reading it from a file, and the comparison flips to False even though the characters are identical. Use == for strings, always; reserve is for None, True and False.
The second is growing a string inside a loop with +=. Because strings are immutable, each round builds an entirely new string and copies everything accumulated so far, so the work grows with the square of the number of pieces. For ten items nobody notices; for a hundred thousand lines it turns a fast script into a slow one. Collect the pieces in a list and call "".join(pieces) once at the end.
The third is forgetting that text which looks numeric is still text. "5" * 3 gives "555", because * repeats a string rather than multiplying it, and "5" + 5 raises a TypeError rather than guessing which you meant. Convert explicitly with int() or float() — and note that int("42.0") fails, because int() parses whole numbers only. Go through float() first when the text may carry a decimal point.
# 1) is compares identity, == compares value
a = "hello"
b = "hello"
print(a == b) # True — same characters
print(a is b) # True — but only because CPython reuses one object here
c = "".join(["hel", "lo"])
print(c == a) # True — still the same characters
print(c is a) # False — a different object, built at run time
# 2) Growing a string in a loop copies it every single time
words = ["python", "for", "placements"]
out = ""
for w in words: # each += builds a whole new string
out += w + " "
print(out.strip())
print(" ".join(words)) # same result, one pass, clearer intent
# 3) Text that looks like a number is still text
print("5" * 3) # 555 — repetition, not multiplication
# print("5" + 5) # TypeError: can only concatenate str to str
print(int("5") + 5) # 10
print(int(" 42 ")) # 42 — surrounding whitespace is allowed
# print(int("42.0")) # ValueError: invalid literal for int()
print(int(float("42.0"))) # 42 - The
isresult above is a detail of how CPython stores short literals, not a language guarantee, and it changes with how the string was created. Anyone who tests strings withiswill eventually hit a pair of identical strings that compare as different.
