Unique, Unordered, and Fast to Search
A set is a collection with two defining properties: every element appears at most once, and there is no order. Both are enforced by Python rather than by you. Adding an element that is already present does nothing at all — no error, no duplicate — and the elements come back in whatever arrangement the implementation finds convenient.
Because there is no order there are no positions, so a set cannot be indexed or sliced. s[0] raises TypeError: 'set' object is not subscriptable. If you need a specific element you are using the wrong container; a set answers "is this in here" and "what is in both", not "what is the third one".
The reason to accept those restrictions is speed. A set stores its elements by hash, exactly as a dictionary stores its keys, so x in s goes straight to the answer regardless of size. The same test on a list walks it from the front. On a hundred names the difference is invisible; inside a loop over a hundred thousand records it is the difference between a script that finishes and one you abandon.
The hash requirement is the price. Elements must be hashable — strings, numbers, tuples of those — for the same reason dictionary keys must be. A list cannot go into a set. Note also the syntax quirk: {1, 2} is a set but {} is an empty dictionary, because dictionaries claimed the braces first. An empty set has to be written set().
fruits = {"apple", "banana", "cherry"}
print(len(fruits)) # 3
# Duplicates simply do not survive
numbers = {1, 2, 2, 3, 3, 3}
print(numbers) # {1, 2, 3}
print(set([1, 2, 2, 3])) # {1, 2, 3} — from any iterable
print(set("hello")) # {'h', 'e', 'l', 'o'} — order will vary
# No order means no positions
# print(fruits[0]) # TypeError: 'set' object is not subscriptable
print(sorted(fruits)) # ['apple', 'banana', 'cherry'] — a LIST
# Membership is the point
print("apple" in fruits) # True
print("mango" in fruits) # False
# Elements must be hashable
ok = {1, "two", (3, 4), None}
# bad = {[1, 2]} # TypeError: unhashable type: 'list'
# The empty-set trap
empty = set()
print(type(empty)) # <class 'set'>
print(type({})) # <class 'dict'> — NOT an empty set {a, b, c}orset(iterable)— duplicates are dropped automaticallyset()for an empty set;{}is an empty dictionary- No order, so no indexing, no slicing and no
.sort() x in sdoes not slow down as the set grows- Elements must be hashable — no lists, dictionaries or other sets inside
sorted(s)gives you an ordered list when you need to display it
- The order you see when printing a set is not random, but it is not something to rely on either: it depends on the hash values and the insertion history, and it can differ between Python versions or between runs for strings. Sort before displaying.
Adding and Removing
add(x) puts one element in. update(iterable) puts in every element of another collection. The distinction mirrors append and extend on lists, and it goes wrong in the same way: s.update("hi") adds the characters 'h' and 'i' separately, because a string is iterable. When you mean one item, use add.
Removal has two spellings, and choosing between them is a small design decision rather than a preference. remove(x) raises KeyError when the element is absent; discard(x) does nothing. Use remove when the element being missing means something has already gone wrong, and discard when "remove it if it is there" is genuinely what you mean.
pop() removes and returns some element, and you do not get to choose which — there is no order, so there is no first. It raises KeyError on an empty set. It is useful for draining a set in a loop and misleading if you expect list-like behaviour. clear() empties the set in place.
One property carries over from lists and matters just as much: sets are mutable, so b = a creates a second name for the same set rather than a copy. Use a.copy() when you want an independent one. And because they are mutable, a set cannot be an element of another set — that is what frozenset is for, later in this lesson.
tags = {"python", "beginner"}
tags.add("placement") # one element
tags.update(["dsa", "resume"]) # every element of an iterable
print(len(tags)) # 5
tags.add("python") # already present — nothing happens
print(len(tags)) # 5
# add vs update, the same trap as append vs extend
s = set()
s.add("hi")
print(s) # {'hi'}
s = set()
s.update("hi")
print(s) # {'h', 'i'} — the string was unpacked
# remove is strict, discard is not
tags.remove("dsa")
# tags.remove("dsa") # KeyError: 'dsa'
tags.discard("dsa") # silent, already gone
# pop takes an arbitrary element
small = {10, 20, 30}
print(small.pop()) # one of them — you cannot choose which
# Assignment is not a copy
a = {1, 2}
b = a
b.add(3)
print(a) # {1, 2, 3}
c = a.copy()
c.add(4)
print(a) # {1, 2, 3} — unchanged .add(x)— one element; a repeat is silently ignored.update(iterable)— many elements; a string unpacks into characters.remove(x)—KeyErrorif absent.discard(x)— no error if absent.pop()— removes an arbitrary element;KeyErroron an empty set.clear(),.copy()— empty it, or make an independent one
- There is no
.append()on a set. If you catch yourself typing it, that is usually a sign the data is ordered and you wanted a list after all.
The Two Jobs Sets Actually Do
Almost every practical use of a set is one of two things: removing duplicates, or testing membership quickly. Everything else is a variation.
Deduplicating is a single call — set(items) — and the caveat is that you lose the order. For a report or an interface that matters, and the fix is a neat idiom: list(dict.fromkeys(items)) keeps the first occurrence of each item in its original position, because dictionaries preserve insertion order and reject duplicate keys. Use set() when order is irrelevant and dict.fromkeys() when it is not.
The membership job is the one that quietly rescues slow programs. If you are checking values against a fixed collection inside a loop, build a set once outside the loop and test against that. The conversion costs one pass; every check after it is free. This is a standard interview answer as well as a standard fix — "the list membership test inside the loop makes it quadratic; convert to a set first".
The same reasoning applies to len(set(items)) for "how many distinct values are there", and to comparing two collections for missing or extra entries, which the next section formalises. What sets deliberately do not do is count: a set knows that a word occurred, not how often. When counts matter, use collections.Counter.
names = ["Asha", "Ravi", "Asha", "Meera", "Ravi"]
print(set(names)) # {'Asha', 'Ravi', 'Meera'} — no order
print(len(set(names))) # 3 distinct people
# Order-preserving deduplication
print(list(dict.fromkeys(names))) # ['Asha', 'Ravi', 'Meera']
# The slow shape
registered = ["CS21B1041", "CS21B1042", "CS21B1043"]
attendance = ["CS21B1042", "CS21B1099"]
for roll in attendance:
if roll in registered: # scans the list every time
print(roll, "present")
# The fix: one conversion, then free lookups
registered_set = set(registered)
for roll in attendance:
if roll in registered_set:
print(roll, "present")
# Sets know 'whether', not 'how many'
words = ["the", "cat", "the"]
print(set(words)) # {'the', 'cat'} — the count is gone
from collections import Counter
print(Counter(words)) # Counter({'the': 2, 'cat': 1}) set(items)only works if every item is hashable, so a list of lists cannot be deduplicated this way. Convert the inner lists to tuples first, then convert back if you need to.
Set Algebra: Union, Intersection, Difference
The four set operations from school mathematics are built into the language, and they replace a surprising amount of loop-and-compare code. Union (|) is everything in either. Intersection (&) is what appears in both. Difference (-) is what is in the first and not the second — and note that it is not symmetric, so a - b and b - a answer different questions. Symmetric difference (^) is everything that is in one but not both.
Each has a method form as well: union, intersection, difference, symmetric_difference. They are not quite interchangeable, and the difference is worth knowing. The operators require both sides to be sets and raise TypeError otherwise; the methods accept any iterable, so a.union([1, 2]) works where a | [1, 2] does not. The operator is shorter when both sides are already sets; the method saves a conversion when the second one is a list.
Once you see these in terms of real questions, they stop feeling academic. Which students registered for both electives? That is an intersection. Which registered students never attended? That is a difference. Which appear in one system but not the other, in either direction? That is a symmetric difference — one line replacing a nested loop.
There are in-place versions too, spelled |=, &=, -= and ^=, which modify the set on the left rather than building a new one. Use them when you are accumulating into a set you own, and remember that anything else holding a reference to that set sees the change.
python_students = {"Asha", "Ravi", "Meera", "Karan"}
java_students = {"Meera", "Karan", "Divya", "Neha"}
print(python_students | java_students) # everyone, once each
print(python_students & java_students) # {'Meera', 'Karan'} — took both
print(python_students - java_students) # {'Asha', 'Ravi'} — Python only
print(java_students - python_students) # {'Divya', 'Neha'} — Java only
print(python_students ^ java_students) # exactly one course, not both
# Method forms accept any iterable; operators need sets on both sides
print(python_students.union(["Zoya"])) # fine
# print(python_students | ["Zoya"]) # TypeError: unsupported operand type
# A real question, answered in one line
registered = {"CS21B1041", "CS21B1042", "CS21B1043"}
attended = {"CS21B1042", "CS21B1043", "CS21B1099"}
print(registered - attended) # {'CS21B1041'} — registered, never attended
print(attended - registered) # {'CS21B1099'} — attended without registering
print(registered & attended) # both lists agree on these
# In-place accumulation
seen = set()
for batch in (["a", "b"], ["b", "c"]):
seen |= set(batch)
print(sorted(seen)) # ['a', 'b', 'c'] a | b/a.union(b)— in eithera & b/a.intersection(b)— in botha - b/a.difference(b)— inaonly; not symmetrica ^ b/a.symmetric_difference(b)— in exactly one of them- Operators need sets on both sides; the methods accept any iterable
|=,&=,-=,^=modify the left-hand set in place
- All four operations return a new set, so the originals are untouched unless you use the in-place form. That makes them safe to chain:
(a | b) - creads exactly as it looks.
Comparisons, Comprehensions and frozenset
Sets answer relationship questions directly. a.issubset(b) asks whether every element of a is in b, and a.isdisjoint(b) asks whether they share nothing. Both have shorter operator forms, and this is where a genuine surprise lives: on sets, <= means "is a subset of", not "is less than". {1, 2} <= {1, 2, 3} is True, and < means a proper subset — a subset that is not the whole thing.
The surprise has a consequence. Because two sets can each fail to be a subset of the other, comparison on sets is only partial: {1, 2} < {2, 3} and {2, 3} < {1, 2} are both False, which is impossible for numbers. Never assume that failing one comparison implies the other, and do not try to sort a list of sets by these operators.
A set comprehension uses the same shape as a list comprehension with braces instead of brackets: {expression for item in iterable if condition}. It builds a set, so duplicates in the result collapse automatically — which is often exactly what you want when extracting a field from a list of records.
Finally, frozenset is an immutable set. It supports everything that reads — membership, union, intersection, subset tests — and nothing that changes. Because it cannot change, it is hashable, so it can be a dictionary key or an element of another set. That is the standard way to represent a group of items as a single value: a set of subject combinations, or the set of tags on an article used as a lookup key.
core = {"Maths", "Physics"}
taken = {"Maths", "Physics", "Chemistry"}
print(core.issubset(taken)) # True
print(core <= taken) # True — subset, NOT "less than"
print(core < taken) # True — proper subset (not equal)
print(taken >= core) # True — superset
print(core.isdisjoint({"Hindi"})) # True — nothing in common
# Set comparison is only partial
print({1, 2} < {2, 3}) # False
print({2, 3} < {1, 2}) # False — both, which numbers never do
# Set comprehension: duplicates collapse for free
records = [("Asha", "CSE"), ("Ravi", "ECE"), ("Meera", "CSE")]
branches = {branch for name, branch in records}
print(branches) # {'CSE', 'ECE'}
print({n % 3 for n in range(10)}) # {0, 1, 2}
# frozenset: immutable, therefore hashable
combo = frozenset({"Maths", "Physics"})
counts = {combo: 12} # a set used as a dictionary key
print(counts[frozenset({"Physics", "Maths"})]) # 12 — order irrelevant
# A set of sets needs frozensets
groups = {frozenset({1, 2}), frozenset({3})}
print(len(groups)) # 2
# {{1, 2}} # TypeError: unhashable type: 'set' frozensethas no.add(),.remove()or.update(). To "change" one, build a new one:combo | {"Chemistry"}returns a new frozenset and leaves the original alone.
When a Set Is the Wrong Choice
Sets are the right answer often enough that it is worth being clear about when they are not. The first case is anything where order matters — a queue, a ranking, the lines of a file, a menu. A set will scramble it, and the scrambling is not reproducible enough to work around.
The second is anything where duplicates carry information. Marks, transactions, votes, log entries: converting these to a set silently deletes data. "Three students scored 87" becomes "87 appeared", and nothing warns you. If you catch yourself calling set() on measurements rather than identifiers, stop and check what you are throwing away.
The third is unhashable data. A list of lists cannot go into a set, and neither can a list of dictionaries — which is exactly the shape that JSON gives you. Convert the parts you want to compare into tuples or extract a hashable identifier first.
Two smaller traps round it off. Changing a set while looping over it raises RuntimeError: Set changed size during iteration, exactly as with a dictionary; loop over a copy if you must delete as you go. And because True == 1 and 1 == 1.0, a set treats them as the same element, so {1, True, 1.0} holds one item. That almost never matters, and when it does, it is very hard to see.
# Order is lost — do not use a set for anything ranked
steps = ["wash", "rinse", "dry"]
print(set(steps)) # arrangement is arbitrary
# Duplicates are data, not noise
marks = [87, 92, 87, 65]
print(sum(marks) / len(marks)) # 82.75 — the real average
print(sum(set(marks)) / len(set(marks))) # 81.33... — wrong, an 87 vanished
# Unhashable elements cannot go in
rows = [{"name": "Asha"}, {"name": "Ravi"}]
# print(set(rows)) # TypeError: unhashable type: 'dict'
print({r["name"] for r in rows}) # extract something hashable instead
# Changing a set while looping over it
s = {1, 2, 3, 4}
# for n in s:
# if n % 2 == 0:
# s.remove(n) # RuntimeError: Set changed size during iteration
for n in list(s): # loop over a snapshot
if n % 2 == 0:
s.remove(n)
print(s) # {1, 3}
# Equal values are one element
print({1, True, 1.0}) # {1}
print(len({0, False})) # 1 - Quick rule for choosing: does order matter, or do repeats matter? If either is yes, use a list. If neither, and you will be testing membership, use a set.
