Lesson 11 of 25

Dictionaries

Lookup by Key, Not by Position

A dictionary stores pairs: a key and the value it maps to. Where a list answers "what is at position 3", a dictionary answers "what is stored under 'name'". That is the whole idea, and it is why dictionaries are the most-used container in real Python code — most data is naturally labelled rather than numbered.

The speed matters as much as the readability. Finding a value by key does not depend on how many entries there are: Python computes the key's hash and goes straight to the right place. Searching a list of a hundred thousand records for a matching roll number means checking them one by one; a dictionary keyed by roll number finds it immediately. Whenever you find yourself scanning a list to match a field, a dictionary keyed on that field is usually the fix.

Keys must be hashable, which in practice means immutable: strings, numbers, booleans, tuples of those, or None. A list can never be a key. Values have no restriction at all — numbers, strings, lists, or other dictionaries, which is how nested structures such as JSON are represented. Keys are also unique: storing under an existing key replaces the old value rather than adding a second entry.

Since Python 3.7, dictionaries keep their insertion order, so looping over one gives you the pairs in the order they were added. Older material describes dictionaries as unordered; that was true before 3.7 and is not any more. Do not confuse order with sorting, though — the order is the order you inserted, not alphabetical or numeric.

Example
student = {
    "name": "Asha",
    "roll": "CS21B1042",
    "marks": 87,
    "subjects": ["Maths", "Physics"],   # values can be anything
}

print(student["name"])       # Asha
print(student["subjects"][0])  # Maths
print(len(student))          # 4

# Keys must be hashable; values need not be
valid = {"text": 1, 42: "answer", (0, 0): "origin", True: "yes"}
# invalid = {[1, 2]: "x"}    # TypeError: unhashable type: 'list'

# Keys are unique — assigning again replaces
student["marks"] = 90
print(student["marks"])      # 90

# Why a dictionary instead of a list of records
rows = [("CS21B1041", 78), ("CS21B1042", 87)]
for roll, marks in rows:          # scans until it finds a match
    if roll == "CS21B1042":
        print(marks)

by_roll = {"CS21B1041": 78, "CS21B1042": 87}
print(by_roll["CS21B1042"])       # straight there, whatever the size

# Insertion order is preserved (Python 3.7+)
d = {}
d["z"] = 1
d["a"] = 2
print(list(d))               # ['z', 'a']  — insertion order, not sorted
  • {key: value, ...} — curly braces, pairs separated by commas
  • d[key] — the value stored under that key
  • Keys must be hashable (immutable) and are unique; values can be anything
  • Lookup by key does not slow down as the dictionary grows
  • Insertion order is preserved from Python 3.7 onwards
  • len(d) — number of pairs; dict() or {} — an empty dictionary
Notes
  • {} creates an empty dictionary, not an empty set. That is a small historical quirk: dictionaries claimed the braces first, so an empty set has to be written set().

Reading Safely: [] versus .get()

Asking for a key that is not there with square brackets raises KeyError and stops the program. That is often the right behaviour — if a configuration file is missing a required setting, you want to know immediately rather than continue with nonsense.

When a key is genuinely optional, .get() is the tool. d.get("email") returns None when the key is absent instead of raising, and d.get("email", "not provided") returns whatever default you name. This is the single most common dictionary method after plain indexing, and using it removes a lot of defensive code.

There is a cost to that safety, and it is worth stating plainly: .get() hides typos. Ask for d.get("nmae") and you get None back with no complaint, and the failure surfaces somewhere far away as a mysterious missing value. Use [] when the key must be there and .get() when it might not be. Choosing .get() everywhere "to be safe" trades a loud error for a quiet one.

in tests keys, not values — a distinction worth fixing now, because it catches people repeatedly. "Asha" in student asks whether "Asha" is a key. To search the values you must say so with in student.values(), and that does scan the whole dictionary, so it has none of the speed advantage of a key lookup.

One more method fills a specific gap. setdefault(key, default) returns the value if the key exists and otherwise inserts the default and returns that. It is how you initialise an entry the first time you touch it, which is exactly what grouping code needs.

Example
student = {"name": "Asha", "marks": 87}

print(student["name"])                 # Asha
# print(student["email"])              # KeyError: 'email'

print(student.get("email"))            # None — no exception
print(student.get("email", "N/A"))     # N/A

# .get() also hides mistakes
print(student.get("nmae"))             # None — a typo, silently

# Check first when you need both safety and a loud failure
if "email" in student:
    print(student["email"])
else:
    print("no email on record")

# in tests KEYS
print("name" in student)               # True
print("Asha" in student)               # False — that is a value
print("Asha" in student.values())      # True — but this scans

# setdefault: read, or create on first touch
settings = {}
settings.setdefault("theme", "light")
settings.setdefault("theme", "dark")   # already there — ignored
print(settings)                        # {'theme': 'light'}
  • d[key] — raises KeyError if absent; use when the key must exist
  • d.get(key) — returns None if absent
  • d.get(key, default) — returns your default if absent
  • key in d — tests keys; value in d.values() tests values and scans
  • d.setdefault(key, default) — read, inserting the default on first use
Notes
  • d.get(key, []) does not store anything in the dictionary — it just hands you a default. If you meant to insert it, use setdefault, or a defaultdict as shown later in this lesson.

Adding, Updating and Removing

One statement does both adding and updating: d[key] = value creates the pair if the key is new and replaces the value if it is not. There is no separate "insert" and "update", and no error either way. That is convenient and it is also the quiet risk — a typo in a key name creates a new entry instead of correcting the existing one, and the dictionary now holds both.

The same rule applies inside a literal. Writing the same key twice in one {...} is not an error; the last one wins and the earlier one vanishes. It is a real source of confusion in a long configuration dictionary where the duplicate is forty lines away from its twin.

For removal there are four options. del d[key] deletes and raises KeyError if the key is absent. d.pop(key) deletes and returns the value, and accepts a default so it can be safe: d.pop(key, None) never raises. d.popitem() removes and returns the last inserted pair, useful for draining a dictionary. d.clear() empties it while keeping the same object.

To combine dictionaries, d.update(other) writes the second one's pairs into the first, overwriting on conflict. {**a, **b} builds a new dictionary from both, and from Python 3.9 the same thing is written a | b. In every case, later values win on a conflicting key — which is precisely what makes this the standard way to layer user settings on top of defaults.

Example
student = {"name": "Asha", "marks": 87}

student["email"] = "asha@college.edu"   # adds
student["marks"] = 90                   # updates — same syntax
print(student)

# The quiet risk: a typo creates instead of correcting
student["marks "] = 95        # trailing space — a different key entirely
print(len(student))           # 4, not 3
del student["marks "]

# A duplicate key in a literal is not an error
config = {"debug": True, "port": 8000, "debug": False}
print(config)                 # {'debug': False, 'port': 8000}

# Four ways to remove
d = {"a": 1, "b": 2, "c": 3, "d": 4}
del d["a"]                    # KeyError if missing
value = d.pop("b")            # removes and returns
print(value)                  # 2
print(d.pop("zz", None))      # None — safe, no exception
last = d.popitem()            # removes the most recently inserted pair
print(last)                   # ('d', 4)
d.clear()
print(d)                      # {}

# Merging: later values win
defaults = {"theme": "light", "size": "medium"}
user = {"theme": "dark"}

print({**defaults, **user})   # {'theme': 'dark', 'size': 'medium'}
print(defaults | user)        # same thing, Python 3.9+

defaults.update(user)         # modifies defaults in place
print(defaults)               # {'theme': 'dark', 'size': 'medium'}
Notes
  • d.update() changes the dictionary in place and returns None, following the same convention as list.sort(). Writing d = d.update(other) throws your dictionary away.

Looping, Views and Comprehensions

Looping over a dictionary gives you its keys. That is the default because keys are what you index with, and it is worth saying out loud because people expect values and are then puzzled by their output. for name in scores: and for name in scores.keys(): are the same thing; the shorter form is preferred.

When you want both parts, .items() yields a (key, value) tuple per entry, which unpacks straight into two loop variables. This is the form you will write most often. .values() gives just the values, which is what you want for sum(), max() and similar.

These three are views, not lists. A view is a live window onto the dictionary: add a pair and any view you are holding reflects it immediately. Views are cheap because nothing is copied, but they cannot be indexed — list(d.keys())[0] if you really need the first key. The liveness is also why changing a dictionary's size while looping over it raises RuntimeError: dictionary changed size during iteration. Loop over list(d) when you intend to delete as you go.

Sorting deserves a mention because dictionaries have no .sort(). Pass .items() to sorted() with a key function that picks the field to order by, and you get a list of pairs back — from which dict() rebuilds a dictionary in that order, since insertion order is preserved.

Finally, a dictionary comprehension builds one in a single expression: {key_expr: value_expr for item in iterable}, with an optional filter. Inverting a dictionary is the classic one-liner, with the classic caveat that duplicate values collapse into one key.

Example
scores = {"Maths": 92, "Physics": 88, "English": 95}

for subject in scores:                 # keys
    print(subject, scores[subject])

for subject, mark in scores.items():   # both — the usual form
    print(f"{subject}: {mark}")

print(sum(scores.values()) / len(scores))   # 91.66666666666667
print(max(scores.values()))                 # 95

# Views are live windows, not copies
keys = scores.keys()
scores["Chemistry"] = 79
print(list(keys))     # ['Maths', 'Physics', 'English', 'Chemistry']

# Which is why deleting during a loop fails
# for s in scores:
#     if scores[s] < 85:
#         del scores[s]      # RuntimeError: dictionary changed size

for s in list(scores):     # a snapshot of the keys
    if scores[s] < 85:
        del scores[s]
print(scores)              # {'Maths': 92, 'English': 95}

# Sorting: sort the items, rebuild if you want a dictionary
scores = {"Maths": 92, "Physics": 88, "English": 95}
by_mark = sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
print(by_mark)             # [('English', 95), ('Maths', 92), ('Physics', 88)]
print(dict(by_mark))

# Dictionary comprehensions
print({n: n ** 2 for n in range(1, 5)})          # {1: 1, 2: 4, 3: 9, 4: 16}
print({s: m for s, m in scores.items() if m >= 90})
print({m: s for s, m in scores.items()})         # inverted
  • for k in d: — keys; the same as d.keys()
  • for k, v in d.items(): — both, unpacked into two names
  • d.values() — the values, for sum(), max() and friends
  • Views are live and uncopied; wrap in list() to index or to delete while looping
  • sorted(d.items(), key=...) — order by key or by value; returns a list of pairs
  • {k: v for ...} — dictionary comprehension, with an optional if
Notes
  • Inverting a dictionary loses data whenever two keys share a value, because the second one overwrites the first. Check with len(inverted) == len(original) if that matters.

Counting and Grouping: What Dictionaries Do Most

Two jobs account for a large share of all dictionary code, and both appear constantly in coding tests and in data work. The first is counting: how many times does each word, character or category occur. The second is grouping: collect all the items that share a property under one key.

Written by hand, both have the same shape — check whether the key exists, initialise it if not, then update it. d[key] = d.get(key, 0) + 1 does the counting version in one line, and d.setdefault(key, []).append(item) does the grouping version in one line. Learn these two idioms; they turn up everywhere.

The standard library then removes even that. collections.Counter counts anything iterable in a single call and adds .most_common(n), which is exactly the "top five words" question interviewers like. collections.defaultdict(list) is a dictionary that creates an empty list the first time you touch a missing key, so the grouping loop becomes a plain append.

One caution about defaultdict: because merely reading a missing key creates it, a stray lookup silently adds an entry. That is usually harmless and occasionally baffling, so if you need to inspect a defaultdict without changing it, use .get().

Example
words = "the quick brown fox jumps over the lazy dog the end".split()

# Counting, written out
counts = {}
for w in words:
    counts[w] = counts.get(w, 0) + 1
print(counts["the"])                 # 3

# Counting, with the standard library
from collections import Counter

tally = Counter(words)
print(tally["the"])                  # 3
print(tally.most_common(2))          # [('the', 3), ('quick', 1)]
print(tally["missing"])              # 0 — no KeyError

# Grouping, written out
students = [("Asha", "CSE"), ("Ravi", "ECE"), ("Meera", "CSE")]

by_branch = {}
for name, branch in students:
    by_branch.setdefault(branch, []).append(name)
print(by_branch)     # {'CSE': ['Asha', 'Meera'], 'ECE': ['Ravi']}

# Grouping, with defaultdict
from collections import defaultdict

by_branch = defaultdict(list)
for name, branch in students:
    by_branch[branch].append(name)   # the empty list appears on demand
print(dict(by_branch))

# The defaultdict caution: reading creates
print(by_branch["MECH"])             # []
print(dict(by_branch))               # 'MECH' is now a real entry
Notes
  • Counter is a dictionary subclass, so everything in this lesson works on it. It adds arithmetic too: Counter(a) - Counter(b) subtracts counts, which is a neat way to compare two texts.

Nested Data, Copying and Three Traps

Real data is nested. An API response or a JSON file becomes a dictionary whose values are lists of dictionaries, and reading it means chaining lookups: data["students"][0]["marks"]["maths"]. Every link in that chain can raise KeyError, and the message names only the missing key, not the path. Chaining .get() with sensible defaults — data.get("students", []) — keeps the failure readable, and json.loads() and json.dumps() convert between text and this structure in one call each.

The first trap is copying. b = a gives a second name for the same dictionary, exactly as with lists, so changing one changes both. a.copy() makes a shallow copy: a new outer dictionary holding the same inner objects. For nested data that is not enough, and copy.deepcopy() is the tool that duplicates every level.

The second trap is a shared mutable default. dict.fromkeys(names, []) looks like a neat way to give every student an empty list, and it gives every student the same empty list — appending for one student appends for all of them. The same mistake appears as a mutable default argument in functions, covered in Lesson 13. A dictionary comprehension fixes it, because the expression is evaluated afresh for each key.

The third trap is keys that are equal without looking it. Python treats True and 1 as equal and they hash the same, so {1: "a", True: "b"} is a one-entry dictionary. The same holds for 1 and 1.0. It rarely matters, but when it does — usually while building a lookup table from mixed data — the missing entry is very hard to explain.

Example
import json, copy

raw = '{"class": "CS21B", "students": [{"name": "Asha", "marks": 87}]}'
data = json.loads(raw)                  # text -> dictionary
print(data["students"][0]["name"])      # Asha

# Chained .get() keeps a missing branch readable
print(data.get("teachers", []))         # []
# print(data["teachers"][0])            # KeyError: 'teachers'

print(json.dumps(data, indent=2)[:40])  # dictionary -> text

# Trap 1: assignment is not a copy
a = {"x": 1}
b = a
b["x"] = 99
print(a)                 # {'x': 99}

nested = {"marks": {"maths": 87}}
shallow = nested.copy()
shallow["marks"]["maths"] = 50
print(nested)            # {'marks': {'maths': 50}}  — inner dict was shared

deep = copy.deepcopy(nested)
deep["marks"]["maths"] = 100
print(nested)            # {'marks': {'maths': 50}}  — safe

# Trap 2: one list shared by every key
names = ["Asha", "Ravi"]
tasks = dict.fromkeys(names, [])
tasks["Asha"].append("lab")
print(tasks)             # {'Asha': ['lab'], 'Ravi': ['lab']}  — both changed

tasks = {n: [] for n in names}     # a fresh list per key
tasks["Asha"].append("lab")
print(tasks)             # {'Asha': ['lab'], 'Ravi': []}

# Trap 3: equal keys collapse
print({1: "a", True: "b"})     # {1: 'b'}
print({1: "a", 1.0: "b"})      # {1: 'b'}
Notes
  • json.dumps() can only write types JSON knows about. Keys become strings — so a dictionary keyed by integers comes back with string keys after a round trip — and values such as set or datetime raise TypeError unless you convert them first.
Ask AI