Data Classes & Structured Data
Topic 7 of 7, with 3 concept checks. dataclasses, namedtuples, and type hints
Give structured records an explicit data contract
Structured records
Use data classes and typed fields to make record shape, defaults, equality, and representation obvious while avoiding shared mutable defaults and accidental identity semantics.
Core lesson 01
@dataclass auto-generates __init__ (from annotated fields), a readable __repr__, and __eq__ (comparing all fields) — eliminating boilerplate for classes that are mainly data containers.
Writing a plain class to hold a few fields usually means hand-writing __init__ to assign each one, __repr__ for a readable printout, and __eq__ to compare instances by value. @dataclass generates all of that automatically from type-annotated class attributes, while still being a normal class you can add methods to.
from dataclasses import dataclass
@dataclass
class Point:
x: int
y: int
p1 = Point(1, 2)
p2 = Point(1, 2)
print(p1) # Point(x=1, y=2) -- auto __repr__
print(p1 == p2) # True -- auto __eq__, compares by valueWhat to remember
What does the @dataclass decorator generate for you automatically?
Common footguns
- Using a mutable default like tags: list = [] directly in a dataclass field — raises a ValueError by design (same root issue as mutable default arguments); use field(default_factory=list) instead.
Core lesson 02
A dataclass instance is mutable by default and supports inheritance/methods naturally. A namedtuple is an immutable tuple subclass — lighter weight, but fields can't be reassigned after creation.
collections.namedtuple (or typing.NamedTuple for the typed version) creates a tuple subclass with named fields — you get indexing, unpacking, and immutability for free, but can't add fields later or mutate them. @dataclass creates a full class — mutable by default (though @dataclass(frozen=True) makes it immutable too), supports default values, methods, and inheritance more naturally, at a slightly higher memory cost than a plain tuple.
from typing import NamedTuple
from dataclasses import dataclass
class PointNT(NamedTuple):
x: int
y: int
@dataclass
class PointDC:
x: int
y: int
p1 = PointNT(1, 2)
# p1.x = 5 # AttributeError -- immutable
p2 = PointDC(1, 2)
p2.x = 5 # fine -- mutable by defaultWhat to remember
What's the difference between a dataclass and a namedtuple?
Common footguns
- Expecting a namedtuple to support adding new attributes later — like any tuple, its structure is fixed at creation.
Core lesson 03
Type hints (def f(x: int) -> str) are optional annotations checked by external static type checkers like mypy — Python itself does not enforce or verify them at runtime.
Annotations like x: int or -> str are stored as metadata (__annotations__) and used by IDEs, linters, and static type checkers (mypy, pyright) to catch mismatches before running the code. But Python's interpreter does not check them at call time — you can pass a string where an int is annotated and Python will happily try to run the function, only failing later if the code actually does something incompatible with that type.
def add(a: int, b: int) -> int:
return a + b
print(add(2, 3)) # 5 -- fine
print(add("2", "3")) # '23' -- runs anyway! str + str is valid, hints aren't checkedWhat to remember
How do type hints work in Python, and are they enforced at runtime?
Common footguns
- Assuming type hints provide runtime safety like a statically typed language — they don't, unless you run a separate type checker in your workflow or add explicit runtime validation (e.g. with pydantic).
Python lab
Browser Python lab
Runtime · idle
Python loads on your first run. Your code stays in this browser.
Best practices
- Use @dataclass for simple data containers instead of hand-writing __init__/__repr__/__eq__.
- Use frozen=True for dataclasses that should behave like immutable value objects.
- Run a static type checker (mypy, pyright) in CI if you rely on type hints for safety.
- Use field(default_factory=...) for mutable default values in dataclasses, never a bare mutable literal.
Apply the concept in Interview practice
Design a LeaderboardmediumLeetCode #1244 · O(n) worst case per top(K)
Keep scores in a dict keyed by player id; for top(K), sort the current values descending and sum the first K — a natural fit for a small dataclass-like Player record.
Open problemDesign TwittermediumLeetCode #355 · O(n log n) per feed fetch
Store each tweet as a (timestamp, tweetId) pair per user, and merge the k most recent lists (own + followees) with a heap to build the news feed.
Open problemInsert Delete GetRandom O(1)mediumLeetCode #380 · O(1) per operation
Combine a list (for O(1) random access) with a dict mapping value to its index; delete by swapping the target with the last element before popping.
Open problemConcept checks
What does the @dataclass decorator generate for you automatically?
Hint
Boilerplate you'd otherwise write by hand in every __init__.
Think __init__, __repr__, and __eq__.
Answer
@dataclass auto-generates __init__ (from annotated fields), a readable __repr__, and __eq__ (comparing all fields) — eliminating boilerplate for classes that are mainly data containers.
Writing a plain class to hold a few fields usually means hand-writing __init__ to assign each one, __repr__ for a readable printout, and __eq__ to compare instances by value. @dataclass generates all of that automatically from type-annotated class attributes, while still being a normal class you can add methods to.
from dataclasses import dataclass
@dataclass
class Point:
x: int
y: int
p1 = Point(1, 2)
p2 = Point(1, 2)
print(p1) # Point(x=1, y=2) -- auto __repr__
print(p1 == p2) # True -- auto __eq__, compares by valueWatch out
- Using a mutable default like tags: list = [] directly in a dataclass field — raises a ValueError by design (same root issue as mutable default arguments); use field(default_factory=list) instead.
What's the difference between a dataclass and a namedtuple?
Hint
One is mutable by default, the other is immutable.
Think of a namedtuple as a lightweight, tuple-based alternative.
Answer
A dataclass instance is mutable by default and supports inheritance/methods naturally. A namedtuple is an immutable tuple subclass — lighter weight, but fields can't be reassigned after creation.
collections.namedtuple (or typing.NamedTuple for the typed version) creates a tuple subclass with named fields — you get indexing, unpacking, and immutability for free, but can't add fields later or mutate them. @dataclass creates a full class — mutable by default (though @dataclass(frozen=True) makes it immutable too), supports default values, methods, and inheritance more naturally, at a slightly higher memory cost than a plain tuple.
from typing import NamedTuple
from dataclasses import dataclass
class PointNT(NamedTuple):
x: int
y: int
@dataclass
class PointDC:
x: int
y: int
p1 = PointNT(1, 2)
# p1.x = 5 # AttributeError -- immutable
p2 = PointDC(1, 2)
p2.x = 5 # fine -- mutable by defaultWatch out
- Expecting a namedtuple to support adding new attributes later — like any tuple, its structure is fixed at creation.
How do type hints work in Python, and are they enforced at runtime?
Hint
They're documentation and tooling support, not a runtime contract.
Nothing stops you from passing the 'wrong' type at runtime by default.
Answer
Type hints (def f(x: int) -> str) are optional annotations checked by external static type checkers like mypy — Python itself does not enforce or verify them at runtime.
Annotations like x: int or -> str are stored as metadata (__annotations__) and used by IDEs, linters, and static type checkers (mypy, pyright) to catch mismatches before running the code. But Python's interpreter does not check them at call time — you can pass a string where an int is annotated and Python will happily try to run the function, only failing later if the code actually does something incompatible with that type.
def add(a: int, b: int) -> int:
return a + b
print(add(2, 3)) # 5 -- fine
print(add("2", "3")) # '23' -- runs anyway! str + str is valid, hints aren't checkedWatch out
- Assuming type hints provide runtime safety like a statically typed language — they don't, unless you run a separate type checker in your workflow or add explicit runtime validation (e.g. with pydantic).