Hypothesis: Property-Based Testing in Python (2026)
Hypothesis: Property-Based Testing in Python (2026) shows how to stop guessing test inputs. Instead of writing five hand-picked examples, you describe the shape of valid data and a rule that must always hold. Hypothesis then generates hundreds of cases, finds the one that breaks your code, and shrinks it to the smallest possible failing input. Every result below comes from a real run.
Hypothesis plugs straight into pytest, so it fits the modern Python testing stack with pytest, Ruff and coverage, works for testing data science code with pytest, and pairs well with FastAPI testing with pytest and TestClient.
TL;DR
pip install hypothesis, then decorate a test with@given(...)and pass strategies such asst.integers(),st.text()orst.lists(...).- Write properties: rules that hold for every input (nothing is lost, decode undoes encode, a discount never raises a price).
- When a property fails, Hypothesis shrinks the input. Our buggy
chunk()was reduced toitems=[0], size=2. - Round-trip tests catch edge cases you would never type by hand: our run-length encoder broke on the one-character string
'0'. - Use
st.builds()for dataclasses,@exampleto pin known cases, andassume()to skip inputs that do not apply. Tested with Hypothesis6.168.5, pytest9.1.1, Python3.13.5.
What is property-based testing?
A normal unit test checks one input against one expected output. A property-based test checks a rule against many generated inputs. You stop asking "does chunk([1, 2, 3], 2) return the right thing?" and start asking "for any list and any chunk size, do I get every item back in order?" The library does the boring part: generating empty lists, huge numbers, Unicode, duplicates and other edge cases, then simplifying whatever fails.
| Piece | What it does | Example |
|---|---|---|
@given | Turns a test into a property test | @given(st.lists(st.integers())) |
| Strategies | Describe valid input data | st.text(), st.integers(min_value=1), st.builds(Order, ...) |
| Shrinking | Reduces a failure to a minimal example | items=[0], size=2 |
@settings | Tune examples, deadlines, determinism | @settings(max_examples=500) |
@example / assume() | Pin a known case / discard irrelevant inputs | @example([], 50) |
Versions tested (2026-10-06)
- Python
3.13.5 hypothesis6.168.5(requires Python 3.10 or newer)pytest9.1.1
python -m venv .venv && source .venv/bin/activate
pip install hypothesis==6.168.5 pytest
python props_demo.py
pytest -q test_orders.py
The demo script calls each Hypothesis test directly and prints a one-line summary so you can see the failing input without a full traceback. In a real project you simply run pytest.
1. Find a bug with your first property
Here is a chunk() helper with an off-by-one bug hidden in the range() stop value. The property is simple: if you flatten the chunks, you must get the original list back.
import sys, hypothesis
from hypothesis import given, settings, strategies as st
print("Python", sys.version.split()[0], "| Hypothesis", hypothesis.__version__)
def report(test):
try:
test()
print(test.__name__, "-> passed")
except (AssertionError, ExceptionGroup) as err:
errs = getattr(err, "exceptions", [err])
err = [e for e in errs if isinstance(e, AssertionError)][0]
note = " ".join(getattr(err, "__notes__", [""])[0].split())
print(test.__name__, "-> FAILED")
print(" ", note.replace("( ", "(").replace(", )", ")"))
# 1) A buggy helper: split a list into chunks of `size`
def chunk(items, size):
return [items[i:i + size] for i in range(0, len(items) - size + 1, size)]
@settings(database=None, derandomize=True)
@given(st.lists(st.integers()), st.integers(min_value=1, max_value=10))
def test_chunk_keeps_every_item(items, size):
flat = [x for part in chunk(items, size) for x in part]
assert flat == items
report(test_chunk_keeps_every_item)
Python 3.13.5 | Hypothesis 6.168.5
test_chunk_keeps_every_item -> FAILED
Failing test case: test_chunk_keeps_every_item(items=[0], size=2)
Hypothesis tried many lists and sizes, found a failure, and then shrank it to the smallest possible case: a one-item list with a chunk size of 2. The last partial chunk is silently dropped. A hand-written test using [1, 2, 3, 4] and size 2 would have passed and hidden the bug. (database=None, derandomize=True only make this demo print the same result on every run; you can leave them out in your own tests.)
2. Fix it and run 500 cases
# 2) Fix the bug and run 500 random cases
def chunk(items, size):
return [items[i:i + size] for i in range(0, len(items), size)]
calls = 0
@settings(max_examples=500, database=None, derandomize=True)
@given(st.lists(st.integers()), st.integers(min_value=1, max_value=10))
def test_chunk_fixed(items, size):
global calls
calls += 1
parts = chunk(items, size)
assert [x for p in parts for x in p] == items
assert all(1 <= len(p) <= size for p in parts)
report(test_chunk_fixed)
print("examples run:", calls)
test_chunk_fixed -> passed
examples run: 500
The fixed version passes 500 generated cases. We also added a second property in the same test: every chunk has between 1 and size items. Stacking several cheap assertions in one property is a great way to pin down behaviour.
3. Round-trip tests catch the weird inputs
"Decode undoes encode" is one of the most useful properties you can write. It works for serializers, parsers, compressors, URL builders and database mappers. Here is a tiny run-length encoder:
# 3) Round-trip property: run-length encoding
def rle_encode(s):
out, i = "", 0
while i < len(s):
j = i
while j < len(s) and s[j] == s[i]:
j += 1
out += f"{j - i}{s[i]}"
i = j
return out
def rle_decode(code):
out, num = "", ""
for ch in code:
if ch.isdigit():
num += ch
else:
out += ch * int(num)
num = ""
return out
@settings(database=None, derandomize=True)
@given(st.text())
def test_rle_round_trip(s):
assert rle_decode(rle_encode(s)) == s
report(test_rle_round_trip)
@settings(database=None, derandomize=True)
@given(st.text(alphabet="abcxyz "))
def test_rle_round_trip_letters(s):
assert rle_decode(rle_encode(s)) == s
report(test_rle_round_trip_letters)
print("rle_encode('aaabcc') =", rle_encode("aaabcc"))
test_rle_round_trip -> FAILED
Failing test case: test_rle_round_trip(s='0')
test_rle_round_trip_letters -> passed
rle_encode('aaabcc') = 3a1b2c
The encoder works for normal letters ('aaabcc' becomes 3a1b2c), but Hypothesis found that the one-character string '0' breaks the round trip: it encodes to 10, which decodes as "ten of nothing". Once you know that, you can either escape digits or, as we did here, restrict the input alphabet and document that limit. Either way, the bug is now a decision instead of a surprise in production.
4. Use Hypothesis with pytest, dataclasses and business rules
In a real project you put properties in an ordinary pytest file. st.builds() creates dataclass instances from field strategies, @example always runs a specific case, and assume() throws away inputs that do not make sense for a property.
from dataclasses import dataclass
from hypothesis import given, example, assume, strategies as st
@dataclass
class Order:
sku: str
qty: int
unit_price_cents: int
def order_total(orders, discount_pct=0):
gross = sum(o.qty * o.unit_price_cents for o in orders)
return gross - gross * discount_pct // 100
orders = st.lists(st.builds(
Order,
sku=st.text(alphabet="ABCDEFGH0123456789", min_size=3, max_size=8),
qty=st.integers(min_value=1, max_value=50),
unit_price_cents=st.integers(min_value=0, max_value=100_000),
))
@given(orders, st.integers(min_value=0, max_value=100))
@example([], 50)
def test_total_never_negative(orders, pct):
assert order_total(orders, pct) >= 0
@given(orders, st.integers(min_value=0, max_value=100))
def test_discount_never_raises_price(orders, pct):
assert order_total(orders, pct) <= order_total(orders)
@given(orders)
def test_order_does_not_matter(orders):
assume(len(orders) >= 2)
assert order_total(orders) == order_total(list(reversed(orders)))
... [100%]
3 passed in 1.33s
Add --hypothesis-show-statistics to see how much work each property did:
test_orders.py::test_total_never_negative:
- 100 passing, 0 failing, and 9 invalid test cases
test_orders.py::test_discount_never_raises_price:
- 100 passing, 0 failing, and 6 invalid test cases
test_orders.py::test_order_does_not_matter:
- 100 passing, 0 failing, and 16 invalid test cases
3 passed in 0.52s
Each property ran 100 generated cases by default. "Invalid" cases are inputs Hypothesis threw away, for example lists shorter than two items rejected by assume(). If a large share of cases are invalid, tighten the strategy instead of filtering.
Good properties to start with
- Round trip:
loads(dumps(x)) == x,decode(encode(s)) == s. - Nothing lost: flattening, splitting, batching and pagination keep every item.
- Invariants: totals are never negative, a sorted list is ordered, IDs stay unique.
- Idempotence:
normalize(normalize(x)) == normalize(x). - Oracle: a fast optimized function matches a slow, obviously correct version.
Common mistakes
- Re-implementing the function in the test. Test a rule, not a copy of the code.
- Over-filtering with
assume(). If Hypothesis keeps discarding inputs, build the right data with strategy arguments such asmin_sizeoralphabet. - Ignoring the example database. Hypothesis saves failing examples in a local
.hypothesis/folder and replays them first on the next run. Keep it out of Git, but do not delete it while you are fixing a bug. - Flaky CI. Use
@settings(derandomize=True)or a settings profile when you need the same inputs on every CI run, and raisedeadlinefor slow tests.
When to use Hypothesis
Use it for pure functions, parsers, serializers, data transformations, money and date maths, and anything with "for all inputs" rules. Keep classic example-based tests for documentation-style cases and for slow integration tests. Most teams get the best results by adding two or three properties to their most critical helpers first, then growing from there.
FAQ
Does Hypothesis replace pytest? No. It is a library that runs inside pytest (or unittest). You keep your fixtures, plugins and CI setup.
How many examples does it run? 100 per test by default. Change it with @settings(max_examples=...), as we did with 500.
Why is the failing example so small? That is shrinking. Hypothesis keeps simplifying the failing input until no smaller version still fails, so you can debug the real cause right away.