Unit Testing Best Practices Cheat Sheet
The Arrange-Act-Assert pattern, mocking external dependencies, parametrized tests, and the FIRST principles for writing reliable unit tests.
Arrange-Act-Assert
Structuring a pytest test into setup, action, and verification.
import pytestdef divide(a, b): if b == 0: raise ValueError("cannot divide by zero") return a / bdef test_divide_returns_quotient(): # Arrange a, b = 10, 2 # Act result = divide(a, b) # Assert assert result == 5def test_divide_by_zero_raises(): with pytest.raises(ValueError, match="cannot divide by zero"): divide(10, 0)
Mocking Dependencies
Isolating the unit under test with unittest.mock.
from unittest.mock import Mock, patchclass EmailService: def send(self, to, body): ... # calls a real SMTP serverdef notify_user(email_service, user_email): email_service.send(user_email, "Welcome!") return Truedef test_notify_user_calls_send(): mock_service = Mock(spec=EmailService) notify_user(mock_service, "[email protected]") mock_service.send.assert_called_once_with("[email protected]", "Welcome!")@patch("mymodule.requests.get") # patch where it's used, not where it's defineddef test_fetch_uses_requests_get(mock_get): mock_get.return_value.status_code = 200 mock_get.return_value.json.return_value = {"ok": True}
Parametrized Tests
Running the same test logic against multiple input/output pairs.
import pytest@pytest.mark.parametrize("a, b, expected", [ (2, 3, 5), (-1, 1, 0), (0, 0, 0),])def test_add(a, b, expected): assert a + b == expected
Testing Best Practices
Habits that keep a test suite fast, trustworthy, and maintainable.
- FIRST principles- Tests should be Fast, Independent, Repeatable, Self-validating, and Timely
- Descriptive names- Name tests after the behavior and expectation, e.g. test_returns_404_when_user_not_found
- One behavior per test- Each test should verify a single behavior so failures pinpoint the exact problem
- Avoid testing internals- Assert on observable outputs/behavior, not private implementation details, so refactors don't break tests
- Mock external dependencies- Isolate the unit under test from databases, network calls, and the filesystem with test doubles
- Use fixtures for setup- Share reusable setup/teardown code (e.g. pytest fixtures) instead of duplicating it in every test
- Cover edge cases- Test empty inputs, boundary values, nulls, and error paths, not just the happy path
- Keep tests deterministic- Avoid depending on real time, random values, network access, or execution order between tests
Test Doubles: Dummy, Stub, Spy, Mock, Fake
Distinguishing the five kinds of test double and when each one is the right tool.
from unittest.mock import Mockclass RealPaymentGateway: def charge(self, amount): ... # hits a real network endpoint# Dummy: passed in but never actually used, just satisfies a signaturedummy_logger = None# Stub: returns canned answers, no behavior verificationclass StubGateway: def charge(self, amount): return {"status": "approved"}# Spy: records how it was called so the test can assert on that afterwardclass SpyGateway: def __init__(self): self.calls = [] def charge(self, amount): self.calls.append(amount) return {"status": "approved"}# Mock: pre-programmed expectations, fails the test if they go unmetmock_gateway = Mock(spec=RealPaymentGateway)mock_gateway.charge.return_value = {"status": "approved"}# Fake: a working, lightweight implementation (e.g. in-memory store) used instead of the real oneclass FakeGateway: def __init__(self): self.ledger = {} def charge(self, amount): self.ledger["last"] = amount return {"status": "approved"}
Property-Based Testing with Hypothesis
Asserting invariants that must hold for all inputs instead of hand-picking individual example cases.
from hypothesis import given, strategies as stdef reverse(lst): return lst[::-1]# Hypothesis generates hundreds of inputs and shrinks failures down to a# minimal reproducing example automatically@given(st.lists(st.integers()))def test_reverse_twice_is_identity(lst): assert reverse(reverse(lst)) == lst@given(st.lists(st.integers(), min_size=1))def test_reverse_preserves_length_and_elements(lst): reversed_lst = reverse(lst) assert len(reversed_lst) == len(lst) assert sorted(reversed_lst) == sorted(lst)
Testing Async Code
Awaiting coroutines under test directly with pytest-asyncio instead of blocking on an event loop manually.
import pytestimport asyncioasync def fetch_user(user_id, db): return await db.get(user_id)@pytest.mark.asyncioasync def test_fetch_user_returns_record(): class FakeDb: async def get(self, user_id): await asyncio.sleep(0) # simulate an await point return {"id": user_id, "name": "Ada"} result = await fetch_user(1, FakeDb()) assert result == {"id": 1, "name": "Ada"}# pyproject.toml:# [tool.pytest.ini_options]# asyncio_mode = "auto" # lets async def tests run without the marker
Fixture Scopes & Dependency Injection
Controlling how often a pytest fixture is rebuilt with scope, and injecting fresh state per test.
import pytestclass Database: def __init__(self): self.records = [] def insert(self, item): self.records.append(item)@pytest.fixture(scope="function") # fresh instance per test (the default)def db(): return Database()@pytest.fixture(scope="session") # built once, shared across the whole rundef expensive_config(): return {"timeout": 30, "retries": 3}def test_insert_adds_record(db): db.insert("a") assert db.records == ["a"]def test_insert_is_isolated(db): # db is a brand-new instance here -- no leakage from the previous test assert db.records == []
Advanced Testing Vocabulary
Concepts that separate a merely-passing test suite from one you can actually trust.
- Test Pyramid- A large base of fast unit tests, fewer integration tests, and a thin layer of slow end-to-end tests
- Mutation Testing- Deliberately injecting small code bugs (mutants) and checking whether the suite fails -- surfaces weak assertions coverage alone misses
- Flaky Test- A test that passes and fails intermittently without code changes, usually from timing, ordering, or shared state
- Snapshot/Golden Testing- Comparing output against a previously approved reference file, useful for large structured outputs
- Contract Testing- Verifying that a consumer and provider of an API agree on the same interface without standing up both services
- Coverage vs Mutation Coverage- Line/branch coverage shows what code ran; mutation coverage shows what was actually verified by an assertion
- Test Isolation- Each test sets up and tears down its own state so tests can run in any order or in parallel
- Golden Master Testing- Capturing legacy system output as a baseline before refactoring, to detect unintended behavior changes
If you find yourself mocking three or four collaborators just to test one function, treat that as a signal the function is doing too much — refactor it before writing more tests around it.