What are common DynamoDB anti-patterns and best practices?
Avoid DynamoDB pitfalls like hot partitions, full scans, and relational modeling. Learn best practices: access-pattern design and single-table design.
Expected Interview Answer
Common DynamoDB anti-patterns include treating it like a relational database (normalized tables with app-side joins), using low-cardinality partition keys that create hot partitions, running frequent full-table Scans, and modeling before knowing access patterns; best practices are the inverse — design around access patterns first, distribute keys evenly, use single-table design with Query, and prefer GSIs over scans.
DynamoDB rewards designing the data model backward from the queries you must serve. Choose high-cardinality partition keys (or add a suffix/sharding for hot keys), overload keys and GSIs to serve multiple entity types from one table, and use Query with composite sort keys instead of scanning. Avoid unbounded item growth, offload large blobs to S3, enable auto-scaling or on-demand capacity, use TTL to expire stale data, and rely on DynamoDB Streams for event-driven propagation rather than polling. Measure with CloudWatch to catch throttling and hot partitions early.
- Consistent low latency by avoiding hot partitions
- Lower cost by avoiding full scans
- Simpler, faster reads via single-table design
- Even capacity utilization across partitions
- Scalable event-driven architecture with Streams
AI Mentor Explanation
A DynamoDB anti-pattern is like setting every fielder in one corner of the ground: all the load lands on a hot spot while the rest of the field sits idle, just as a low-cardinality partition key overloads one partition. Good captains spread fielders to cover every scoring zone evenly. Best practice is that field placement — distribute keys so traffic spreads across partitions, and plan each over to a known plan rather than reacting ball by ball with a wasteful full-ground scan.
Step-by-Step Explanation
Step 1
Design from access patterns
Write down every query the app must serve before creating tables; the model is derived from those patterns, not from entity diagrams.
Step 2
Choose high-cardinality partition keys
Pick keys with many distinct, evenly hit values; add a sharding suffix for known hot keys to spread traffic.
Step 3
Prefer Query over Scan
Use partition + sort key conditions or GSIs to fetch exactly what you need; reserve Scan for rare, bounded maintenance tasks.
Step 4
Adopt single-table design where it fits
Overload keys and GSIs so one table serves multiple entity types and relationships without app-side joins.
Step 5
Handle large and expiring data
Offload big blobs to S3, cap item growth, and use TTL to auto-delete stale items instead of manual cleanup.
Step 6
Scale and observe
Enable on-demand or auto-scaling, use Streams for event propagation, and watch CloudWatch for throttling and hot partitions.
What Interviewer Expects
- Recognizing relational thinking as the top anti-pattern
- Explaining hot partitions and key sharding
- Why Scans are costly and how to avoid them
- Understanding single-table design and key overloading
- Using GSIs, TTL, Streams, and auto-scaling appropriately
- Designing from access patterns first
Common Mistakes
- Normalizing data and doing joins in application code
- Using a low-cardinality partition key (e.g., a status flag) that creates hot partitions
- Relying on Scan for routine reads
- Modeling the schema before knowing access patterns
- Letting items grow unbounded instead of offloading to S3 or splitting
Best Answer (HR Friendly)
“The biggest mistake with DynamoDB is treating it like a traditional SQL database and searching the whole table for data, which is slow and expensive. The best practice is to design around how you'll actually look up your data, spread that data evenly so no single part gets overloaded, and fetch it with targeted queries instead of full scans.”
Code Example
// BAD: reads every item, then filters client-side — slow and costly
const res = await ddb.scan({
TableName: 'Orders',
FilterExpression: 'customerId = :c',
ExpressionAttributeValues: { ':c': { S: customerId } },
}).promise()
// FilterExpression runs AFTER the full scan; you pay for all items read// GOOD: targeted Query on the partition key
const res = await ddb.query({
TableName: 'Orders',
KeyConditionExpression: 'customerId = :c AND createdAt > :d',
ExpressionAttributeValues: {
':c': { S: customerId },
':d': { S: '2026-01-01' },
},
}).promise()
// Sharding a hot partition key (e.g., a popular product's events)
const shard = Math.floor(Math.random() * 10)
const pk = `PRODUCT#${productId}#${shard}` // spreads writes across 10 partitionsFollow-up Questions
- How do you detect and fix a hot partition in DynamoDB?
- What is single-table design and when should you avoid it?
- When is a Scan actually acceptable in DynamoDB?
- How do GSIs differ from LSIs, and when do you use each?
- How can DynamoDB Streams support event-driven architectures?
MCQ Practice
1. Which is a common DynamoDB anti-pattern?
Frequent full-table Scans read and bill for every item; targeted Queries on well-chosen keys are the best practice instead.
2. What causes a hot partition in DynamoDB?
Low-cardinality keys funnel most requests to one partition, exceeding its throughput and causing throttling.
3. Which technique helps spread load across partitions for a hot key?
Appending a shard suffix distributes writes and reads for a hot key across multiple physical partitions.
Flash Cards
Top DynamoDB anti-pattern? — Treating it like a relational DB — normalized tables with joins done in application code.
What is a hot partition? — A partition receiving disproportionate traffic due to a low-cardinality key, causing throttling.
Why avoid Scan for routine reads? — Scan reads (and bills for) every item; filters apply after reading, unlike a targeted Query.
What is single-table design? — Storing multiple entity types in one table via overloaded keys and GSIs to serve queries without joins.
How to fix a hot key? — Add a sharding suffix to the partition key to spread traffic across many partitions.