How does Elasticsearch handle updates and why are documents immutable?
Learn why Elasticsearch documents are immutable, how updates become delete-and-reindex, how segment merges reclaim space, and how concurrency control works.
Expected Interview Answer
Elasticsearch documents are immutable because they live in Lucene segments that are written once and never modified; an update is really a delete-and-reindex — the old document is marked deleted and a new version is written, with the space reclaimed later during segment merges.
Underneath, Elasticsearch stores data in immutable Lucene segments. When you update a document, Elasticsearch fetches the current version, applies your change, writes a brand-new document with an incremented _version, and flags the previous copy as deleted in the segment's live-docs bitmap — it is not erased in place. Those tombstoned documents keep consuming space until background segment merges rewrite the live documents into new segments and drop the deleted ones. Immutability is what makes segments cacheable, safe for concurrent reads without locking, and cheap to replicate, while optimistic concurrency via _seq_no and _primary_term prevents lost updates.
- Immutable segments are lock-free and safe for concurrent reads
- Segments can be aggressively cached at the OS and file level
- Simplifies replication since files never change after writing
- Optimistic concurrency prevents conflicting updates being lost
- Append-only writes are fast and crash-recoverable via the translog
AI Mentor Explanation
It is like an official scorebook written in ink: you never erase an entry, so to correct a score you strike the old line, note it as void, and write a fresh corrected line below, and only when the book is recopied clean are the struck-out lines finally left out entirely.
Step-by-Step Explanation
Step 1
Read current version
An update first retrieves the existing document along with its _seq_no and _primary_term.
Step 2
Apply and reindex
The change is applied and a new document is written with an incremented _version.
Step 3
Tombstone the old copy
The previous document is flagged as deleted in the segment's live-docs bitmap, not erased in place.
Step 4
Write to segment and translog
The new document lands in an in-memory buffer and the translog for durability, then a refresh makes it searchable.
Step 5
Reclaim via merge
Background segment merges rewrite live documents into new segments and drop the tombstoned ones, freeing space.
What Interviewer Expects
- That Lucene segments are immutable and write-once
- That an update is a delete-and-reindex, not an in-place edit
- How deleted documents are tombstoned and later merged away
- The role of optimistic concurrency (_seq_no, _primary_term)
- Why immutability enables caching, lock-free reads, and easy replication
Common Mistakes
- Believing Elasticsearch edits documents in place
- Thinking a delete immediately frees disk space
- Ignoring that frequent updates bloat segments until merges run
- Confusing document _version with optimistic concurrency control
- Assuming updates avoid rewriting the whole document
Best Answer (HR Friendly)
“Elasticsearch never edits a stored document in place. When you update one, it quietly writes a fresh copy and marks the old one as deleted, and a background cleanup later reclaims the space. Keeping stored data unchangeable is what lets Elasticsearch cache it, read it without locking, and copy it safely across the cluster.”
Code Example
# Partial update — internally a delete + reindex
POST /products/_update/42
{
"doc": { "price": 199 }
}
# Safe concurrent update using seq_no / primary_term:
PUT /products/_doc/42?if_seq_no=17&if_primary_term=3
{
"title": "Wireless Headphones",
"price": 199
}
# Fails with 409 if another update changed the doc first,
# preventing a lost update. The old copy is tombstoned
# and reclaimed during a later segment merge.Follow-up Questions
- What is a Lucene segment and why is it immutable?
- How do segment merges reclaim space from deleted documents?
- What roles do _seq_no and _primary_term play in updates?
- How does the translog provide durability before a refresh?
- Why can heavy update workloads hurt Elasticsearch performance?
MCQ Practice
1. What actually happens when you update an Elasticsearch document?
Segments are immutable, so an update tombstones the old document and writes a new version.
2. When is the space from a deleted document reclaimed?
Tombstoned documents keep space until a segment merge rewrites live docs and drops the deleted ones.
3. Which fields provide optimistic concurrency control for updates?
if_seq_no and if_primary_term let an update fail with a conflict if the document changed meanwhile.
Flash Cards
Why are Elasticsearch documents immutable? — They live in write-once Lucene segments that are never modified after being written.
What is an update really? — A delete-and-reindex: the old copy is tombstoned and a new version is written.
When is deleted-document space freed? — During background segment merges that rewrite live documents into new segments.
How are lost updates prevented? — Optimistic concurrency via _seq_no and _primary_term rejects conflicting writes with a 409.