15 KiB
ADR-001: Generational MVCC storage and the immutable Db API
Status: Accepted (ships in 8.0) Date: 2026-06-10
Context
Before 8.0, Brainy carried two overlapping version-control subsystems: a
copy-on-write branching layer (fork/checkout/commit/branches) and a
separate versioning subsystem (versions.save/list/compare/restore/prune),
plus a read-only historical adapter for commit-based time travel. Together
they were ~5,100 LOC of mechanism for one product need: read a consistent
past state while the store keeps moving, and snapshot/restore cheaply.
Neither subsystem gave a precise isolation guarantee. Reads raced in-place JSON overwrites, so a "snapshot" was only as immutable as the bytes it happened to share with the live store.
8.0 replaces both with one mechanism: generational MVCC over immutable,
generation-stamped records, exposed through a Datomic-style immutable
database value (Db). The same model is implemented natively by versioned
index providers (LSM snapshots), so semantics are identical on the pure-JS
path and the native path.
Decision
The model
- A monotonic u64 generation counter is the store's logical clock. It
advances once per committed
transact()batch and once per single-operation write (add/update/remove/relate/…), sobrain.generation()is always a meaningful watermark. It is persisted in_system/generation.jsonand never reissued for anything durable. brain.now()pins the current generation in O(1) and returns aDb— an immutable view. Pins are refcounted;db.release()(with aFinalizationRegistrybackstop for leaked values) ends the pin.brain.transact(ops, { meta, ifAtGeneration })commits a declarative batch atomically as exactly one generation, with whole-store compare-and-swap (ifAtGeneration→GenerationConflictError) and reified transaction metadata appended to_system/tx-log.jsonl.brain.asOf(generation | Date | snapshotPath)opens past state;db.with(ops)layers a speculative in-memory overlay (never touching disk, the counter, or index providers);db.persist(path)cuts an instant snapshot;brain.restore(path, { confirm: true })replaces state from one;Brainy.load(path)opens a snapshot read-only with the full query surface.
Persisted layout
All paths are storage-root-relative:
_system/generation.json { generation, updatedAt } atomic tmp+rename
_system/manifest.json { version, generation, atomic tmp+rename
committedAt, horizon } (the commit point)
_system/tx-log.jsonl one line per committed append-only
transact: { generation,
timestamp, meta? }
_generations/<N>/tx.json the generation-N delta: immutable
touched noun/verb ids + meta
_generations/<N>/prev/<id>.json before-image of <id> as of immutable
commit N (raw stored bytes;
null parts = file was absent)
Why per-generation deltas instead of a global id → latest generation
map in the manifest: a global map makes every commit O(all ids) — the
whole map must be rewritten to swap it atomically. The delta layout makes a
commit O(ids touched) and keeps the manifest a fixed-size watermark, while
point-in-time resolution stays correct (see "Read resolution" below). The
trade is that resolution at a pinned generation scans the deltas of later
commits — bounded by the number of commits since the pin, which is exactly
the window compaction keeps short.
Before-images are deliberately the only per-id records. They serve both roles the layer needs — the crash-recovery undo log and the point-in-time read source. After-images would duplicate state that is already readable (the canonical entity files hold the latest bytes; earlier states resolve from later before-images) and would double record I/O per commit.
Commit protocol (durability)
transact() commits under a store-wide mutex:
- CAS check. A stale
ifAtGenerationthrowsGenerationConflictErrorbefore anything is staged. - Reserve generation
N(counter increment). - Stage the undo log: write the before-image of every touched id plus
tx.jsonunder_generations/N/, then fsync the files and their directories. From this point, any crash is recoverable to the exact pre-transaction bytes. - Execute the planned batch through the TransactionManager (which has its own operation-level rollback for in-flight failures).
- Commit point: persist the counter, then write
_system/manifest.jsonvia atomic tmp+rename and fsync it. The rename is the commit: a generation directory is committed if and only ifN ≤ manifest.generation. - Append the tx-log line (advisory metadata — a crash between 5 and 6 keeps the transaction).
Crash recovery (on open): any _generations/<N> directory with
N > manifest.generation is an uncommitted transaction. Its before-images
are restored to the canonical paths (idempotently — recovery itself can
crash and rerun) and the directory is removed. Because recovery runs before
any index is built, and a recovery that rolled something back forces a full
index rebuild, derived indexes never observe rolled-back state. Reader-mode
instances skip recovery (readers never write; the next writer repairs).
A failed (non-crash) transaction takes the same staging directory down the abort path: the TransactionManager rolls back applied operations, the staging directory is removed, and the generation reservation is returned — a failed batch leaves the generation counter unchanged.
Read resolution at a pinned generation
The state of id X at pinned generation G is:
- the before-image stored by the first committed generation after G that touched X, or
- the live canonical bytes, when nothing after G touched X.
While nothing has committed past G, every read on the Db delegates to
the live fast paths untouched — now() adds no read overhead until history
actually moves.
Two read paths, one result set. get(), metadata-level find(), and
filter-based related() resolve directly through the record layer at any
reachable pinned generation — no extra cost beyond scanning the deltas of
later commits. Index-accelerated dimensions (semantic/vector search, graph
traversal, cursors, aggregation) are served by at-generation index
materialization: the first such query on a historical Db copies the
exact at-G record set (live bytes for ids untouched since the pin,
before-images for the rest; a final reconciliation pass runs under the
commit mutex so transactions racing the copy cannot skew it) into an
ephemeral in-memory store and opens a read-only engine over it — the same
vector/metadata/graph index classes the live brain uses, sharing the host's
embedder and aggregate definitions. The handle is cached on the Db and
freed by release().
Cost, stated plainly: materialization is O(n at G) time and memory,
once per Db. That is the open-core price of historical index queries. A
native VersionedIndexProvider (isGenerationVisible() + pins over
retained LSM segments) serves the same reads with no rebuild at all — the
materializer is the correctness baseline, the provider is the accelerator.
The one remaining boundary. Speculative with() overlays throw
SpeculativeOverlayError for index-accelerated queries and persist():
overlay entities carry no embeddings (with() never invokes the embedder),
so a "full" index query over an overlay would silently exclude the
overlay's own entities. Commit with transact() to get the full surface.
History granularity. Generation records are written per transact()
batch only. Single-operation writes advance the counter (so watermarks and
CAS stay sound) but do not stage before-images: they remain visible through
earlier pins and are not reported by db.since(). Code that needs pinned
isolation across its writes uses transact() — that is the documented 8.0
contract, stated rather than papered over.
Pinning, retention, compaction
- Each live
Dbholds one refcounted pin on its generation (plus apin(generation)on every registeredVersionedIndexProvider, whose explicit pin lifetime overrides any time-based snapshot retention the provider has). compactHistory({ retainGenerations?, retainMs? })reclaims record-sets. A record-setNis reclaimed only whenNis at or below every live pin — deletingNcan only break readers pinned belowN, because resolution reads before-images from generations strictly greater than the pin. With both retention options supplied, both must allow the reclaim.- The manifest records the horizon (highest reclaimed generation).
Generations below the horizon are unreachable;
asOf()on them throwsGenerationCompactedError. The horizon itself stays reachable, resolved from the record-sets above it. To keep a generation readable forever,persist()it first — snapshots are self-contained.
Snapshots and restore
db.persist(path) flushes indexes, then cuts the snapshot under the
store's commit mutex (no commit, compaction, or counter write can
interleave). On filesystem storage it is a hard-link farm: every data
file is immutable-by-rename, so linking is safe — later rewrites swap
inodes and the snapshot keeps the old bytes. The two exceptions are handled
explicitly: the append-in-place tx-log is byte-copied, and process-local
lock state is excluded. Cross-device targets (and filesystems that refuse
links) fall back to per-file byte copies. In-memory stores serialize to the
same directory layout, so persisting a memory brain produces a real,
durable, loadable store.
persist() requires the view to still be the store's latest generation
(a snapshot captures current bytes); a view that history has moved past
throws rather than persisting the wrong state.
restore(path, { confirm: true }) replaces the store's contents from a
snapshot via byte copy (never links — the snapshot stays independent),
reloads all adapter-internal derived state, rebuilds all indexes, and
floors the generation counter at its pre-restore value so observed
generation numbers are never reissued. Live pins do not survive a restore;
a warning is logged when any exist.
Versioned index providers
Native index providers may implement the optional 4-method
VersionedIndexProvider capability (generation(),
isGenerationVisible(), pin(), release() — BigInt generations at the
boundary). The locked consistency model: providers are post-commit
appliers. The storage-record commit is the source of truth; provider
index state is derived. On open, a provider behind the committed watermark
replays the gap from storage (or requests a rebuild) — there are no
provider rollback hooks, because uncommitted transactions are repaired at
the storage layer before any index opens. Speculative with() overlays
never reach providers.
Guarantees (and their proofs)
Each stated guarantee has a test that proves it, not merely exercises it
(tests/integration/db-mvcc.test.ts, plus
tests/unit/db/generationStore.test.ts for the record layer in isolation):
| Guarantee | Proof |
|---|---|
Snapshot isolation: a pinned Db reads exactly its pinned state, forever |
proof 1 (200 mutations, including deletes, against a pinned view) |
| Atomicity: a failing batch applies nothing; generation unchanged | proofs 2a/2b/2c (plan-time failure, injected execution-phase storage failure, ifRev conflict) |
| Whole-store CAS | proof 3 (ifAtGeneration success + conflict with exact expected/actual) |
| Snapshot integrity under source mutation (hard-link safety) | proofs 4a/4b/4c |
| Compaction never breaks a pinned read; release enables reclaim | proof 5 |
with() overlays touch nothing durable |
proof 6 |
| Generation monotonicity across close/reopen | proof 7 |
| Crash before the manifest rename recovers to exact pre-transaction state through the real recovery path | proof 8 (fault injection that skips abort cleanup, exactly as a dead process would) |
| Balanced provider pin/release lockstep | proof 9 |
One deliberate softness: single-operation generation bumps persist the counter coalesced (per write burst), not per write. Durable artifacts — records, manifests, snapshots — always persist the counter synchronously at their own commit points, so a crash inside the coalescing window can lose only counter values that nothing durable ever referenced.
Failure modes
| Failure | Outcome |
|---|---|
| Crash before staging completes | Partial staging directory > manifest watermark → removed on next open; canonical state untouched. |
| Crash after staging, before/during batch execution | Before-images restored on next open; indexes rebuilt; byte-identical pre-transaction state. |
| Crash after execution, before manifest rename | Same as above — the rename is the only commit point. |
| Crash after manifest rename, before tx-log append | Transaction kept (committed); tx-log misses one advisory line; asOf(Date) resolution for that commit falls back to neighboring entries. |
| Batch fails mid-execution (no crash) | TransactionManager operation rollback + staging-directory removal + reservation return; generation unchanged. |
asOf() below the compaction horizon |
GenerationCompactedError — explicit, never partial data. |
Index-accelerated query on a with() overlay |
SpeculativeOverlayError — explicit, never silently-incomplete results (overlay entities carry no embeddings). |
persist() of a view history has moved past |
GenerationConflictError — a snapshot captures current bytes; persist before further writes. |
| Torn trailing tx-log line (crashed append) | Tolerated; unparseable lines are skipped by readers. |
Lineage
The design is an assembly of well-understood prior art, chosen for being boring where it counts:
- Datomic — the database-as-a-value: an immutable
Dbyou query, withwith()for speculation and reified transaction metadata instead of commit messages. - LMDB — reader pins: readers never block writers; a reader's view stays valid because nothing overwrites the pages (here: records) it references; reclamation waits for the last reader.
- LSM trees / Cassandra — immutable segments make snapshots hard links and make compaction a retention policy instead of a locking problem.
Consequences
- One mechanism replaces the COW and versioning subsystems (their removal is the companion change to this ADR).
- In-place branch switching (
checkout) is gone by design; the replacement is opening a persisted snapshot as a separate instance — a name→path mapping where a product needs named branches. - Every commit pays O(ids touched) extra writes (before-images + delta + manifest). Single-operation writes pay only an in-memory counter bump with coalesced persistence.
- The full query surface works at every reachable pinned generation.
Record-path reads (
get, metadatafind, filterrelated) are effectively free; index-accelerated historical queries pay a one-time O(n at G) materialization perDbon the open-core path (freed onrelease()), and run rebuild-free on a nativeVersionedIndexProvider.