263 lines
14 KiB
Markdown
263 lines
14 KiB
Markdown
|
|
# ADR-001: Generational MVCC storage and the immutable Db API
|
||
|
|
|
||
|
|
**Status:** Accepted (ships in 8.0)
|
||
|
|
**Date:** 2026-06-10
|
||
|
|
|
||
|
|
## Context
|
||
|
|
|
||
|
|
Before 8.0, Brainy carried two overlapping version-control subsystems: a
|
||
|
|
copy-on-write branching layer (`fork`/`checkout`/`commit`/`branches`) and a
|
||
|
|
separate versioning subsystem (`versions.save/list/compare/restore/prune`),
|
||
|
|
plus a read-only historical adapter for commit-based time travel. Together
|
||
|
|
they were ~5,100 LOC of mechanism for one product need: *read a consistent
|
||
|
|
past state while the store keeps moving, and snapshot/restore cheaply.*
|
||
|
|
|
||
|
|
Neither subsystem gave a precise isolation guarantee. Reads raced in-place
|
||
|
|
JSON overwrites, so a "snapshot" was only as immutable as the bytes it
|
||
|
|
happened to share with the live store.
|
||
|
|
|
||
|
|
8.0 replaces both with **one mechanism**: generational MVCC over immutable,
|
||
|
|
generation-stamped records, exposed through a Datomic-style immutable
|
||
|
|
database value (`Db`). The same model is implemented natively by versioned
|
||
|
|
index providers (LSM snapshots), so semantics are identical on the pure-JS
|
||
|
|
path and the native path.
|
||
|
|
|
||
|
|
## Decision
|
||
|
|
|
||
|
|
### The model
|
||
|
|
|
||
|
|
- A **monotonic u64 generation counter** is the store's logical clock. It
|
||
|
|
advances once per committed `transact()` batch and once per
|
||
|
|
single-operation write (`add`/`update`/`delete`/`relate`/…), so
|
||
|
|
`brain.generation()` is always a meaningful watermark. It is persisted in
|
||
|
|
`_system/generation.json` and never reissued for anything durable.
|
||
|
|
- `brain.now()` **pins** the current generation in O(1) and returns a `Db` —
|
||
|
|
an immutable view. Pins are refcounted; `db.release()` (with a
|
||
|
|
`FinalizationRegistry` backstop for leaked values) ends the pin.
|
||
|
|
- `brain.transact(ops, { meta, ifAtGeneration })` commits a declarative
|
||
|
|
batch atomically as **exactly one generation**, with whole-store
|
||
|
|
compare-and-swap (`ifAtGeneration` → `GenerationConflictError`) and
|
||
|
|
reified transaction metadata appended to `_system/tx-log.jsonl`.
|
||
|
|
- `brain.asOf(generation | Date | snapshotPath)` opens past state;
|
||
|
|
`db.with(ops)` layers a speculative in-memory overlay (never touching
|
||
|
|
disk, the counter, or index providers); `db.persist(path)` cuts an
|
||
|
|
instant snapshot; `brain.restore(path, { confirm: true })` replaces state
|
||
|
|
from one; `Brainy.load(path)` opens a snapshot read-only with the full
|
||
|
|
query surface.
|
||
|
|
|
||
|
|
### Persisted layout
|
||
|
|
|
||
|
|
All paths are storage-root-relative:
|
||
|
|
|
||
|
|
```
|
||
|
|
_system/generation.json { generation, updatedAt } atomic tmp+rename
|
||
|
|
_system/manifest.json { version, generation, atomic tmp+rename
|
||
|
|
committedAt, horizon } (the commit point)
|
||
|
|
_system/tx-log.jsonl one line per committed append-only
|
||
|
|
transact: { generation,
|
||
|
|
timestamp, meta? }
|
||
|
|
_generations/<N>/tx.json the generation-N delta: immutable
|
||
|
|
touched noun/verb ids + meta
|
||
|
|
_generations/<N>/prev/<id>.json before-image of <id> as of immutable
|
||
|
|
commit N (raw stored bytes;
|
||
|
|
null parts = file was absent)
|
||
|
|
```
|
||
|
|
|
||
|
|
**Why per-generation deltas instead of a global `id → latest generation`
|
||
|
|
map in the manifest:** a global map makes every commit O(all ids) — the
|
||
|
|
whole map must be rewritten to swap it atomically. The delta layout makes a
|
||
|
|
commit O(ids touched) and keeps the manifest a fixed-size watermark, while
|
||
|
|
point-in-time resolution stays correct (see "Read resolution" below). The
|
||
|
|
trade is that resolution at a pinned generation scans the deltas of later
|
||
|
|
commits — bounded by the number of commits since the pin, which is exactly
|
||
|
|
the window compaction keeps short.
|
||
|
|
|
||
|
|
Before-images are deliberately the *only* per-id records. They serve both
|
||
|
|
roles the layer needs — the crash-recovery undo log and the point-in-time
|
||
|
|
read source. After-images would duplicate state that is already readable
|
||
|
|
(the canonical entity files hold the latest bytes; earlier states resolve
|
||
|
|
from later before-images) and would double record I/O per commit.
|
||
|
|
|
||
|
|
### Commit protocol (durability)
|
||
|
|
|
||
|
|
`transact()` commits under a store-wide mutex:
|
||
|
|
|
||
|
|
1. **CAS check.** A stale `ifAtGeneration` throws `GenerationConflictError`
|
||
|
|
before anything is staged.
|
||
|
|
2. **Reserve** generation `N` (counter increment).
|
||
|
|
3. **Stage the undo log:** write the before-image of every touched id plus
|
||
|
|
`tx.json` under `_generations/N/`, then **fsync** the files and their
|
||
|
|
directories. From this point, any crash is recoverable to the exact
|
||
|
|
pre-transaction bytes.
|
||
|
|
4. **Execute** the planned batch through the TransactionManager (which has
|
||
|
|
its own operation-level rollback for in-flight failures).
|
||
|
|
5. **Commit point:** persist the counter, then write `_system/manifest.json`
|
||
|
|
via atomic tmp+rename and fsync it. The rename *is* the commit: a
|
||
|
|
generation directory is committed if and only if `N ≤
|
||
|
|
manifest.generation`.
|
||
|
|
6. Append the tx-log line (advisory metadata — a crash between 5 and 6
|
||
|
|
keeps the transaction).
|
||
|
|
|
||
|
|
**Crash recovery (on open):** any `_generations/<N>` directory with
|
||
|
|
`N > manifest.generation` is an uncommitted transaction. Its before-images
|
||
|
|
are restored to the canonical paths (idempotently — recovery itself can
|
||
|
|
crash and rerun) and the directory is removed. Because recovery runs before
|
||
|
|
any index is built, and a recovery that rolled something back forces a full
|
||
|
|
index rebuild, derived indexes never observe rolled-back state. Reader-mode
|
||
|
|
instances skip recovery (readers never write; the next writer repairs).
|
||
|
|
|
||
|
|
A failed (non-crash) transaction takes the same staging directory down the
|
||
|
|
abort path: the TransactionManager rolls back applied operations, the
|
||
|
|
staging directory is removed, and the generation reservation is returned —
|
||
|
|
a failed batch leaves the generation counter unchanged.
|
||
|
|
|
||
|
|
### Read resolution at a pinned generation
|
||
|
|
|
||
|
|
The state of id X at pinned generation G is:
|
||
|
|
|
||
|
|
- the before-image stored by the **first committed generation after G that
|
||
|
|
touched X**, or
|
||
|
|
- the live canonical bytes, when nothing after G touched X.
|
||
|
|
|
||
|
|
While nothing has committed past G, *every* read on the `Db` delegates to
|
||
|
|
the live fast paths untouched — `now()` adds no read overhead until history
|
||
|
|
actually moves.
|
||
|
|
|
||
|
|
**Honesty boundary.** `get()`, metadata-level `find()`, and `related()` are
|
||
|
|
fully correct at any reachable pinned generation. Index-accelerated
|
||
|
|
dimensions (semantic/vector search, graph traversal, cursors, aggregation)
|
||
|
|
are served by live indexes that represent only the current generation, so
|
||
|
|
at a historical generation they throw
|
||
|
|
`NotYetSupportedAtHistoricalGenerationError` — never silently-wrong
|
||
|
|
results. The escape hatch is `persist()` + `Brainy.load()`, which rebuilds
|
||
|
|
indexes from the snapshot and supports the full query surface. Rebuilding
|
||
|
|
indexes from a pinned record set in-place (the standard LSM answer) is the
|
||
|
|
documented follow-up; versioned native providers already expose
|
||
|
|
`isGenerationVisible()` for provider-accelerated historical reads.
|
||
|
|
|
||
|
|
**History granularity.** Generation *records* are written per `transact()`
|
||
|
|
batch only. Single-operation writes advance the counter (so watermarks and
|
||
|
|
CAS stay sound) but do not stage before-images: they remain visible through
|
||
|
|
earlier pins and are not reported by `db.since()`. Code that needs pinned
|
||
|
|
isolation across its writes uses `transact()` — that is the documented 8.0
|
||
|
|
contract, stated rather than papered over.
|
||
|
|
|
||
|
|
### Pinning, retention, compaction
|
||
|
|
|
||
|
|
- Each live `Db` holds one refcounted pin on its generation (plus a
|
||
|
|
`pin(generation)` on every registered `VersionedIndexProvider`, whose
|
||
|
|
explicit pin lifetime overrides any time-based snapshot retention the
|
||
|
|
provider has).
|
||
|
|
- `compactHistory({ retainGenerations?, retainMs? })` reclaims record-sets.
|
||
|
|
A record-set `N` is reclaimed only when `N` is at or below **every** live
|
||
|
|
pin — deleting `N` can only break readers pinned *below* `N`, because
|
||
|
|
resolution reads before-images from generations strictly greater than the
|
||
|
|
pin. With both retention options supplied, both must allow the reclaim.
|
||
|
|
- The manifest records the **horizon** (highest reclaimed generation).
|
||
|
|
Generations below the horizon are unreachable; `asOf()` on them throws
|
||
|
|
`GenerationCompactedError`. The horizon itself stays reachable, resolved
|
||
|
|
from the record-sets above it. To keep a generation readable forever,
|
||
|
|
`persist()` it first — snapshots are self-contained.
|
||
|
|
|
||
|
|
### Snapshots and restore
|
||
|
|
|
||
|
|
`db.persist(path)` flushes indexes, then cuts the snapshot under the
|
||
|
|
store's commit mutex (no commit, compaction, or counter write can
|
||
|
|
interleave). On filesystem storage it is a **hard-link farm**: every data
|
||
|
|
file is immutable-by-rename, so linking is safe — later rewrites swap
|
||
|
|
inodes and the snapshot keeps the old bytes. The two exceptions are handled
|
||
|
|
explicitly: the append-in-place tx-log is byte-copied, and process-local
|
||
|
|
lock state is excluded. Cross-device targets (and filesystems that refuse
|
||
|
|
links) fall back to per-file byte copies. In-memory stores serialize to the
|
||
|
|
same directory layout, so persisting a memory brain produces a real,
|
||
|
|
durable, loadable store.
|
||
|
|
|
||
|
|
`persist()` requires the view to still be the store's latest generation
|
||
|
|
(a snapshot captures current bytes); a view that history has moved past
|
||
|
|
throws rather than persisting the wrong state.
|
||
|
|
|
||
|
|
`restore(path, { confirm: true })` replaces the store's contents from a
|
||
|
|
snapshot via byte copy (never links — the snapshot stays independent),
|
||
|
|
reloads all adapter-internal derived state, rebuilds all indexes, and
|
||
|
|
floors the generation counter at its pre-restore value so observed
|
||
|
|
generation numbers are never reissued. Live pins do not survive a restore;
|
||
|
|
a warning is logged when any exist.
|
||
|
|
|
||
|
|
### Versioned index providers
|
||
|
|
|
||
|
|
Native index providers may implement the optional 4-method
|
||
|
|
`VersionedIndexProvider` capability (`generation()`,
|
||
|
|
`isGenerationVisible()`, `pin()`, `release()` — BigInt generations at the
|
||
|
|
boundary). The locked consistency model: providers are **post-commit
|
||
|
|
appliers**. The storage-record commit is the source of truth; provider
|
||
|
|
index state is derived. On open, a provider behind the committed watermark
|
||
|
|
replays the gap from storage (or requests a rebuild) — there are no
|
||
|
|
provider rollback hooks, because uncommitted transactions are repaired at
|
||
|
|
the storage layer before any index opens. Speculative `with()` overlays
|
||
|
|
never reach providers.
|
||
|
|
|
||
|
|
## Guarantees (and their proofs)
|
||
|
|
|
||
|
|
Each stated guarantee has a test that proves it, not merely exercises it
|
||
|
|
(`tests/integration/db-mvcc.test.ts`, plus
|
||
|
|
`tests/unit/db/generationStore.test.ts` for the record layer in isolation):
|
||
|
|
|
||
|
|
| Guarantee | Proof |
|
||
|
|
|---|---|
|
||
|
|
| Snapshot isolation: a pinned `Db` reads exactly its pinned state, forever | proof 1 (200 mutations, including deletes, against a pinned view) |
|
||
|
|
| Atomicity: a failing batch applies nothing; generation unchanged | proofs 2a/2b/2c (plan-time failure, injected execution-phase storage failure, `ifRev` conflict) |
|
||
|
|
| Whole-store CAS | proof 3 (`ifAtGeneration` success + conflict with exact expected/actual) |
|
||
|
|
| Snapshot integrity under source mutation (hard-link safety) | proofs 4a/4b/4c |
|
||
|
|
| Compaction never breaks a pinned read; release enables reclaim | proof 5 |
|
||
|
|
| `with()` overlays touch nothing durable | proof 6 |
|
||
|
|
| Generation monotonicity across close/reopen | proof 7 |
|
||
|
|
| Crash before the manifest rename recovers to exact pre-transaction state through the real recovery path | proof 8 (fault injection that skips abort cleanup, exactly as a dead process would) |
|
||
|
|
| Balanced provider pin/release lockstep | proof 9 |
|
||
|
|
|
||
|
|
One deliberate softness: single-operation generation bumps persist the
|
||
|
|
counter coalesced (per write burst), not per write. Durable artifacts —
|
||
|
|
records, manifests, snapshots — always persist the counter synchronously at
|
||
|
|
their own commit points, so a crash inside the coalescing window can lose
|
||
|
|
only counter values that nothing durable ever referenced.
|
||
|
|
|
||
|
|
## Failure modes
|
||
|
|
|
||
|
|
| Failure | Outcome |
|
||
|
|
|---|---|
|
||
|
|
| Crash before staging completes | Partial staging directory > manifest watermark → removed on next open; canonical state untouched. |
|
||
|
|
| Crash after staging, before/during batch execution | Before-images restored on next open; indexes rebuilt; byte-identical pre-transaction state. |
|
||
|
|
| Crash after execution, before manifest rename | Same as above — the rename is the only commit point. |
|
||
|
|
| Crash after manifest rename, before tx-log append | Transaction kept (committed); tx-log misses one advisory line; `asOf(Date)` resolution for that commit falls back to neighboring entries. |
|
||
|
|
| Batch fails mid-execution (no crash) | TransactionManager operation rollback + staging-directory removal + reservation return; generation unchanged. |
|
||
|
|
| `asOf()` below the compaction horizon | `GenerationCompactedError` — explicit, never partial data. |
|
||
|
|
| Index-accelerated query at a historical pin | `NotYetSupportedAtHistoricalGenerationError` — explicit, never silently-wrong results. |
|
||
|
|
| Torn trailing tx-log line (crashed append) | Tolerated; unparseable lines are skipped by readers. |
|
||
|
|
|
||
|
|
## Lineage
|
||
|
|
|
||
|
|
The design is an assembly of well-understood prior art, chosen for being
|
||
|
|
boring where it counts:
|
||
|
|
|
||
|
|
- **Datomic** — the database-as-a-value: an immutable `Db` you query, with
|
||
|
|
`with()` for speculation and reified transaction metadata instead of
|
||
|
|
commit messages.
|
||
|
|
- **LMDB** — reader pins: readers never block writers; a reader's view
|
||
|
|
stays valid because nothing overwrites the pages (here: records) it
|
||
|
|
references; reclamation waits for the last reader.
|
||
|
|
- **LSM trees / Cassandra** — immutable segments make snapshots hard links
|
||
|
|
and make compaction a retention policy instead of a locking problem.
|
||
|
|
|
||
|
|
## Consequences
|
||
|
|
|
||
|
|
- One mechanism replaces the COW and versioning subsystems (their removal
|
||
|
|
is the companion change to this ADR).
|
||
|
|
- In-place branch switching (`checkout`) is gone by design; the replacement
|
||
|
|
is opening a persisted snapshot as a separate instance — a name→path
|
||
|
|
mapping where a product needs named branches.
|
||
|
|
- Every commit pays O(ids touched) extra writes (before-images + delta +
|
||
|
|
manifest). Single-operation writes pay only an in-memory counter bump
|
||
|
|
with coalesced persistence.
|
||
|
|
- Historical index-accelerated queries are an explicit error until the
|
||
|
|
asOf-rebuild ships; correctness-critical historical reads (`get`,
|
||
|
|
metadata `find`, `related`) work today and forever at any reachable pin.
|