docs(8.0): consistency-model concept + snapshots guide — Db API replaces branching docs
This commit is contained in:
parent
e5feae4104
commit
cc8037db10
23 changed files with 1053 additions and 1871 deletions
282
docs/concepts/consistency-model.md
Normal file
282
docs/concepts/consistency-model.md
Normal file
|
|
@ -0,0 +1,282 @@
|
|||
---
|
||||
title: Consistency Model
|
||||
slug: concepts/consistency-model
|
||||
public: true
|
||||
category: concepts
|
||||
template: concept
|
||||
order: 4
|
||||
description: The exact guarantees behind Brainy's Db API — snapshot isolation, atomic transactions, two levels of compare-and-swap, time travel, retention, snapshots, and crash recovery.
|
||||
next:
|
||||
- guides/snapshots-and-time-travel
|
||||
- guides/optimistic-concurrency
|
||||
---
|
||||
|
||||
# Consistency Model
|
||||
|
||||
Brainy 8.0's consistency story rests on one mechanism: **generational MVCC**
|
||||
— multi-version concurrency control over immutable, generation-stamped
|
||||
records. It is exposed through a single value type, the **`Db`**: an
|
||||
immutable, point-in-time view of the whole store that you query like the
|
||||
live brain.
|
||||
|
||||
```typescript
|
||||
const db = brain.now() // pin the current state — O(1), no I/O
|
||||
|
||||
await brain.transact([
|
||||
{ op: 'update', id: invoiceId, metadata: { status: 'paid' } }
|
||||
])
|
||||
|
||||
await db.get(invoiceId) // still 'pending' — pinned, forever
|
||||
await brain.get(invoiceId) // 'paid' — live
|
||||
await db.release() // unpin when done
|
||||
```
|
||||
|
||||
This page states the guarantees precisely — what is promised, what it costs,
|
||||
and where the honest limits are. The design record is
|
||||
[ADR-001](../ADR-001-generational-mvcc.md); every guarantee below is proven
|
||||
by a dedicated test in `tests/integration/db-mvcc.test.ts`.
|
||||
|
||||
## The generation clock
|
||||
|
||||
A **monotonic generation counter** is the store's logical clock:
|
||||
|
||||
- It advances **once per committed `transact()` batch** and once per
|
||||
single-operation write (`add`/`update`/`delete`/`relate`/…).
|
||||
- `brain.generation()` reads it; it is persisted in the data directory and
|
||||
**never reissued** — not across restarts, and not across `restore()`
|
||||
(the counter is floored at its pre-restore value).
|
||||
|
||||
Every `Db` is pinned at one generation. `db.generation` and `db.timestamp`
|
||||
identify the view; `newerDb.since(olderDb)` returns exactly the entity and
|
||||
relationship ids that committed transactions touched between two views.
|
||||
|
||||
## Snapshot isolation for reads
|
||||
|
||||
**Guarantee:** a `Db` reads exactly the state at its pinned generation, no
|
||||
matter what commits afterwards — including deletes. There are no torn reads,
|
||||
no partially applied batches, and no drift over time.
|
||||
|
||||
- `brain.now()` pins the current generation in O(1).
|
||||
- `brain.transact()` returns a `Db` pinned at the freshly committed
|
||||
generation.
|
||||
- `brain.asOf(generation | Date | snapshotPath)` pins past state.
|
||||
|
||||
While nothing has committed past the pin, reads delegate to the live fast
|
||||
paths — pinning is free until history actually moves. Once later
|
||||
transactions commit, the view keeps serving the **full query surface** at
|
||||
its generation (see "Reading the past" below).
|
||||
|
||||
Writers are never blocked by readers and readers never block writers: a
|
||||
pinned view stays valid because nothing overwrites the immutable records it
|
||||
resolves from (the LMDB reader-pin model).
|
||||
|
||||
## Transaction atomicity
|
||||
|
||||
`brain.transact(ops)` executes a declarative batch — `add`, `update`,
|
||||
`remove`, `relate`, `unrelate` — **atomically as exactly one generation**:
|
||||
|
||||
```typescript
|
||||
const db = await brain.transact([
|
||||
{ op: 'add', id: orderId, type: NounType.Document, subtype: 'order', data: 'Order #1042' },
|
||||
{ op: 'add', id: itemId, type: NounType.Thing, subtype: 'line-item', data: 'Widget x3' },
|
||||
{ op: 'relate', from: orderId, to: itemId, type: VerbType.Contains, subtype: 'order-line' }
|
||||
], { meta: { author: 'order-service', requestId: 'req-9f2' } })
|
||||
|
||||
db.receipt.ids // resolved id per operation, in input order
|
||||
```
|
||||
|
||||
Either every operation applies, or none do and the store is byte-identical
|
||||
to its pre-transaction state. Operation semantics mirror the corresponding
|
||||
single-operation methods — validation, subtype enforcement, relationship
|
||||
deduplication, delete cascades — and later operations may reference ids
|
||||
created earlier in the same batch.
|
||||
|
||||
**The commit point is one atomic rename.** The durability protocol:
|
||||
|
||||
1. Before-images of every touched id are staged into an immutable
|
||||
generation directory and **fsynced**.
|
||||
2. The batch executes through the transaction manager (which has its own
|
||||
operation-level rollback for non-crash failures).
|
||||
3. The store manifest is replaced via atomic temp-file rename and fsynced.
|
||||
**The rename is the commit** — a generation is committed if and only if
|
||||
the manifest says so.
|
||||
|
||||
**Crash recovery:** on the next open, any staged generation above the
|
||||
manifest watermark is an uncommitted transaction; its before-images are
|
||||
restored (idempotently — recovery can itself crash and rerun) and derived
|
||||
indexes never observe the rolled-back state. A crash anywhere before the
|
||||
rename rolls back to the exact pre-transaction bytes; a crash after it keeps
|
||||
the transaction.
|
||||
|
||||
Transaction metadata (`meta`) is reified Datomic-style: recorded in an
|
||||
append-only transaction log readable via `brain.transactionLog()` — audit
|
||||
fields live in the database, not in commit messages.
|
||||
|
||||
## Two levels of compare-and-swap
|
||||
|
||||
Concurrent `transact()` calls commit serially (snapshot-isolated batches).
|
||||
For lost-update protection across a read–modify–write cycle, Brainy offers
|
||||
CAS at two granularities:
|
||||
|
||||
| Granularity | Mechanism | Conflict error | Use when |
|
||||
|---|---|---|---|
|
||||
| **Per entity** | `_rev` + `{ op: 'update', ifRev }` (also on `brain.update()`) | `RevisionConflictError` | "This entity must not have changed since I read it." |
|
||||
| **Whole store** | `transact(ops, { ifAtGeneration })` | `GenerationConflictError` | "*Nothing* may have committed since I read." |
|
||||
|
||||
```typescript
|
||||
const view = brain.now()
|
||||
const order = await view.get(orderId)
|
||||
|
||||
try {
|
||||
await brain.transact(
|
||||
[{ op: 'update', id: orderId, metadata: { total: recompute(order) }, ifRev: order._rev }],
|
||||
{ ifAtGeneration: view.generation }
|
||||
)
|
||||
} catch (err) {
|
||||
if (err instanceof GenerationConflictError) {
|
||||
// Something committed since the pin — re-read and retry.
|
||||
}
|
||||
} finally {
|
||||
await view.release()
|
||||
}
|
||||
```
|
||||
|
||||
An `ifRev` conflict on any operation rejects the **whole batch**; an
|
||||
`ifAtGeneration` conflict is detected before anything is staged. Both leave
|
||||
the store untouched and the generation counter unchanged. See
|
||||
[Optimistic concurrency with `_rev`](../guides/optimistic-concurrency.md)
|
||||
for the per-entity pattern in depth.
|
||||
|
||||
## Reading the past
|
||||
|
||||
`brain.asOf()` accepts a generation number, a `Date` (resolved through the
|
||||
transaction log to the newest generation committed at or before it), or a
|
||||
snapshot directory path. Historical views serve the **full query surface**
|
||||
— `get()`, `find()` in every mode, semantic search, graph traversal,
|
||||
cursors, aggregation — through two complementary paths:
|
||||
|
||||
- **Record path** (free): `get()`, metadata-level `find()`, and
|
||||
filter-based `related()` resolve directly through the immutable record
|
||||
layer. Ids untouched since the pin still ride the live fast paths.
|
||||
- **Index path** (paid once): index-accelerated queries — semantic/vector
|
||||
search, graph traversal, cursors, aggregation — are served by an
|
||||
**at-generation index materialization** built lazily on first use:
|
||||
Brainy reconstructs in-memory indexes over the exact record set at that
|
||||
generation. This costs O(n at the pinned generation) time and memory,
|
||||
**once per `Db`**, cached until `release()`. That is the open-core price
|
||||
of historical index queries, stated plainly.
|
||||
|
||||
A native index provider implementing the optional
|
||||
`VersionedIndexProvider` plugin capability serves the same historical reads
|
||||
from its retained index segments **without any rebuild** — the materializer
|
||||
is the correctness baseline, the provider is the accelerator. Semantics are
|
||||
identical on both paths.
|
||||
|
||||
### History granularity — the honest limit
|
||||
|
||||
Generation *records* are written per `transact()` batch only.
|
||||
Single-operation writes (`add`/`update`/`delete`/`relate`/… outside
|
||||
`transact()`) advance the generation counter — so watermarks and CAS stay
|
||||
sound — but do **not** stage before-images: they remain visible through
|
||||
earlier pins and are not reported by `db.since()`. Code that needs pinned
|
||||
isolation across its own writes uses `transact()`. This is the documented
|
||||
8.0 contract, not an accident.
|
||||
|
||||
## Speculative writes: `db.with()`
|
||||
|
||||
`db.with(ops)` returns a new `Db` whose reads see the operations applied
|
||||
**in memory, on top of the view** — Datomic's `with`. Nothing touches disk,
|
||||
the generation counter, or index providers:
|
||||
|
||||
```typescript
|
||||
const current = brain.now()
|
||||
const whatIf = await current.with([
|
||||
{ op: 'update', id: employeeId, metadata: { team: 'platform' } }
|
||||
])
|
||||
|
||||
await whatIf.find({ where: { team: 'platform' } }) // sees the change
|
||||
await brain.get(employeeId) // unchanged — nothing committed
|
||||
```
|
||||
|
||||
**The one boundary:** overlay entities carry no embeddings (`with()` never
|
||||
invokes the embedder), so index-accelerated queries and `persist()` on a
|
||||
speculative view throw `SpeculativeOverlayError` rather than returning
|
||||
silently incomplete results. `get()`, metadata-filter `find()`, and
|
||||
filter-based `related()` work fully on overlays. To get the full surface,
|
||||
commit the same operations with `brain.transact()`.
|
||||
|
||||
## Retention and compaction
|
||||
|
||||
Historical records cost disk space, so retention is explicit:
|
||||
|
||||
- Every live `Db` holds a refcounted **pin**; a record-set is never
|
||||
reclaimed while any pin could need it — pinned reads stay correct across
|
||||
compaction, always.
|
||||
- `brain.compactHistory({ retainGenerations?, retainMs? })` reclaims
|
||||
everything no retention rule and no pin protects, and records the
|
||||
**horizon** — `asOf()` below it throws `GenerationCompactedError`,
|
||||
explicitly, never partial data.
|
||||
- To keep a state readable forever, `persist()` it first: snapshots are
|
||||
self-contained and unaffected by compaction of the source store.
|
||||
|
||||
Release `Db` values you do not keep (including the ones `transact()`
|
||||
returns). A `FinalizationRegistry` backstop releases leaked pins at garbage
|
||||
collection, but explicit `release()` is what makes compaction
|
||||
deterministic.
|
||||
|
||||
## Durability: snapshots and restore
|
||||
|
||||
`db.persist(path)` cuts a **self-contained snapshot** under the store's
|
||||
commit mutex, so no commit or compaction can interleave. On filesystem
|
||||
storage it is built from **hard links**: because every data file is
|
||||
immutable-by-rename, linking is safe — the snapshot is created without
|
||||
copying entity data, shares disk space with the source, and later writes to
|
||||
the source can never alter it (rewrites swap inodes; the snapshot keeps the
|
||||
old bytes). Cross-device targets fall back to byte copies; in-memory stores
|
||||
serialize to the same directory layout, producing a real, durable store.
|
||||
|
||||
Two rules keep snapshots honest:
|
||||
|
||||
- `persist()` requires the view to still be the store's **latest**
|
||||
generation (a snapshot captures current bytes); a view that history has
|
||||
moved past throws `GenerationConflictError` instead of persisting the
|
||||
wrong state.
|
||||
- `brain.restore(path, { confirm: true })` replaces the store's entire
|
||||
state from a snapshot via byte copy (never links — the snapshot stays
|
||||
independent), rebuilds all indexes, and floors the generation counter so
|
||||
observed generation numbers are never reissued. Live pins do not survive
|
||||
a restore — release them first (a warning is logged when any exist).
|
||||
|
||||
`Brainy.load(path)` (or `brain.asOf(path)`) opens a snapshot as a
|
||||
self-contained **read-only** store with the full query surface, including
|
||||
vector search.
|
||||
|
||||
## What is not guaranteed
|
||||
|
||||
Stated plainly, so nothing surprises you in production:
|
||||
|
||||
- **Single-writer.** Brainy is a single-writer, many-reader database
|
||||
([multi-process model](./multi-process.md)). Transactions are atomic
|
||||
within one writer process — there is no distributed or cross-process
|
||||
transaction coordination.
|
||||
- **History granularity.** Only `transact()` batches produce historical
|
||||
records; single-operation writes between commits stay visible through
|
||||
earlier pins (see above).
|
||||
- **Compacted history is gone.** `asOf()` below the compaction horizon
|
||||
fails explicitly; persist what you must keep.
|
||||
- **Counter persistence is coalesced for single-operation writes.** Durable
|
||||
artifacts (records, manifests, snapshots) always persist the counter at
|
||||
their own commit points, so a crash inside the coalescing window can lose
|
||||
only counter values nothing durable ever referenced.
|
||||
- **Speculative overlays are metadata-only readers.** Index-accelerated
|
||||
queries on `with()` views throw rather than guess.
|
||||
|
||||
## Where to go next
|
||||
|
||||
- [Snapshots & Time Travel](../guides/snapshots-and-time-travel.md) — the
|
||||
recipes: backup, restore, time-travel debugging, what-if analysis, audit
|
||||
trails.
|
||||
- [Optimistic concurrency with `_rev`](../guides/optimistic-concurrency.md)
|
||||
— the per-entity CAS pattern.
|
||||
- [ADR-001: Generational MVCC](../ADR-001-generational-mvcc.md) — the full
|
||||
design record, including the persisted layout and the proof table.
|
||||
Loading…
Add table
Add a link
Reference in a new issue