docs(8.0): consistency-model concept + snapshots guide — Db API replaces branching docs

This commit is contained in:
David Snelling 2026-06-11 08:37:26 -07:00
parent e5feae4104
commit cc8037db10
23 changed files with 1053 additions and 1871 deletions

View file

@ -134,8 +134,8 @@ console.log(health.overall)
await reader.close()
```
Every mutation method (`add`, `update`, `delete`, `relate`, `commit`,
`fork`, `branch`, ...) throws on a read-only instance with a clear message.
Every mutation method (`add`, `update`, `delete`, `relate`, `transact`,
`restore`, ...) throws on a read-only instance with a clear message.
## Backups
@ -148,9 +148,10 @@ brainy inspect backup /data/brain /backups/brain-2026-05-15.tar
```
For periodic backups (hourly, daily), schedule this via cron or your
container scheduler. For point-in-time recovery, use Brainy's COW
`commit()` API — the snapshots there are content-addressed and never
overwritten.
container scheduler. For point-in-time recovery, use the Db API's
`db.persist(path)` — a self-contained hard-link snapshot that later writes
can never alter, restorable with `brain.restore(path, { confirm: true })`.
See [Snapshots & Time Travel](./snapshots-and-time-travel.md).
## Comparing two stores

View file

@ -7,8 +7,8 @@ template: guide
order: 8
description: Use the per-entity `_rev` counter and `update({ ifRev })` to coordinate concurrent writes safely. Covers the lock pattern, idempotent inserts with `ifAbsent`, and recovery on conflict.
next:
- concepts/consistency-model
- guides/find-limits
- api/README
---
# Optimistic concurrency with `_rev`
@ -25,7 +25,7 @@ Brainy 7.31.0 adds a per-entity revision counter so multiple writers can coordin
| `add({ id, ifAbsent: true })` | By-ID idempotent insert. Returns the existing `id` if one is already present; no throw, no overwrite. |
| `addMany({ items, ifAbsent: true })` | Applies `ifAbsent` to every item. Per-item `ifAbsent` overrides the batch flag. |
`_rev` is independent of `brain.versions.*` (named snapshots), of `brain.fork()` / branches (COW), and of VFS file versioning. They all coexist; `_rev` is the per-write counter used for CAS.
`_rev` is the **per-entity** counter. Its store-wide counterpart is the generation counter behind the [Db API](../concepts/consistency-model.md): `brain.transact(ops, { ifAtGeneration })` is CAS over the whole store, `update({ ifRev })` (and `ifRev` on `transact()` update operations) is CAS over one entity.
## The lock pattern
@ -134,42 +134,49 @@ await brain.addIfMissing({ // ← not a real API
})
```
It's race-prone outside a transaction: two concurrent imports both see "not found," both insert, you get duplicates. Without a unique-index primitive (which Brainy doesn't have today), this pattern needs to live inside a transaction:
It's race-prone as a plain read-then-write: two concurrent imports both see "not found," both insert, you get duplicates. Without a unique-index primitive (which Brainy doesn't have today), close the race with whole-store CAS — read at a pinned generation, then commit only if nothing moved:
```ts
// The race-safe shape, available once brain.transact() ships in 8.0.
await brain.transact(async tx => {
const existing = await tx.find({
type: 'Person',
where: { email: 'x@y.com' },
limit: 1
})
if (existing.length === 0) {
await tx.add({ type: 'Person', data: '...', metadata: { email: 'x@y.com' } })
import { GenerationConflictError } from '@soulcraft/brainy'
async function addIfMissingByEmail(email: string, data: string) {
for (let attempt = 0; attempt < 5; attempt++) {
const db = brain.now()
try {
const existing = await db.find({
type: NounType.Person,
where: { email },
limit: 1
})
if (existing.length > 0) return existing[0].id
const committed = await brain.transact(
[{ op: 'add', type: NounType.Person, subtype: 'customer', data, metadata: { email } }],
{ ifAtGeneration: db.generation } // rejects if ANYTHING committed since the read
)
return committed.receipt!.ids[0]
} catch (err) {
if (err instanceof GenerationConflictError) continue // world moved — re-read + retry
throw err
} finally {
await db.release()
}
}
})
throw new Error('addIfMissingByEmail conflict after 5 attempts')
}
```
For 7.31.0, lean on `ifAbsent` when you control the ID, and accept the inherent dedup race when you don't. The 8.0 `brain.transact()` is the natural home for the attribute-based variant.
`ifAtGeneration` is deliberately coarse — *any* committed write invalidates it — so keep the retry bound. When you control the ID, `ifAbsent` stays the cheaper tool.
## How `_rev` interacts with the other versioning systems
## How `_rev` relates to generations
Brainy has several persistence primitives that all touch the word "version" in different ways. `_rev` is independent of every one of them:
Brainy 8.0 has exactly two write-coordination counters, at two granularities:
| System | What it tracks | When it advances |
|---|---|---|
| **`_rev`** (new in 7.31.0) | Per-entity write counter | Auto, on every successful `update()` |
| **`brain.versions.save()`** | Named snapshots per entity | Explicit — you call `save()` with a tag |
| **`brain.fork()` / branches** | Whole-brain copy-on-write | Explicit — you call `fork(name)` |
| **VFS file versioning** | Per-VFS-file snapshots (Document entity) | Same as `brain.versions.save()` |
| Counter | Scope | What it tracks | CAS surface | Conflict error |
|---|---|---|---|---|
| **`_rev`** | One entity | Per-entity write count, bumped on every successful update | `update({ ifRev })`, `{ op: 'update', ifRev }` in `transact()` | `RevisionConflictError` |
| **Generation** | Whole store | One tick per committed `transact()` batch or single-operation write | `transact(ops, { ifAtGeneration })` | `GenerationConflictError` |
If you're on a branch, each branch has its own copy of every entity (that's what COW means), so each branch has its own `_rev` per entity — same as every other field. Snapshots taken via `brain.versions.save()` capture the entity at that moment including its `_rev` at that time; the snapshot's own `version: number` is the snapshot index, distinct from `_rev`.
They compose: a `transact()` batch can carry per-entity `ifRev` checks *and* a whole-store `ifAtGeneration`; any failed check rejects the entire batch before anything is staged. Generations also power snapshots and time travel (`brain.now()`, `brain.asOf()`, `db.persist()`) — see the [consistency model](../concepts/consistency-model.md) and [Snapshots & Time Travel](./snapshots-and-time-travel.md).
## What's coming in 8.0
Brainy 8.0 ships a Datomic-style immutable `Db` API where `brain.transact(tx)` is the public multi-write atomicity primitive. Two things change for code written against 7.31.0's surface:
- `_rev` survives unchanged. It's part of the 8.0 entity shape; the auto-bump moves to the new `transact()` write path. Code that uses `update({ ifRev })` keeps working.
- A new `brain.transact(tx, { ifAtGeneration: prev.generation })` adds generation-based CAS for whole-transaction atomicity, layered on top of `_rev`. Per-entity for single-record patterns (the lock above), generation-based for "did the world move under me."
If you adopt `_rev` + `ifRev` in 7.31.0, no migration work is needed for 8.0.
A snapshot or historical view captures each entity *including* its `_rev` at that moment, so reading the past and writing back with `ifRev` against the live state works exactly as you'd hope: the write fails if the entity moved since the state you copied from.

View file

@ -1,6 +1,6 @@
# Schema Migrations
Brainy includes a built-in migration system for transforming entity and verb metadata across storage versions. Migrations are pure functions that run once per storage instance, with automatic backup, resume support, and error tracking.
Brainy includes a built-in migration system for transforming entity and verb metadata across storage versions. Migrations are pure functions that run once per storage instance, with optional snapshot backup (`backupTo`), resume support, and error tracking.
---
@ -48,15 +48,13 @@ When `brain.init()` runs:
When `brain.migrate()` runs:
1. **Backup** — creates an instant COW branch (`pre-migration-7.17.0`) tagged with `system:backup` metadata. Rollback is possible by switching to this branch.
1. **Backup (optional)** — with `backupTo`, a hard-link snapshot of the current generation is persisted before any transform runs. Rollback is `brain.restore(backupPath, { confirm: true })`.
2. **Transform main branch** — iterates all nouns/verbs in paginated batches. For each entity, calls the `transform` function. If it returns a new object, saves it. If it returns `null`, skips. Vectors are never touched.
2. **Transform** — iterates all nouns/verbs in paginated batches. For each entity, calls the `transform` function. If it returns a new object, saves it. If it returns `null`, skips. Vectors are never touched.
3. **Transform other branches** — switches to each user branch, runs the same transforms. Inherited (already-migrated) entities return `null` and are skipped automatically.
3. **Save state** — records each completed migration ID so it never re-runs.
4. **Save state** — records each completed migration ID so it never re-runs.
5. **Rebuild indexes** — if any entities were modified, rebuilds the MetadataIndex.
4. **Rebuild indexes** — if any entities were modified, rebuilds the MetadataIndex.
---
@ -78,7 +76,7 @@ interface Migration {
- **Return a new object** to modify the entity's metadata.
- **Return `null`** to skip (no change needed).
- **Must be idempotent** — running the same transform twice on the same data should produce the same result (or return `null` the second time). This is required because branch iterations may re-encounter inherited entities.
- **Must be idempotent** — running the same transform twice on the same data should produce the same result (or return `null` the second time). This is required because interrupted runs resume and re-encounter already-migrated entities.
- **Must be pure** — no side effects, no async, no external state.
- Transforms only see metadata. Vectors, embeddings, and the `data` field stored inside metadata are available as properties on the metadata object.
@ -110,9 +108,9 @@ const preview = await brain.migrate({ dryRun: true })
// preview.sampleChanges — up to 5 before/after samples
// preview.estimatedTime — rough time estimate string
// Apply migrations
const result = await brain.migrate()
// result.backupBranch — name of COW backup branch, or null
// Apply migrations (optionally with a pre-migration snapshot)
const result = await brain.migrate({ backupTo: '/backups/pre-migration' })
// result.backupPath — snapshot path, or null when no backupTo was supplied
// result.migrationsApplied — array of migration IDs that ran
// result.entitiesProcessed — total entities scanned
// result.entitiesModified — entities actually changed
@ -164,20 +162,20 @@ const result = await brain.migrate({ maxErrors: 10000 })
## Backup and Rollback
Before modifying any data, `brain.migrate()` calls `brain.fork()` to create a COW snapshot. This is instant regardless of dataset size — it's a pointer copy, not a data copy.
The backup branch is named `pre-migration-{version}` and tagged with metadata:
- `type: 'system:backup'`
- `migrationVersion: '7.17.0'`
- `author: 'brainy-migration'`
To roll back, switch to the backup branch:
Pass `backupTo` and `brain.migrate()` persists a snapshot of the current generation **before any transform runs**. On filesystem storage the snapshot is a hard-link farm — created without copying entity data, and immune to later writes (see [Snapshots & Time Travel](./snapshots-and-time-travel.md)):
```typescript
await brain.checkout('pre-migration-7.17.0')
const result = await brain.migrate({ backupTo: '/backups/pre-migration-8.0' })
console.log(result.backupPath) // '/backups/pre-migration-8.0' (null when no backupTo)
```
Old backup branches from previous migrations are cleaned up automatically before each new migration run.
To roll back, restore the snapshot wholesale:
```typescript
await brain.restore('/backups/pre-migration-8.0', { confirm: true })
```
Without `backupTo`, no backup is taken — transforms are idempotent (they return `null` when already applied), but a pre-migration snapshot is the cheap insurance for anything destructive.
---

View file

@ -0,0 +1,277 @@
---
title: Snapshots & Time Travel
slug: guides/snapshots-and-time-travel
public: true
category: guides
template: guide
order: 9
description: Recipes for the Db API — instant backups with persist(), restore, time-travel debugging with asOf(), persist-before-migrate, what-if analysis with with(), and audit trails via transaction metadata.
next:
- concepts/consistency-model
- guides/optimistic-concurrency
---
# Snapshots & Time Travel
Brainy 8.0 treats the database as a **value**: `brain.now()` pins the
current state as an immutable `Db`, `brain.transact()` commits an atomic
batch and hands you the resulting value, `brain.asOf()` opens past state,
and `db.persist()` cuts a self-contained snapshot. This guide is the recipe
book. The precise guarantees behind every recipe live in the
[consistency model](../concepts/consistency-model.md).
## Instant backup
Pin the current state, persist it, release:
```typescript
const db = brain.now()
try {
await db.persist('/backups/2026-06-11')
} finally {
await db.release()
}
```
On filesystem storage the snapshot is built from **hard links**: every data
file in Brainy is immutable-by-rename, so the snapshot is created without
copying entity data and shares disk space with the live store. Later writes
can never alter it — a rewrite swaps the inode, the snapshot keeps the old
bytes. Cross-device targets fall back to per-file byte copies, and
persisting an in-memory brain serializes it to the same directory layout —
a real, durable store.
Two things to know:
- `persist()` requires the view to still be the store's **latest**
generation. If something committed after your pin, it throws
`GenerationConflictError` instead of snapshotting the wrong state — pin
and persist before further writes, or retry with a fresh `brain.now()`.
- The target directory must be empty or absent.
For scheduled backups, this loop is the whole job:
```typescript
const db = brain.now()
try {
await db.persist(`/backups/${new Date().toISOString().slice(0, 10)}`)
} finally {
await db.release()
}
```
## Restore
`restore()` replaces the store's **entire** current state from a snapshot —
entities, relationships, indexes, history. It is deliberately loud about it:
```typescript
await brain.restore('/backups/2026-06-11', { confirm: true })
```
- `{ confirm: true }` is mandatory — current state is destroyed.
- The snapshot is copied in (never linked), so it stays independent and can
be restored again later.
- All indexes are rebuilt from the restored records.
- The generation counter is floored at its pre-restore value, so generation
numbers you observed before the restore are never reissued.
- Live `Db` pins do not survive a restore — release them first.
## Open a snapshot read-only
You do not have to restore to look inside a snapshot. `Brainy.load()` opens
it as a self-contained read-only store with the **full query surface**,
including vector search:
```typescript
const db = await Brainy.load('/backups/2026-06-11')
const hits = await db.search('unpaid invoices from the spring campaign')
const orders = await db.find({ type: NounType.Document, subtype: 'order' })
await db.release() // closes the underlying read-only instance
```
`brain.asOf('/backups/2026-06-11')` does the same from an existing brain.
This is also the 8.0 answer to "named branches": a branch is a name → path
mapping your application keeps, where each path is a persisted snapshot.
Need a writable copy? Restore the snapshot into a fresh data directory and
open a writer on it — instead of switching a shared store between branches
in place, every line of code always sees exactly the store it opened.
## Time-travel debugging
When production data looks wrong, query the past directly — by wall-clock
time or by generation:
```typescript
// What did this order look like yesterday?
const yesterday = await brain.asOf(new Date(Date.now() - 86_400_000))
const before = await yesterday.get(orderId)
// Full queries work at any reachable generation — search, graph, filters:
const thenActive = await yesterday.find({
type: NounType.Document,
subtype: 'order',
where: { status: 'active' }
})
await yesterday.release()
```
Pin two points in time and diff them:
```typescript
const before = await brain.asOf(1041)
const after = brain.now()
const changed = await after.since(before)
changed.nouns // entity ids touched by transactions in between
changed.verbs // relationship ids touched in between
await before.release()
await after.release()
```
Three things to remember:
- History granularity is `transact()` commits — single-operation writes
advance the clock but do not produce historical records (see the
[consistency model](../concepts/consistency-model.md)). Use `transact()`
for writes you want to travel back through.
- The first index-accelerated query (semantic search, traversal, cursors,
aggregation) at a historical generation builds an in-memory index
materialization — O(n at that generation), once per `Db`, freed on
`release()`. Metadata-level reads are free.
- Generations reclaimed by `compactHistory()` throw
`GenerationCompactedError` — persist anything you need to keep forever.
## Safe schema migration
`brain.migrate()` integrates with snapshots directly: pass `backupTo` and a
hard-link snapshot of the current generation is persisted **before any
transform runs**:
```typescript
const result = await brain.migrate({ backupTo: '/backups/pre-migration-8.0' })
console.log(result.migrationsApplied, result.backupPath)
// If the migration went wrong, roll the whole store back:
await brain.restore('/backups/pre-migration-8.0', { confirm: true })
```
The same persist-before-mutate pattern works for any risky bulk operation,
not just migrations:
```typescript
const pin = brain.now()
try {
await pin.persist('/backups/pre-bulk-edit')
} finally {
await pin.release()
}
await runRiskyBulkEdit(brain)
```
## What-if analysis
`db.with(ops)` applies a transaction **speculatively, in memory** — nothing
touches disk, the generation counter, or the indexes. Ask "what would the
store look like if…", then commit the same operations for real:
```typescript
const ops = [
{ op: 'update', id: employeeId, metadata: { team: 'platform' } },
{ op: 'relate', from: employeeId, to: milestoneId, type: VerbType.ParticipatesIn, subtype: 'assignment' }
]
const base = brain.now()
const whatIf = await base.with(ops)
await whatIf.get(employeeId) // sees the change
await whatIf.find({ where: { team: 'platform' } }) // metadata finds work
await whatIf.related(employeeId) // overlay relations included
await whatIf.release()
await base.release()
// Looks right — make it real, atomically:
await brain.transact(ops)
```
**The boundary:** speculative entities carry no embeddings (`with()` never
invokes the embedder), so semantic search, traversal, cursors, aggregation,
and `persist()` throw `SpeculativeOverlayError` on overlay views instead of
returning silently incomplete results. `get()`, metadata-filter `find()`,
and filter-based `related()` are fully supported. Overlays chain — calling
`with()` on an overlay stacks another layer.
## Audit trails
`transact()` reifies transaction metadata: whatever you pass as `meta` is
recorded durably alongside the committed generation and timestamp, readable
via `brain.transactionLog()`:
```typescript
await brain.transact(
[{ op: 'update', id: invoiceId, metadata: { status: 'approved' } }],
{ meta: { author: 'approvals-service', actor: 'jane@example.com', reason: 'PO-7741' } }
)
const log = await brain.transactionLog({ limit: 20 }) // newest first
// [{ generation: 1042, timestamp: 1765432100000, meta: { author: 'approvals-service', ... } }]
```
Combine the log with `asOf()` to reconstruct exactly what any transaction
did:
```typescript
const [entry] = await brain.transactionLog({ limit: 1 })
const after = await brain.asOf(entry.generation)
const before = await brain.asOf(entry.generation - 1)
const touched = await after.since(before)
for (const id of touched.nouns) {
console.log(id, await before.get(id), '→', await after.get(id))
}
await before.release()
await after.release()
```
For per-entity write coordination (rather than whole-store history), the
`_rev` counter and `ifRev` CAS remain the right tool — see
[optimistic concurrency](./optimistic-concurrency.md).
## Keeping history bounded
Historical records cost disk space. Reclaim what no live pin protects:
```typescript
await brain.compactHistory({
retainGenerations: 100, // keep the 100 most recent commits
retainMs: 7 * 24 * 60 * 60 * 1000 // and everything from the last 7 days
})
```
Compaction never breaks a pinned read — record-sets are reclaimed only when
no live `Db` could need them. Release views you are done with (including the
ones `transact()` returns), and `persist()` any generation you want to keep
beyond the retention window: snapshots are self-contained and unaffected by
compaction.
## From branches to values
If you used the pre-8.0 `fork`/`checkout`/`commit`/`versions` surface, every
use case maps to a sharper tool:
| Pre-8.0 habit | 8.0 recipe |
|---|---|
| `fork()` to experiment safely | `db.with(ops)` for speculation in memory; a restored snapshot in a fresh directory for a long-lived writable copy |
| `commit()` checkpoints | `transact(ops, { meta })` — every batch is an atomic, logged, time-travelable commit |
| `checkout()` to switch branches | Open the snapshot you want — `Brainy.load(path)` read-only, or restore into its own directory. No in-place switching: every handle always sees one unambiguous store. |
| `getHistory()` | `brain.transactionLog()` + `db.since(priorDb)` |
| `versions.save()` per-entity snapshots | A pinned `Db` or persisted snapshot captures *every* entity at that moment; `asOf()` reads any entity's past state |
| `versions.restore()` | `brain.restore(snapshot, { confirm: true })` for the whole store, or read the old entity via `asOf()` and write it back with `transact()` |
| Backup branches | `db.persist(path)` — instant, hard-link-shared, self-contained |

View file

@ -5,7 +5,7 @@ public: true
category: guides
template: guide
order: 2
description: "Two adapters cover every deployment: in-memory for tests + ephemeral workloads, filesystem for everything that needs to persist. Both support copy-on-write branching. Cloud backup is operator tooling, not a built-in adapter."
description: "Two adapters cover every deployment: in-memory for tests + ephemeral workloads, filesystem for everything that needs to persist. Both share one on-disk contract, including generational history and snapshots. Cloud backup is operator tooling, not a built-in adapter."
next:
- guides/plugins
- concepts/zero-config
@ -20,8 +20,10 @@ Brainy 8.0 ships **two storage adapters**:
- **`MemoryStorage`** — in-memory only. The right choice for tests, ephemeral
workloads, and short-lived demos.
Both implement the same `StorageAdapter` interface, support copy-on-write
branching, and use the same on-disk layout (memory's "disk" is a JS Map).
Both implement the same `StorageAdapter` interface, support the full Db API
(generational history, snapshots, restore — see the
[consistency model](../concepts/consistency-model.md)), and use the same
on-disk layout (memory's "disk" is a JS Map).
## Quick start
@ -44,7 +46,7 @@ const brainAuto = new Brainy({ storage: { type: 'auto' } })
| Use case | Adapter | Why |
|---|---|---|
| Production app | `filesystem` | Durable, branchable, mmap-able |
| Production app | `filesystem` | Durable, snapshot-able, mmap-able |
| Tests, CI | `memory` | No disk teardown; fast |
| Short-lived data pipeline | `memory` | No persistence needed |
| In-browser demo | `memory` | Filesystem unavailable in browsers |
@ -70,20 +72,20 @@ azcopy sync /var/lib/brainy "https://account.blob.core.windows.net/brainy?sv=...
Brainy's filesystem layout is sync-friendly:
- Atomic writes (temp + rename) — readers never see torn files
- Per-shard files — `rsync`-style incremental sync works well
- Content-addressed blobs (`_blobs/`) — immutable, cache-friendly
- Branch refs live under `_cow/` — pick up branches automatically
- Immutable generation records (`_generations/`) — append-only, cache-friendly
For point-in-time backups, take a filesystem snapshot (ZFS, btrfs, LVM, EBS,
etc.) or use `brain.persist(path)` to write a self-contained snapshot you
can sync independently of the live brain.
etc.) or use `brain.now().persist(path)` to write a self-contained snapshot
you can sync independently of the live brain — see
[Snapshots & Time Travel](./snapshots-and-time-travel.md).
## Why no cloud adapters in 8.0?
Cloud storage adapters lived in Brainy 4.x-7.x. They were dropped in 8.0
per **BR-BRAINY-80-STORAGE-SIMPLIFY** because:
because:
- Zero production consumers used them at scale. Every Soulcraft consumer
ran on local filesystem.
- Zero production consumers used them at scale — every known production
deployment ran on local filesystem.
- Cloud-storage HNSW / DiskANN doesn't perform — vector indexes need
low-latency random reads that S3 / GCS / R2 / Azure can't provide
consistently.
@ -97,23 +99,21 @@ Brainy 8.0 is smaller, faster to install, and clearer about what it does.
## Configuration
```ts
interface StorageOptions {
// The adapter type. Defaults to 'auto'.
type?: 'auto' | 'memory' | 'filesystem'
// BrainyConfig['storage'] — either a config object or a pre-constructed adapter:
storage?:
| {
// The adapter type. Defaults to 'auto'
// (filesystem on Node-like runtimes, memory otherwise).
type: 'auto' | 'memory' | 'filesystem'
// Force a specific adapter regardless of type.
forceMemoryStorage?: boolean
forceFileSystemStorage?: boolean
// Root directory for filesystem storage. Passed through to storage
// factories, including plugin-provided ones.
rootDirectory?: string
// Filesystem only.
rootDirectory?: string
// COW branch to open. Defaults to 'main'.
branch?: string
// COW compression toggle. Defaults to true.
enableCompression?: boolean
}
// Adapter-specific options.
options?: any
}
| StorageAdapter // e.g. storage: new MemoryStorage()
```
## Direct construction
@ -143,6 +143,7 @@ backup tooling. The recipe:
4. Set up an operator backup job using `gsutil` / `aws s3` / `rclone` /
`azcopy` on a cron — hourly or whatever your RPO requires. Point it at
the brainy data dir.
5. For point-in-time backups, use filesystem snapshots or `brain.persist()`.
5. For point-in-time backups, use filesystem snapshots or
`brain.now().persist(path)`.
Same data, same APIs, no library-side cloud code.