brainy/docs/guides/snapshots-and-time-travel.md

438 lines
17 KiB
Markdown
Raw Permalink Normal View History

---
title: Snapshots & Time Travel
slug: guides/snapshots-and-time-travel
public: true
category: guides
template: guide
order: 9
feat(8.0): temporal range verbs — diff, history, since(gen|Date), asOf{exclusive}, transactionLog window asOf() answers "state AT a point"; these answer "what happened BETWEEN two points" and "one entity's whole history", all on the existing generation records. - Extract resolveAsOfGeneration() as the shared gen|Date→{generation,timestamp} chokepoint (reachability + Date semantics identical everywhere). asOf()'s inclusive path is byte-identical — db-mvcc still 25/25. - asOf(target, { exclusive }) — strict-before (resolved−1, clamped at gen 0, never a RangeError). - db.since(Db | generation | Date) — overload; number/Date resolve via the shared resolver (new DbHost.resolveGeneration); EXCLUSIVE lower bound; since(db) === since(db.generation). Same-store guard via Db.belongsToStore. - brain.diff(a, b) → { added, removed, modified } split by nouns/verbs. EARNS its name: candidate set is ONLY changedBetween(gLow,gHigh), each classified by existence at both endpoints + a key-order-insensitive value compare (new src/db/stableEqual.ts) — a touched-but-reverted / born-and-died id is in no bucket. - brain.history(id, { from, to }) → every distinct version oldest→newest, each value === asOf(version.generation).get(id); null = removal; kind auto-detected (throws on UUID-space collision). New store helper generationsTouching(). - brain.transactionLog({ from, to, limit }) — INCLUSIVE generation/Date window (contrast since's exclusive lower bound); limit applied last. - Compaction policy locked: diff/since THROW GenerationCompactedError below the horizon; history TRUNCATES to it. Types DiffResult/HistoryVersion/EntityHistory exported from the package root. Tests: tests/integration/db-temporal.test.ts (8 proofs incl. diff-earns-its-name, history↔asOf cross-check, the composition proof, granularity, compaction contrast) + tests/unit/db/stableEqual.test.ts (6). docs/guides/snapshots-and-time-travel.md + RELEASES updated. 1477 unit + db-temporal 8 + db-mvcc 25 green.
2026-06-19 13:21:02 -07:00
description: Recipes for the Db API — instant backups with persist(), restore, time-travel debugging with asOf(), range queries over history (diff, history, since, log windows), persist-before-migrate, what-if analysis with with(), and audit trails via transaction metadata.
next:
- concepts/consistency-model
- guides/optimistic-concurrency
---
# Snapshots & Time Travel
Brainy 8.0 treats the database as a **value**: `brain.now()` pins the
current state as an immutable `Db`, `brain.transact()` commits an atomic
batch and hands you the resulting value, `brain.asOf()` opens past state,
and `db.persist()` cuts a self-contained snapshot. This guide is the recipe
book. The precise guarantees behind every recipe live in the
[consistency model](../concepts/consistency-model.md).
## Instant backup
Pin the current state, persist it, release:
```typescript
const db = brain.now()
try {
await db.persist('/backups/2026-06-11')
} finally {
await db.release()
}
```
On filesystem storage the snapshot is built from **hard links**: every data
file in Brainy is immutable-by-rename, so the snapshot is created without
copying entity data and shares disk space with the live store. Later writes
can never alter it — a rewrite swaps the inode, the snapshot keeps the old
bytes. Cross-device targets fall back to per-file byte copies, and
persisting an in-memory brain serializes it to the same directory layout —
a real, durable store.
Two things to know:
- `persist()` requires the view to still be the store's **latest**
generation. If something committed after your pin, it throws
`GenerationConflictError` instead of snapshotting the wrong state — pin
and persist before further writes, or retry with a fresh `brain.now()`.
- The target directory must be empty or absent.
For scheduled backups, this loop is the whole job:
```typescript
const db = brain.now()
try {
await db.persist(`/backups/${new Date().toISOString().slice(0, 10)}`)
} finally {
await db.release()
}
```
## Restore
`restore()` replaces the store's **entire** current state from a snapshot —
entities, relationships, indexes, history. It is deliberately loud about it:
```typescript
await brain.restore('/backups/2026-06-11', { confirm: true })
```
- `{ confirm: true }` is mandatory — current state is destroyed.
- The snapshot is copied in (never linked), so it stays independent and can
be restored again later.
- All indexes are rebuilt from the restored records.
- The generation counter is floored at its pre-restore value, so generation
numbers you observed before the restore are never reissued.
- Live `Db` pins do not survive a restore — release them first.
## Open a snapshot read-only
You do not have to restore to look inside a snapshot. `Brainy.load()` opens
it as a self-contained read-only store with the **full query surface**,
including vector search:
```typescript
const db = await Brainy.load('/backups/2026-06-11')
feat(8.0): API simplification — remove neural()/Db.search, one storage `path` key, integration→0 8.0 RC cleanup toward "one place per thing, zero-config, no deprecation": - Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})` / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor, NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept. - Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams). Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` → now `find({ query, limit })`. - Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native storage provider resolves the identical root (no split-brain). - Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now applied as a post-filter on `result.score` (the documented way to bound semantic results). - Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()` and failed dimension validation; they are metadata-only updates now. - Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an in-place path change that preserves the blob and the entity id, for files and directories. - Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability. Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk flushes and emits `progress.queryable`, so imported data is queryable during the import. - Sweep all docs, comments, and JSDoc for the removed/changed APIs. Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
2026-06-20 13:31:11 -07:00
const hits = await db.find({ query: 'unpaid invoices from the spring campaign' })
const orders = await db.find({ type: NounType.Document, subtype: 'order' })
await db.release() // closes the underlying read-only instance
```
`brain.asOf('/backups/2026-06-11')` does the same from an existing brain.
This is also the 8.0 answer to "named branches": a branch is a name → path
mapping your application keeps, where each path is a persisted snapshot.
Need a writable copy? Restore the snapshot into a fresh data directory and
open a writer on it — instead of switching a shared store between branches
in place, every line of code always sees exactly the store it opened.
## Time-travel debugging
When production data looks wrong, query the past directly — by wall-clock
time or by generation:
```typescript
// What did this order look like yesterday?
const yesterday = await brain.asOf(new Date(Date.now() - 86_400_000))
const before = await yesterday.get(orderId)
// Full queries work at any reachable generation — search, graph, filters:
const thenActive = await yesterday.find({
type: NounType.Document,
subtype: 'order',
where: { status: 'active' }
})
await yesterday.release()
```
Pin two points in time and diff them:
```typescript
const before = await brain.asOf(1041)
const after = brain.now()
const changed = await after.since(before)
changed.nouns // entity ids touched by transactions in between
changed.verbs // relationship ids touched in between
await before.release()
await after.release()
```
Three things to remember:
feat(8.0): Model-B per-write generation-stamping + adaptive retention knob Every write — transact() AND single-op add/update/remove/relate — is now its own immutable generation (Model-B), so a now() pin always freezes and asOf/since/diff/history/transactionLog reflect single-ops exactly like transacts. Closes the Model-A hole where pins did not freeze against single-op writes. Generation-stamping: - GenerationStore.commitSingleOp: a one-operation commitTransaction with deferred durability. Wired into add/update/remove/relate/updateRelation/ unrelate + removeMany (the *Many and VFS paths delegate to these). - Async group-commit (flushPendingSingleOps): the live write is acknowledged immediately; its before-image is buffered in an in-memory pending tier that resolveAt/chains/changedBetween/tx-log read like on-disk generations, so the synchronous now() freezes with no forced flush. One fsync per window (triggers: size / 50ms timer / flush / close / transact / compactHistory). - Crash recovery is drop-without-restore for group-commit generations (marked groupCommit:true): a crash mid-flush discards the partial generation and never restores its before-images, which would otherwise revert the already-acknowledged live write. - Init-time infrastructure (the VFS root) is the un-versioned generation-0 baseline: a fresh brain reports generation()===0 and an empty transactionLog(); the first user write is generation 1. - Historical find()/related() overlay bound is the full reserved watermark (generation()), so un-flushed single-op writes are overlaid too. Retention knob: - config `history` -> `retention`: 'all' | 'adaptive' | { maxGenerations?, maxAge?, maxBytes?, budgetBytes?, autoCompact? }; unset -> adaptive (disk/RAM pressure, zero-config). CompactHistoryOptions floors -> caps (retainGenerations->maxGenerations, retainMs->maxAge, +maxBytes): reclaim oldest-unpinned while ANY cap is exceeded; pins always exempt. - brain.setRetentionBudget(bytes) drives the adaptive byte budget at runtime (a coordinator's fair-share input). Per-generation bytes recorded in each delta enable historyBytes() introspection without a storage size API. Tests: per-write generation resolution, pin freeze vs add/update/remove, drop-without-restore corruption-trap (fault injector), clean-reopen replay, maxBytes/maxAge/no-cap reclamation, retention-then-reopen. 107 db/generation/ temporal tests green, tsc clean. Docs (ADR-001, consistency-model, snapshots guide, api reference, RELEASES) updated to per-write granularity + retention.
2026-06-22 15:19:58 -07:00
- History granularity is per-write: EVERY write — `transact()` AND a
single-operation `add`/`update`/`remove`/`relate` — is its own immutable
generation, so a pin always freezes against later writes and every write is
individually addressable via `asOf()` (see the
[consistency model](../concepts/consistency-model.md)). Use `transact()` when
you want several operations to share ONE atomic generation.
- The first index-accelerated query (semantic search, traversal, cursors,
aggregation) at a historical generation builds an in-memory index
materialization — O(n at that generation), once per `Db`, freed on
`release()`. Metadata-level reads are free.
- Generations reclaimed by `compactHistory()` throw
`GenerationCompactedError` — persist anything you need to keep forever.
feat(8.0): temporal range verbs — diff, history, since(gen|Date), asOf{exclusive}, transactionLog window asOf() answers "state AT a point"; these answer "what happened BETWEEN two points" and "one entity's whole history", all on the existing generation records. - Extract resolveAsOfGeneration() as the shared gen|Date→{generation,timestamp} chokepoint (reachability + Date semantics identical everywhere). asOf()'s inclusive path is byte-identical — db-mvcc still 25/25. - asOf(target, { exclusive }) — strict-before (resolved−1, clamped at gen 0, never a RangeError). - db.since(Db | generation | Date) — overload; number/Date resolve via the shared resolver (new DbHost.resolveGeneration); EXCLUSIVE lower bound; since(db) === since(db.generation). Same-store guard via Db.belongsToStore. - brain.diff(a, b) → { added, removed, modified } split by nouns/verbs. EARNS its name: candidate set is ONLY changedBetween(gLow,gHigh), each classified by existence at both endpoints + a key-order-insensitive value compare (new src/db/stableEqual.ts) — a touched-but-reverted / born-and-died id is in no bucket. - brain.history(id, { from, to }) → every distinct version oldest→newest, each value === asOf(version.generation).get(id); null = removal; kind auto-detected (throws on UUID-space collision). New store helper generationsTouching(). - brain.transactionLog({ from, to, limit }) — INCLUSIVE generation/Date window (contrast since's exclusive lower bound); limit applied last. - Compaction policy locked: diff/since THROW GenerationCompactedError below the horizon; history TRUNCATES to it. Types DiffResult/HistoryVersion/EntityHistory exported from the package root. Tests: tests/integration/db-temporal.test.ts (8 proofs incl. diff-earns-its-name, history↔asOf cross-check, the composition proof, granularity, compaction contrast) + tests/unit/db/stableEqual.test.ts (6). docs/guides/snapshots-and-time-travel.md + RELEASES updated. 1477 unit + db-temporal 8 + db-mvcc 25 green.
2026-06-19 13:21:02 -07:00
## Range queries over history
`asOf()` answers "what was the state AT a point". Four range verbs answer
"what happened BETWEEN two points" and "what is one entity's whole history".
They all build on the same generation records — no extra bookkeeping.
### `diff(a, b)` — what changed, classified
`since()` gives you the raw set of *touched* ids. `diff()` goes further: it
resolves each touched id at both endpoints and classifies it as **added**,
**removed**, or **modified** — split by entities (`nouns`) and relationships
(`verbs`). An id that was touched but ended up identical (changed then
reverted, or created and deleted within the interval) lands in **none** of the
buckets. Endpoints are a generation, a `Date`, or a `Db`, in either order:
```typescript
const d = await brain.diff(1041, brain.generation())
d.added.nouns // entity ids created between the two states
d.removed.nouns // entity ids deleted
d.modified.nouns // entity ids whose stored value actually changed
d.added.verbs // …relationships, the same three ways
```
Orientation is `a → b`: `added` means "exists at `b`, not at `a`". The
comparison behind `modified` is key-order-insensitive, so a no-op re-write of
the same fields never shows up as a change.
### `history(id, range?)` — one entity, every version
`asOf()` is per-*generation*; `history()` is per-*entity*. It returns every
distinct version of one id over a range, oldest first — each `value` is the
materialized state at that version (and `null` marks a removal):
```typescript
const h = await brain.history(invoiceId)
for (const v of h.versions) {
console.log(v.generation, v.value?.metadata?.status ?? '(deleted)')
}
// 1041 'draft'
// 1043 'approved'
// 1050 'paid'
```
Every version ties to the trusted `asOf()` path — `v.value` equals
`(await brain.asOf(v.generation)).get(id)`. Pass `{ from, to }` (generation or
`Date`) to bound the range; a `from` below the compaction horizon is quietly
truncated to it rather than throwing (history is best-effort over surviving
records).
### `since()` and `transactionLog()` take ranges too
`since()` accepts a `Db`, a generation number, or a `Date` — all equivalent,
all an **exclusive** lower bound (`db.since(prior)` equals
`db.since(prior.generation)`):
```typescript
await brain.now().since(1041) // ids changed after generation 1041
await brain.now().since(new Date(Date.now() - 3_600_000)) // …in the last hour
```
`transactionLog({ from, to })` windows the commit log **inclusively** on both
ends (a log window names the commits it spans — the deliberate contrast to
`since`'s exclusive lower bound); `limit` applies after the window, newest
first:
```typescript
const window = await brain.transactionLog({ from: 1041, to: 1050 }) // commits 1041…1050
const recent = await brain.transactionLog({ from: lastHour, limit: 20 })
```
### Composing them
"Which orders changed in this window?" is `diff` ids intersected with an
`asOf` query — the two agree by construction:
```typescript
const changed = await brain.diff(g1, g2)
const atG2 = await brain.asOf(g2)
const changedOrders = (await atG2.find({ type: NounType.Document, subtype: 'order' }))
.map(r => r.id)
.filter(id => changed.added.nouns.includes(id) || changed.modified.nouns.includes(id))
await atG2.release()
```
One contrast to keep straight: `diff` and `since` **throw**
`GenerationCompactedError` for a bound below the horizon, while `history`
**truncates** to the horizon — diffs must be exact, history is best-effort.
## Safe schema migration
`brain.migrate()` integrates with snapshots directly: pass `backupTo` and a
hard-link snapshot of the current generation is persisted **before any
transform runs**:
```typescript
const result = await brain.migrate({ backupTo: '/backups/pre-migration-8.0' })
console.log(result.migrationsApplied, result.backupPath)
// If the migration went wrong, roll the whole store back:
await brain.restore('/backups/pre-migration-8.0', { confirm: true })
```
The same persist-before-mutate pattern works for any risky bulk operation,
not just migrations:
```typescript
const pin = brain.now()
try {
await pin.persist('/backups/pre-bulk-edit')
} finally {
await pin.release()
}
await runRiskyBulkEdit(brain)
```
## What-if analysis
`db.with(ops)` applies a transaction **speculatively, in memory** — nothing
touches disk, the generation counter, or the indexes. Ask "what would the
store look like if…", then commit the same operations for real:
```typescript
const ops = [
{ op: 'update', id: employeeId, metadata: { team: 'platform' } },
{ op: 'relate', from: employeeId, to: milestoneId, type: VerbType.ParticipatesIn, subtype: 'assignment' }
]
const base = brain.now()
const whatIf = await base.with(ops)
await whatIf.get(employeeId) // sees the change
await whatIf.find({ where: { team: 'platform' } }) // metadata finds work
await whatIf.related(employeeId) // overlay relations included
await whatIf.release()
await base.release()
// Looks right — make it real, atomically:
await brain.transact(ops)
```
**The boundary:** speculative entities carry no embeddings (`with()` never
invokes the embedder), so semantic search, traversal, cursors, aggregation,
and `persist()` throw `SpeculativeOverlayError` on overlay views instead of
returning silently incomplete results. `get()`, metadata-filter `find()`,
and filter-based `related()` are fully supported. Overlays chain — calling
`with()` on an overlay stacks another layer.
## Audit trails
`transact()` reifies transaction metadata: whatever you pass as `meta` is
recorded durably alongside the committed generation and timestamp, readable
via `brain.transactionLog()`:
```typescript
await brain.transact(
[{ op: 'update', id: invoiceId, metadata: { status: 'approved' } }],
{ meta: { author: 'approvals-service', actor: 'jane@example.com', reason: 'PO-7741' } }
)
const log = await brain.transactionLog({ limit: 20 }) // newest first
// [{ generation: 1042, timestamp: 1765432100000, meta: { author: 'approvals-service', ... } }]
```
Combine the log with `asOf()` to reconstruct exactly what any transaction
did:
```typescript
const [entry] = await brain.transactionLog({ limit: 1 })
const after = await brain.asOf(entry.generation)
const before = await brain.asOf(entry.generation - 1)
const touched = await after.since(before)
for (const id of touched.nouns) {
console.log(id, await before.get(id), '→', await after.get(id))
}
await before.release()
await after.release()
```
For per-entity write coordination (rather than whole-store history), the
`_rev` counter and `ifRev` CAS remain the right tool — see
[optimistic concurrency](./optimistic-concurrency.md).
## Keeping history bounded
feat(8.0): Model-B per-write generation-stamping + adaptive retention knob Every write — transact() AND single-op add/update/remove/relate — is now its own immutable generation (Model-B), so a now() pin always freezes and asOf/since/diff/history/transactionLog reflect single-ops exactly like transacts. Closes the Model-A hole where pins did not freeze against single-op writes. Generation-stamping: - GenerationStore.commitSingleOp: a one-operation commitTransaction with deferred durability. Wired into add/update/remove/relate/updateRelation/ unrelate + removeMany (the *Many and VFS paths delegate to these). - Async group-commit (flushPendingSingleOps): the live write is acknowledged immediately; its before-image is buffered in an in-memory pending tier that resolveAt/chains/changedBetween/tx-log read like on-disk generations, so the synchronous now() freezes with no forced flush. One fsync per window (triggers: size / 50ms timer / flush / close / transact / compactHistory). - Crash recovery is drop-without-restore for group-commit generations (marked groupCommit:true): a crash mid-flush discards the partial generation and never restores its before-images, which would otherwise revert the already-acknowledged live write. - Init-time infrastructure (the VFS root) is the un-versioned generation-0 baseline: a fresh brain reports generation()===0 and an empty transactionLog(); the first user write is generation 1. - Historical find()/related() overlay bound is the full reserved watermark (generation()), so un-flushed single-op writes are overlaid too. Retention knob: - config `history` -> `retention`: 'all' | 'adaptive' | { maxGenerations?, maxAge?, maxBytes?, budgetBytes?, autoCompact? }; unset -> adaptive (disk/RAM pressure, zero-config). CompactHistoryOptions floors -> caps (retainGenerations->maxGenerations, retainMs->maxAge, +maxBytes): reclaim oldest-unpinned while ANY cap is exceeded; pins always exempt. - brain.setRetentionBudget(bytes) drives the adaptive byte budget at runtime (a coordinator's fair-share input). Per-generation bytes recorded in each delta enable historyBytes() introspection without a storage size API. Tests: per-write generation resolution, pin freeze vs add/update/remove, drop-without-restore corruption-trap (fault injector), clean-reopen replay, maxBytes/maxAge/no-cap reclamation, retention-then-reopen. 107 db/generation/ temporal tests green, tsc clean. Docs (ADR-001, consistency-model, snapshots guide, api reference, RELEASES) updated to per-write granularity + retention.
2026-06-22 15:19:58 -07:00
Under Model-B every write is a generation, so history can grow quickly —
Brainy auto-compacts on every `flush()`/`close()` under the **`retention`**
knob (configured on the constructor):
```typescript
feat(8.0): Model-B per-write generation-stamping + adaptive retention knob Every write — transact() AND single-op add/update/remove/relate — is now its own immutable generation (Model-B), so a now() pin always freezes and asOf/since/diff/history/transactionLog reflect single-ops exactly like transacts. Closes the Model-A hole where pins did not freeze against single-op writes. Generation-stamping: - GenerationStore.commitSingleOp: a one-operation commitTransaction with deferred durability. Wired into add/update/remove/relate/updateRelation/ unrelate + removeMany (the *Many and VFS paths delegate to these). - Async group-commit (flushPendingSingleOps): the live write is acknowledged immediately; its before-image is buffered in an in-memory pending tier that resolveAt/chains/changedBetween/tx-log read like on-disk generations, so the synchronous now() freezes with no forced flush. One fsync per window (triggers: size / 50ms timer / flush / close / transact / compactHistory). - Crash recovery is drop-without-restore for group-commit generations (marked groupCommit:true): a crash mid-flush discards the partial generation and never restores its before-images, which would otherwise revert the already-acknowledged live write. - Init-time infrastructure (the VFS root) is the un-versioned generation-0 baseline: a fresh brain reports generation()===0 and an empty transactionLog(); the first user write is generation 1. - Historical find()/related() overlay bound is the full reserved watermark (generation()), so un-flushed single-op writes are overlaid too. Retention knob: - config `history` -> `retention`: 'all' | 'adaptive' | { maxGenerations?, maxAge?, maxBytes?, budgetBytes?, autoCompact? }; unset -> adaptive (disk/RAM pressure, zero-config). CompactHistoryOptions floors -> caps (retainGenerations->maxGenerations, retainMs->maxAge, +maxBytes): reclaim oldest-unpinned while ANY cap is exceeded; pins always exempt. - brain.setRetentionBudget(bytes) drives the adaptive byte budget at runtime (a coordinator's fair-share input). Per-generation bytes recorded in each delta enable historyBytes() introspection without a storage size API. Tests: per-write generation resolution, pin freeze vs add/update/remove, drop-without-restore corruption-trap (fault injector), clean-reopen replay, maxBytes/maxAge/no-cap reclamation, retention-then-reopen. 107 db/generation/ temporal tests green, tsc clean. Docs (ADR-001, consistency-model, snapshots guide, api reference, RELEASES) updated to per-write granularity + retention.
2026-06-22 15:19:58 -07:00
// Zero-config: ADAPTIVE — keep as much history as free disk/RAM allows,
// reclaiming oldest-first under pressure. (This is the default.)
new Brainy({ /* retention unset */ })
// Unbounded — never reclaim history (opt in explicitly):
new Brainy({ retention: 'all' })
// Explicit CAPS — reclaim oldest-unpinned generations while ANY cap is exceeded:
new Brainy({ retention: { maxGenerations: 1000, maxAge: 7 * 86_400_000, maxBytes: 512 * 1024 ** 2 } })
```
Reclaim manually at any time (the same caps):
```typescript
await brain.compactHistory({ maxGenerations: 100, maxAge: 7 * 24 * 60 * 60 * 1000 })
```
Compaction never breaks a pinned read — record-sets are reclaimed only when
feat(8.0): Model-B per-write generation-stamping + adaptive retention knob Every write — transact() AND single-op add/update/remove/relate — is now its own immutable generation (Model-B), so a now() pin always freezes and asOf/since/diff/history/transactionLog reflect single-ops exactly like transacts. Closes the Model-A hole where pins did not freeze against single-op writes. Generation-stamping: - GenerationStore.commitSingleOp: a one-operation commitTransaction with deferred durability. Wired into add/update/remove/relate/updateRelation/ unrelate + removeMany (the *Many and VFS paths delegate to these). - Async group-commit (flushPendingSingleOps): the live write is acknowledged immediately; its before-image is buffered in an in-memory pending tier that resolveAt/chains/changedBetween/tx-log read like on-disk generations, so the synchronous now() freezes with no forced flush. One fsync per window (triggers: size / 50ms timer / flush / close / transact / compactHistory). - Crash recovery is drop-without-restore for group-commit generations (marked groupCommit:true): a crash mid-flush discards the partial generation and never restores its before-images, which would otherwise revert the already-acknowledged live write. - Init-time infrastructure (the VFS root) is the un-versioned generation-0 baseline: a fresh brain reports generation()===0 and an empty transactionLog(); the first user write is generation 1. - Historical find()/related() overlay bound is the full reserved watermark (generation()), so un-flushed single-op writes are overlaid too. Retention knob: - config `history` -> `retention`: 'all' | 'adaptive' | { maxGenerations?, maxAge?, maxBytes?, budgetBytes?, autoCompact? }; unset -> adaptive (disk/RAM pressure, zero-config). CompactHistoryOptions floors -> caps (retainGenerations->maxGenerations, retainMs->maxAge, +maxBytes): reclaim oldest-unpinned while ANY cap is exceeded; pins always exempt. - brain.setRetentionBudget(bytes) drives the adaptive byte budget at runtime (a coordinator's fair-share input). Per-generation bytes recorded in each delta enable historyBytes() introspection without a storage size API. Tests: per-write generation resolution, pin freeze vs add/update/remove, drop-without-restore corruption-trap (fault injector), clean-reopen replay, maxBytes/maxAge/no-cap reclamation, retention-then-reopen. 107 db/generation/ temporal tests green, tsc clean. Docs (ADR-001, consistency-model, snapshots guide, api reference, RELEASES) updated to per-write granularity + retention.
2026-06-22 15:19:58 -07:00
no live `Db` could need them (live pins are ALWAYS exempt). Release views you
are done with (including the ones `transact()` returns), and `persist()` any
generation you want to keep beyond the retention window: snapshots are
self-contained and unaffected by compaction.
feat: temporal VFS — file content joins the Model-B immutability model The temporal model had a hole exactly where files were concerned: every entity write is an immutable generation with before-images, but VFS content BYTES lived under an eager refCount GC left over from the pre-8.0 design — unlink could physically destroy bytes that in-window history still referenced, and overwrite never released the old hash at all (an unbounded silent leak whose accidental byproduct was the only thing "preserving" history). Reading the past could therefore return a stale field, a dangling hash, or nothing, depending on luck. Fix: blob reclamation becomes a HISTORY decision instead of a LIVENESS decision. Each blob's metadata now carries historyRefCount alongside the live refCount: - The commit seam counts one history reference per persisted before-image record carrying a content hash (commitTransaction staging and the group-commit flush), recorded BEFORE the record-set persists and carried in the generation delta (blobHashes — always present on new deltas, so compaction only falls back to reading records for pre-contract generations). An aborted transaction compensates best-effort. - unlink/rmdir/overwrite drop ONLY the live reference (BlobStorage.delete → release; overwrite finally releases the superseded hash — cancelling the dedup increment on same-content rewrites and closing the leak), and only AFTER the canonical mutation commits, so a failed delete can never leave a live file whose bytes compaction might reclaim. - History compaction is the ONE reclamation point: after deleting a generation's record-set it releases that set's references and physically reclaims any hash at zero live AND zero history references. Pins are exempt automatically. Crash ordering is over-count-only in every path (record before persist, release after delete), so a crash can leak until the scrub recounts but can never reclaim bytes a retained generation needs. scrubBlobHistoryRefCounts() restores exactness; existing stores get a one-time marker-gated backfill on open, failing into leak-safe mode (reclamation disabled) rather than guessing. On top of the protected history, the temporal API the generational model always implied: - vfs.readFile(path, { asOf }) — the exact bytes as of a generation or Date, materialized from the history (pinned view released so compaction is never blocked by a read). - vfs.history(path) — FileVersion[] ascending ({ generation, timestamp, hash, size, mimeType? }), the newest entry being the live state. - Overwrites now refresh the file entity's data/embedding text — semantic search and the data field previously served the FIRST version's text forever (the stale-field defect a consumer's incident recovery depended on by luck). Integration suite (temporal-vfs.test.ts): per-version exact reads + history listing, leak-fix + history protection on overwrite, rm keeps bytes readable, compaction reclaims past-window bytes and preserves in-window (including the cross-file dedup case where an old file's history and a newer file's removal share one hash), data freshness, and scrub exactness.
2026-07-10 16:43:48 -07:00
## Time travel for files (the VFS)
Since 8.2.0, time travel covers Virtual Filesystem **content**, not just
entity records. File bytes are retention-protected: a content blob referenced
by any generation inside the retention window is never reclaimed, so reading
the past always returns the exact bytes — never a stale field or a
dangling hash.
**`vfs.readFile(path, { asOf })`** takes a generation number or a `Date` and
returns the file's exact bytes as they stood then. It resolves the path's
current entity, then materializes its state at the target generation — so it
answers *"what did the file at this path hold at that point?"* It bypasses
the content cache; the `encoding` option still applies. Asking about a
generation before the file existed throws the usual not-found error, and
asking past the retention window's compaction horizon throws a
compacted-generation error.
**`vfs.history(path)`** returns the file's versions inside the retention
window, oldest first — one `FileVersion` per generation that wrote the file,
the newest entry being the current state:
```typescript
// A CMS page evolves…
await brain.vfs.writeFile('/pages/home.json', '{"title":"Launch"}')
await brain.vfs.writeFile('/pages/home.json', '{"title":"Launch v2"}')
await brain.vfs.writeFile('/pages/home.json', '{"title":""}') // bad deploy!
// Every version is listed and readable:
const versions = await brain.vfs.history('/pages/home.json')
// → [{ generation, timestamp, hash, size, mimeType? }, …] ascending
const good = versions[versions.length - 2]
const bytes = await brain.vfs.readFile('/pages/home.json', {
asOf: good.generation
})
// Restore = write the old bytes back. This is a NEW write (a new
// generation) — history is never rewritten, so the bad version stays
// visible in the audit trail.
await brain.vfs.writeFile('/pages/home.json', bytes)
```
Two lifecycle consequences worth stating plainly:
- **Deleting or overwriting a file no longer frees its bytes immediately.**
Old content lives until history compaction reclaims the generations that
reference it — the same `retention` budget that bounds all Model-B history
(and pinned views are exempt, exactly as above). Size your `retention` for
the file-version depth you want; `retention: 'all'` keeps every version of
every file forever.
- **After `compactHistory()` reclaims a generation, its file versions are
gone** and their bytes are physically reclaimed. (This also fixed a
pre-8.2.0 defect where overwritten content was never reclaimed at all — an
unbounded silent leak.)
## From branches to values
If you used the pre-8.0 `fork`/`checkout`/`commit`/`versions` surface, every
use case maps to a sharper tool:
| Pre-8.0 habit | 8.0 recipe |
|---|---|
| `fork()` to experiment safely | `db.with(ops)` for speculation in memory; a restored snapshot in a fresh directory for a long-lived writable copy |
| `commit()` checkpoints | `transact(ops, { meta })` — every batch is an atomic, logged, time-travelable commit |
| `checkout()` to switch branches | Open the snapshot you want — `Brainy.load(path)` read-only, or restore into its own directory. No in-place switching: every handle always sees one unambiguous store. |
| `getHistory()` | `brain.transactionLog()` + `db.since(priorDb)` |
| `versions.save()` per-entity snapshots | A pinned `Db` or persisted snapshot captures *every* entity at that moment; `asOf()` reads any entity's past state |
| `versions.restore()` | `brain.restore(snapshot, { confirm: true })` for the whole store, or read the old entity via `asOf()` and write it back with `transact()` |
| Backup branches | `db.persist(path)` — instant, hard-link-shared, self-contained |