open-brainy/src/db/types.ts
David Snelling 38b0041464 feat: generation fact log — after-image commit records, dual-written at every commit point
Every committed generation now also appends a FACT — an after-image
commit record (what each touched entity/relationship became, or a
body-less tombstone for a removal) — to an append-only, crc32c-framed
segment log under _generations/facts/. The before-image history and the
canonical tree remain authoritative; the fact log gives consumers ONE
sequential, self-verifying stream (index heals, incremental replays)
in place of a per-entity directory walk.

- Wire format: positional msgpack facts [generation, timestamp, ops,
  meta, blobHashes]; op = [kind u8, id bin16, record | nil tombstone];
  32-byte segment header (magic, formatVersion, firstGeneration,
  zeroed+verified reserved); length+crc32c frame per fact; zero-padded
  segment names so lexicographic order == generation order; JSON
  manifest with an atomic rename flip, manifest-first rotation.
- Commit protocol: facts append+fsync BEFORE the commit point inside
  the existing durability window, so a crash can only leave the log
  AHEAD of committed truth — open() truncates back (torn tails detected
  by CRC). Absent generation = never committed; a scan can never see an
  uncommitted fact. transact() facts are durable-on-return; single-op
  facts ride the group-commit flush exactly like buffered history. A
  fact-append failure fails the write, loudly — a silent gap would be a
  lie a later replay discovers.
- New public surface: brain.scanFacts() (sequential batches with heal
  telemetry: head/segments/approx up front, per-batch generation range
  + bytes + segment id, loud abort on gaps, summary cross-check) and
  brain.factSegmentPaths() (immutable sealed segments for zero-copy
  consumers; the mutable tail excluded). Exported types CommitFact,
  FactOp, FactScanBatch, FactScanHandle.
- Storage: optional binary raw-byte primitives (appendRawBytes,
  readRawBytes, writeRawBytes, rawByteSize) on StorageAdapter —
  feature-detected; filesystem + memory adapters implement them; an
  adapter without them hosts no fact log. Fact segments are byte-copied
  (never hard-linked) into snapshots. The _generations/facts/ namespace
  is registered as a protected family (rebuildable: false): no sweeper
  or GC may delete under it.
- New crc32c (Castagnoli) utility with RFC known-answer tests.
2026-07-15 10:49:02 -07:00

474 lines
20 KiB
TypeScript

/**
* @module db/types
* @description Shared types for the 8.0 generational-MVCC Db API: the
* declarative transaction-operation union consumed by `brain.transact()`,
* the options bags for `transact()` / `compactHistory()`, the persisted
* record shapes of the generational record layer, and the narrow storage
* contract (`GenerationStorage`) the record layer needs from an adapter.
*
* Persisted layout (all paths relative to the storage root):
*
* ```
* _system/generation.json — { generation, updatedAt } (atomic tmp+rename)
* _system/manifest.json — { version, generation, committedAt, horizon }
* (atomic tmp+rename — the rename IS the commit point)
* _system/tx-log.jsonl — one JSON line per committed transact:
* { generation, timestamp, meta? } (append-only)
* _generations/<N>/tx.json — the generation-N delta: touched noun/verb ids + meta
* _generations/<N>/prev/<id>.json — immutable before-image of <id> as of commit N
* ```
*
* Design note: the manifest holds the committed watermark + compaction
* horizon rather than a full `id → latest record generation` map. The id
* mapping is distributed across the per-generation `tx.json` deltas, which
* makes each commit O(ids touched) instead of O(all ids) — see
* `docs/ADR-001-generational-mvcc.md` for the full justification.
*/
import type { AddParams, UpdateParams, RelateParams, Entity, Relation } from '../types/brainy.types.js'
// ============================================================================
// Transaction operations (brain.transact input)
// ============================================================================
/**
* @description Create an entity. Carries the same parameters as
* `brain.add()`; supply `id` to choose the entity id (otherwise a UUID v4 is
* generated and returned in the operation's planning, exactly as `add()`
* does).
*/
export interface TxAddOperation<T = any> extends AddParams<T> {
/** Discriminator. */
op: 'add'
}
/**
* @description Update an entity. Carries the same parameters as
* `brain.update()` including per-entity CAS via `ifRev` — an `ifRev`
* conflict rejects the whole transaction before anything is applied.
*/
export interface TxUpdateOperation<T = any> extends UpdateParams<T> {
/** Discriminator. */
op: 'update'
}
/**
* @description Remove an entity (and, exactly like `brain.remove()`, every
* relationship where it is source or target — the cascade is part of the
* same atomic batch).
*/
export interface TxRemoveOperation {
/** Discriminator. */
op: 'remove'
/** Id of the entity to remove. */
id: string
}
/**
* @description Create a relationship. Carries the same parameters as
* `brain.relate()` (including `bidirectional`). Duplicate relationships
* (same from/to/type) are deduplicated exactly as `relate()` does — the
* operation becomes a no-op that resolves to the existing relationship id.
*/
export interface TxRelateOperation<T = any> extends RelateParams<T> {
/** Discriminator. */
op: 'relate'
}
/**
* @description Delete a relationship by id (mirror of `brain.unrelate()`).
*/
export interface TxUnrelateOperation {
/** Discriminator. */
op: 'unrelate'
/** Id of the relationship to delete. */
id: string
}
/**
* @description One declarative operation inside a `brain.transact()` batch.
* The batch executes atomically: either every operation applies and the
* store advances exactly one generation, or none apply and the store is
* byte-identical to its pre-transaction state.
*/
export type TxOperation<T = any> =
| TxAddOperation<T>
| TxUpdateOperation<T>
| TxRemoveOperation
| TxRelateOperation<T>
| TxUnrelateOperation
/**
* @description Options for `brain.transact()`.
*/
export interface TransactOptions {
/**
* Transaction metadata, reified Datomic-style: recorded in
* `_system/tx-log.jsonl` alongside the generation number and commit
* timestamp. Use it for audit fields (author, reason, request id) —
* it replaces commit messages from the pre-8.0 versioning API.
*/
meta?: Record<string, unknown>
/**
* Whole-store compare-and-swap. When provided, the transaction commits
* only if the store's current generation equals this value; otherwise it
* throws `GenerationConflictError` (with `expected`/`actual`) before any
* record is staged.
*/
ifAtGeneration?: number
}
/**
* @description Per-operation results of a committed transaction, in input
* order. `add` resolves to the created entity id, `relate` to the created
* (or deduplicated existing) relationship id; `update`/`remove`/`unrelate`
* resolve to the id they acted on.
*/
export interface TransactReceipt {
/** The generation this transaction committed as. */
generation: number
/** Commit timestamp (ms since epoch) recorded in the tx-log. */
timestamp: number
/** Resolved id per input operation, in input order. */
ids: string[]
}
// ============================================================================
// Compaction
// ============================================================================
/**
* @description Options for `brain.compactHistory()` — explicit retention
* **caps** (Model-B `retention` knob). Each supplied field is an upper bound:
* the oldest unpinned generations are reclaimed while **ANY** supplied cap is
* exceeded (predictable ops — the more you constrain, the more is reclaimed).
* Live `Db` pins are ALWAYS exempt. With no options, every generation not
* protected by a live pin is eligible.
*
* (8.0 rename: the pre-release `retainGenerations`/`retainMs` *floors* became
* `maxGenerations`/`maxAge` *caps* — the `max*` naming signals the semantics.)
*/
export interface CompactHistoryOptions {
/**
* Keep at most this many of the most recent committed generations — reclaim
* the oldest unpinned generations while the count exceeds it (`0` = reclaim
* everything not protected by a live pin).
*/
maxGenerations?: number
/**
* Keep only generations committed within the last `maxAge` milliseconds —
* reclaim unpinned generations older than the window.
*/
maxAge?: number
/**
* Keep total generational-history bytes at or below this — reclaim the oldest
* unpinned generations while the total exceeds it. Byte accounting is the sum
* of each surviving generation's serialized record set (`GenerationDelta.bytes`).
*/
maxBytes?: number
}
/**
* @description Result of `brain.compactHistory()`.
*/
export interface CompactHistoryResult {
/** Number of generation record-sets reclaimed by this call. */
removedGenerations: number
/**
* The new compaction horizon — the highest generation whose record-set has
* been reclaimed. Generations BELOW the horizon are no longer reachable via
* `asOf()` (their reads would need the reclaimed before-images); the
* horizon itself stays reachable, resolved from the record-sets above it.
*/
horizon: number
}
// ============================================================================
// Db surfaces
// ============================================================================
/**
* @description Result of `db.since(priorDb)` — the ids touched by committed
* generations in the interval `(priorDb.generation, db.generation]`. Under
* Model-B EVERY write is its own generation, so single-operation writes
* (`add`/`update`/`remove`/`relate` outside `transact()`) ARE reflected here,
* exactly like `transact()` commits.
*/
export interface ChangedIds {
/** Entity (noun) ids touched in the interval, sorted. */
nouns: string[]
/** Relationship (verb) ids touched in the interval, sorted. */
verbs: string[]
/** Exclusive lower bound of the interval. */
fromGeneration: number
/** Inclusive upper bound of the interval. */
toGeneration: number
}
/**
* @description The result of {@link Brainy.diff}: ids that came into existence,
* went away, or changed value between two states, each split by kind. Unlike the
* raw touched-id set of {@link ChangedIds}, `diff` EARNS its name — it resolves
* each touched id at BOTH endpoints and classifies by existence (added/removed)
* and stable value comparison (modified). An id touched within the interval but
* whose endpoint states are identical (touched-then-reverted) appears in NONE of
* the three buckets. Orientation is `a → b`: `added` exists at `b` not `a`.
*/
export interface DiffResult {
/** Ids absent at `a`, present at `b`. */
added: { nouns: string[]; verbs: string[] }
/** Ids present at `a`, absent at `b`. */
removed: { nouns: string[]; verbs: string[] }
/** Ids present at both endpoints with a changed stored value. */
modified: { nouns: string[]; verbs: string[] }
/** The resolved generation of endpoint `a`. */
fromGeneration: number
/** The resolved generation of endpoint `b`. */
toGeneration: number
}
/**
* @description One version of a single entity or relationship in its
* {@link EntityHistory} — the materialized state at a committed generation that
* changed it. `value` is `null` when the id did not exist at that version
* (created-after / removed). Granularity is per-write (every `transact()` AND
* single-op write is its own generation — Model-B).
*/
export interface HistoryVersion<T = any> {
/** The committed generation that produced this version. */
generation: number
/** Commit timestamp of that generation (ms since epoch). */
timestamp: number
/** Materialized state at this version, or `null` if the id did not exist then. */
value: Entity<T> | Relation<T> | null
}
/**
* @description The result of {@link Brainy.history}: every distinct version of
* ONE id over a generation range, oldest → newest. Consecutive identical states
* are collapsed (one entry per distinct state). `kind` is auto-detected from the
* id's presence in the noun vs verb spaces.
*/
export interface EntityHistory<T = any> {
/** The id whose history this is. */
id: string
/** Whether `id` is an entity (`'noun'`) or relationship (`'verb'`). */
kind: 'noun' | 'verb'
/** Distinct versions, oldest first. */
versions: HistoryVersion<T>[]
}
// ============================================================================
// Persisted record shapes
// ============================================================================
/**
* @description An immutable before-image of one entity or relationship,
* persisted under `_generations/<N>/prev/<id>.json` — the state of the id
* immediately before commit `N` touched it. `metadata`/`vector` hold the
* *raw stored objects* (exact bytes that live at the entity's canonical
* storage paths); `null` means the corresponding file did not exist (for
* before-images of created ids, both are `null`). Before-images are both
* the crash-recovery undo log and the point-in-time read source: the state
* of id X at pinned generation G is the before-image of the first committed
* generation after G that touched X.
*/
export interface GenerationRecord {
/** Whether this record images an entity (noun) or a relationship (verb). */
kind: 'noun' | 'verb'
/** Raw stored metadata object, or `null` if the metadata file was absent. */
metadata: unknown | null
/** Raw stored vector object, or `null` if the vector file was absent. */
vector: unknown | null
}
/**
* @description The per-generation delta persisted at
* `_generations/<N>/tx.json`. For a `transact()` generation it is written and
* fsynced *before* any operation of the transaction executes, so crash
* recovery always knows exactly which ids to restore from before-images. For a
* single-operation (Model-B) generation it is written at group-commit flush
* time with `groupCommit: true` — see that flag.
*/
export interface GenerationDelta {
/** The generation this delta belongs to. */
generation: number
/** Commit timestamp (ms since epoch). */
timestamp: number
/** Transaction metadata (mirrored into the tx-log on commit). */
meta?: Record<string, unknown>
/** Entity ids touched by this generation. */
nouns: string[]
/** Relationship ids touched by this generation. */
verbs: string[]
/**
* Content-blob hashes referenced by this generation's before-image records
* (a MULTISET — one entry per referencing record occurrence), captured at
* persist time so compaction can release the exact history references it
* reclaims without re-reading the records. Always present (possibly empty)
* on deltas written under the temporal-blob contract; absent only on
* pre-contract generations, for which compaction falls back to reading the
* record-set itself.
*/
blobHashes?: string[]
/**
* `true` for a Model-B single-operation generation persisted by the async
* group-commit flush (`GenerationStore.flushPendingSingleOps`). It marks the
* **drop-without-restore** crash-recovery contract: the live write already
* mutated canonical storage and was acknowledged to the caller BEFORE the
* flush, so an uncommitted (crashed-mid-flush) generation with this flag set
* must be DISCARDED — its before-images must NEVER be restored, or recovery
* would silently revert acknowledged user writes. Absent/`false` on a
* `transact()` generation, whose before-images ARE the durable undo log and
* are restored on rollback (the before-image + execute are one atomic unit).
*/
groupCommit?: boolean
/**
* Serialized byte size of this generation's record set (the `prev/<id>.json`
* before-images plus this delta), recorded at write time so `maxBytes` /
* adaptive retention can bound total history without a storage size API. Read
* back lazily by `GenerationStore.historyBytes()`. Absent on generations
* written before byte accounting existed (treated as 0).
*/
bytes?: number
}
/**
* @description Shape of `_system/manifest.json`. Replaced via atomic
* tmp+rename on every commit — the rename is the commit point: a generation
* directory is committed if and only if its number is ≤ `generation`.
*/
export interface GenerationManifest {
/** Schema version of the manifest file. */
version: 1
/** Committed-transaction watermark. */
generation: number
/** ISO timestamp of the most recent commit (or compaction). */
committedAt: string
/** Compaction horizon — record-sets for generations ≤ this are reclaimed. */
horizon: number
}
/**
* @description One line of `_system/tx-log.jsonl`.
*/
export interface TxLogEntry {
/** The committed generation. */
generation: number
/** Commit timestamp (ms since epoch). */
timestamp: number
/** Transaction metadata, when supplied to `transact()`. */
meta?: Record<string, unknown>
}
// ============================================================================
// Storage contract
// ============================================================================
/**
* @description The narrow storage contract the generational record layer
* needs. `BaseStorage` implements it structurally (both shipped adapters —
* filesystem and memory — inherit the implementation), so any
* `StorageAdapter` produced by the 8.0 storage factory satisfies it.
*
* Raw-object methods bypass branch scoping and operate on storage-root-
* relative paths (the record layer owns `_system/` + `_generations/`);
* entity-raw methods read/write the canonical entity files (branch-scoped,
* write-cache coherent) so before/after images capture exactly the bytes
* the live read paths see.
*/
export interface GenerationStorage {
/** Read a raw object at a storage-root-relative path (`null` if absent). */
readRawObject(path: string): Promise<any | null>
/** Write a raw object at a storage-root-relative path (atomic tmp+rename on disk). */
writeRawObject(path: string, data: any): Promise<void>
/** Delete a raw object (no-op if absent). */
deleteRawObject(path: string): Promise<void>
/** List raw object paths under a prefix (normalized, `.gz`-stripped). */
listRawObjects(prefix: string): Promise<string[]>
/** Remove every object under a prefix (and the directory itself on disk). */
removeRawPrefix(prefix: string): Promise<void>
/** Durability barrier: fsync the given object paths (no-op in memory). */
syncRawObjects(paths: string[]): Promise<void>
/**
* OPTIONAL transaction durability barrier. `commitTransaction` calls
* {@link beginWriteBarrier} immediately before running the planned operations
* and {@link flushWriteBarrier} immediately after — BEFORE the generation
* counter and manifest are advanced. An adapter whose canonical writes are
* not synchronously durable (the filesystem adapter's tmp+rename lands in the
* page cache) MUST implement these so a transaction reported "committed" is
* durable before its generation stamp: otherwise a hard kill can leave the
* fsync'd counter ahead of the still-buffered entity bytes (phantom progress
* for any generation-based consumer). Adapters whose writes are already
* durable per-call (cloud object PUT) may leave these undefined — the store
* treats them as no-ops. `beginWriteBarrier` also resets any tracking left by
* an aborted prior transaction.
*/
beginWriteBarrier?(): void
/** @see beginWriteBarrier — fsync every canonical write since begin. */
flushWriteBarrier?(): Promise<void>
/** Read an entity's raw stored metadata+vector objects. */
readNounRaw(id: string): Promise<{ metadata: any | null; vector: any | null }>
/** Restore an entity's raw stored objects (`null` part ⇒ delete that file). */
writeNounRaw(id: string, record: { metadata: any | null; vector: any | null }): Promise<void>
/** Read a relationship's raw stored metadata+vector objects. */
readVerbRaw(id: string): Promise<{ metadata: any | null; vector: any | null }>
/** Restore a relationship's raw stored objects (`null` part ⇒ delete that file). */
writeVerbRaw(id: string, record: { metadata: any | null; vector: any | null }): Promise<void>
/** Append one line to `_system/tx-log.jsonl`. */
appendTxLogLine(line: string): Promise<void>
/** Read all lines of `_system/tx-log.jsonl` (empty array if absent). */
readTxLogLines(): Promise<string[]>
/**
* OPTIONAL binary raw-byte primitives — the substrate for the generation
* fact log's append-only CRC-framed segments. Feature-detected: a storage
* layer that omits them hosts no fact log (dual-write is skipped; readers
* fall back to canonical enumeration). Paths are used VERBATIM (no
* suffixing). Append durability rides `syncRawObjects` at the commit
* barrier, exactly like the staged history files.
*/
appendRawBytes?(path: string, bytes: Uint8Array): Promise<void>
/** Read a raw binary file whole; absent → null; a real fault throws. */
readRawBytes?(path: string): Promise<Uint8Array | null>
/** Replace a raw binary file atomically (tmp → fsync → rename). */
writeRawBytes?(path: string, bytes: Uint8Array): Promise<void>
/** Byte size of a raw binary file, or null when absent. */
rawByteSize?(path: string): Promise<number | null>
/**
* OPTIONAL temporal-blob contract (implemented by blob-aware storage; the
* generation store treats the hashes as opaque strings). Extract the
* content-blob hashes a record-set references — a pure multiset extraction,
* no side effects.
*/
extractBlobHashesFromRecords?(records: GenerationRecord[]): string[]
/**
* OPTIONAL: record history references for the given hashes (one increment
* per occurrence). Called BEFORE the referencing record-set is persisted —
* a crash between the two can only over-count (a leak the scrub repairs),
* never under-count (which would risk reclaiming bytes history still needs).
*/
recordHistoryBlobReferences?(hashes: string[]): Promise<void>
/**
* OPTIONAL: release history references for the given hashes (one decrement
* per occurrence) and physically reclaim any hash left with zero live AND
* zero history references. Called AFTER the referencing record-set is
* deleted by compaction — same over-count-only crash ordering.
*/
releaseHistoryBlobReferences?(hashes: string[]): Promise<void>
/**
* Register the generation-bump hook invoked on every entity-visible
* single-operation write (see `BaseStorage.setGenerationBumpHook`).
*/
setGenerationBumpHook(hook: (() => void) | undefined): void
/** Snapshot the entire store to a directory (hard-link farm on disk). */
snapshotToDirectory(targetPath: string): Promise<void>
/** Replace the entire store's contents from a snapshot directory. */
restoreFromDirectory(sourcePath: string): Promise<void>
}