Compare commits
No commits in common. "main" and "v8.3.1" have entirely different histories.
66 changed files with 275 additions and 7636 deletions
78
CHANGELOG.md
78
CHANGELOG.md
|
|
@ -2,84 +2,6 @@
|
|||
|
||||
All notable changes to this project will be documented in this file. See [standard-version](https://github.com/conventional-changelog/standard-version) for commit guidelines.
|
||||
|
||||
### [8.9.0](https://github.com/soulcraftlabs/brainy/compare/v8.8.2...v8.9.0) (2026-07-19)
|
||||
|
||||
- docs: measured performance envelopes v1 (per-op p50/p95 at 1k and 10k, pure-JS floor) (5cabd78)
|
||||
- fix: release drains in-flight writer-lock heartbeat — no phantom lock after unlink (70e4bc8)
|
||||
- feat: flush() never compacts — history maintenance moves to close() with bounded passes (300d9f2)
|
||||
|
||||
|
||||
### [8.8.2](https://github.com/soulcraftlabs/brainy/compare/v8.8.1...v8.8.2) (2026-07-19)
|
||||
|
||||
- fix: one field-resolution law across aggregation hooks, source.where, removeMany, and find() spellings (945d92d)
|
||||
- chore: push public docs to the soulcraft.com ingest door on release (42037d0)
|
||||
|
||||
|
||||
### [8.8.1](https://github.com/soulcraftlabs/brainy/compare/v8.8.0...v8.8.1) (2026-07-18)
|
||||
|
||||
- fix: O(1) adaptive retention accounting + historyStats fleet audit (6207e48)
|
||||
- fix: import dedup off-switch honesty + brain-owned lifecycle for the background pass (4fcef7b)
|
||||
|
||||
|
||||
### [8.8.0](https://github.com/soulcraftlabs/brainy/compare/v8.7.1...v8.8.0) (2026-07-17)
|
||||
|
||||
- feat: OS-limit detection for pool-scale deployments (16a73b8)
|
||||
|
||||
|
||||
### [8.7.1](https://github.com/soulcraftlabs/brainy/compare/v8.7.0...v8.7.1) (2026-07-17)
|
||||
|
||||
- fix: race-proof writer-lock acquisition + machine-readable conflict through init (01a3b46)
|
||||
|
||||
|
||||
### [8.7.0](https://github.com/soulcraftlabs/brainy/compare/v8.6.0...v8.7.0) (2026-07-17)
|
||||
|
||||
- feat: scaled transact budgets + labeled timeout diagnostics + envelope docs (6ef9fcb)
|
||||
|
||||
|
||||
### [8.6.0](https://github.com/soulcraftlabs/brainy/compare/v8.5.2...v8.6.0) (2026-07-17)
|
||||
|
||||
- feat: brain.auditGraph() — read-only graph-truth audit (2a03fae)
|
||||
|
||||
|
||||
### [8.5.2](https://github.com/soulcraftlabs/brainy/compare/v8.5.1...v8.5.2) (2026-07-17)
|
||||
|
||||
- fix: exception-safe aggregation backfill + generation-verified adoption + loud open-path guards (a77b064)
|
||||
|
||||
|
||||
### [8.5.1](https://github.com/soulcraftlabs/brainy/compare/v8.5.0...v8.5.1) (2026-07-17)
|
||||
|
||||
- fix: aggregation state adoption on reopen + single-flight backfill + query-cap ratchet removal (da55be7)
|
||||
- docs: external-backups/sparse-storage guide + generation fact log concept (593bb8b)
|
||||
|
||||
|
||||
### [8.5.0](https://github.com/soulcraftlabs/brainy/compare/v8.4.0...v8.5.0) (2026-07-15)
|
||||
|
||||
- test: tolerant timing assertion in the execution-time measure test (4dc0a92)
|
||||
- feat: committedGeneration capability + pinned durability/stability contracts (d1ecee1)
|
||||
- docs: RELEASES.md entry for 8.5.0 (provider fact-log access + shared verifier) (e4f37cd)
|
||||
- feat: provider access to the fact log + shared stamp verifier via internals (352e356)
|
||||
|
||||
|
||||
### [8.4.0](https://github.com/soulcraftlabs/brainy/compare/v8.3.3...v8.4.0) (2026-07-15)
|
||||
|
||||
- docs: RELEASES.md entry for 8.4.0 (generation fact log + family stamp) (4a60b43)
|
||||
- feat: entity-tree family stamp — sourceGeneration + rollup coherence at open (2888ae6)
|
||||
- feat: generation fact log — after-image commit records, dual-written at every commit point (38b0041)
|
||||
|
||||
|
||||
### [8.3.3](https://github.com/soulcraftlabs/brainy/compare/v8.3.2...v8.3.3) (2026-07-15)
|
||||
|
||||
- docs: RELEASES.md entry for 8.3.3 (rename containment fix + repair) (c3feafd)
|
||||
- test: lens-consistency regression — combined vs subtype-only vs canonical ground truth (4fb41f9)
|
||||
- fix: VFS rename moves the containment edge — no ghost in the old directory (af8c179)
|
||||
|
||||
|
||||
### [8.3.2](https://github.com/soulcraftlabs/brainy/compare/v8.3.1...v8.3.2) (2026-07-14)
|
||||
|
||||
- docs: RELEASES.md entry for 8.3.2 (honest counters) (0932ecd)
|
||||
- fix: honest counters — removal never re-reads the removed record + repairIndex recounts and persists all rollups (2e2ba9c)
|
||||
|
||||
|
||||
### [8.3.1](https://github.com/soulcraftlabs/brainy/compare/v8.3.0...v8.3.1) (2026-07-14)
|
||||
|
||||
- docs: RELEASES.md entry for 8.3.1 (full-removal deletes + family-scoped gate) (c0c68ac)
|
||||
|
|
|
|||
384
RELEASES.md
384
RELEASES.md
|
|
@ -8,392 +8,8 @@ Full auto-generated changelog: `CHANGELOG.md` · Releases: https://github.com/so
|
|||
- Debugging data, query, or storage behaviour
|
||||
- A new Brainy feature is available that you want to adopt
|
||||
|
||||
## Removed APIs — 7.x → 8.x (the complete ledger)
|
||||
|
||||
Every public API removed at the 8.0 major, with its sanctioned replacement. If your code
|
||||
still calls a left-column name on 8.x it throws (or the config key is rejected) — the
|
||||
replacement is always a one-line change. (Standing contract from 8.9.0 forward: removals
|
||||
happen only at majors, after ≥1 minor of loud runtime deprecation naming the replacement.)
|
||||
|
||||
| Removed (7.x) | Replacement (8.x) |
|
||||
|---|---|
|
||||
| `brain.search(query, k)` | `find({ query })` — semantic; `find({ query, searchMode })` for hybrid |
|
||||
| `brain.getRelations({...})` | `related(id, opts)` for adjacency; `find({ connected: {...} })` for scoped traversal |
|
||||
| `brain.neural()` clustering | `find({ vector })` + aggregation `GROUP BY` |
|
||||
| `Db.search()` | `db.find({ vector })` |
|
||||
| Pre-8.0 storage path aliases (`directory`, `basePath`, …) | one `storage.path` key (old aliases throw) |
|
||||
| Reserved keys inside `metadata` bags (silently remapped in 7.x) | top-level params (`subtype`, `visibility`, `confidence`, `weight`, …) — reserved-in-bag throws |
|
||||
| 7.x COW branches layout (`branches/main/`) | generational MVCC (`asOf()`, `now()`, `db.persist(path)`) — on-disk migration is automatic at first 8.x open |
|
||||
|
||||
The fork/snapshot family (`brain.snapshot()`, `createSnapshot()`, `restoreSnapshot()`)
|
||||
is sometimes cited as a 7.x removal — those methods never existed on 7.x; the 8.0 Db API
|
||||
(`asOf`/`persist`/`restore({confirm})`) is their first real implementation.
|
||||
|
||||
---
|
||||
|
||||
## v8.9.0 — 2026-07-19 (flush is durability-only: history maintenance moves to close())
|
||||
|
||||
The write path stops paying maintenance costs — the last structural piece of the
|
||||
flush-storm class (a production deployment measured single writes blocked 25–191s behind
|
||||
history reclaim running inline on flush under memory pressure):
|
||||
|
||||
- **`flush()` never compacts history.** It persists the current window's deltas and
|
||||
nothing else — its cost no longer depends on history backlog or retention mode, in any
|
||||
configuration. **`close()` is the auto-compaction site** (time-bounded per pass, ~5s;
|
||||
an early stop is a consistent prefix and the next pass resumes).
|
||||
- **`compactHistory()` gains `timeBudgetMs`** — bound your own maintenance windows; the
|
||||
same resumable-prefix guarantee applies.
|
||||
- **The documented trade**: a long-lived writer that never closes accumulates history
|
||||
until its next explicit `compactHistory()`. Predictable writes, explicit maintenance.
|
||||
If you run bounded retention on an always-on service, schedule a periodic
|
||||
`compactHistory({ ...caps, timeBudgetMs })` in your maintenance window.
|
||||
- **New public doc: `docs/performance-envelopes.md`** — measured per-op envelopes
|
||||
(p50/p95 at stated scales, hardware, and backend, with the measuring script cited).
|
||||
Refresh rule going forward: any release touching a measured path re-runs that op's
|
||||
benchmark and updates the envelope in the same release.
|
||||
- **New in this file: the Removed APIs 7.x→8.x table** (top of this document) — every
|
||||
removal with its sanctioned replacement, one place, per the engine-currency contract.
|
||||
Standing from here: removals only at majors, after ≥1 minor of loud runtime deprecation.
|
||||
|
||||
## v8.8.2 — 2026-07-19 (one field-resolution law: reserved-field aggregates stop drifting)
|
||||
|
||||
Four fixes from a consumer conformance audit, all rooted in the same disease — two field-resolution
|
||||
regimes where there must be one:
|
||||
|
||||
- **Aggregates grouped by a RESERVED field (`subtype`, `visibility`, …) now decrement on
|
||||
delete.** The delete/update hooks fed the aggregation engine a partial entity view (type,
|
||||
service, data, metadata only), so a reserved-field `groupBy` resolved to a nonexistent group
|
||||
on the way DOWN — counts drifted upward forever after any delete, and updates that moved an
|
||||
entity between reserved-field groups double-counted it. The hooks now pass the full-fidelity
|
||||
entity view (every reserved field top-level, the same shape the add path uses). If your
|
||||
deployment derives stats from reserved-field aggregates, re-define those aggregates once
|
||||
after upgrading (a changed definition triggers one rescan) or run them fresh — the drifted
|
||||
persisted counts do not self-heal retroactively.
|
||||
- **Aggregation `source.where` on reserved fields now filters** instead of silently matching
|
||||
nothing: the matcher resolves fields through the same resolver `groupBy` uses (top-level
|
||||
standard fields + custom metadata), so `where: { subtype: 'note' }` means what it says.
|
||||
- **`removeMany()` refuses empty/invalid selectors loudly.** A bare array passed positionally
|
||||
(`removeMany([id])` instead of `removeMany({ ids: [id] })`), an empty params object, or
|
||||
`ids: []` used to resolve successfully having deleted nothing. All three now throw.
|
||||
- **`find()` accepts both where-key spellings.** Metadata is flattened at index time
|
||||
(`metadata.entry.title` indexes as `entry.title`); a `metadata.`-prefixed where key now
|
||||
falls back to its flattened spelling when the prefixed one isn't indexed — the
|
||||
"unindexed field(s), returning []" confusion for storage-shaped spellings is gone. (A
|
||||
literal nested custom key named `metadata` still wins when indexed as spelled.)
|
||||
|
||||
## v8.8.1 — 2026-07-18 (flush no longer walks the whole generation history + the import dedup off-switch is now honest)
|
||||
|
||||
### The flush-storm fix (production incident, reported by a long-running deployment)
|
||||
|
||||
Under the default adaptive retention, **every `flush()` re-walked the entire committed
|
||||
generation history** to compute total history bytes for the budget check — O(all
|
||||
generations) with disk re-reads past the 4,096-entry delta-cache bound. On a brain with
|
||||
70,000+ accumulated generations that turned every write into a full-tail scan (60-100s
|
||||
writes), even though the budget (free-RAM-based) never tripped and nothing was ever
|
||||
reclaimed. Fixed:
|
||||
|
||||
- `historyBytes()` now maintains a **running total**: seeded by one walk on first use,
|
||||
then updated incrementally at every commit and reclaim — the adaptive retention check
|
||||
on every flush is O(1). Invariant regression-pinned (running total ≡ fresh walk through
|
||||
both commit paths and compaction).
|
||||
- New **`brain.historyStats()`** (read-only, exported `HistoryStats`): generation count,
|
||||
total on-disk bytes, generation/timestamp range, compaction horizon, retention mode,
|
||||
and the effective adaptive budget — the one-call fleet-audit for sizing retention
|
||||
exposure per brain.
|
||||
- Interim guidance for keep-everything deployments already affected: `retention: 'all'`
|
||||
skips the adaptive accounting entirely (and is the correct policy if you never want
|
||||
history reclaimed). The accumulated files are harmless at rest; this release removes
|
||||
the per-write cost of their existence.
|
||||
|
||||
### The import dedup off-switch (lifecycle honesty)
|
||||
|
||||
The post-import background deduplication pass (a merge-DELETE writer that runs ~5 minutes
|
||||
after an import, merging entities judged duplicates by id / name / vector similarity) had
|
||||
three lifecycle defects, all fixed:
|
||||
|
||||
- **`enableDeduplication: false` now actually disables it.** The background pass was
|
||||
scheduled unconditionally — an import that explicitly opted out could still have
|
||||
entities auto-removed 5 minutes later. The flag now gates BOTH the inline merge and
|
||||
the background pass (regression-pinned).
|
||||
- **One deduplicator per brain, owned by the brain.** Each `import()` call constructed its
|
||||
own coordinator + deduplicator, so the "debounced" timer never actually debounced across
|
||||
imports (N imports = N delete timers). The brain now owns a single instance — the
|
||||
debounce genuinely spans imports — and `close()` cancels pending work, so a delete pass
|
||||
can never fire against a closed brain.
|
||||
- **The 5-minute timer is unref'd** — a pending pass no longer holds the process open
|
||||
(the exit-hang class; this timer had escaped the earlier sweep).
|
||||
|
||||
Retention note for keep-everything deployments: with `enableDeduplication: false` on
|
||||
import calls and `retention: 'all'` in config, no engine path removes records
|
||||
automatically.
|
||||
|
||||
## v8.8.0 — 2026-07-17 (OS-limit detection for pool-scale deployments)
|
||||
|
||||
Small minor: brains now detect the two OS limits that bite at pool scale and warn **before**
|
||||
the incident instead of during it.
|
||||
|
||||
- At open (once per process, Linux-only, measurement-only), Brainy reads `RLIMIT_NOFILE`
|
||||
(soft/hard, from `/proc/self/limits`) and `vm.max_map_count`, and warns loudly when either
|
||||
sits below the pool-scale floors (soft NOFILE < 65 536; max_map_count < 262 144) — with the
|
||||
exact raise commands (`ulimit -n` / `LimitNOFILE=` / `sysctl vm.max_map_count`). On stock
|
||||
defaults the failure otherwise arrives as `EMFILE` or a failed mmap deep inside an index
|
||||
open, long after the real cause stopped being visible. An unreadable limit produces **no**
|
||||
warning — no measurement, no claim (non-Linux platforms stay silent).
|
||||
- Exported for ops doors: `checkOsLimits()` returns the full `OsLimitsReport`
|
||||
(values + warnings) programmatically, with the floors exported as constants.
|
||||
|
||||
## v8.7.1 — 2026-07-17 (writer-lock acquisition is race-proof + machine-readable through init)
|
||||
|
||||
Two hardenings of the multi-process writer lock (the `locks/_writer.lock` lease that makes a
|
||||
second writer on the same brain directory fail loudly):
|
||||
|
||||
- **Lock acquisition claims atomically.** The acquire path used read-then-write, leaving a
|
||||
window where two processes racing an *absent* lock could both "succeed" — and the loser
|
||||
kept running unlocked, silently. The claim is now an atomic create-exclusive write
|
||||
(`O_EXCL`): exactly one racer wins; the loser re-evaluates and either fails loudly with
|
||||
the winner's details or performs a verified stale-takeover. Bounded retries; contention
|
||||
beyond them fails loudly rather than degrading into a lockless open.
|
||||
- **`BRAINY_WRITER_LOCKED` survives `init()`.** The conflict error documents a
|
||||
machine-readable contract (`err.code`, `err.lockInfo` with the holder's pid/host/
|
||||
heartbeat), but init's error wrapping silently stripped both, leaving consumers a message
|
||||
to regex against. The error now passes through unwrapped.
|
||||
|
||||
Measured while verifying (for operators sizing audits): `brain.auditGraph()` at a
|
||||
production-consumer scale of ~2,600 relationships / 800 entities costs ~0.1 s warm and
|
||||
~0.5 s cold, with exact scar counting across reopen.
|
||||
|
||||
## v8.7.0 — 2026-07-17 (bulk-transact ergonomics: scaled budgets + timeout telemetry)
|
||||
|
||||
The bulk-import ergonomics release, from a consumer's measured production incident (a serial
|
||||
import on network-attached storage at ~2 s/op met a flat 30 s transaction budget):
|
||||
|
||||
- **The transact apply budget now scales with the batch** — `max(30 s, opCount × 2 s)` — or
|
||||
is exactly what you pass as the new `TransactOptions.timeoutMs`. A flat 30 s cap silently
|
||||
limited honest bulk work to ~15 operations on slow disks while looking generous for small
|
||||
batches. Internal batch paths (e.g. `removeMany` chunks) get the same scaling.
|
||||
- **`TransactionTimeoutError` is a diagnosis, not just a failure**: it now reports the
|
||||
operation it stopped at as `i/N` with the operation's name, elapsed vs budgeted time, and
|
||||
states the batch rolled back atomically and is retryable. Its `context` carries the same
|
||||
fields programmatically.
|
||||
- **The transact envelope is documented** — batch sizing, budget math, chunking with
|
||||
`ifAbsent` idempotency, and the precompute pattern (`embedBatch` + per-op `vector`) that
|
||||
keeps model inference out of the commit path. Guide: `docs/guides/optimistic-concurrency.md`.
|
||||
- Note: `brain.embed()` / `brain.embedBatch()` (the precompute APIs) already ship — public,
|
||||
with native-provider passthrough, verified end-to-end (batch and single paths produce
|
||||
bit-identical vectors; a vector-supplied `add` is fully searchable). Honest measurement:
|
||||
on the default WASM engine, batch throughput ≈ sequential (~160 ms/text) — the precompute
|
||||
win is keeping inference out of the budgeted commit path, not raw embedding speed.
|
||||
|
||||
## v8.6.0 — 2026-07-17 (brain.auditGraph — the graph-truth verification instrument)
|
||||
|
||||
A minor release adding one new public API, from the fleet's graph-trust program: a read-only
|
||||
audit that **proves whether relationship reads return stored truth** on a given brain.
|
||||
|
||||
- **`brain.auditGraph(options?)`** walks every canonical relationship record, queries the same
|
||||
read path applications use (`related()` / VFS `readdir`) with all visibility tiers included,
|
||||
and classifies every discrepancy into its failure family: `missingFromReads` (records the
|
||||
read path omits — a stale adjacency index), `danglingEndpoints` (relationships whose endpoint
|
||||
entity no longer exists — the historical partial-delete scar class), and `readOnlyVerbIds`
|
||||
(read-path edges with no stored record — ghosts). Design-hidden internal/system edges are
|
||||
counted separately so intentional hiding is never misclassified as loss. Counts are exact;
|
||||
example lists cap at `maxExamples` with an explicit `truncatedExamples` flag; the result is
|
||||
narrated loudly on incoherence. Mutates nothing — safe on a live brain.
|
||||
- The operational pairing: audit → if incoherent, `repairIndex()` → audit again. A `coherent`
|
||||
report after the repair is the verified all-clear. Run it after any engine upgrade, restore,
|
||||
or migration. Guide: `docs/guides/inspection.md`.
|
||||
- Also: `getNounIds` pagination now refuses an undecodable resume cursor loudly (the same
|
||||
contract `getNouns`/`getVerbs` gained in 8.5.2 — the third and final walk brought under it).
|
||||
- Types exported: `GraphAuditReport`, `GraphAuditDiscrepancy`.
|
||||
|
||||
## v8.5.2 — 2026-07-17 (aggregation backfill: exception-safe, generation-verified, and loud)
|
||||
|
||||
Hardening patch from a migration incident (a byte-copied store on new hardware; the service
|
||||
entered a silent full-CPU loop at boot). Four changes, all in the aggregation engine's
|
||||
backfill/adoption path:
|
||||
|
||||
- **Backfill walks are exception-safe and non-destructive.** A rescan now builds into a
|
||||
staging map and swaps in atomically on completion; a mid-walk failure drops the staging map,
|
||||
keeps the previous live state serving, and surfaces the storage error to the failing query.
|
||||
Previously the walk wiped live state *before* a scan that could throw, never cleared the
|
||||
pending flag on failure, and re-ran a full walk on every subsequent query — a silent
|
||||
wipe/walk/throw loop at the caller's retry rate.
|
||||
- **Failed walks are latched.** After a walk fails, retries within a 30-second cooldown rethrow
|
||||
the recorded error instantly instead of re-walking — a tight caller-side retry loop now costs
|
||||
one loud error per query, never a full store walk per query.
|
||||
- **Adoption is generation-verified.** Persisted aggregation state is stamped with the store's
|
||||
committed generation at flush; reopen adoption requires the stamp to equal the current
|
||||
watermark. Stale state (unclean shutdown) or over-counting state (a fact-log truncation on a
|
||||
copied store pulled the watermark back) triggers exactly one loud rescan — never a silent
|
||||
adopt. Pre-8.5.2 state on generation-aware stores rescans once after upgrade, then is stamped.
|
||||
- **The path narrates.** Adoption decisions, no-adoptable-state outcomes, walk start/finish
|
||||
(entity count + duration), and walk failures all log by default; a non-advancing storage
|
||||
pagination cursor aborts the walk loudly instead of looping forever.
|
||||
|
||||
Plus three guards from a full audit of every loop in the open/init path:
|
||||
|
||||
- **Invalid pagination cursors fail loudly.** A supplied-but-undecodable resume token to
|
||||
`getNouns`/`getVerbs` used to silently restart the walk at offset 0 — to a `while(hasMore)`
|
||||
caller that re-serves page 1 forever (an unbounded silent CPU loop). It now throws with a
|
||||
clear message instead.
|
||||
- **The graph cold-load verb walk has a stall guard.** `hasMore=true` with a missing or
|
||||
non-advancing cursor aborts loudly instead of re-reading the same page forever.
|
||||
- **A derived index AHEAD of the store is named at open.** Brainy already surfaced a provider
|
||||
generation *behind* the committed watermark; the *ahead* direction (the signature of a
|
||||
byte-copy of a live service, or a log truncation during crash recovery) now logs a loud
|
||||
warning explaining what happened and that `brain.repairIndex()` forces a heal — instead of
|
||||
passing unnamed into whatever the derived index does next.
|
||||
|
||||
## v8.5.1 — 2026-07-17 (aggregation state survives restarts + the query-cap ratchet removed)
|
||||
|
||||
Patch release from a production incident (aggregate/count paths taking 40–90 s on an idle box
|
||||
while vector search stayed fast, and every `find({ limit: 5000 })` suddenly failing against an
|
||||
"auto-configured query limit of 1000"). Three fixes, one cosmetic:
|
||||
|
||||
- **Aggregation state is actually adopted on reopen.** The boot pattern `defineAggregate()` →
|
||||
query raced the engine's async state load: the synchronous define always won, flagged a
|
||||
backfill, and the first query then wiped the just-loaded persisted state and re-walked the
|
||||
entire store — every restart, forever. Reopening with an unchanged definition now adopts the
|
||||
persisted state directly (zero scans); a backfill runs only on a real definition change, a
|
||||
missing/failed state load, or a write that landed before adoption (exactness wins). Apps that
|
||||
rely on persisted definitions without re-defining at boot also no longer race a spurious
|
||||
"Aggregate not defined".
|
||||
- **Backfills are single-flight and batched.** Concurrent queries on a cold aggregate used to
|
||||
each wipe the others' partial state and start their own full store walk — under steady query
|
||||
arrival the store never converged (the 40–90 s loop). Now all concurrent queries share one
|
||||
walk, and one walk fills every aggregate pending backfill (M aggregates ≠ M scans).
|
||||
- **The query-cap "learning" ratchet is removed.** `maxLimit` is a memory-protection bound, but
|
||||
a hidden tuner shrank it 20 % per recorded query while the lifetime-average query time
|
||||
exceeded 1 s — down to a floor of 1000, below the documented 10 000 auto floor, with no
|
||||
practical recovery, and the resulting error blamed "available free memory" (stale basis
|
||||
label). The cap now comes from its construction-time basis (or your explicit
|
||||
`maxQueryLimit` / `reservedQueryMemory`) alone and never changes at runtime; query timing is
|
||||
recorded for diagnostics only.
|
||||
- **Cosmetic:** the engine's own persistence keys (`__aggregation_*`, `brainy:entityIdMapper`)
|
||||
no longer log `[Storage] Unknown key format` at boot — they were always routed correctly;
|
||||
they're now recognized before the warning fires.
|
||||
|
||||
Operationally: if a host was bitten, upgrading and restarting is the whole fix — no repair
|
||||
ritual needed. Setting `maxQueryLimit` explicitly remains the valve that bypasses auto-detection
|
||||
entirely.
|
||||
|
||||
## v8.5.0 — 2026-07-15 (provider access to the fact log + the shared stamp verifier)
|
||||
|
||||
Small additive follow-up to 8.4.0, from the native accelerator's first consumption pass:
|
||||
|
||||
- **Index providers can now reach the fact log through the storage adapter** — new optional
|
||||
capability `storage.scanFacts()` / `storage.factLogHeadGeneration()` / `storage.factSegmentPaths()`,
|
||||
wired by the host brain at init as a closure over its live log. Providers hold only `storage` and
|
||||
must never construct their own fact-log reader (the log's open path is writer-side); `null` means
|
||||
"no fact log here — use the enumeration walk."
|
||||
- **The family-stamp verifier is shared** via `@soulcraft/brainy/internals`
|
||||
(`readFamilyStamp` / `writeFamilyStamp` / `verifyFamilyStamp` + types) so first-party native
|
||||
providers run literally the same verification function, never a synchronized copy.
|
||||
- **Rollup stamp invariants accept strings** (`number | string`) — content fingerprints such as a
|
||||
per-tree SHA-256 are valid invariant values; a type mismatch reads as incoherence, never a pass.
|
||||
- **`storage.committedGeneration()`** — the committed watermark as a capability, so a provider
|
||||
compares its stamp's `sourceGeneration` against the store's truth without parsing the private
|
||||
manifest format.
|
||||
- **Two durability/stability contracts pinned in the suite:** fsync-before-ack (holds for
|
||||
`transact()` today; the single-op path is pinned as the documented future target — group commit
|
||||
becomes latency batching, never durability skipping) and scan-stability-under-rotation (a scan
|
||||
snapshot yields exactly its facts — no gaps, duplicates, or bleed-in — while segments rotate
|
||||
beneath it).
|
||||
|
||||
No behavior change for applications; all additions.
|
||||
|
||||
## v8.4.0 — 2026-07-15 (the generation fact log — a sequential, self-verifying commit stream)
|
||||
|
||||
A minor release, fully backward-compatible (all additions; no behavior change for existing APIs).
|
||||
This is infrastructure: it changes nothing about how you query today, and lays the substrate that
|
||||
makes index heals and incremental catch-up sequential-read problems instead of directory walks.
|
||||
|
||||
- **Every committed write now also appends a "fact" — an after-image commit record.** Alongside the
|
||||
existing before-image history, each committed generation (single-op and `transact()` alike) appends
|
||||
what each touched entity/relationship *became* — or a body-less tombstone for a removal — to an
|
||||
append-only, checksummed segment log under `_generations/facts/`. Crash-safe by construction: a
|
||||
torn tail is detected and ignored; on open the log is reconciled to committed truth, so an absent
|
||||
generation always means "never committed." `transact()` facts are durable when `transact()`
|
||||
returns; single-op facts ride the same group-commit flush as their history.
|
||||
|
||||
- **New: `brain.scanFacts()`** — stream committed facts in commit order, in batches, with heal-grade
|
||||
telemetry (total scope up front; per-batch generation range, byte size, and segment; a summary
|
||||
cross-check at the end; loud abort on any gap — never a silent skip). **`brain.factSegmentPaths()`**
|
||||
hands zero-copy consumers the immutable sealed segment files directly. New exported types:
|
||||
`CommitFact`, `FactOp`, `FactScanBatch`, `FactScanHandle`. Facts accumulate from the first write
|
||||
after upgrading — pre-existing history is not retroactively converted (enumeration remains the
|
||||
fallback for old data).
|
||||
|
||||
- **New: the entity-tree family stamp.** At every flush/close, brainy stamps which committed
|
||||
generation the canonical entity files reflect plus the rollup invariants (entity/relationship
|
||||
counts) that verify the tree whole. At open, coherence is a comparison — a genuine divergence is
|
||||
loud and names the failing invariant; `repairIndex()` recounts from canonical and re-stamps. New
|
||||
exports: `readFamilyStamp`, `verifyFamilyStamp`, `ENTITY_TREE_STAMP_PATH`, `FamilyStamp` types.
|
||||
|
||||
- **Storage adapters** gain optional binary raw-byte primitives (`appendRawBytes`, `readRawBytes`,
|
||||
`writeRawBytes`, `rawByteSize`) — feature-detected; the filesystem and memory adapters implement
|
||||
them; an adapter without them simply hosts no fact log. The fact-log namespace is registered as a
|
||||
protected family: no sweeper or GC can delete under it.
|
||||
|
||||
No API breaks. 24 new tests; the full commit-path regression suite is green.
|
||||
|
||||
## v8.3.3 — 2026-07-15 (rename moves the containment edge — no ghost in the old directory)
|
||||
|
||||
One production-reported fix plus a repair path, completing the delete/move hygiene arc (8.3.1 fixed
|
||||
deletes, 8.3.2 fixed counters, this fixes moves).
|
||||
|
||||
- **A cross-directory `vfs.rename()` now MOVES the containment edge instead of accumulating one per
|
||||
parent.** The old parent's `Contains` edge was never removed on a move, leaving the entity a child
|
||||
of **both** directories: `readdir(oldDir)` kept listing it after the move, re-creating the old path
|
||||
showed the same name twice, and any tree-walking consumer (sync engines, file browsers) saw the
|
||||
file in two places. The old edge is now removed by edge id, resolved from the graph's own adjacency
|
||||
— a removal never requires reading the thing being removed. Bonus fix in the same seam: a move **to
|
||||
the root** now gets its containment edge (it was previously skipped, orphaning the file out of
|
||||
`readdir('/')`).
|
||||
|
||||
- **`repairIndex()` now also reconciles VFS containment** (new `vfs.repairContainment()`): every VFS
|
||||
entity's containment edges are checked against its canonical `metadata.path` — stale old-parent
|
||||
ghosts and duplicate edges are removed, a missing expected edge is restored, and user
|
||||
knowledge-graph edges are never touched (only `vfs-contains` edges are candidates). Loud per
|
||||
repair. Stores that performed cross-directory renames under ≤8.3.2 should run `brain.repairIndex()`
|
||||
once after upgrading — the same single ritual now heals orphan directories, counters, **and**
|
||||
containment edges.
|
||||
|
||||
- Also ships a permanent lens-consistency regression suite (combined type+subtype vs subtype-only vs
|
||||
canonical ground truth, id-for-id, warm and after a cold reopen), ported from the field
|
||||
investigation that closed the historical lens-drop report.
|
||||
|
||||
No API changes beyond the new optional `vfs.repairContainment()` (also invoked by `repairIndex()`).
|
||||
|
||||
## v8.3.2 — 2026-07-14 (honest counters — the recount + removal-without-re-reading)
|
||||
|
||||
Completes 8.3.1's delete-hygiene story at the counter layer, from a production proof chain reported
|
||||
by a downstream deployment: persisted entity totals were permanently **inflated** — deletes whose
|
||||
count decrement was silently skipped — and because paginated `totalCount` serves
|
||||
`Math.max(persistedTotal, scanned)`, the inflated number always won and **no disk cleanup could ever
|
||||
lower it**.
|
||||
|
||||
- **A removal's count decrement no longer requires re-reading the record being removed.** The
|
||||
decrement was sourced from re-reading the entity's metadata inside the delete; if that read
|
||||
returned `null` (a replace race, or a ghost left by a pre-8.3.1 partial delete) the decrement was
|
||||
silently skipped while the paired add had counted — minting drift on every write→delete→re-create
|
||||
cycle. The caller's pre-delete read now rides through the whole delete path
|
||||
(`remove()`/`removeMany()` → the delete operation → `deleteNoun`/`deleteVerb` →
|
||||
`deleteNounMetadata`/`deleteVerbMetadata`, both sides symmetric): a null internal read falls back
|
||||
to the known prior record instead of skipping. The `StorageAdapter` signatures gain an optional
|
||||
`priorMetadata` parameter (additive; existing adapters unaffected).
|
||||
|
||||
- **`repairIndex()` is the sanctioned counter recount — unconditional, and it actually persists.**
|
||||
`rebuildTypeCounts()` previously rebuilt only the type-statistics arrays and computed the total
|
||||
*just to log it* — the persisted scalar (`counts.json`) survived every "rebuild" untouched, so an
|
||||
already-inflated brain could never be corrected. One canonical walk now rebuilds **every** counter
|
||||
rollup — scalar totals, per-type maps, and type statistics — and persists them, and `repairIndex()`
|
||||
runs it unconditionally (not only when orphan directories are found: counters can be inflated over
|
||||
perfectly clean shelves). Brains with delete history should run `brain.repairIndex()` once after
|
||||
upgrading; the correction survives reopen.
|
||||
|
||||
No API breaks (optional-parameter additions only). Regression tests cover the drift cycle, the
|
||||
null-read decrement fallback, and the persisted recount across reopen.
|
||||
|
||||
## v8.3.1 — 2026-07-14 (full-removal deletes + family-scoped migration gate)
|
||||
|
||||
Two production-reported fixes in the write/index spine, plus an operator repair path. No API changes;
|
||||
|
|
|
|||
|
|
@ -1,117 +0,0 @@
|
|||
---
|
||||
title: The Generation Fact Log
|
||||
slug: concepts/generation-fact-log
|
||||
public: true
|
||||
category: concepts
|
||||
template: concept
|
||||
order: 6
|
||||
description: Every committed write also appends a self-verifying "fact" — an after-image commit record — to an append-only log. What facts are, the crash-safety model, the scanFacts() streaming surface, family stamps, and how index providers consume the log for sequential heals.
|
||||
next:
|
||||
- concepts/consistency-model
|
||||
- guides/snapshots-and-time-travel
|
||||
---
|
||||
|
||||
# The Generation Fact Log
|
||||
|
||||
Since 8.4.0, every committed generation also appends a **fact** — a compact record of what each
|
||||
touched entity or relationship *became* — to an append-only, checksummed log under
|
||||
`_generations/facts/`. Where the generational history answers *"what did things look like
|
||||
before?"* (before-images, powering `asOf()` and rollback), the fact log answers *"what happened,
|
||||
in order?"* — one sequential, self-verifying stream of the store's present being written.
|
||||
|
||||
Nothing about querying changes. The fact log exists for three consumers:
|
||||
|
||||
1. **Index heals and rebuilds** — one sequential read in commit order replaces a per-entity
|
||||
directory walk over millions of files.
|
||||
2. **Incremental catch-up** — a derived index that knows which generation it reflects reads *just
|
||||
the gap*, instead of rebuilding from scratch.
|
||||
3. **Replay and audit tooling** — anything that wants the store's committed timeline as a stream.
|
||||
|
||||
## What a fact is
|
||||
|
||||
One fact per committed generation:
|
||||
|
||||
- **`generation`** and **`timestamp`** — which commit, when.
|
||||
- **`ops`** — every write in that commit: `{ kind: 'noun' | 'verb', id, record }` where `record`
|
||||
holds the entity's full after-image (both stored legs), or **`null` for a tombstone** — a
|
||||
removal carries no body, by design.
|
||||
- **`meta`** — the transaction metadata `transact()` was submitted with, when present.
|
||||
- **`blobHashes`** — content-blob references, for exact reclamation accounting.
|
||||
|
||||
Facts accumulate **from the first write after upgrading** — pre-existing history is not
|
||||
retroactively converted, and consumers fall back to the enumeration walk when no log exists.
|
||||
|
||||
## Crash safety, in one paragraph
|
||||
|
||||
Facts are appended and fsynced **inside the same durability window as the commit itself**, before
|
||||
the commit point — so after a crash, the log can only ever be *ahead* of committed truth, never
|
||||
behind it with a hole. On open, the store reconciles the log back to the committed watermark:
|
||||
torn tails are detected by per-record checksums and cut; whole records beyond the watermark are
|
||||
truncated. The invariant every reader can rely on: **an absent generation was never committed; a
|
||||
present fact was.** `transact()` facts are durable the moment `transact()` returns; single-op
|
||||
facts share the same group-commit flush as the rest of their generation, so a hard kill loses the
|
||||
fact and the generation *together* — never a torn state.
|
||||
|
||||
## Reading the log
|
||||
|
||||
```typescript
|
||||
const scan = brain.scanFacts({ fromGeneration: 1 })
|
||||
if (scan) {
|
||||
// Telemetry up front — progress bars get a denominator from second zero.
|
||||
console.log(scan.headGeneration, scan.segmentCount, scan.approxFactCount)
|
||||
|
||||
for await (const batch of scan.batches()) {
|
||||
// Each batch: { facts, firstGeneration, lastGeneration, factCount, byteSize, segmentId }
|
||||
for (const fact of batch.facts) {
|
||||
for (const op of fact.ops) {
|
||||
if (op.record === null) {
|
||||
// a tombstone: op.id was removed in this generation
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
console.log(scan.summary()) // { factsYielded, segmentsRead } — the cross-check
|
||||
}
|
||||
```
|
||||
|
||||
- `scanFacts()` returns `null` when the store hosts no fact log (older store, or a storage adapter
|
||||
without binary append support) — fall back to enumerating entities.
|
||||
- Scans run against a **snapshot**: facts appended after the scan opens never bleed in, each fact
|
||||
is yielded exactly once, and a detected gap aborts loudly — never a silent skip.
|
||||
- `brain.factSegmentPaths()` returns the immutable, *sealed* segment files for zero-copy consumers
|
||||
(the append-mutable tail is excluded — read it through `scanFacts()`).
|
||||
|
||||
## Family stamps: how a projection proves it's current
|
||||
|
||||
Anything derived from the store — an index, the entity file tree itself — carries a **family
|
||||
stamp**: a small JSON record of *which committed generation the projection reflects*
|
||||
(`sourceGeneration`) plus the invariants that verify it whole (exact per-file byte sizes for
|
||||
bounded families; rollup invariants like entity counts for unbounded ones). At open, coherence is
|
||||
a **comparison**, not a walk:
|
||||
|
||||
- stamp equals the committed watermark and invariants hold → serve;
|
||||
- stamp is behind → the projection reads just the gap from the fact log;
|
||||
- invariants fail → loud, named divergence — `brain.repairIndex()` rebuilds from canonical and
|
||||
re-stamps.
|
||||
|
||||
The verifier is exported (`verifyFamilyStamp`) so every projection — TypeScript or native — runs
|
||||
literally the same check.
|
||||
|
||||
## For plugin authors: the storage capability
|
||||
|
||||
Index providers receive the storage adapter, not the brain — so the host wires the log onto it.
|
||||
Feature-detect and prefer the stream; fall back to enumeration:
|
||||
|
||||
```typescript
|
||||
const scan = storage.scanFacts?.({ fromGeneration: stamp.sourceGeneration + 1 })
|
||||
if (scan) {
|
||||
// sequential catch-up from the log
|
||||
} else {
|
||||
// enumeration walk (older store or adapter)
|
||||
}
|
||||
const committed = storage.committedGeneration?.() // the watermark stamps compare against
|
||||
```
|
||||
|
||||
Providers must never construct their own reader over the log's files — the open path belongs to
|
||||
the single writer (it reconciles the log at open); the capability is the sanctioned seam.
|
||||
|
|
@ -11,13 +11,7 @@ No batch jobs. No scheduled recalculations. Aggregates stay current with every w
|
|||
**Defining over existing data:** if you define an aggregate on a store that already holds
|
||||
matching entities, Brainy backfills it from those entities on the first query (a one-time scan,
|
||||
then purely incremental). So `defineAggregate()` behaves the same whether you define it before
|
||||
or after the data exists.
|
||||
|
||||
**Reopening a persisted brain:** aggregate state persists across restarts. Re-defining the
|
||||
same aggregate at boot (the normal declarative pattern) adopts the persisted state directly —
|
||||
no rescan. A backfill scan runs only when the definition actually changed, when no persisted
|
||||
state exists, or when the state failed to load; and however many aggregates need backfilling,
|
||||
they share a single scan.
|
||||
or after the data exists — including when a persisted brain reopens already populated.
|
||||
|
||||
## Quick Start
|
||||
|
||||
|
|
|
|||
|
|
@ -1,99 +0,0 @@
|
|||
---
|
||||
title: External Backups & Sparse Storage
|
||||
slug: guides/external-backups
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 10
|
||||
description: How to back up a brain directory with external tools (tar, rsync, cp) without exploding sparse files — why a store can show 100+ GB "apparent" size on a small disk, which files are sparse, and how persist()/restore() handle it for you.
|
||||
next:
|
||||
- guides/snapshots-and-time-travel
|
||||
- concepts/storage-adapters
|
||||
---
|
||||
|
||||
# External Backups & Sparse Storage
|
||||
|
||||
The built-in snapshot path — [`db.persist()` and `brain.restore()`](/docs/guides/snapshots-and-time-travel) —
|
||||
already handles everything on this page for you. Read this when you back up a brain directory with
|
||||
**external tools**: `tar`, `rsync`, `cp`, `scp`, or a filesystem-level backup agent.
|
||||
|
||||
## The one-sentence rule
|
||||
|
||||
> **Always use the sparse-aware flag**: `tar czSf` (capital `S`), `rsync --sparse`,
|
||||
> `cp --sparse=always`. A naive copy can turn a 2 GB store into a 100+ GB one — or fail
|
||||
> the disk entirely.
|
||||
|
||||
## Why: some files are sparse
|
||||
|
||||
When a native accelerator plugin is active, parts of the index live in **memory-mapped files**
|
||||
created at a large fixed virtual size — the file's *apparent* size — while the filesystem only
|
||||
allocates blocks that were actually written. A brand-new id-mapper file can report tens of
|
||||
gigabytes in `ls -l` while occupying a few megabytes on disk.
|
||||
|
||||
Check the difference yourself:
|
||||
|
||||
```bash
|
||||
ls -lh brain-data/_id_mapper/ # APPARENT size (can be huge)
|
||||
du -sh brain-data/ # ALLOCATED size (the real footprint)
|
||||
```
|
||||
|
||||
The sparse candidates in a brain directory:
|
||||
|
||||
| Path | What it is |
|
||||
|---|---|
|
||||
| `_id_mapper/` | The native id-mapper's mmap files (large fixed virtual size) |
|
||||
| `_blobs/` | Native index files (vector base, segments) — may be mmap-backed |
|
||||
|
||||
Everything else (entities, `_system`, `_generations`, `_cas` content blobs) is ordinary dense data.
|
||||
|
||||
## Doing it right
|
||||
|
||||
**tar** — the `S` flag detects holes and stores only real data:
|
||||
|
||||
```bash
|
||||
tar czSf brain-backup.tgz /data/brain
|
||||
# restore preserves the holes:
|
||||
tar xzSf brain-backup.tgz -C /data/
|
||||
```
|
||||
|
||||
**rsync**:
|
||||
|
||||
```bash
|
||||
rsync -a --sparse /data/brain/ backup-host:/backups/brain/
|
||||
```
|
||||
|
||||
**cp**:
|
||||
|
||||
```bash
|
||||
cp -a --sparse=always /data/brain /backups/brain
|
||||
```
|
||||
|
||||
**What goes wrong without the flag:** the copy *materializes* every hole as real zero bytes.
|
||||
A store whose apparent size exceeds the target disk fails with `ENOSPC` partway through — and a
|
||||
copy that *does* fit silently costs the full apparent size in storage and transfer time.
|
||||
|
||||
## What the built-in paths do (so you don't have to)
|
||||
|
||||
- **`db.persist(path)`** snapshots via **hard links** — instant and space-shared, since every data
|
||||
file is immutable-by-rename. The handful of append-in-place files (the transaction log, the
|
||||
commit fact log's tail segment) and mmap-mutated directories (`_id_mapper/`) are **byte-copied**
|
||||
instead, so a post-snapshot write can never reach through a shared inode into your backup.
|
||||
- **`brain.restore(path, { confirm: true })`** is **non-destructive and sparse-aware**: the snapshot
|
||||
is copied into a staging area *before* any live data is touched (all-zero blocks stay holes), and
|
||||
only after the copy fully succeeds does an atomic swap move it into place. A failed copy —
|
||||
including `ENOSPC` — leaves the live store exactly as it was. A crash mid-swap completes forward
|
||||
on the next open.
|
||||
|
||||
## Live-store caveats for external tools
|
||||
|
||||
1. **Prefer snapshotting a `persist()` output, not the live directory.** `persist()` produces a
|
||||
crash-consistent, immutable snapshot; running `tar` against a live, actively-written directory
|
||||
can capture a torn mid-write state. If you must archive live, stop writes first (or accept that
|
||||
the archive is only as consistent as the moment's flush state).
|
||||
2. **Never prune or "clean up" files inside a brain directory.** Index files that look stale or
|
||||
redundant are load-bearing; the store protects its declared index families from in-process
|
||||
deletion, but an external `rm` bypasses that fence. If space is the concern, `du -sh` first —
|
||||
the allocated size is usually far smaller than it looks.
|
||||
3. **Verify restores by opening them.** `Brainy.load(path)` opens any snapshot or restored
|
||||
directory read-only — the store verifies its own coherence at open and reports loudly if
|
||||
anything is missing or torn.
|
||||
|
|
@ -40,10 +40,6 @@ Brainy picks `maxLimit` from the first of these that's available:
|
|||
|
||||
Worked example: a 4 GB Cloud Run container picks priority 3 → `floor(4 GB × 0.25 / 25 KB) = floor(40 960) = 40 000` results. A 900 MB free-memory box on priority 4 gets `floor(900 MB / 25 KB) = ~36 000`.
|
||||
|
||||
The cap is fixed at construction and never changes at runtime. Query timing is recorded
|
||||
for diagnostics only — a burst of slow queries cannot silently shrink the cap, and the
|
||||
auto-detected tiers (3 and 4) never go below a floor of 10 000.
|
||||
|
||||
> **Calibration note.** Pre-7.30.2 used 100 KB per result instead of 25 KB, which produced caps that were 4× too tight for typical workloads (an 8 KB / result reality). 7.30.2 recalibrated to match observed entity sizes; existing `limit: 10_000` safety patterns now pass silently on any reasonably-sized box.
|
||||
|
||||
## What happens when you exceed the cap
|
||||
|
|
|
|||
|
|
@ -300,10 +300,7 @@ await brain.import(data, {
|
|||
// Deduplication
|
||||
enableDeduplication: true, // Check for duplicate entities (default: true)
|
||||
deduplicationThreshold: 0.85, // Similarity threshold for duplicates (0-1, default: 0.85)
|
||||
// Notes: false disables BOTH the inline merge and the background pass that
|
||||
// runs ~5 min after the last import (merged duplicates are deleted).
|
||||
// The inline pass auto-disables for imports >100 entities (O(n²) cost);
|
||||
// the background pass still covers those unless the flag is false.
|
||||
// Note: Auto-disabled for imports >100 entities
|
||||
|
||||
// Performance
|
||||
chunkSize: 100, // Batch size for processing (default: varies by operation)
|
||||
|
|
|
|||
|
|
@ -86,22 +86,11 @@ await brain.import(file, {
|
|||
|
||||
```typescript
|
||||
await brain.import(file, {
|
||||
enableDeduplication: true, // Check for duplicates (default: true)
|
||||
enableDeduplication: true, // Check for duplicates (default: false)
|
||||
deduplicationThreshold: 0.85 // Similarity threshold (default: 0.85)
|
||||
})
|
||||
```
|
||||
|
||||
Deduplication merges entities judged duplicates — the non-primary records are
|
||||
**deleted**. Set `enableDeduplication: false` to disable it entirely: the flag
|
||||
gates both the inline merge during import and the background pass that runs
|
||||
about 5 minutes after the last import.
|
||||
|
||||
```typescript
|
||||
await brain.import(file, {
|
||||
enableDeduplication: false // No merging, inline or background
|
||||
})
|
||||
```
|
||||
|
||||
### Import Tracking
|
||||
|
||||
Track and organize imports by project:
|
||||
|
|
|
|||
|
|
@ -166,34 +166,6 @@ brainy inspect diff /data/brain-prod /data/brain-staging
|
|||
Sample-based — for a full diff, dump both with `inspect dump` and compare
|
||||
the JSONL.
|
||||
|
||||
## Auditing graph-read truth
|
||||
|
||||
`brain.auditGraph()` (8.6.0+) proves — or disproves — that relationship reads
|
||||
return canonical truth on a given brain, without mutating anything. It walks
|
||||
every stored relationship record, asks the same read path your application
|
||||
uses (`related()`, VFS `readdir`) with every visibility tier included, and
|
||||
classifies every discrepancy:
|
||||
|
||||
```typescript
|
||||
const report = await brain.auditGraph()
|
||||
|
||||
report.coherent // true = related()/readdir can be trusted on this brain
|
||||
report.missingFromReadsCount // records the read path omits — stale index
|
||||
report.danglingEndpointsCount // relationships whose endpoint entity is gone
|
||||
report.readOnlyCount // read-path edges with NO stored record — ghosts
|
||||
report.visibilityHiddenCount // internal/system edges hidden by design (not a fault)
|
||||
```
|
||||
|
||||
Counts are always exact; the example lists (`missingFromReads`,
|
||||
`danglingEndpoints`, `readOnlyVerbIds`) are capped at `maxExamples`
|
||||
(default 100) and `truncatedExamples` says so when they are.
|
||||
|
||||
Run it after any engine upgrade, restore, or migration. If it reports
|
||||
discrepancies, run `brain.repairIndex()` and audit again — a `coherent`
|
||||
report after the repair is the verified statement that the heal worked.
|
||||
Cost: one relationship-record walk plus one indexed read per distinct
|
||||
source entity — safe on a live brain.
|
||||
|
||||
## Repairing a corrupted store
|
||||
|
||||
If invariants fail and you suspect index corruption, `inspect repair`
|
||||
|
|
|
|||
|
|
@ -180,35 +180,3 @@ Brainy 8.0 has exactly two write-coordination counters, at two granularities:
|
|||
They compose: a `transact()` batch can carry per-entity `ifRev` checks *and* a whole-store `ifAtGeneration`; any failed check rejects the entire batch before anything is staged. Generations also power snapshots and time travel (`brain.now()`, `brain.asOf()`, `db.persist()`) — see the [consistency model](../concepts/consistency-model.md) and [Snapshots & Time Travel](./snapshots-and-time-travel.md).
|
||||
|
||||
A snapshot or historical view captures each entity *including* its `_rev` at that moment, so reading the past and writing back with `ifRev` against the live state works exactly as you'd hope: the write fails if the entity moved since the state you copied from.
|
||||
|
||||
## The transact envelope: batch size, budget, and bulk imports
|
||||
|
||||
`transact()` applies its batch atomically under one commit — which means the whole batch
|
||||
shares one **apply budget**. Since 8.7.0 the budget scales with the batch:
|
||||
`max(30 s, opCount × 2 s)`, or exactly what you pass as `timeoutMs`. A tripped budget rolls
|
||||
the entire batch back (nothing partial survives) and throws a retryable
|
||||
`TransactionTimeoutError` that names the operation it stopped at, the batch size, and the
|
||||
elapsed vs budgeted time — a diagnosis, not just a failure:
|
||||
|
||||
```
|
||||
Transaction timed out at operation 41/120 ('add') — 246012ms elapsed, budget 240000ms.
|
||||
The batch rolled back atomically; retry with a higher timeoutMs or a smaller batch.
|
||||
```
|
||||
|
||||
Practical envelope guidance for bulk work:
|
||||
|
||||
1. **Precompute embeddings outside the commit path.** Embedding inside `transact()` spends
|
||||
the budget on model inference. Use `brain.embedBatch(texts)` and pass each vector via
|
||||
the op's `vector` field — the commit then pays only storage costs, and a retried batch
|
||||
never re-pays inference. (The win is *where* the inference happens, not raw embedding
|
||||
throughput: on the default WASM engine, batch and sequential embedding measure
|
||||
comparably, ~160 ms/text; native embedding providers may batch faster.)
|
||||
2. **Chunk very large imports** into batches of a few hundred ops with one `transact()`
|
||||
each. You lose whole-import atomicity but keep per-chunk atomicity, bounded memory, and
|
||||
resumability — pair with `ifAbsent` upserts so a retried chunk is idempotent.
|
||||
3. **Slow disks change the math, not the contract.** On network-attached storage a single
|
||||
op can cost ~2 s (canonical write + fsync + index maintenance). The scaled default
|
||||
absorbs that; pass an explicit `timeoutMs` only when you know better than the scale.
|
||||
4. **`addMany`/`relateMany` are the convenience tier** — they chunk and batch-embed for
|
||||
you, with per-item error reporting instead of batch atomicity. Choose by what you need:
|
||||
atomic-all-or-nothing → `transact()`; resilient bulk load → `addMany`.
|
||||
|
|
|
|||
|
|
@ -9,7 +9,6 @@ description: Recipes for the Db API — instant backups with persist(), restore,
|
|||
next:
|
||||
- concepts/consistency-model
|
||||
- guides/optimistic-concurrency
|
||||
- guides/external-backups
|
||||
---
|
||||
|
||||
# Snapshots & Time Travel
|
||||
|
|
@ -42,10 +41,6 @@ bytes. Cross-device targets fall back to per-file byte copies, and
|
|||
persisting an in-memory brain serializes it to the same directory layout —
|
||||
a real, durable store.
|
||||
|
||||
> Archiving a brain directory with **external tools** (`tar`, `rsync`, `cp`)?
|
||||
> Some index files are sparse and can explode to their apparent size under a
|
||||
> naive copy — see [External Backups & Sparse Storage](/docs/guides/external-backups).
|
||||
|
||||
Two things to know:
|
||||
|
||||
- `persist()` requires the view to still be the store's **latest**
|
||||
|
|
@ -344,12 +339,8 @@ For per-entity write coordination (rather than whole-store history), the
|
|||
## Keeping history bounded
|
||||
|
||||
Under Model-B every write is a generation, so history can grow quickly —
|
||||
Brainy auto-compacts at `close()` (time-bounded per pass) under the
|
||||
**`retention`** knob (configured on the constructor). Since 8.9.0, `flush()`
|
||||
never compacts: flushing is durability work and costs only what the current
|
||||
window's writes cost, regardless of history backlog. A long-lived writer that
|
||||
never closes keeps its history until its next explicit `compactHistory()` —
|
||||
schedule one in your maintenance window if you run bounded retention:
|
||||
Brainy auto-compacts on every `flush()`/`close()` under the **`retention`**
|
||||
knob (configured on the constructor):
|
||||
|
||||
```typescript
|
||||
// Zero-config: ADAPTIVE — keep as much history as free disk/RAM allows,
|
||||
|
|
@ -363,13 +354,10 @@ new Brainy({ retention: 'all' })
|
|||
new Brainy({ retention: { maxGenerations: 1000, maxAge: 7 * 86_400_000, maxBytes: 512 * 1024 ** 2 } })
|
||||
```
|
||||
|
||||
Reclaim manually at any time (the same caps, plus an optional per-pass time
|
||||
budget for maintenance windows — an early stop is a consistent prefix and the
|
||||
next pass resumes):
|
||||
Reclaim manually at any time (the same caps):
|
||||
|
||||
```typescript
|
||||
await brain.compactHistory({ maxGenerations: 100, maxAge: 7 * 24 * 60 * 60 * 1000 })
|
||||
await brain.compactHistory({ maxBytes: 512 * 1024 ** 2, timeBudgetMs: 10_000 })
|
||||
```
|
||||
|
||||
Compaction never breaks a pinned read — record-sets are reclaimed only when
|
||||
|
|
|
|||
|
|
@ -1,83 +0,0 @@
|
|||
---
|
||||
title: Performance Envelopes
|
||||
slug: guides/performance-envelopes
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 40
|
||||
description: Measured per-operation latency envelopes at stated scales — what to expect, on what hardware, and exactly how each number was produced.
|
||||
next:
|
||||
- guides/find-limits
|
||||
---
|
||||
|
||||
# Performance Envelopes
|
||||
|
||||
Every number on this page is **measured, never projected** — produced by the script
|
||||
cited at the bottom, against the built package (the artifact you install), on the stated
|
||||
hardware. Each entry says what was measured, at what scale, on which storage backend.
|
||||
When a release touches a measured path, that operation is re-measured and this page
|
||||
updates in the same release.
|
||||
|
||||
Two scopes to keep straight:
|
||||
|
||||
- **These envelopes are the pure-JS engine** (no native accelerator registered) on
|
||||
filesystem storage. This is the floor every deployment gets from `npm install` alone.
|
||||
- **Accelerated deployments** (the optional native provider) publish their own numbers —
|
||||
this page never claims them.
|
||||
|
||||
## Read operations
|
||||
|
||||
Reads are where the architecture pays off: after the write path has done its indexing
|
||||
work, queries answer from purpose-built indexes without scanning.
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `get(id)` (warm) | p50 < 0.1ms | p50 < 0.1ms | served from cache/metadata index |
|
||||
| `find` (metadata: indexed equality + range, limit 100) | p50 1.0ms · p95 1.8ms | p50 7.0ms · p95 8.9ms | column-store bitmap paths |
|
||||
| `related(id)` (per-node adjacency) | p50 < 0.1ms · p95 0.2ms | p50 < 0.1ms | LSM adjacency index — O(degree), scale-independent |
|
||||
| `find` (semantic: embed + HNSW, 1k docs) | p50 178ms · p95 393ms | — | dominated by WASM query embedding (measured on a machine under concurrent load — treat the p95 as an upper bound); the vector search itself is single-digit ms |
|
||||
|
||||
## Write operations
|
||||
|
||||
Under Model-B **every write is its own durable generation** — a single-op `add` pays
|
||||
serialization, before-image staging, and fsync before it acks. That durability is priced
|
||||
into the write path visibly, by design:
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `add` (single-op) | p50 167ms · p95 171ms | p50 165ms · p95 172ms | full durable generation per write — flat across scale |
|
||||
| `addMany` (bulk) | ~163ms/entity | ~187ms/entity | **currently per-item commits** — see the honest note below |
|
||||
| `relateMany` | ~0.8ms/edge | ~0.9ms/edge | edges batch efficiently today |
|
||||
| `flush` (steady-state, 1 pending write) | p50 8ms · p95 10ms | p50 45ms · p95 52ms | durability-only since 8.9.0 — cost no longer depends on history backlog or retention mode |
|
||||
|
||||
**The honest note on bulk writes:** `addMany` today commits each item as its own
|
||||
generation (the same durability as single-op `add`, serialized by the single-writer
|
||||
lock), so bulk-load cost is N × single-op cost. Batched chunk commits (one generation
|
||||
and one fsync window per chunk, as `removeMany` already does) are designed into the
|
||||
unified-commit work on the current roadmap. Until that ships, size bulk imports
|
||||
accordingly — 10k entities is minutes, not seconds, on filesystem storage.
|
||||
|
||||
## Open / close
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `open` (empty store) | ~560ms | ~190ms | includes embedder initialization |
|
||||
| `open` (warm, populated, clean shutdown) | 763ms | 4.9s | pure-JS vector index load dominates and grows with entity count; the native accelerator exists precisely to remove this |
|
||||
| `close` | bounded | bounded | auto-compaction pass is time-bounded (~5s max) since 8.9.0 |
|
||||
|
||||
A store that was NOT cleanly closed pays index rebuilds on top of the warm-open
|
||||
number (tens of seconds at 10k) — clean shutdown is worth engineering for.
|
||||
|
||||
## How these were produced
|
||||
|
||||
- **Hardware**: Intel Core i9-14900HX (32 threads), 62GB RAM, NVMe, Linux, Node v22.
|
||||
- **Backend**: `storage: { type: 'filesystem' }`, pure JS (no native providers).
|
||||
- **Embeddings**: deterministic stub for non-semantic ops (isolates engine cost);
|
||||
the real WASM embedder for the semantic row (that's what you'll run).
|
||||
- **Method**: p50/p95 over 50–200 samples per op against the built `dist/`;
|
||||
the measuring script ships in the repo history and re-runs per release.
|
||||
|
||||
Numbers on different hardware will differ; the *shape* (sub-2ms indexed reads,
|
||||
~160ms embedding-bound semantic queries, durability-priced writes) is the envelope
|
||||
you should hold your deployment against. If your measurements diverge from these
|
||||
shapes by an order of magnitude, something is wrong — file it.
|
||||
4
package-lock.json
generated
4
package-lock.json
generated
|
|
@ -1,12 +1,12 @@
|
|||
{
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "8.9.0",
|
||||
"version": "8.3.1",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "8.9.0",
|
||||
"version": "8.3.1",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@msgpack/msgpack": "^3.1.2",
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
{
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "8.9.0",
|
||||
"version": "8.3.1",
|
||||
"description": "Universal Knowledge Protocol™ - World's first Triple Intelligence database unifying vector, graph, and document search in one API. Stage 3 CANONICAL: 42 nouns × 127 verbs covering 96-97% of all human knowledge.",
|
||||
"main": "dist/index.js",
|
||||
"module": "dist/index.js",
|
||||
|
|
|
|||
|
|
@ -1,116 +0,0 @@
|
|||
#!/usr/bin/env node
|
||||
/**
|
||||
* @module scripts/push-docs
|
||||
* @description Push this repo's PUBLIC docs to the soulcraft.com docs ingest
|
||||
* door after an npm publish (VENUE-DOCS-RELEASE-PUSH — retires the old
|
||||
* build-time docs sync).
|
||||
*
|
||||
* Contract (mirrors the reference implementation on the serving side):
|
||||
* POST {base}/api/docs/ingest
|
||||
* headers: x-service-secret: $DOCS_INGEST_SECRET, Content-Type: application/json
|
||||
* body: { docs: [{ slug, title, markdown, nav: { order, section } }] }
|
||||
* batches of 10, idempotent per slug.
|
||||
*
|
||||
* A doc is public iff its frontmatter has `public: true` AND a `slug`. The
|
||||
* frontmatter is stripped; `category` → nav.section, `order` → nav.order.
|
||||
*
|
||||
* Deliberately NOT pushed: the combined /docs landing index. It spans BOTH
|
||||
* engine corpora (this repo's and the native accelerator's), so a per-repo
|
||||
* push would clobber the union — the index is authored on the serving side.
|
||||
*
|
||||
* Env: DOCS_INGEST_SECRET (required), DOCS_INGEST_BASE (default
|
||||
* https://soulcraft.com). Exits 0 with a LOUD warning when the secret is
|
||||
* absent (the npm publish has already happened; the serving side runs its
|
||||
* interim sync on request) and exits 1 when a push actually fails — the docs
|
||||
* site would silently trail npm otherwise, and that must be visible.
|
||||
*/
|
||||
import * as fs from 'node:fs'
|
||||
import * as path from 'node:path'
|
||||
|
||||
const BASE = (process.env.DOCS_INGEST_BASE || 'https://soulcraft.com').replace(/\/+$/, '')
|
||||
const SECRET = process.env.DOCS_INGEST_SECRET
|
||||
const DOCS_DIR = path.join(path.dirname(new URL(import.meta.url).pathname), '..', 'docs')
|
||||
const BATCH = 10
|
||||
|
||||
if (!SECRET) {
|
||||
console.warn(
|
||||
'⚠️ DOCS PUSH SKIPPED: DOCS_INGEST_SECRET is not set.\n' +
|
||||
' soulcraft.com/docs now TRAILS this npm release until docs are pushed.\n' +
|
||||
' Either export DOCS_INGEST_SECRET and re-run `node scripts/push-docs.js`,\n' +
|
||||
' or ping venue on VENUE-DOCS-RELEASE-PUSH for the interim sync.'
|
||||
)
|
||||
process.exit(0)
|
||||
}
|
||||
|
||||
/** Minimal frontmatter split — returns [meta, body] or [null, raw]. */
|
||||
function parseFrontmatter(raw) {
|
||||
const m = raw.match(/^---\n([\s\S]*?)\n---\n([\s\S]*)$/)
|
||||
if (!m) return [null, raw]
|
||||
const meta = {}
|
||||
for (const line of m[1].split('\n')) {
|
||||
const kv = line.match(/^(\w[\w-]*):\s*(.*)$/)
|
||||
if (kv) meta[kv[1]] = kv[2].trim().replace(/^["']|["']$/g, '')
|
||||
}
|
||||
return [meta, m[2]]
|
||||
}
|
||||
|
||||
const docs = []
|
||||
;(function walk(dir) {
|
||||
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
|
||||
const full = path.join(dir, entry.name)
|
||||
if (entry.isDirectory()) walk(full)
|
||||
else if (entry.name.endsWith('.md')) {
|
||||
const [meta, body] = parseFrontmatter(fs.readFileSync(full, 'utf-8'))
|
||||
if (!meta || meta.public !== 'true' || !meta.slug) continue
|
||||
docs.push({
|
||||
slug: meta.slug,
|
||||
title: meta.title || meta.slug,
|
||||
markdown: body.trim(),
|
||||
nav: {
|
||||
order: Number.parseInt(meta.order || '99', 10) || 99,
|
||||
section: meta.category || 'guides'
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
})(DOCS_DIR)
|
||||
|
||||
if (docs.length === 0) {
|
||||
console.error('❌ DOCS PUSH FAILED: zero public docs collected — refusing to push an empty corpus.')
|
||||
process.exit(1)
|
||||
}
|
||||
docs.sort((a, b) => a.slug.localeCompare(b.slug))
|
||||
console.log(`Pushing ${docs.length} public docs to ${BASE}/api/docs/ingest …`)
|
||||
|
||||
let failed = false
|
||||
for (let i = 0; i < docs.length; i += BATCH) {
|
||||
const batch = docs.slice(i, i + BATCH)
|
||||
try {
|
||||
const res = await fetch(`${BASE}/api/docs/ingest`, {
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'x-service-secret': SECRET,
|
||||
'Content-Type': 'application/json',
|
||||
'User-Agent': 'brainy-docs-push/1.0'
|
||||
},
|
||||
body: JSON.stringify({ docs: batch }),
|
||||
signal: AbortSignal.timeout(120_000)
|
||||
})
|
||||
if (!res.ok) {
|
||||
throw new Error(`HTTP ${res.status}: ${(await res.text()).slice(0, 300)}`)
|
||||
}
|
||||
console.log(` batch ${i / BATCH + 1}: ${batch.map((d) => d.slug).join(', ')} → ok`)
|
||||
} catch (err) {
|
||||
failed = true
|
||||
console.error(` batch ${i / BATCH + 1} FAILED: ${err instanceof Error ? err.message : err}`)
|
||||
}
|
||||
}
|
||||
|
||||
if (failed) {
|
||||
console.error(
|
||||
'❌ DOCS PUSH INCOMPLETE — soulcraft.com/docs may trail npm. ' +
|
||||
'Re-run `node scripts/push-docs.js` or ping venue on VENUE-DOCS-RELEASE-PUSH.'
|
||||
)
|
||||
process.exit(1)
|
||||
}
|
||||
console.log('✅ Docs pushed.')
|
||||
|
|
@ -196,18 +196,6 @@ else
|
|||
fi
|
||||
echo -e "${GREEN}✅ GitHub release created${NC}\n"
|
||||
|
||||
# Step 12: Push public docs to the soulcraft.com docs ingest door
|
||||
# (VENUE-DOCS-RELEASE-PUSH). Skips with a loud warning when
|
||||
# DOCS_INGEST_SECRET is unset; fails loudly (without undoing the publish —
|
||||
# that already happened) when a push errors, so the docs site never
|
||||
# silently trails npm.
|
||||
echo -e "${BLUE}1️⃣2️⃣ Pushing public docs to soulcraft.com/docs...${NC}"
|
||||
if node scripts/push-docs.js; then
|
||||
echo -e "${GREEN}✅ Docs push step done${NC}\n"
|
||||
else
|
||||
echo -e "${RED}❌ Docs push FAILED — soulcraft.com/docs trails npm until re-run or interim sync${NC}\n"
|
||||
fi
|
||||
|
||||
echo -e "${GREEN}━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━${NC}"
|
||||
echo -e "${GREEN}🎉 Release ${NEW_VERSION} complete!${NC}"
|
||||
echo -e "${GREEN}━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━${NC}"
|
||||
|
|
|
|||
|
|
@ -29,7 +29,6 @@ import { matchesMetadataFilter } from '../utils/metadataFilter.js'
|
|||
import { compareCodePoints } from '../utils/collation.js'
|
||||
import { bucketTimestamp } from './timeWindows.js'
|
||||
import { NounType } from '../types/graphTypes.js'
|
||||
import { prodLog } from '../utils/logger.js'
|
||||
|
||||
/** Persistence key for aggregate definitions */
|
||||
const DEFINITIONS_KEY = '__aggregation_definitions__'
|
||||
|
|
@ -88,18 +87,10 @@ function matchesSource(entity: Record<string, unknown>, source: AggregateDefinit
|
|||
if (entity.service !== source.service) return false
|
||||
}
|
||||
|
||||
// Where filter — resolve each filtered field through resolveEntityField,
|
||||
// the SAME single source of truth groupBy uses (top-level standard fields
|
||||
// + custom metadata). Matching only the metadata sub-object made
|
||||
// where:{subtype}/{visibility}/… a silent no-op: reserved fields never
|
||||
// live in the custom bag, so those filters could never match anything.
|
||||
// Metadata where filter — match against the entity's metadata sub-object
|
||||
if (source.where && Object.keys(source.where).length > 0) {
|
||||
const e = entity as unknown as HNSWNounWithMetadata
|
||||
const resolved: Record<string, unknown> = {}
|
||||
for (const key of Object.keys(source.where)) {
|
||||
resolved[key] = resolveEntityField(e, key)
|
||||
}
|
||||
if (!matchesMetadataFilter(resolved, source.where)) return false
|
||||
const metadata = (entity.metadata ?? entity) as Record<string, unknown>
|
||||
if (!matchesMetadataFilter(metadata, source.where)) return false
|
||||
}
|
||||
|
||||
return true
|
||||
|
|
@ -336,30 +327,6 @@ export class AggregationIndex {
|
|||
/** Track aggregates with stale MIN/MAX (need lazy recompute) */
|
||||
private staleMinMax = new Map<string, Set<string>>()
|
||||
|
||||
/** Resolves when init() has finished loading persisted definitions/state. */
|
||||
private initPromise: Promise<void> | null = null
|
||||
|
||||
/** True once init() has settled (success or failure). */
|
||||
private initDone = false
|
||||
|
||||
/**
|
||||
* Aggregates registered by the app before init() finished loading persisted
|
||||
* state, awaiting reconciliation: init() adopts the persisted state when the
|
||||
* definition hash matches; anything left unadopted when init settles resolves
|
||||
* to a backfill. Deciding backfill eagerly at define time was the boot-order
|
||||
* bug that wiped valid persisted state on every restart — the synchronous
|
||||
* defineAggregate() always beats the async init().
|
||||
*/
|
||||
private pendingAdopt = new Set<string>()
|
||||
|
||||
/**
|
||||
* In-flight rescan targets. While a name has a staging map, ALL
|
||||
* contributions (the walk's and concurrent write hooks') land there instead
|
||||
* of the live map; the live map keeps serving until {@link finishBackfill}
|
||||
* swaps the staging map in atomically.
|
||||
*/
|
||||
private backfillStaging = new Map<string, Map<string, AggregateGroupState>>()
|
||||
|
||||
constructor(storage: StorageAdapter, nativeProvider?: AggregationProvider) {
|
||||
this.storage = storage
|
||||
this.nativeProvider = nativeProvider
|
||||
|
|
@ -369,128 +336,21 @@ export class AggregationIndex {
|
|||
|
||||
/**
|
||||
* Initialize: load persisted definitions and state, detect changes, rebuild stale.
|
||||
*
|
||||
* Idempotent — repeated calls return the same promise. Definitions registered
|
||||
* *before* this completes (the normal boot order: `defineAggregate()` is
|
||||
* synchronous and always beats this async load) are reconciled rather than
|
||||
* clobbered: the app's definition wins, and its persisted state is adopted
|
||||
* when the definition hash matches — backfill happens only on a real change.
|
||||
*/
|
||||
init(): Promise<void> {
|
||||
if (!this.initPromise) {
|
||||
this.initPromise = this.loadPersisted().finally(() => {
|
||||
this.resolvePendingAdoptToBackfill()
|
||||
this.initDone = true
|
||||
})
|
||||
}
|
||||
return this.initPromise
|
||||
}
|
||||
|
||||
/**
|
||||
* Await the persisted-state load (if one was started) and settle every
|
||||
* pending adoption decision. After this resolves, `getPendingBackfills()`
|
||||
* is authoritative: a name is listed iff it genuinely needs a rescan.
|
||||
* Query paths must await this before consulting backfill state.
|
||||
*/
|
||||
async ready(): Promise<void> {
|
||||
if (this.initPromise) {
|
||||
try {
|
||||
await this.initPromise
|
||||
} catch {
|
||||
// The owner already surfaced the load failure loudly; backfill covers.
|
||||
}
|
||||
}
|
||||
this.resolvePendingAdoptToBackfill()
|
||||
}
|
||||
|
||||
/**
|
||||
* Any definition still awaiting state adoption has no persisted state to
|
||||
* adopt (or init never ran / failed) — it must backfill.
|
||||
*/
|
||||
private resolvePendingAdoptToBackfill(): void {
|
||||
if (this.pendingAdopt.size > 0) {
|
||||
prodLog.info(
|
||||
`[Aggregation] no adoptable persisted state for: ${Array.from(this.pendingAdopt).join(', ')} — flagged for backfill`
|
||||
)
|
||||
}
|
||||
for (const name of this.pendingAdopt) this.needsBackfill.add(name)
|
||||
this.pendingAdopt.clear()
|
||||
}
|
||||
|
||||
/**
|
||||
* May this persisted state be ADOPTED? When the store exposes its committed
|
||||
* watermark, the state's `sourceGeneration` must EQUAL it: behind means
|
||||
* later writes are missing from the state (unclean shutdown); ahead means
|
||||
* it counts writes that no longer exist (e.g. a fact-log truncation on a
|
||||
* copied store pulled the watermark back). Either way: one exact rescan,
|
||||
* said out loud — never a silent adopt. Stores without the capability (and
|
||||
* pre-stamp state on them) fall back to hash-only adoption.
|
||||
*/
|
||||
private stateGenerationAdoptable(name: string, stateData: unknown): boolean {
|
||||
const committed = this.storage.committedGeneration?.() ?? null
|
||||
if (committed === null) return true
|
||||
const raw = (stateData as Record<string, unknown>).sourceGeneration
|
||||
const stamped = typeof raw === 'number' ? raw : null
|
||||
if (stamped === committed) return true
|
||||
prodLog.warn(
|
||||
`[Aggregation] '${name}': persisted state is at generation ${stamped ?? 'unstamped'} ` +
|
||||
`but the store's committed generation is ${committed} — rescanning instead of adopting`
|
||||
)
|
||||
return false
|
||||
}
|
||||
|
||||
private async loadPersisted(): Promise<void> {
|
||||
async init(): Promise<void> {
|
||||
// Load persisted definitions
|
||||
const savedDefs = await this.storage.getMetadata(DEFINITIONS_KEY)
|
||||
if (savedDefs && typeof savedDefs === 'object' && savedDefs.definitions) {
|
||||
const defs = savedDefs.definitions as Array<AggregateDefinition & { _hash?: string }>
|
||||
|
||||
for (const def of defs) {
|
||||
const savedHash = def._hash || ''
|
||||
|
||||
if (this.definitions.has(def.name)) {
|
||||
// The app re-registered this aggregate before the load finished.
|
||||
// The app's definition wins — never clobber it with the persisted
|
||||
// copy. Adopt the persisted state when the definition is unchanged
|
||||
// AND no write has landed for it yet (a landed write would be lost
|
||||
// by adoption; the hook flips such names to backfill).
|
||||
const appHash = this.definitionHashes.get(def.name) || ''
|
||||
if (appHash === savedHash && this.pendingAdopt.has(def.name)) {
|
||||
const stateData = await this.storage.getMetadata(`${STATE_KEY_PREFIX}${def.name}__`)
|
||||
if (
|
||||
stateData &&
|
||||
stateData.groups &&
|
||||
this.stateGenerationAdoptable(def.name, stateData)
|
||||
) {
|
||||
const groupMap = new Map<string, AggregateGroupState>()
|
||||
for (const group of stateData.groups as AggregateGroupState[]) {
|
||||
groupMap.set(serializeGroupKey(group.groupKey), group)
|
||||
}
|
||||
this.states.set(def.name, groupMap)
|
||||
this.pendingAdopt.delete(def.name)
|
||||
this.needsBackfill.delete(def.name)
|
||||
prodLog.info(
|
||||
`[Aggregation] '${def.name}': adopted persisted state (${groupMap.size} groups) — no rescan`
|
||||
)
|
||||
}
|
||||
// No/invalid persisted state: stays in pendingAdopt and resolves
|
||||
// to backfill when init settles.
|
||||
}
|
||||
continue
|
||||
}
|
||||
|
||||
// Not registered this session — restore definition + state from
|
||||
// persistence.
|
||||
this.definitions.set(def.name, def)
|
||||
const currentHash = hashDefinition(def)
|
||||
const savedHash = def._hash || ''
|
||||
|
||||
// Load persisted state
|
||||
const stateData = await this.storage.getMetadata(`${STATE_KEY_PREFIX}${def.name}__`)
|
||||
if (
|
||||
stateData &&
|
||||
stateData.groups &&
|
||||
savedHash === currentHash &&
|
||||
this.stateGenerationAdoptable(def.name, stateData)
|
||||
) {
|
||||
if (stateData && stateData.groups && savedHash === currentHash) {
|
||||
// Definition unchanged — load state
|
||||
const groupMap = new Map<string, AggregateGroupState>()
|
||||
for (const group of stateData.groups as AggregateGroupState[]) {
|
||||
|
|
@ -498,10 +358,6 @@ export class AggregationIndex {
|
|||
groupMap.set(serialized, group)
|
||||
}
|
||||
this.states.set(def.name, groupMap)
|
||||
this.needsBackfill.delete(def.name)
|
||||
prodLog.info(
|
||||
`[Aggregation] '${def.name}': restored definition + adopted persisted state (${groupMap.size} groups)`
|
||||
)
|
||||
} else {
|
||||
// Definition changed or no saved state — start fresh and backfill from
|
||||
// existing entities (the owner drains needsBackfill on first query).
|
||||
|
|
@ -542,22 +398,14 @@ export class AggregationIndex {
|
|||
}))
|
||||
await this.storage.saveMetadata(DEFINITIONS_KEY, { definitions: defsToSave })
|
||||
|
||||
// Persist dirty states, stamped with the committed generation they
|
||||
// reflect. The stamp is what makes reopen-adoption verifiable: state at a
|
||||
// different generation than the store's committed watermark is stale (an
|
||||
// unclean shutdown after later writes) or over-counts (a fact-log
|
||||
// truncation on a copied store pulled the watermark BACK below the
|
||||
// stamp) — either way the answer is one exact rescan, never a silent
|
||||
// adopt. Read the generation after collecting groups so any racing
|
||||
// commit resolves toward rescan, not wrong-adopt.
|
||||
// Persist dirty states
|
||||
for (const name of this.dirty) {
|
||||
const stateMap = this.states.get(name)
|
||||
if (stateMap) {
|
||||
const groups = Array.from(stateMap.values())
|
||||
const sourceGeneration = this.storage.committedGeneration?.() ?? null
|
||||
await this.storage.saveMetadata(
|
||||
`${STATE_KEY_PREFIX}${name}__`,
|
||||
sourceGeneration === null ? { groups } : { groups, sourceGeneration }
|
||||
{ groups }
|
||||
)
|
||||
}
|
||||
}
|
||||
|
|
@ -604,19 +452,10 @@ export class AggregationIndex {
|
|||
this.definitions.set(def.name, def)
|
||||
this.definitionHashes.set(def.name, newHash)
|
||||
|
||||
// First sight this session, before init() settled: defer the backfill
|
||||
// decision — init() adopts the persisted state on hash match, and anything
|
||||
// left unadopted resolves to backfill. Deciding eagerly here wiped valid
|
||||
// persisted state on every restart.
|
||||
if (!this.states.has(def.name) && !this.initDone) {
|
||||
this.states.set(def.name, new Map())
|
||||
this.pendingAdopt.add(def.name)
|
||||
}
|
||||
// Reset state if definition changed or doesn't exist yet, and flag it for
|
||||
// backfill so already-stored entities are counted (write-time hooks only see
|
||||
// future writes). The owner drains this on the next query via getPendingBackfills().
|
||||
else if (!this.states.has(def.name) || (oldHash && oldHash !== newHash)) {
|
||||
this.pendingAdopt.delete(def.name)
|
||||
if (!this.states.has(def.name) || (oldHash && oldHash !== newHash)) {
|
||||
this.states.set(def.name, new Map())
|
||||
this.needsBackfill.add(def.name)
|
||||
}
|
||||
|
|
@ -637,8 +476,6 @@ export class AggregationIndex {
|
|||
this.definitionHashes.delete(name)
|
||||
this.states.delete(name)
|
||||
this.staleMinMax.delete(name)
|
||||
this.pendingAdopt.delete(name)
|
||||
this.needsBackfill.delete(name)
|
||||
|
||||
// Notify native provider
|
||||
if (this.nativeProvider?.removeAggregate) {
|
||||
|
|
@ -676,17 +513,9 @@ export class AggregationIndex {
|
|||
return Array.from(this.needsBackfill)
|
||||
}
|
||||
|
||||
/**
|
||||
* Begin a rescan into a STAGING map. The live state is not touched — it
|
||||
* keeps serving (possibly stale, but flagged pending) until the rescan
|
||||
* completes and swaps in atomically. A mid-walk failure drops the staging
|
||||
* map via {@link abortBackfill} and loses nothing: wiping live state before
|
||||
* a scan that could throw was the destructive-before-durable defect.
|
||||
* Contributions (walk + concurrent write hooks) land in staging while it
|
||||
* exists, so the swapped-in result reflects writes that raced the walk.
|
||||
*/
|
||||
/** Clear an aggregate's state so a full rescan cannot double-count. */
|
||||
beginBackfill(name: string): void {
|
||||
this.backfillStaging.set(name, new Map())
|
||||
this.states.set(name, new Map())
|
||||
// Reset native provider state for this aggregate too, if present.
|
||||
const def = this.definitions.get(name)
|
||||
if (def && this.nativeProvider?.removeAggregate && this.nativeProvider?.defineAggregate) {
|
||||
|
|
@ -695,15 +524,6 @@ export class AggregationIndex {
|
|||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Abandon an in-flight rescan after a failure: drop the staging map, keep
|
||||
* the live state serving, leave the aggregate flagged as pending so a later
|
||||
* attempt rescans. The failure itself must be surfaced loudly by the owner.
|
||||
*/
|
||||
abortBackfill(name: string): void {
|
||||
this.backfillStaging.delete(name)
|
||||
}
|
||||
|
||||
/** Feed one already-stored entity into a single aggregate during backfill. */
|
||||
backfillEntity(name: string, entity: Record<string, unknown>): void {
|
||||
if (isAggregateEntity(entity)) return
|
||||
|
|
@ -717,33 +537,14 @@ export class AggregationIndex {
|
|||
}
|
||||
}
|
||||
|
||||
/** Swap the rebuilt staging state in atomically; persists on next flush(). */
|
||||
/** Mark an aggregate's backfill complete; rebuilt state persists on next flush(). */
|
||||
finishBackfill(name: string): void {
|
||||
const staged = this.backfillStaging.get(name)
|
||||
if (staged) {
|
||||
this.states.set(name, staged)
|
||||
this.backfillStaging.delete(name)
|
||||
}
|
||||
this.needsBackfill.delete(name)
|
||||
this.dirty.add(name)
|
||||
}
|
||||
|
||||
// ============= Write-Time Hooks =============
|
||||
|
||||
/**
|
||||
* A write is landing for an aggregate whose persisted-state adoption is still
|
||||
* pending — adopting after this write would lose its contribution. Settle the
|
||||
* decision now: an exact rescan instead of adoption. The window is the few
|
||||
* milliseconds between a boot-time defineAggregate() and init() completing,
|
||||
* so this rarely fires; when it does, correctness wins over the walk.
|
||||
*/
|
||||
private resolveAdoptOnWrite(name: string): void {
|
||||
if (this.pendingAdopt.has(name)) {
|
||||
this.pendingAdopt.delete(name)
|
||||
this.needsBackfill.add(name)
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Called when an entity is added. Updates all matching aggregates.
|
||||
*/
|
||||
|
|
@ -752,7 +553,6 @@ export class AggregationIndex {
|
|||
|
||||
for (const [name, def] of this.definitions) {
|
||||
if (!matchesSource(entity, def.source)) continue
|
||||
this.resolveAdoptOnWrite(name)
|
||||
|
||||
if (this.nativeProvider) {
|
||||
const results = this.nativeProvider.incrementalUpdate(name, def, entity, 'add')
|
||||
|
|
@ -779,10 +579,6 @@ export class AggregationIndex {
|
|||
const oldMatches = matchesSource(oldEntity, def.source)
|
||||
const newMatches = matchesSource(newEntity, def.source)
|
||||
|
||||
if (oldMatches || newMatches) {
|
||||
this.resolveAdoptOnWrite(name)
|
||||
}
|
||||
|
||||
if (this.nativeProvider && (oldMatches || newMatches)) {
|
||||
const results = this.nativeProvider.incrementalUpdate(name, def, newEntity, 'update', oldEntity)
|
||||
this.applyNativeResults(name, results)
|
||||
|
|
@ -809,7 +605,6 @@ export class AggregationIndex {
|
|||
|
||||
for (const [name, def] of this.definitions) {
|
||||
if (!matchesSource(entity, def.source)) continue
|
||||
this.resolveAdoptOnWrite(name)
|
||||
|
||||
if (this.nativeProvider) {
|
||||
const results = this.nativeProvider.incrementalUpdate(name, def, entity, 'delete')
|
||||
|
|
@ -962,7 +757,7 @@ export class AggregationIndex {
|
|||
def: AggregateDefinition,
|
||||
entity: Record<string, unknown>
|
||||
): void {
|
||||
const stateMap = (this.backfillStaging.get(aggName) ?? this.states.get(aggName))!
|
||||
const stateMap = this.states.get(aggName)!
|
||||
|
||||
// Fan out: an unnest dimension makes one entity contribute to several groups.
|
||||
for (const groupKey of computeGroupKeys(entity, def.groupBy)) {
|
||||
|
|
@ -1020,7 +815,7 @@ export class AggregationIndex {
|
|||
def: AggregateDefinition,
|
||||
entity: Record<string, unknown>
|
||||
): void {
|
||||
const stateMap = (this.backfillStaging.get(aggName) ?? this.states.get(aggName))!
|
||||
const stateMap = this.states.get(aggName)!
|
||||
|
||||
// Fan out: reverse the entity's contribution from every group it joined.
|
||||
for (const groupKey of computeGroupKeys(entity, def.groupBy)) {
|
||||
|
|
@ -1076,7 +871,7 @@ export class AggregationIndex {
|
|||
* Apply results from native provider back into the state maps.
|
||||
*/
|
||||
private applyNativeResults(aggName: string, results: AggregateGroupState[]): void {
|
||||
const stateMap = (this.backfillStaging.get(aggName) ?? this.states.get(aggName))!
|
||||
const stateMap = this.states.get(aggName)!
|
||||
for (const group of results) {
|
||||
const serialized = serializeGroupKey(group.groupKey)
|
||||
stateMap.set(serialized, group)
|
||||
|
|
|
|||
785
src/brainy.ts
785
src/brainy.ts
File diff suppressed because it is too large
Load diff
|
|
@ -848,17 +848,7 @@ export interface StorageAdapter {
|
|||
*/
|
||||
getNounsByNounType(nounType: string): Promise<HNSWNounWithMetadata[]>
|
||||
|
||||
/**
|
||||
* Delete a noun — FULL canonical removal (both legs + the entity container).
|
||||
*
|
||||
* @param id The entity id.
|
||||
* @param priorMetadata OPTIONAL already-known metadata of the entity being
|
||||
* removed (the caller's pre-delete read). The count decrement must never
|
||||
* REQUIRE re-reading the record being removed: when the internal read
|
||||
* returns `null` (replace race, or a ghost left by an earlier version) the
|
||||
* decrement falls back to this record instead of being silently skipped.
|
||||
*/
|
||||
deleteNoun(id: string, priorMetadata?: NounMetadata | null): Promise<void>
|
||||
deleteNoun(id: string): Promise<void>
|
||||
|
||||
/**
|
||||
* Save verb - Pure HNSW verb with core fields only
|
||||
|
|
@ -915,15 +905,7 @@ export interface StorageAdapter {
|
|||
*/
|
||||
getVerbsByType(type: string): Promise<HNSWVerbWithMetadata[]>
|
||||
|
||||
/**
|
||||
* Delete a verb — FULL canonical removal (both legs + the container).
|
||||
*
|
||||
* @param id The relationship id.
|
||||
* @param priorMetadata OPTIONAL already-known metadata of the edge being
|
||||
* removed (the caller's pre-delete read); keeps the count decrement honest
|
||||
* when the internal read returns `null` (see `deleteNoun`).
|
||||
*/
|
||||
deleteVerb(id: string, priorMetadata?: VerbMetadata | null): Promise<void>
|
||||
deleteVerb(id: string): Promise<void>
|
||||
|
||||
/**
|
||||
* Save metadata
|
||||
|
|
@ -1127,77 +1109,6 @@ export interface StorageAdapter {
|
|||
*/
|
||||
listDerivedFamilies?(): Promise<DerivedFamilyDeclaration[]>
|
||||
|
||||
/**
|
||||
* @description OPTIONAL binary raw-byte primitives — the substrate for
|
||||
* append-only log-structured files (the generation fact log's CRC-framed
|
||||
* segments). Feature-detected: an adapter that omits them simply hosts no
|
||||
* fact log (readers fall back to canonical enumeration). Paths are
|
||||
* storage-root-relative and used VERBATIM (no `.gz`/`.bin` suffixing —
|
||||
* unlike the JSON object and blob primitives).
|
||||
*
|
||||
* Append to a raw binary file, creating it (and parent directories) when
|
||||
* absent. Append durability is the CALLER's job via `syncRawObjects` —
|
||||
* matching the commit protocol, which batches fsyncs at its barrier.
|
||||
*/
|
||||
appendRawBytes?(path: string, bytes: Uint8Array): Promise<void>
|
||||
|
||||
/**
|
||||
* Read a raw binary file whole. Absent → `null`; a real IO fault throws
|
||||
* (never masked as absence).
|
||||
*/
|
||||
readRawBytes?(path: string): Promise<Uint8Array | null>
|
||||
|
||||
/**
|
||||
* Replace a raw binary file atomically (write-new → fsync → rename) — the
|
||||
* reconcile primitive (e.g. truncating a fact-log tail back to committed
|
||||
* truth after a crash).
|
||||
*/
|
||||
writeRawBytes?(path: string, bytes: Uint8Array): Promise<void>
|
||||
|
||||
/**
|
||||
* Byte size of a raw binary file, or `null` when absent.
|
||||
*/
|
||||
rawByteSize?(path: string): Promise<number | null>
|
||||
|
||||
/**
|
||||
* @description OPTIONAL fact-scan capability — how an index provider that
|
||||
* holds only `storage` reaches the generation fact log (the host brain
|
||||
* wires it at init; providers must never construct their own fact-log
|
||||
* reader — the log's open path is writer-side). Present ⟺ this store hosts
|
||||
* a fact log AND the host wired the capability. Returns a scan handle over
|
||||
* committed facts (heal-grade telemetry included), or `null` when no fact
|
||||
* log exists — callers fall back to the canonical enumeration walk.
|
||||
*/
|
||||
scanFacts?(options?: {
|
||||
fromGeneration?: number
|
||||
toGeneration?: number
|
||||
kinds?: Array<'noun' | 'verb'>
|
||||
batchSize?: number
|
||||
}): import('./db/factLog.js').FactScanHandle | null
|
||||
|
||||
/**
|
||||
* @description OPTIONAL (rides the fact-scan capability): the fact log's
|
||||
* head generation — the replay target for `stamp.sourceGeneration + 1`
|
||||
* catch-ups. `null` when no fact log exists.
|
||||
*/
|
||||
factLogHeadGeneration?(): number | null
|
||||
|
||||
/**
|
||||
* @description OPTIONAL (rides the fact-scan capability): the COMMITTED
|
||||
* generation watermark — the manifest truth a projection's
|
||||
* `sourceGeneration` compares against. Exposed as a capability so a
|
||||
* provider never parses the store's private manifest format. `null` when
|
||||
* the capability is unwired.
|
||||
*/
|
||||
committedGeneration?(): number | null
|
||||
|
||||
/**
|
||||
* @description OPTIONAL (rides the fact-scan capability): the immutable,
|
||||
* sealed fact-segment file paths covering `fromGeneration` — the zero-copy
|
||||
* handoff. The mutable tail is excluded (read it via `scanFacts`).
|
||||
*/
|
||||
factSegmentPaths?(options?: { fromGeneration?: number }): string[]
|
||||
|
||||
/**
|
||||
* Save statistics data
|
||||
* @param statistics The statistics data to save
|
||||
|
|
|
|||
|
|
@ -945,13 +945,6 @@ export class Db<T = any> {
|
|||
* {@link SpeculativeOverlayError} (commit them with `brain.transact()`
|
||||
* first).
|
||||
*
|
||||
* SPARSE FILES: a native accelerator's mmap index files can be sparse —
|
||||
* huge apparent size, small allocated size. `persist()` handles them
|
||||
* correctly (hard links share the allocation). But if you then archive the
|
||||
* snapshot with EXTERNAL tools, use the sparse-aware flags (`tar czSf`,
|
||||
* `rsync --sparse`, `cp --sparse=always`) or the copy materializes every
|
||||
* hole — see docs/guides/external-backups-and-sparse-storage.md.
|
||||
*
|
||||
* @param path - Absolute directory for the snapshot (created; must be
|
||||
* empty or absent).
|
||||
* @throws GenerationConflictError when this view is no longer the latest
|
||||
|
|
|
|||
|
|
@ -1,721 +0,0 @@
|
|||
/**
|
||||
* @module db/factLog
|
||||
* @description The generation FACT LOG — an append-only, CRC-framed record of
|
||||
* every committed generation as an AFTER-IMAGE "fact": what each touched
|
||||
* entity/relationship BECAME (or a body-less tombstone when it was removed).
|
||||
* This is the dual-write half of the log-canonical transition: today the
|
||||
* before-image history + canonical tree remain authoritative; the fact log is
|
||||
* appended at the same commit points and reconciled to committed truth at
|
||||
* open, so consumers (index heals, replays, scans) can read one sequential,
|
||||
* self-verifying stream instead of walking the entity tree.
|
||||
*
|
||||
* ## Wire format (frozen; additive-only within a major)
|
||||
*
|
||||
* Fact (msgpack, POSITIONAL array — the segment header's formatVersion
|
||||
* governs the schema):
|
||||
*
|
||||
* fact := [ generation:u64, timestamp:u64, ops, meta|nil, blobHashes|nil ]
|
||||
* op := [ kind:u8 (0=noun, 1=verb), id:bin16 (raw uuid bytes),
|
||||
* record:[metaLeg, vecLeg] | nil ] // nil = TOMBSTONE
|
||||
*
|
||||
* Segment file (`_generations/facts/seg-<firstGeneration, zero-padded 20>.bfl`):
|
||||
*
|
||||
* header := magic "BFACTS\0\0" (8B) | formatVersion:u32 LE |
|
||||
* firstGeneration:u64 LE | reserved 12B (ZEROED, verified)
|
||||
* frame := length:u32 LE | crc32c:u32 LE (of payload) | payload
|
||||
*
|
||||
* A fact is never split across segments; a torn tail (length overrun or CRC
|
||||
* mismatch) terminates that segment's scan — everything before it is intact.
|
||||
* Zero-padded names make lexicographic order == generation order.
|
||||
*
|
||||
* ## Invariant
|
||||
*
|
||||
* After {@link FactLog.open}, the log contains EXACTLY the committed prefix:
|
||||
* facts are appended BEFORE the commit point (inside the same durability
|
||||
* window), so a crash can only leave the log AHEAD of committed truth — open
|
||||
* truncates any fact beyond the committed generation. Absent generation =
|
||||
* never committed; present = committed. A scan can never see an uncommitted
|
||||
* fact.
|
||||
*
|
||||
* The manifest (`_generations/facts/manifest.json`, JSON — forensics stay
|
||||
* terminal-readable) is the single source of truth for the segment SET;
|
||||
* rotation flips it atomically (write-new → fsync → rename) BEFORE the new
|
||||
* tail's first byte exists, so no segment file is ever unaccounted for.
|
||||
*/
|
||||
import { encode as defaultEncode, decode as defaultDecode } from '@msgpack/msgpack'
|
||||
import { crc32c } from '../utils/crc32c.js'
|
||||
import { prodLog } from '../utils/logger.js'
|
||||
|
||||
// Swappable msgpack implementation — defaults to the JS codec; a native
|
||||
// provider (registered via the plugin registry's 'msgpack' key) may replace
|
||||
// it. Byte-compatibility is the contract (positional arrays, bin16 ids).
|
||||
let msgpackEncode: (value: unknown) => Uint8Array = defaultEncode
|
||||
let msgpackDecode: (bytes: Uint8Array) => unknown = defaultDecode
|
||||
|
||||
/** Replace the msgpack encode/decode implementation at runtime. */
|
||||
export function setFactCodec(impl: {
|
||||
encode: (value: unknown) => Uint8Array
|
||||
decode: (bytes: Uint8Array) => unknown
|
||||
}): void {
|
||||
msgpackEncode = impl.encode
|
||||
msgpackDecode = impl.decode
|
||||
}
|
||||
|
||||
/** Storage-root-relative home of the fact log. */
|
||||
export const FACTS_PREFIX = '_generations/facts'
|
||||
/** The facts manifest path (JSON). */
|
||||
export const FACTS_MANIFEST_PATH = `${FACTS_PREFIX}/manifest.json`
|
||||
/** Current segment format version (header field; additive-only within a major). */
|
||||
export const FACTS_FORMAT_VERSION = 1
|
||||
/** Rotation threshold: seal the tail segment once it exceeds this many bytes. */
|
||||
const SEGMENT_ROTATE_BYTES = 8 * 1024 * 1024
|
||||
/** Segment header: magic(8) + formatVersion(4) + firstGeneration(8) + reserved(12). */
|
||||
const HEADER_BYTES = 32
|
||||
const MAGIC = new Uint8Array([0x42, 0x46, 0x41, 0x43, 0x54, 0x53, 0x00, 0x00]) // "BFACTS\0\0"
|
||||
/** Frame prefix: length(4) + crc32c(4). */
|
||||
const FRAME_PREFIX_BYTES = 8
|
||||
|
||||
/** One write inside a fact: what the id became (or a tombstone). */
|
||||
export interface FactOp {
|
||||
kind: 'noun' | 'verb'
|
||||
id: string
|
||||
/** The AFTER-IMAGE legs, or `null` for a tombstone (the id was removed). */
|
||||
record: { metadata: unknown | null; vector: unknown | null } | null
|
||||
}
|
||||
|
||||
/** One committed generation, as scanned back out of the log. */
|
||||
export interface CommitFact {
|
||||
generation: number
|
||||
timestamp: number
|
||||
ops: FactOp[]
|
||||
meta?: Record<string, unknown>
|
||||
blobHashes?: string[]
|
||||
}
|
||||
|
||||
/** The telemetry a scan batch carries (frozen shape). */
|
||||
export interface FactScanBatch {
|
||||
facts: CommitFact[]
|
||||
firstGeneration: number
|
||||
lastGeneration: number
|
||||
factCount: number
|
||||
byteSize: number
|
||||
segmentId: string
|
||||
}
|
||||
|
||||
/**
|
||||
* Liveness bound on a scan's FIRST batch (Stage-2 co-freeze, D1 contract):
|
||||
* `batches()` must yield its first batch — or fail loudly — within this many
|
||||
* ms of the first pull. A backlogged or damaged store may be SLOW, but it may
|
||||
* never be SILENT: a consumer awaiting the first batch is otherwise
|
||||
* indistinguishable from a wedge (the exact failure shape a production heal
|
||||
* hit against a generations-backlogged brain).
|
||||
*/
|
||||
export const SCANFACTS_FIRST_BATCH_MS = 10_000
|
||||
|
||||
/** The telemetry a scan OPEN returns (frozen shape). */
|
||||
export interface FactScanHandle {
|
||||
headGeneration: number
|
||||
segmentCount: number
|
||||
approxFactCount: number
|
||||
/**
|
||||
* Ordered batches; a detected gap aborts LOUDLY, never a silent skip.
|
||||
* Liveness contract: the FIRST batch resolves or rejects within
|
||||
* {@link SCANFACTS_FIRST_BATCH_MS} of the first pull — never a silent hang.
|
||||
*/
|
||||
batches: () => AsyncGenerator<FactScanBatch>
|
||||
/** Close telemetry — the invariant cross-check, valid after iteration ends. */
|
||||
summary: () => { factsYielded: number; segmentsRead: number }
|
||||
}
|
||||
|
||||
/** Manifest entry for a sealed segment. */
|
||||
interface SegmentEntry {
|
||||
file: string
|
||||
firstGeneration: number
|
||||
lastGeneration: number
|
||||
facts: number
|
||||
bytes: number
|
||||
}
|
||||
|
||||
/** The facts manifest (JSON on disk). */
|
||||
interface FactsManifest {
|
||||
formatVersion: number
|
||||
segments: SegmentEntry[]
|
||||
/** The append target. Its true content is established by scanning (crash tolerance). */
|
||||
tailSegment: string | null
|
||||
updatedAt: string
|
||||
}
|
||||
|
||||
/** The narrow byte-level storage surface the fact log rides. */
|
||||
export interface FactLogStorage {
|
||||
appendRawBytes(path: string, bytes: Uint8Array): Promise<void>
|
||||
readRawBytes(path: string): Promise<Uint8Array | null>
|
||||
writeRawBytes(path: string, bytes: Uint8Array): Promise<void>
|
||||
rawByteSize(path: string): Promise<number | null>
|
||||
readRawObject(path: string): Promise<any | null>
|
||||
writeRawObject(path: string, data: any): Promise<void>
|
||||
syncRawObjects(paths: string[]): Promise<void>
|
||||
deleteRawObject(path: string): Promise<void>
|
||||
}
|
||||
|
||||
/** True when the storage adapter exposes every primitive the fact log needs. */
|
||||
export function storageSupportsFactLog(storage: unknown): storage is FactLogStorage {
|
||||
const s = storage as Record<string, unknown>
|
||||
return (
|
||||
typeof s.appendRawBytes === 'function' &&
|
||||
typeof s.readRawBytes === 'function' &&
|
||||
typeof s.writeRawBytes === 'function' &&
|
||||
typeof s.rawByteSize === 'function'
|
||||
)
|
||||
}
|
||||
|
||||
/** uuid string → 16 raw bytes (bin16 on the wire). */
|
||||
function uuidToBytes(id: string): Uint8Array {
|
||||
const hex = id.replace(/-/g, '')
|
||||
if (hex.length !== 32) {
|
||||
// Non-uuid ids (legacy/natural keys) ride as UTF-8 with a length prefix
|
||||
// marker impossible for uuids: we refuse instead — the write API has
|
||||
// guaranteed uuid ids since 8.0, so anything else is a corruption signal.
|
||||
throw new Error(`fact log: id is not a uuid: ${id}`)
|
||||
}
|
||||
const bytes = new Uint8Array(16)
|
||||
for (let i = 0; i < 16; i++) {
|
||||
bytes[i] = parseInt(hex.slice(i * 2, i * 2 + 2), 16)
|
||||
}
|
||||
return bytes
|
||||
}
|
||||
|
||||
/** 16 raw bytes → canonical lowercase uuid string. */
|
||||
function bytesToUuid(bytes: Uint8Array): string {
|
||||
let hex = ''
|
||||
for (let i = 0; i < 16; i++) hex += bytes[i].toString(16).padStart(2, '0')
|
||||
return `${hex.slice(0, 8)}-${hex.slice(8, 12)}-${hex.slice(12, 16)}-${hex.slice(16, 20)}-${hex.slice(20)}`
|
||||
}
|
||||
|
||||
/** Zero-padded segment filename: lexicographic order == generation order. */
|
||||
function segmentFileName(firstGeneration: number): string {
|
||||
return `seg-${String(firstGeneration).padStart(20, '0')}.bfl`
|
||||
}
|
||||
|
||||
/** Build a segment header. Reserved bytes are ZEROED (and verified on open). */
|
||||
function buildHeader(firstGeneration: number): Uint8Array {
|
||||
const header = new Uint8Array(HEADER_BYTES)
|
||||
header.set(MAGIC, 0)
|
||||
const view = new DataView(header.buffer)
|
||||
view.setUint32(8, FACTS_FORMAT_VERSION, true)
|
||||
view.setBigUint64(12, BigInt(firstGeneration), true)
|
||||
// bytes 20..31 stay zero (reserved)
|
||||
return header
|
||||
}
|
||||
|
||||
/** Encode one fact into a framed record (length + crc32c + msgpack payload). */
|
||||
function encodeFrame(fact: CommitFact): Uint8Array {
|
||||
const payload = msgpackEncode([
|
||||
fact.generation,
|
||||
fact.timestamp,
|
||||
fact.ops.map((op) => [
|
||||
op.kind === 'noun' ? 0 : 1,
|
||||
uuidToBytes(op.id),
|
||||
op.record === null ? null : [op.record.metadata, op.record.vector]
|
||||
]),
|
||||
fact.meta ?? null,
|
||||
fact.blobHashes && fact.blobHashes.length > 0 ? fact.blobHashes : null
|
||||
])
|
||||
const frame = new Uint8Array(FRAME_PREFIX_BYTES + payload.length)
|
||||
const view = new DataView(frame.buffer)
|
||||
view.setUint32(0, payload.length, true)
|
||||
view.setUint32(4, crc32c(payload), true)
|
||||
frame.set(payload, FRAME_PREFIX_BYTES)
|
||||
return frame
|
||||
}
|
||||
|
||||
/** Decode one msgpack payload back into a CommitFact. */
|
||||
function decodeFact(payload: Uint8Array): CommitFact {
|
||||
const raw = msgpackDecode(payload) as unknown[]
|
||||
const [generation, timestamp, ops, meta, blobHashes] = raw as [
|
||||
number,
|
||||
number,
|
||||
Array<[number, Uint8Array, [unknown, unknown] | null]>,
|
||||
Record<string, unknown> | null,
|
||||
string[] | null
|
||||
]
|
||||
return {
|
||||
generation: Number(generation),
|
||||
timestamp: Number(timestamp),
|
||||
ops: ops.map(([kind, idBytes, record]) => ({
|
||||
kind: kind === 0 ? ('noun' as const) : ('verb' as const),
|
||||
id: bytesToUuid(idBytes),
|
||||
record: record === null ? null : { metadata: record[0] ?? null, vector: record[1] ?? null }
|
||||
})),
|
||||
...(meta ? { meta } : {}),
|
||||
...(blobHashes && blobHashes.length > 0 ? { blobHashes } : {})
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a segment's bytes: verify the header, then walk frames until the end
|
||||
* or a torn tail (length overrun / CRC mismatch), which terminates the walk —
|
||||
* everything before it is intact. Returns the decoded facts plus the byte
|
||||
* length of the VALID prefix (header + intact frames), which reconciliation
|
||||
* uses to cut a torn tail without re-encoding.
|
||||
*/
|
||||
function parseSegment(
|
||||
file: string,
|
||||
bytes: Uint8Array
|
||||
): { facts: CommitFact[]; validBytes: number } {
|
||||
if (bytes.length < HEADER_BYTES) {
|
||||
prodLog.warn(`[FactLog] segment ${file} shorter than its header — treating as empty`)
|
||||
return { facts: [], validBytes: 0 }
|
||||
}
|
||||
for (let i = 0; i < MAGIC.length; i++) {
|
||||
if (bytes[i] !== MAGIC[i]) {
|
||||
throw new Error(`fact log: segment ${file} has a bad magic — not a fact segment`)
|
||||
}
|
||||
}
|
||||
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength)
|
||||
const version = view.getUint32(8, true)
|
||||
if (version !== FACTS_FORMAT_VERSION) {
|
||||
throw new Error(
|
||||
`fact log: segment ${file} has formatVersion ${version}; this build reads ${FACTS_FORMAT_VERSION}`
|
||||
)
|
||||
}
|
||||
for (let i = 20; i < HEADER_BYTES; i++) {
|
||||
if (bytes[i] !== 0) {
|
||||
// Non-zero reserved bytes = a future format this build cannot verify.
|
||||
throw new Error(`fact log: segment ${file} has non-zero reserved header bytes — unverifiable`)
|
||||
}
|
||||
}
|
||||
|
||||
const facts: CommitFact[] = []
|
||||
let offset = HEADER_BYTES
|
||||
while (offset + FRAME_PREFIX_BYTES <= bytes.length) {
|
||||
const length = view.getUint32(offset, true)
|
||||
const expectedCrc = view.getUint32(offset + 4, true)
|
||||
const start = offset + FRAME_PREFIX_BYTES
|
||||
const end = start + length
|
||||
if (end > bytes.length) break // torn tail: frame length overruns the file
|
||||
const payload = bytes.subarray(start, end)
|
||||
if (crc32c(payload) !== expectedCrc) break // torn tail: payload CRC mismatch
|
||||
facts.push(decodeFact(payload))
|
||||
offset = end
|
||||
}
|
||||
return { facts, validBytes: offset }
|
||||
}
|
||||
|
||||
/**
|
||||
* The generation fact log. One instance per open store; every method assumes
|
||||
* the single-writer discipline the generation store already enforces (calls
|
||||
* arrive under its commit mutex).
|
||||
*/
|
||||
export class FactLog {
|
||||
private readonly storage: FactLogStorage
|
||||
/** Rotation threshold (bytes); tests may lower it to exercise rotation. */
|
||||
private readonly rotateBytes: number
|
||||
private manifest: FactsManifest = {
|
||||
formatVersion: FACTS_FORMAT_VERSION,
|
||||
segments: [],
|
||||
tailSegment: null,
|
||||
updatedAt: new Date(0).toISOString()
|
||||
}
|
||||
/** Decoded facts of the TAIL segment (bounded by the rotation threshold). */
|
||||
private tailFacts: CommitFact[] = []
|
||||
/** Byte size of the tail segment file (valid prefix). */
|
||||
private tailBytes = 0
|
||||
/** Highest generation in the log (0 = empty). */
|
||||
private head = 0
|
||||
/** Segment paths appended since the last sync (the fsync batch). */
|
||||
private readonly dirtySegments = new Set<string>()
|
||||
|
||||
constructor(storage: FactLogStorage, options?: { rotateBytes?: number }) {
|
||||
this.storage = storage
|
||||
this.rotateBytes = options?.rotateBytes ?? SEGMENT_ROTATE_BYTES
|
||||
}
|
||||
|
||||
/** The highest committed generation the log holds (0 = empty). */
|
||||
headGeneration(): number {
|
||||
return this.head
|
||||
}
|
||||
|
||||
/**
|
||||
* Open the log and reconcile it to committed truth: read the manifest,
|
||||
* establish the tail's intact content (torn-tail scan), then TRUNCATE any
|
||||
* fact with `generation > committedGeneration` — those never committed (a
|
||||
* crash between fact-append and the commit point). After open, the log is
|
||||
* exactly the committed prefix.
|
||||
*/
|
||||
async open(committedGeneration: number): Promise<void> {
|
||||
const stored = (await this.storage.readRawObject(FACTS_MANIFEST_PATH)) as FactsManifest | null
|
||||
if (stored && typeof stored === 'object' && Array.isArray(stored.segments)) {
|
||||
if (stored.formatVersion !== FACTS_FORMAT_VERSION) {
|
||||
throw new Error(
|
||||
`fact log: manifest formatVersion ${stored.formatVersion}; this build reads ${FACTS_FORMAT_VERSION}`
|
||||
)
|
||||
}
|
||||
this.manifest = stored
|
||||
}
|
||||
|
||||
// Drop sealed segments that sit ENTIRELY beyond committed truth (a crash
|
||||
// right after a rotation whose facts never committed), newest first.
|
||||
while (this.manifest.segments.length > 0) {
|
||||
const last = this.manifest.segments[this.manifest.segments.length - 1]
|
||||
if (last.firstGeneration > committedGeneration) {
|
||||
prodLog.warn(
|
||||
`[FactLog] dropping sealed segment ${last.file} (generations ${last.firstGeneration}..` +
|
||||
`${last.lastGeneration} never committed)`
|
||||
)
|
||||
await this.storage.deleteRawObject(`${FACTS_PREFIX}/${last.file}`)
|
||||
this.manifest.segments.pop()
|
||||
await this.persistManifest()
|
||||
} else if (last.lastGeneration > committedGeneration) {
|
||||
// A sealed segment STRADDLING committed truth: cut it back.
|
||||
await this.truncateSegmentTo(last.file, committedGeneration)
|
||||
const cut = await this.reloadSegmentEntry(last.file)
|
||||
this.manifest.segments[this.manifest.segments.length - 1] = cut
|
||||
await this.persistManifest()
|
||||
break
|
||||
} else {
|
||||
break
|
||||
}
|
||||
}
|
||||
|
||||
// Establish the tail: scan its intact prefix, then truncate beyond
|
||||
// committed truth (the common crash shape: buffered single-op facts whose
|
||||
// counter never went durable).
|
||||
if (this.manifest.tailSegment) {
|
||||
const tailPath = `${FACTS_PREFIX}/${this.manifest.tailSegment}`
|
||||
const bytes = await this.storage.readRawBytes(tailPath)
|
||||
if (bytes === null) {
|
||||
// Manifest named a tail whose first byte never landed — an empty tail.
|
||||
this.tailFacts = []
|
||||
this.tailBytes = 0
|
||||
} else {
|
||||
const { facts, validBytes } = parseSegment(this.manifest.tailSegment, bytes)
|
||||
const kept = facts.filter((f) => f.generation <= committedGeneration)
|
||||
if (kept.length !== facts.length || validBytes !== bytes.length) {
|
||||
const dropped = facts.length - kept.length
|
||||
if (dropped > 0) {
|
||||
prodLog.warn(
|
||||
`[FactLog] truncating ${dropped} uncommitted fact(s) beyond generation ` +
|
||||
`${committedGeneration} from the tail (never committed)`
|
||||
)
|
||||
}
|
||||
await this.rewriteTail(kept)
|
||||
} else {
|
||||
this.tailFacts = facts
|
||||
this.tailBytes = validBytes
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
this.head = this.computeHead()
|
||||
}
|
||||
|
||||
/**
|
||||
* Append one committed generation's fact. NOT durable until {@link sync} —
|
||||
* the caller batches durability at its commit barrier (transact syncs in
|
||||
* the same call; Model-B group-commit syncs at flush).
|
||||
*/
|
||||
async append(fact: CommitFact): Promise<void> {
|
||||
if (fact.generation <= this.head) {
|
||||
throw new Error(
|
||||
`fact log: non-monotonic append (generation ${fact.generation} ≤ head ${this.head})`
|
||||
)
|
||||
}
|
||||
if (this.manifest.tailSegment === null) {
|
||||
await this.startTail(fact.generation)
|
||||
} else if (this.tailBytes >= this.rotateBytes) {
|
||||
await this.rotate(fact.generation)
|
||||
}
|
||||
const frame = encodeFrame(fact)
|
||||
const tailPath = `${FACTS_PREFIX}/${this.manifest.tailSegment}`
|
||||
await this.storage.appendRawBytes(tailPath, frame)
|
||||
this.tailFacts.push(fact)
|
||||
this.tailBytes += frame.length
|
||||
this.head = fact.generation
|
||||
this.dirtySegments.add(tailPath)
|
||||
}
|
||||
|
||||
/** Fsync every segment appended since the last sync. */
|
||||
async sync(): Promise<void> {
|
||||
if (this.dirtySegments.size === 0) return
|
||||
const paths = [...this.dirtySegments]
|
||||
this.dirtySegments.clear()
|
||||
await this.storage.syncRawObjects(paths)
|
||||
}
|
||||
|
||||
/**
|
||||
* Open a scan over committed facts. The scan runs against a MANIFEST
|
||||
* SNAPSHOT (sealed segments + the tail's decoded facts at open) — exactly-
|
||||
* once per fact, inclusive bounds, stable under concurrent appends. Gaps
|
||||
* abort LOUDLY: a missing generation inside a segment's declared range is
|
||||
* corruption, never silently skipped.
|
||||
*/
|
||||
scanFacts(options?: {
|
||||
fromGeneration?: number
|
||||
toGeneration?: number
|
||||
kinds?: Array<'noun' | 'verb'>
|
||||
batchSize?: number
|
||||
/** Test override for the first-batch liveness bound (default {@link SCANFACTS_FIRST_BATCH_MS}). */
|
||||
firstBatchTimeoutMs?: number
|
||||
}): FactScanHandle {
|
||||
const from = options?.fromGeneration ?? 1
|
||||
const to = options?.toGeneration ?? this.head
|
||||
const kinds = options?.kinds
|
||||
const batchSize = Math.max(1, options?.batchSize ?? 256)
|
||||
|
||||
// Snapshot: the segment list + tail content as of NOW.
|
||||
const segments = this.manifest.segments.filter(
|
||||
(s) => s.lastGeneration >= from && s.firstGeneration <= to
|
||||
)
|
||||
const tailSnapshot = this.tailFacts.filter((f) => f.generation >= from && f.generation <= to)
|
||||
const tailId = this.manifest.tailSegment ?? 'tail'
|
||||
const approxFactCount =
|
||||
segments.reduce((sum, s) => sum + s.facts, 0) + tailSnapshot.length
|
||||
|
||||
let factsYielded = 0
|
||||
let segmentsRead = 0
|
||||
const storage = this.storage
|
||||
|
||||
async function* batches(this: void): AsyncGenerator<FactScanBatch> {
|
||||
let expectedNext = 0 // gap detection: generations are monotonic, not necessarily dense
|
||||
const emit = (facts: CommitFact[], segmentId: string, byteSize: number): FactScanBatch => ({
|
||||
facts,
|
||||
firstGeneration: facts[0].generation,
|
||||
lastGeneration: facts[facts.length - 1].generation,
|
||||
factCount: facts.length,
|
||||
byteSize,
|
||||
segmentId
|
||||
})
|
||||
const filterOps = (fact: CommitFact): CommitFact =>
|
||||
kinds
|
||||
? { ...fact, ops: fact.ops.filter((op) => kinds.includes(op.kind)) }
|
||||
: fact
|
||||
|
||||
for (const entry of segments) {
|
||||
const bytes = await storage.readRawBytes(`${FACTS_PREFIX}/${entry.file}`)
|
||||
if (bytes === null) {
|
||||
throw new Error(
|
||||
`fact log: sealed segment ${entry.file} is MISSING — the log is damaged; aborting scan`
|
||||
)
|
||||
}
|
||||
const { facts } = parseSegment(entry.file, bytes)
|
||||
segmentsRead++
|
||||
const inRange = facts.filter((f) => f.generation >= from && f.generation <= to)
|
||||
for (const f of inRange) {
|
||||
if (f.generation <= expectedNext - 1) {
|
||||
throw new Error(`fact log: out-of-order fact ${f.generation} in ${entry.file} — aborting scan`)
|
||||
}
|
||||
expectedNext = f.generation + 1
|
||||
}
|
||||
for (let i = 0; i < inRange.length; i += batchSize) {
|
||||
const slice = inRange.slice(i, i + batchSize).map(filterOps)
|
||||
if (slice.length === 0) continue
|
||||
factsYielded += slice.length
|
||||
yield emit(slice, entry.file, slice.reduce((n, f) => n + encodeFrame(f).length, 0))
|
||||
}
|
||||
}
|
||||
|
||||
if (tailSnapshot.length > 0) {
|
||||
segmentsRead++
|
||||
for (const f of tailSnapshot) {
|
||||
if (f.generation <= expectedNext - 1) {
|
||||
throw new Error(`fact log: out-of-order fact ${f.generation} in the tail — aborting scan`)
|
||||
}
|
||||
expectedNext = f.generation + 1
|
||||
}
|
||||
for (let i = 0; i < tailSnapshot.length; i += batchSize) {
|
||||
const slice = tailSnapshot.slice(i, i + batchSize).map(filterOps)
|
||||
factsYielded += slice.length
|
||||
yield emit(slice, tailId, slice.reduce((n, f) => n + encodeFrame(f).length, 0))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Liveness wrapper: the FIRST pull races the contract deadline. Only the
|
||||
// first — the bound is time-to-first-batch (proof the producer is alive),
|
||||
// not per-batch pacing; and it runs only while a pull is actually pending,
|
||||
// so consumer think-time between pulls never counts against the producer.
|
||||
const firstBatchTimeoutMs = options?.firstBatchTimeoutMs ?? SCANFACTS_FIRST_BATCH_MS
|
||||
async function* batchesWithLiveness(this: void): AsyncGenerator<FactScanBatch> {
|
||||
const inner = batches()
|
||||
let timer: NodeJS.Timeout | undefined
|
||||
try {
|
||||
const deadline = new Promise<never>((_, reject) => {
|
||||
timer = setTimeout(
|
||||
() =>
|
||||
reject(
|
||||
new Error(
|
||||
`fact log: scanFacts produced no first batch within ${firstBatchTimeoutMs}ms ` +
|
||||
`(liveness contract) — the store is wedged or unreadably slow; aborting scan LOUDLY ` +
|
||||
`instead of hanging the consumer.`
|
||||
)
|
||||
),
|
||||
firstBatchTimeoutMs
|
||||
)
|
||||
timer.unref?.()
|
||||
})
|
||||
const first = await Promise.race([inner.next(), deadline])
|
||||
if (first.done) return
|
||||
yield first.value
|
||||
} finally {
|
||||
clearTimeout(timer)
|
||||
}
|
||||
yield* inner
|
||||
}
|
||||
|
||||
return {
|
||||
headGeneration: this.head,
|
||||
segmentCount: segments.length + (tailSnapshot.length > 0 ? 1 : 0),
|
||||
approxFactCount,
|
||||
batches: batchesWithLiveness,
|
||||
summary: () => ({ factsYielded, segmentsRead })
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The mmap fast path (capability handoff): the immutable sealed segment
|
||||
* files covering `fromGeneration`, in order. The TAIL is deliberately NOT
|
||||
* included — it is append-mutable; consumers read it via {@link scanFacts}.
|
||||
*/
|
||||
segmentPaths(options?: { fromGeneration?: number }): string[] {
|
||||
const from = options?.fromGeneration ?? 1
|
||||
return this.manifest.segments
|
||||
.filter((s) => s.lastGeneration >= from)
|
||||
.map((s) => `${FACTS_PREFIX}/${s.file}`)
|
||||
}
|
||||
|
||||
/**
|
||||
* Drop every fact with `generation > keepThrough` — the in-session abort
|
||||
* compensation: a transact appends its fact BEFORE the commit point, so a
|
||||
* real (non-crash) abort after the append must take the fact back out. The
|
||||
* dropped facts can only live in the TAIL (they were just appended); the
|
||||
* rewrite is atomic and bounded by the rotation threshold.
|
||||
*/
|
||||
async dropAbove(keepThrough: number): Promise<void> {
|
||||
if (this.head <= keepThrough) return
|
||||
const kept = this.tailFacts.filter((f) => f.generation <= keepThrough)
|
||||
if (kept.length === this.tailFacts.length) {
|
||||
throw new Error(
|
||||
`fact log: dropAbove(${keepThrough}) found no droppable facts in the tail ` +
|
||||
`(head ${this.head}) — the fact to drop was already sealed; the log needs reopen`
|
||||
)
|
||||
}
|
||||
await this.rewriteTail(kept)
|
||||
this.head = this.computeHead()
|
||||
}
|
||||
|
||||
// -- internals -------------------------------------------------------------
|
||||
|
||||
private computeHead(): number {
|
||||
if (this.tailFacts.length > 0) return this.tailFacts[this.tailFacts.length - 1].generation
|
||||
const sealed = this.manifest.segments
|
||||
if (sealed.length > 0) return sealed[sealed.length - 1].lastGeneration
|
||||
return 0
|
||||
}
|
||||
|
||||
/** Create the very first tail segment (manifest-first, then header bytes). */
|
||||
private async startTail(firstGeneration: number): Promise<void> {
|
||||
const file = segmentFileName(firstGeneration)
|
||||
this.manifest.tailSegment = file
|
||||
await this.persistManifest()
|
||||
await this.storage.appendRawBytes(`${FACTS_PREFIX}/${file}`, buildHeader(firstGeneration))
|
||||
this.tailFacts = []
|
||||
this.tailBytes = HEADER_BYTES
|
||||
}
|
||||
|
||||
/**
|
||||
* Seal the tail into the manifest and start a new one. Manifest-first: the
|
||||
* flip both seals the old tail AND names the new one atomically, so no
|
||||
* segment file ever exists unaccounted for.
|
||||
*/
|
||||
private async rotate(nextGeneration: number): Promise<void> {
|
||||
const sealedFile = this.manifest.tailSegment
|
||||
if (!sealedFile) return
|
||||
// Seal what the tail actually holds.
|
||||
await this.sync() // sealed segments are always fully durable
|
||||
const entry: SegmentEntry = {
|
||||
file: sealedFile,
|
||||
firstGeneration: this.tailFacts[0]?.generation ?? nextGeneration,
|
||||
lastGeneration: this.tailFacts[this.tailFacts.length - 1]?.generation ?? nextGeneration - 1,
|
||||
facts: this.tailFacts.length,
|
||||
bytes: this.tailBytes
|
||||
}
|
||||
const newFile = segmentFileName(nextGeneration)
|
||||
this.manifest.segments.push(entry)
|
||||
this.manifest.tailSegment = newFile
|
||||
await this.persistManifest()
|
||||
await this.storage.appendRawBytes(`${FACTS_PREFIX}/${newFile}`, buildHeader(nextGeneration))
|
||||
this.tailFacts = []
|
||||
this.tailBytes = HEADER_BYTES
|
||||
}
|
||||
|
||||
/** Atomically persist the manifest (write-new → fsync → rename downstream). */
|
||||
private async persistManifest(): Promise<void> {
|
||||
this.manifest.updatedAt = new Date().toISOString()
|
||||
await this.storage.writeRawObject(FACTS_MANIFEST_PATH, this.manifest)
|
||||
await this.storage.syncRawObjects([FACTS_MANIFEST_PATH])
|
||||
}
|
||||
|
||||
/** Rewrite the tail segment to hold exactly `facts` (atomic replace). */
|
||||
private async rewriteTail(facts: CommitFact[]): Promise<void> {
|
||||
const file = this.manifest.tailSegment
|
||||
if (!file) return
|
||||
const first = facts[0]?.generation ?? this.segmentFirstGenerationFromName(file)
|
||||
const parts: Uint8Array[] = [buildHeader(first)]
|
||||
for (const f of facts) parts.push(encodeFrame(f))
|
||||
const total = parts.reduce((n, p) => n + p.length, 0)
|
||||
const merged = new Uint8Array(total)
|
||||
let offset = 0
|
||||
for (const p of parts) {
|
||||
merged.set(p, offset)
|
||||
offset += p.length
|
||||
}
|
||||
await this.storage.writeRawBytes(`${FACTS_PREFIX}/${file}`, merged)
|
||||
this.tailFacts = facts
|
||||
this.tailBytes = total
|
||||
}
|
||||
|
||||
/** Cut a SEALED segment back to `committedGeneration` (atomic replace). */
|
||||
private async truncateSegmentTo(file: string, committedGeneration: number): Promise<void> {
|
||||
const path = `${FACTS_PREFIX}/${file}`
|
||||
const bytes = await this.storage.readRawBytes(path)
|
||||
if (bytes === null) return
|
||||
const { facts } = parseSegment(file, bytes)
|
||||
const kept = facts.filter((f) => f.generation <= committedGeneration)
|
||||
prodLog.warn(
|
||||
`[FactLog] truncating sealed segment ${file} to generation ${committedGeneration} ` +
|
||||
`(${facts.length - kept.length} uncommitted fact(s) dropped)`
|
||||
)
|
||||
const first = kept[0]?.generation ?? this.segmentFirstGenerationFromName(file)
|
||||
const parts: Uint8Array[] = [buildHeader(first)]
|
||||
for (const f of kept) parts.push(encodeFrame(f))
|
||||
const total = parts.reduce((n, p) => n + p.length, 0)
|
||||
const merged = new Uint8Array(total)
|
||||
let offset = 0
|
||||
for (const p of parts) {
|
||||
merged.set(p, offset)
|
||||
offset += p.length
|
||||
}
|
||||
await this.storage.writeRawBytes(path, merged)
|
||||
}
|
||||
|
||||
/** Re-derive a sealed segment's manifest entry from its actual bytes. */
|
||||
private async reloadSegmentEntry(file: string): Promise<SegmentEntry> {
|
||||
const bytes = await this.storage.readRawBytes(`${FACTS_PREFIX}/${file}`)
|
||||
const { facts, validBytes } = bytes
|
||||
? parseSegment(file, bytes)
|
||||
: { facts: [] as CommitFact[], validBytes: 0 }
|
||||
return {
|
||||
file,
|
||||
firstGeneration: facts[0]?.generation ?? this.segmentFirstGenerationFromName(file),
|
||||
lastGeneration: facts[facts.length - 1]?.generation ?? 0,
|
||||
facts: facts.length,
|
||||
bytes: validBytes
|
||||
}
|
||||
}
|
||||
|
||||
/** Parse the zero-padded firstGeneration back out of a segment filename. */
|
||||
private segmentFirstGenerationFromName(file: string): number {
|
||||
const match = /^seg-(\d{20})\.bfl$/.exec(file)
|
||||
return match ? Number(match[1]) : 0
|
||||
}
|
||||
}
|
||||
|
|
@ -1,152 +0,0 @@
|
|||
/**
|
||||
* @module db/familyStamp
|
||||
* @description The generalized FAMILY STAMP — one JSON shape that declares,
|
||||
* for any derived projection, WHICH source state it reflects and HOW to verify
|
||||
* it is whole. The entity tree (canonical current-state files) carries the
|
||||
* first brainy-side stamp; native index families carry the same shape. One
|
||||
* verifier reads both member modes:
|
||||
*
|
||||
* - `enumerated` — bounded families: exact byte size per member file,
|
||||
* verified at open.
|
||||
* - `rollup` — unbounded families (the entity tree: millions of files):
|
||||
* the verified surface is a small set of rollup invariants (entity/
|
||||
* relationship counts) plus `sourceGeneration`.
|
||||
*
|
||||
* `sourceGeneration` is the generation of the source-of-truth log this
|
||||
* projection reflects — open-time coherence becomes a COMPARISON (stamp vs
|
||||
* log head), not a walk:
|
||||
*
|
||||
* - equal + invariants hold → coherent, serve.
|
||||
* - behind → the projection missed the tail (crash between commit and stamp);
|
||||
* for the Stage-1 tree this is benign by construction (the tree is written
|
||||
* BY the commit), so the stamp refreshes; a DERIVED projection would replay
|
||||
* the gap instead.
|
||||
* - invariants FAIL at equal generation → genuine incoherence: loud, and the
|
||||
* repair ritual (`repairIndex()`, whose recount rebuilds the rollups from a
|
||||
* canonical walk) heals it.
|
||||
*
|
||||
* Stamps are JSON on purpose — every incident gets debugged by reading a
|
||||
* stamp in a terminal.
|
||||
*/
|
||||
|
||||
/** Storage-root-relative directory holding family stamps. */
|
||||
export const FAMILY_STAMPS_PREFIX = '_system/family-stamps'
|
||||
|
||||
/** The entity tree's stamp path. */
|
||||
export const ENTITY_TREE_STAMP_PATH = `${FAMILY_STAMPS_PREFIX}/entity-tree.json`
|
||||
|
||||
/** One enumerated member: a file and its exact expected byte size. */
|
||||
export interface EnumeratedMember {
|
||||
path: string
|
||||
bytes: number
|
||||
}
|
||||
|
||||
/**
|
||||
* The stamp's verified surface, in one of the two member modes. Rollup
|
||||
* invariant values may be numbers (counts, byte sizes) or strings (content
|
||||
* fingerprints, e.g. a per-tree SHA-256) — the verifier compares by strict
|
||||
* equality either way, so a type mismatch reads as incoherence, never a pass.
|
||||
*/
|
||||
export type StampMembers =
|
||||
| { mode: 'enumerated'; files: EnumeratedMember[] }
|
||||
| { mode: 'rollup'; invariants: Record<string, number | string> }
|
||||
|
||||
/** The generalized family stamp (one shape, one verifier, both engines). */
|
||||
export interface FamilyStamp {
|
||||
/** Which projection this stamps (e.g. `'entity-tree'`). */
|
||||
family: string
|
||||
/** Monotonic per-family stamp generation — bumps on every committed stamp. */
|
||||
generation: number
|
||||
/** ISO timestamp of the stamp write. */
|
||||
committedAt: string
|
||||
/** The source-of-truth generation this projection reflects. */
|
||||
sourceGeneration: number
|
||||
/** The verified surface. */
|
||||
members: StampMembers
|
||||
}
|
||||
|
||||
/** The verdict of an open-time stamp verification. */
|
||||
export type StampVerdict =
|
||||
| { state: 'coherent' }
|
||||
| { state: 'absent' } // legacy store — first stamp writes at the next flush
|
||||
| { state: 'behind'; stampSource: number; head: number }
|
||||
| { state: 'incoherent'; failures: string[] }
|
||||
| { state: 'unverifiable'; reason: string } // a FAULT reading the stamp — never conflated with absence
|
||||
|
||||
/** The narrow storage surface stamps ride (JSON objects + fsync). */
|
||||
export interface StampStorage {
|
||||
readRawObject(path: string): Promise<any | null>
|
||||
writeRawObject(path: string, data: any): Promise<void>
|
||||
syncRawObjects(paths: string[]): Promise<void>
|
||||
}
|
||||
|
||||
/** Read a family's stamp; `null` when none was ever written. */
|
||||
export async function readFamilyStamp(
|
||||
storage: StampStorage,
|
||||
path: string
|
||||
): Promise<FamilyStamp | null> {
|
||||
const stored = (await storage.readRawObject(path)) as FamilyStamp | null
|
||||
if (!stored || typeof stored !== 'object' || typeof stored.family !== 'string') return null
|
||||
return stored
|
||||
}
|
||||
|
||||
/** Write a family's stamp durably (atomic object write + fsync). */
|
||||
export async function writeFamilyStamp(
|
||||
storage: StampStorage,
|
||||
path: string,
|
||||
stamp: Omit<FamilyStamp, 'generation' | 'committedAt'> & { generation?: number }
|
||||
): Promise<void> {
|
||||
const prior = await readFamilyStamp(storage, path)
|
||||
const full: FamilyStamp = {
|
||||
...stamp,
|
||||
generation: (prior?.generation ?? 0) + 1,
|
||||
committedAt: new Date().toISOString()
|
||||
}
|
||||
await storage.writeRawObject(path, full)
|
||||
await storage.syncRawObjects([path])
|
||||
}
|
||||
|
||||
/**
|
||||
* The ONE verifier, both member modes. `actual` supplies the observed rollup
|
||||
* values (rollup mode) or file sizes (enumerated mode, keyed by path);
|
||||
* `head` is the source-of-truth generation now.
|
||||
*/
|
||||
export function verifyFamilyStamp(
|
||||
stamp: FamilyStamp | null,
|
||||
head: number,
|
||||
actual: Record<string, number | string>
|
||||
): StampVerdict {
|
||||
if (stamp === null) return { state: 'absent' }
|
||||
if (stamp.sourceGeneration > head) {
|
||||
// A stamp AHEAD of the log claims state that never committed — the
|
||||
// projection was stamped against truth that a crash rolled back.
|
||||
return {
|
||||
state: 'incoherent',
|
||||
failures: [`sourceGeneration ${stamp.sourceGeneration} is ahead of the log head ${head}`]
|
||||
}
|
||||
}
|
||||
if (stamp.sourceGeneration < head) {
|
||||
return { state: 'behind', stampSource: stamp.sourceGeneration, head }
|
||||
}
|
||||
const failures: string[] = []
|
||||
if (stamp.members.mode === 'rollup') {
|
||||
for (const [name, expected] of Object.entries(stamp.members.invariants)) {
|
||||
const observed = actual[name]
|
||||
if (observed === undefined) {
|
||||
failures.push(`rollup invariant '${name}' has no observed value`)
|
||||
} else if (observed !== expected) {
|
||||
failures.push(`rollup invariant '${name}': stamped ${expected}, observed ${observed}`)
|
||||
}
|
||||
}
|
||||
} else {
|
||||
for (const member of stamp.members.files) {
|
||||
const observed = actual[member.path]
|
||||
if (observed === undefined) {
|
||||
failures.push(`member '${member.path}' is missing`)
|
||||
} else if (observed !== member.bytes) {
|
||||
failures.push(`member '${member.path}': stamped ${member.bytes} bytes, observed ${observed}`)
|
||||
}
|
||||
}
|
||||
}
|
||||
return failures.length > 0 ? { state: 'incoherent', failures } : { state: 'coherent' }
|
||||
}
|
||||
|
|
@ -1,459 +0,0 @@
|
|||
/**
|
||||
* @module db/generationSegments
|
||||
* @description The generation-segment store — Stage-2 D1+D3+repacking's file
|
||||
* format (co-frozen 2026-07-19; design: the d1-d3-repacking spec).
|
||||
*
|
||||
* Packs CONSECUTIVE cold generations' record-sets (before-images + delta)
|
||||
* into append-once segment files with derived sidecar indexes, so history
|
||||
* scales in SEGMENTS (tens) instead of FILES-PER-GENERATION (hundreds of
|
||||
* thousands), and cold-open reads ONE manifest instead of listing the
|
||||
* backlog. Layout under `_generations/segments/`:
|
||||
*
|
||||
* - `seg-<firstGen, zero-padded 20>.bgs` — magic "BGS1", then one frame per
|
||||
* generation: `u32 payloadLen | u32 crc32c | msgpack payload`. Payload is
|
||||
* POSITIONAL: `[generation, timestamp, delta, records[], flags]` with
|
||||
* records `[kindByte, id, record]`. `flags` reserves encoding evolution
|
||||
* (bit 0 = compressed payload — v1 always 0; a future writer upgrade,
|
||||
* never a format break). Sealed segments are IMMUTABLE — the fact log's
|
||||
* own law, generalized.
|
||||
* - `seg-<firstGen>.idx` — DERIVED sidecar (msgpack): per-generation frame
|
||||
* offsets (point reads = one ranged read, never a listing) + per-id
|
||||
* generation postings (per-id chain rebuilds read only what they need).
|
||||
* Corrupt/missing → rebuilt from its segment in one sequential read,
|
||||
* loudly.
|
||||
* - `manifest.json` — the segment catalogue + `compactedBelow` (D3's
|
||||
* horizon marker). Cold-open reads THIS; the packed backlog is never
|
||||
* listed.
|
||||
*
|
||||
* D3 semantics carried here: bounded-retention reclaim drops WHOLE segments
|
||||
* at boundaries (O(1) per segment, no rewrite); under the archival profile
|
||||
* (`retention: 'all'`) nothing here is ever dropped — folding is the only
|
||||
* transform (re-representation, never deletion).
|
||||
*/
|
||||
|
||||
import { encode as msgpackEncode, decode as msgpackDecode } from '@msgpack/msgpack'
|
||||
import { crc32c } from '../utils/crc32c.js'
|
||||
import type { FactLogStorage } from './factLog.js'
|
||||
import { prodLog } from '../utils/logger.js'
|
||||
|
||||
/** Directory for segment files + manifest, under the generations prefix. */
|
||||
export const SEGMENTS_PREFIX = '_generations/segments'
|
||||
|
||||
/** Target sealed-segment size (co-freeze proposal; tunable on evidence). */
|
||||
export const SEGMENT_TARGET_BYTES = 64 * 1024 * 1024
|
||||
|
||||
const MAGIC = new TextEncoder().encode('BGS1')
|
||||
const FRAME_PREFIX_BYTES = 8 // u32 payloadLen + u32 crc32c
|
||||
const MANIFEST_PATH = `${SEGMENTS_PREFIX}/manifest.json`
|
||||
|
||||
/** One generation's fold input — exactly what the live tier holds for it. */
|
||||
export interface FoldGeneration {
|
||||
generation: number
|
||||
timestamp: number
|
||||
/** The tx.json delta object, carried verbatim. */
|
||||
delta: unknown
|
||||
/** The before-image record-set (empty for record-less generations). */
|
||||
records: Array<{ kind: 'noun' | 'verb'; id: string; record: unknown }>
|
||||
}
|
||||
|
||||
/** Manifest entry for one sealed segment. */
|
||||
export interface SegmentMeta {
|
||||
file: string
|
||||
firstGeneration: number
|
||||
lastGeneration: number
|
||||
frames: number
|
||||
bytes: number
|
||||
/** crc32c of the full segment byte stream — the digest chain's link. */
|
||||
checksum: number
|
||||
}
|
||||
|
||||
interface SegmentManifest {
|
||||
version: 1
|
||||
compactedBelow: number
|
||||
segments: SegmentMeta[]
|
||||
}
|
||||
|
||||
interface SidecarIndex {
|
||||
version: 1
|
||||
/** [generation, frameOffset, frameLen] ascending by generation. */
|
||||
generations: Array<[number, number, number]>
|
||||
/** `${kindByte}:${id}` → ascending generations holding a record for it. */
|
||||
ids: Record<string, number[]>
|
||||
}
|
||||
|
||||
const segmentFileName = (firstGeneration: number): string =>
|
||||
`seg-${String(firstGeneration).padStart(20, '0')}.bgs`
|
||||
const sidecarFileName = (firstGeneration: number): string =>
|
||||
`seg-${String(firstGeneration).padStart(20, '0')}.idx`
|
||||
|
||||
/**
|
||||
* The generation-segment store. Owns the packed tier ONLY — the live
|
||||
* per-generation tier and the routing between tiers belong to
|
||||
* `GenerationStore`. All mutating entry points here are called under the
|
||||
* generation store's commit mutex.
|
||||
*/
|
||||
export class GenerationSegmentStore {
|
||||
private readonly storage: FactLogStorage
|
||||
private manifest: SegmentManifest = { version: 1, compactedBelow: 0, segments: [] }
|
||||
/** Sidecar cache — segments are immutable, so entries never invalidate. */
|
||||
private readonly sidecars = new Map<string, SidecarIndex>()
|
||||
|
||||
constructor(storage: FactLogStorage) {
|
||||
this.storage = storage
|
||||
}
|
||||
|
||||
/** Load the manifest (ONE read — never a directory listing). */
|
||||
async open(): Promise<void> {
|
||||
const raw = (await this.storage.readRawObject(MANIFEST_PATH)) as SegmentManifest | null
|
||||
if (raw) {
|
||||
if (raw.version !== 1) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] manifest version ${String(raw.version)} is newer than this ` +
|
||||
`engine understands — refusing to serve partial history. Upgrade the engine.`
|
||||
)
|
||||
}
|
||||
this.manifest = raw
|
||||
}
|
||||
}
|
||||
|
||||
/** The packed tier's catalogue (ascending, immutable snapshot). */
|
||||
segments(): readonly SegmentMeta[] {
|
||||
return this.manifest.segments
|
||||
}
|
||||
|
||||
/** D3's horizon marker: generations below this were reclaimed (bounded profiles only). */
|
||||
compactedBelow(): number {
|
||||
return this.manifest.compactedBelow
|
||||
}
|
||||
|
||||
/** The covering sealed segment for `gen`, or null if it lives outside the packed tier. */
|
||||
private coveringSegment(gen: number): SegmentMeta | null {
|
||||
// Manifest is ascending and ranges never overlap — binary search.
|
||||
const segs = this.manifest.segments
|
||||
let lo = 0
|
||||
let hi = segs.length - 1
|
||||
while (lo <= hi) {
|
||||
const mid = (lo + hi) >> 1
|
||||
const s = segs[mid]
|
||||
if (gen < s.firstGeneration) hi = mid - 1
|
||||
else if (gen > s.lastGeneration) lo = mid + 1
|
||||
else return s
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
/** True when `gen` is packed (readable from this tier). */
|
||||
hasGeneration(gen: number): boolean {
|
||||
return this.coveringSegment(gen) !== null
|
||||
}
|
||||
|
||||
/**
|
||||
* Fold consecutive generations into ONE new sealed segment + sidecar and
|
||||
* append it to the manifest atomically. Caller guarantees: `gens` is
|
||||
* ascending, contiguous with the packed tier (first = last packed + 1 when
|
||||
* segments exist), and already durable in the live tier. Crash between the
|
||||
* segment write and the caller's live-tier delete leaves a DUPLICATE
|
||||
* representation — resolved live-tier-wins by the reader; never a gap.
|
||||
*/
|
||||
async fold(gens: FoldGeneration[]): Promise<SegmentMeta> {
|
||||
if (gens.length === 0) {
|
||||
throw new Error('[GenerationSegments] fold() requires at least one generation')
|
||||
}
|
||||
for (let i = 1; i < gens.length; i++) {
|
||||
if (gens[i].generation <= gens[i - 1].generation) {
|
||||
throw new Error('[GenerationSegments] fold() input must be strictly ascending')
|
||||
}
|
||||
}
|
||||
const last = this.manifest.segments[this.manifest.segments.length - 1]
|
||||
if (last && gens[0].generation <= last.lastGeneration) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] fold() overlaps the packed tier: ${gens[0].generation} ≤ ` +
|
||||
`sealed ${last.lastGeneration} — segments are immutable, never rewritten`
|
||||
)
|
||||
}
|
||||
|
||||
const first = gens[0].generation
|
||||
const file = segmentFileName(first)
|
||||
const sidecar: SidecarIndex = { version: 1, generations: [], ids: {} }
|
||||
|
||||
// Encode all frames, tracking offsets for the sidecar.
|
||||
const parts: Uint8Array[] = [MAGIC]
|
||||
let offset = MAGIC.length
|
||||
for (const g of gens) {
|
||||
const payload = msgpackEncode([
|
||||
g.generation,
|
||||
g.timestamp,
|
||||
g.delta,
|
||||
g.records.map((r) => [r.kind === 'noun' ? 0 : 1, r.id, r.record]),
|
||||
0 // flags: v1 = uncompressed
|
||||
])
|
||||
const frame = new Uint8Array(FRAME_PREFIX_BYTES + payload.length)
|
||||
const view = new DataView(frame.buffer)
|
||||
view.setUint32(0, payload.length, true)
|
||||
view.setUint32(4, crc32c(payload), true)
|
||||
frame.set(payload, FRAME_PREFIX_BYTES)
|
||||
sidecar.generations.push([g.generation, offset, frame.length])
|
||||
for (const r of g.records) {
|
||||
const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}`
|
||||
;(sidecar.ids[key] ??= []).push(g.generation)
|
||||
}
|
||||
parts.push(frame)
|
||||
offset += frame.length
|
||||
}
|
||||
const total = parts.reduce((n, p) => n + p.length, 0)
|
||||
const bytes = new Uint8Array(total)
|
||||
let at = 0
|
||||
for (const p of parts) {
|
||||
bytes.set(p, at)
|
||||
at += p.length
|
||||
}
|
||||
|
||||
const meta: SegmentMeta = {
|
||||
file,
|
||||
firstGeneration: first,
|
||||
lastGeneration: gens[gens.length - 1].generation,
|
||||
frames: gens.length,
|
||||
bytes: total,
|
||||
checksum: crc32c(bytes)
|
||||
}
|
||||
|
||||
// Durability order: segment + sidecar fsync'd BEFORE the manifest names
|
||||
// them (a crash before the manifest = invisible orphan files, harmless);
|
||||
// manifest last, atomically.
|
||||
const segPath = `${SEGMENTS_PREFIX}/${file}`
|
||||
const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(first)}`
|
||||
await this.storage.writeRawBytes(segPath, bytes)
|
||||
await this.storage.writeRawBytes(idxPath, msgpackEncode(sidecar))
|
||||
await this.storage.syncRawObjects([segPath, idxPath])
|
||||
const next: SegmentManifest = {
|
||||
...this.manifest,
|
||||
segments: [...this.manifest.segments, meta]
|
||||
}
|
||||
await this.storage.writeRawObject(MANIFEST_PATH, next)
|
||||
await this.storage.syncRawObjects([MANIFEST_PATH])
|
||||
this.manifest = next
|
||||
this.sidecars.set(file, sidecar)
|
||||
return meta
|
||||
}
|
||||
|
||||
/** Load (or rebuild, loudly) a segment's sidecar. */
|
||||
private async sidecarFor(meta: SegmentMeta): Promise<SidecarIndex> {
|
||||
const cached = this.sidecars.get(meta.file)
|
||||
if (cached) return cached
|
||||
const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(meta.firstGeneration)}`
|
||||
const raw = await this.storage.readRawBytes(idxPath)
|
||||
if (raw) {
|
||||
try {
|
||||
const idx = msgpackDecode(raw) as SidecarIndex
|
||||
if (idx.version === 1) {
|
||||
this.sidecars.set(meta.file, idx)
|
||||
return idx
|
||||
}
|
||||
} catch {
|
||||
// fall through to rebuild
|
||||
}
|
||||
}
|
||||
// Sidecars are DERIVED: rebuild from the segment, loudly — never serve
|
||||
// wrong offsets silently.
|
||||
prodLog.warn(
|
||||
`[GenerationSegments] sidecar for ${meta.file} missing or unreadable — rebuilding from the segment`
|
||||
)
|
||||
const rebuilt = await this.rebuildSidecar(meta)
|
||||
await this.storage.writeRawBytes(idxPath, msgpackEncode(rebuilt))
|
||||
this.sidecars.set(meta.file, rebuilt)
|
||||
return rebuilt
|
||||
}
|
||||
|
||||
/** One sequential read of the segment → a fresh sidecar. Verifies every frame CRC. */
|
||||
private async rebuildSidecar(meta: SegmentMeta): Promise<SidecarIndex> {
|
||||
const frames = await this.readAllFrames(meta)
|
||||
const idx: SidecarIndex = { version: 1, generations: [], ids: {} }
|
||||
for (const f of frames) {
|
||||
idx.generations.push([f.generation, f.offset, f.frameLen])
|
||||
for (const r of f.records) {
|
||||
const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}`
|
||||
;(idx.ids[key] ??= []).push(f.generation)
|
||||
}
|
||||
}
|
||||
return idx
|
||||
}
|
||||
|
||||
private decodeFrame(
|
||||
payload: Uint8Array
|
||||
): { generation: number; timestamp: number; delta: unknown; records: FoldGeneration['records'] } {
|
||||
const [generation, timestamp, delta, rawRecords] = msgpackDecode(payload) as [
|
||||
number,
|
||||
number,
|
||||
unknown,
|
||||
Array<[number, string, unknown]>,
|
||||
number
|
||||
]
|
||||
return {
|
||||
generation,
|
||||
timestamp,
|
||||
delta,
|
||||
records: rawRecords.map(([kindByte, id, record]) => ({
|
||||
kind: kindByte === 0 ? ('noun' as const) : ('verb' as const),
|
||||
id,
|
||||
record
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
private async readAllFrames(meta: SegmentMeta): Promise<
|
||||
Array<ReturnType<GenerationSegmentStore['decodeFrame']> & { offset: number; frameLen: number }>
|
||||
> {
|
||||
const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`)
|
||||
if (!bytes) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] sealed segment ${meta.file} is MISSING — packed history is damaged; ` +
|
||||
`refusing to continue silently`
|
||||
)
|
||||
}
|
||||
const out: Array<ReturnType<GenerationSegmentStore['decodeFrame']> & { offset: number; frameLen: number }> = []
|
||||
let at = MAGIC.length
|
||||
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength)
|
||||
while (at + FRAME_PREFIX_BYTES <= bytes.length) {
|
||||
const payloadLen = view.getUint32(at, true)
|
||||
const crc = view.getUint32(at + 4, true)
|
||||
const payload = bytes.subarray(at + FRAME_PREFIX_BYTES, at + FRAME_PREFIX_BYTES + payloadLen)
|
||||
if (payload.length !== payloadLen || crc32c(payload) !== crc) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] frame CRC mismatch in ${meta.file} at offset ${at} — ` +
|
||||
`packed history is damaged; refusing to serve it`
|
||||
)
|
||||
}
|
||||
out.push({ ...this.decodeFrame(payload), offset: at, frameLen: FRAME_PREFIX_BYTES + payloadLen })
|
||||
at += FRAME_PREFIX_BYTES + payloadLen
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
/** Read one packed generation's frame via its sidecar offset (one ranged read). */
|
||||
private async readFrame(
|
||||
gen: number
|
||||
): Promise<ReturnType<GenerationSegmentStore['decodeFrame']> | null> {
|
||||
const meta = this.coveringSegment(gen)
|
||||
if (!meta) return null
|
||||
const idx = await this.sidecarFor(meta)
|
||||
// generations ascending → binary search.
|
||||
const gens = idx.generations
|
||||
let lo = 0
|
||||
let hi = gens.length - 1
|
||||
while (lo <= hi) {
|
||||
const mid = (lo + hi) >> 1
|
||||
if (gens[mid][0] < gen) lo = mid + 1
|
||||
else if (gens[mid][0] > gen) hi = mid - 1
|
||||
else {
|
||||
const [, offset, frameLen] = gens[mid]
|
||||
const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`)
|
||||
if (!bytes) {
|
||||
throw new Error(`[GenerationSegments] sealed segment ${meta.file} is MISSING`)
|
||||
}
|
||||
const frame = bytes.subarray(offset, offset + frameLen)
|
||||
const view = new DataView(frame.buffer, frame.byteOffset, frame.byteLength)
|
||||
const payloadLen = view.getUint32(0, true)
|
||||
const crc = view.getUint32(4, true)
|
||||
const payload = frame.subarray(FRAME_PREFIX_BYTES, FRAME_PREFIX_BYTES + payloadLen)
|
||||
if (payload.length !== payloadLen || crc32c(payload) !== crc) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] frame CRC mismatch for generation ${gen} in ${meta.file} — ` +
|
||||
`packed history is damaged; refusing to serve it`
|
||||
)
|
||||
}
|
||||
return this.decodeFrame(payload)
|
||||
}
|
||||
}
|
||||
// In the covering range but not present: the packed tier is dense by
|
||||
// construction (fold packs every generation it is handed, including
|
||||
// record-less ones) — absence inside a sealed range is damage.
|
||||
throw new Error(
|
||||
`[GenerationSegments] generation ${gen} is inside sealed segment ${meta.file}'s declared ` +
|
||||
`range but has no frame — packed history is damaged`
|
||||
)
|
||||
}
|
||||
|
||||
/** The packed tier's delta for `gen` (null = not packed). */
|
||||
async readDelta(gen: number): Promise<{ delta: unknown; timestamp: number } | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
return frame ? { delta: frame.delta, timestamp: frame.timestamp } : null
|
||||
}
|
||||
|
||||
/** The packed tier's full record-set for `gen` (null = not packed). */
|
||||
async readRecords(gen: number): Promise<FoldGeneration['records'] | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
return frame ? frame.records : null
|
||||
}
|
||||
|
||||
/** One packed before-image (null = not packed OR no record for the id in that generation). */
|
||||
async readRecord(gen: number, kind: 'noun' | 'verb', id: string): Promise<unknown | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
if (!frame) return null
|
||||
const hit = frame.records.find((r) => r.kind === kind && r.id === id)
|
||||
return hit ? hit.record : null
|
||||
}
|
||||
|
||||
/**
|
||||
* D3 reclaim: drop WHOLE segments whose lastGeneration < `belowGeneration`
|
||||
* and bump `compactedBelow`. Partial segments are never dropped — the
|
||||
* boundary waits. NEVER called under the archival profile (the caller
|
||||
* enforces retention semantics; this method only executes boundary drops).
|
||||
*/
|
||||
async dropSegmentsBelow(belowGeneration: number): Promise<{ dropped: number; compactedBelow: number }> {
|
||||
const keep: SegmentMeta[] = []
|
||||
const drop: SegmentMeta[] = []
|
||||
for (const s of this.manifest.segments) {
|
||||
;(s.lastGeneration < belowGeneration ? drop : keep).push(s)
|
||||
}
|
||||
if (drop.length === 0) {
|
||||
return { dropped: 0, compactedBelow: this.manifest.compactedBelow }
|
||||
}
|
||||
const compactedBelow = Math.max(
|
||||
this.manifest.compactedBelow,
|
||||
drop[drop.length - 1].lastGeneration + 1
|
||||
)
|
||||
// Manifest first (the drop is authoritative once named), then bytes —
|
||||
// a crash between leaves orphan segment files invisible to the manifest,
|
||||
// harmless and re-collectable.
|
||||
const next: SegmentManifest = { ...this.manifest, compactedBelow, segments: keep }
|
||||
await this.storage.writeRawObject(MANIFEST_PATH, next)
|
||||
await this.storage.syncRawObjects([MANIFEST_PATH])
|
||||
this.manifest = next
|
||||
for (const s of drop) {
|
||||
await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${s.file}`)
|
||||
await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${sidecarFileName(s.firstGeneration)}`)
|
||||
this.sidecars.delete(s.file)
|
||||
}
|
||||
return { dropped: drop.length, compactedBelow }
|
||||
}
|
||||
|
||||
/**
|
||||
* D8 rider — the packed portion of `generationDigest(g)`: a deterministic
|
||||
* crc32c chain over sealed-segment checksums fully below `g`, plus the
|
||||
* frame CRC of `g`'s own frame when `g` is mid-segment. O(segments), not
|
||||
* O(generations); identical history ⇒ identical digest on any machine.
|
||||
* The live-tier portion is composed by the caller.
|
||||
*/
|
||||
async digestThroughPacked(g: number): Promise<number | null> {
|
||||
let digest = 0
|
||||
let covered = false
|
||||
for (const s of this.manifest.segments) {
|
||||
if (s.lastGeneration <= g) {
|
||||
digest = crc32c(new TextEncoder().encode(`${digest}:${s.checksum}`))
|
||||
if (s.lastGeneration === g) covered = true
|
||||
} else if (s.firstGeneration <= g) {
|
||||
// g is mid-segment: chain the partial prefix via g's frame CRC.
|
||||
const frame = await this.readFrame(g)
|
||||
if (frame === null) return null
|
||||
const idx = await this.sidecarFor(s)
|
||||
const upTo = idx.generations.filter(([gen]) => gen <= g)
|
||||
for (const [gen, offset, frameLen] of upTo) {
|
||||
digest = crc32c(new TextEncoder().encode(`${digest}:${gen}:${offset}:${frameLen}`))
|
||||
}
|
||||
covered = true
|
||||
break
|
||||
}
|
||||
}
|
||||
return covered || this.manifest.segments.length > 0 ? digest : null
|
||||
}
|
||||
}
|
||||
|
|
@ -45,9 +45,6 @@ import type {
|
|||
GenerationStorage,
|
||||
TxLogEntry
|
||||
} from './types.js'
|
||||
import { FactLog, storageSupportsFactLog, type CommitFact, type FactOp } from './factLog.js'
|
||||
import { GenerationSegmentStore, type FoldGeneration } from './generationSegments.js'
|
||||
import { crc32c } from '../utils/crc32c.js'
|
||||
|
||||
/**
|
||||
* The byte-identical before-images of every id a commit touches, read UNDER
|
||||
|
|
@ -124,16 +121,6 @@ export interface GenerationStoreOpenResult {
|
|||
export class GenerationStore {
|
||||
private readonly storage: GenerationStorage
|
||||
|
||||
/**
|
||||
* The generation FACT LOG (dual-write transition) — an append-only,
|
||||
* CRC-framed record of every committed generation as an AFTER-IMAGE fact.
|
||||
* `null` when the storage layer lacks the binary raw-byte primitives.
|
||||
* Appends ride the same commit protocol: a fact-append failure FAILS the
|
||||
* write (loud — a silent fact gap would make the log a lie that a later
|
||||
* replay discovers), and open() reconciles the log to committed truth.
|
||||
*/
|
||||
private factLog: FactLog | null = null
|
||||
|
||||
/** Latest reserved/observed generation (≥ {@link committed}). */
|
||||
private counter = 0
|
||||
/** Committed-transaction watermark (manifest generation). */
|
||||
|
|
@ -259,30 +246,6 @@ export class GenerationStore {
|
|||
*/
|
||||
private deltaCacheMax = 4096
|
||||
|
||||
/**
|
||||
* Running total of on-disk history bytes across committed generations —
|
||||
* `null` until {@link historyBytes} pays its one seeding walk. Maintained
|
||||
* incrementally at commit/reclaim so the adaptive retention check on every
|
||||
* flush() is O(1), never a tail re-walk. Never updated by cache re-reads
|
||||
* ({@link setDelta} inserts are cache population, not new history).
|
||||
*/
|
||||
private historyBytesTotal: number | null = null
|
||||
|
||||
/**
|
||||
* The packed tier (D1+D3): sealed segments holding folded cold
|
||||
* generations. Null until {@link open} wires it (and on storage adapters
|
||||
* without raw-byte primitives — the live tier then carries everything,
|
||||
* exactly as before the packed tier existed).
|
||||
*/
|
||||
private segments: GenerationSegmentStore | null = null
|
||||
|
||||
/**
|
||||
* Live-tier window: generations newer than `committed - REPACK_LIVE_WINDOW`
|
||||
* are never folded — the hot tail stays in the per-generation layout the
|
||||
* write path owns. Matches the resident chain window's scale.
|
||||
*/
|
||||
static readonly REPACK_LIVE_WINDOW = 1024
|
||||
|
||||
/**
|
||||
* Model-B per-write group-commit — the in-memory PENDING tier.
|
||||
*
|
||||
|
|
@ -437,46 +400,6 @@ export class GenerationStore {
|
|||
|
||||
this.opened = true
|
||||
|
||||
// Generation FACT LOG (dual-write transition): when the storage layer
|
||||
// exposes the binary raw-byte primitives, open the after-image fact log
|
||||
// and reconcile it to committed truth — facts are appended BEFORE the
|
||||
// commit point, so a crash can only leave the log AHEAD; open truncates
|
||||
// any fact beyond `committed`. Storage without the primitives simply
|
||||
// hosts no fact log (readers fall back to canonical enumeration).
|
||||
if (storageSupportsFactLog(this.storage)) {
|
||||
this.factLog = new FactLog(this.storage)
|
||||
await this.factLog.open(this.committed)
|
||||
} else {
|
||||
this.factLog = null
|
||||
}
|
||||
|
||||
// PACKED TIER (D1+D3): same capability gate as the fact log. Opening
|
||||
// reads ONE manifest — never a listing of the packed backlog — and seeds
|
||||
// committedRanges with the sealed ranges so packed generations resolve
|
||||
// exactly like live ones.
|
||||
if (storageSupportsFactLog(this.storage)) {
|
||||
this.segments = new GenerationSegmentStore(this.storage)
|
||||
await this.segments.open()
|
||||
const packedRanges = this.segments
|
||||
.segments()
|
||||
.map((s): [number, number] => [s.firstGeneration, Math.min(s.lastGeneration, this.committed)])
|
||||
.filter(([lo, hi]) => lo <= hi)
|
||||
if (packedRanges.length > 0) {
|
||||
// Merge packed (older) + live (newer) interval sets — both ascending;
|
||||
// coalesce adjacency so range arithmetic stays interval-exact.
|
||||
const merged: Array<[number, number]> = []
|
||||
for (const r of [...packedRanges, ...this.committedRanges].sort((a, b) => a[0] - b[0])) {
|
||||
const last = merged[merged.length - 1]
|
||||
if (last && r[0] <= last[1] + 1) last[1] = Math.max(last[1], r[1])
|
||||
else merged.push([r[0], r[1]])
|
||||
}
|
||||
this.committedRanges = merged
|
||||
}
|
||||
this.horizonGen = Math.max(this.horizonGen, this.segments.compactedBelow() - 1)
|
||||
} else {
|
||||
this.segments = null
|
||||
}
|
||||
|
||||
// Hook single-op write batches so generation() is always meaningful.
|
||||
// Suppressed while a transact batch executes (the batch is ONE generation).
|
||||
if (!options?.readOnly) {
|
||||
|
|
@ -536,85 +459,6 @@ export class GenerationStore {
|
|||
return this.horizonGen
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Read-only history footprint for fleet audits: how much
|
||||
* generational history this store holds on disk. `bytes` pays (and seeds)
|
||||
* the one-time {@link historyBytes} walk on first call — subsequent calls
|
||||
* are O(1). The oldest/newest timestamps come from those generations'
|
||||
* deltas (cache-bounded reads).
|
||||
* @returns Counts, bytes, generation range, and the compaction horizon.
|
||||
*/
|
||||
/**
|
||||
* @description D8 (gate-to-generation provenance): a deterministic content
|
||||
* digest of the generation log THROUGH `g` — identical history ⇒ identical
|
||||
* digest on any machine; any divergence (different records, different
|
||||
* order, reclaimed range) ⇒ different digest. Composed from the packed
|
||||
* tier's sealed-segment checksum chain (O(segments)) plus the live tier's
|
||||
* per-generation delta digests (O(live window at most)). Release gates pin
|
||||
* {generation, digest} and verify both at execution time.
|
||||
* @param g - The generation to digest through (≤ committed).
|
||||
* @returns A hex digest string, stable across reopen and repacking states
|
||||
* ONLY for fully-packed prefixes — repacking changes representation, so
|
||||
* the composed digest is defined over CONTENT: live-tier gens hash their
|
||||
* delta + record ids, packed gens hash via frame CRCs. A gate should pin
|
||||
* after a repack pass for long-term stability, or re-pin on repack.
|
||||
*/
|
||||
async generationDigest(g: number): Promise<string> {
|
||||
if (!Number.isInteger(g) || g < 1 || g > this.committed) {
|
||||
throw new RangeError(
|
||||
`generationDigest(): generation ${g} is out of range [1, ${this.committed}]`
|
||||
)
|
||||
}
|
||||
if (g <= this.horizonGen) {
|
||||
throw new GenerationCompactedError(g, this.horizonGen)
|
||||
}
|
||||
let digest = 0
|
||||
const enc = new TextEncoder()
|
||||
if (this.segments) {
|
||||
const packed = await this.segments.digestThroughPacked(g)
|
||||
if (packed !== null) digest = packed
|
||||
}
|
||||
// Live-tier composition: every committed gen ≤ g not covered by a sealed
|
||||
// segment hashes its delta content in ascending order.
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
if (gen > g) break
|
||||
if (this.segments?.hasGeneration(gen)) continue
|
||||
const delta = await this.getDelta(gen)
|
||||
digest = crc32c(
|
||||
enc.encode(
|
||||
`${digest}:${gen}:${delta.timestamp}:${[...delta.nouns].sort().join(',')}:${[...delta.verbs].sort().join(',')}`
|
||||
)
|
||||
)
|
||||
}
|
||||
return digest.toString(16).padStart(8, '0')
|
||||
}
|
||||
|
||||
async historyStats(): Promise<{
|
||||
generations: number
|
||||
bytes: number
|
||||
oldestGeneration: number | null
|
||||
newestGeneration: number | null
|
||||
oldestTimestamp: number | null
|
||||
newestTimestamp: number | null
|
||||
horizon: number
|
||||
}> {
|
||||
let oldest: number | null = null
|
||||
let newest: number | null = null
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
if (oldest === null) oldest = gen
|
||||
newest = gen
|
||||
}
|
||||
return {
|
||||
generations: this.committedCount(),
|
||||
bytes: await this.historyBytes(),
|
||||
oldestGeneration: oldest,
|
||||
newestGeneration: newest,
|
||||
oldestTimestamp: oldest !== null ? (await this.getDelta(oldest)).timestamp : null,
|
||||
newestTimestamp: newest !== null ? (await this.getDelta(newest)).timestamp : null,
|
||||
horizon: this.horizonGen
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Read one generation's persisted before-image records — the
|
||||
* compaction fallback for generations written before deltas carried
|
||||
|
|
@ -627,17 +471,14 @@ export class GenerationStore {
|
|||
try {
|
||||
paths = await this.storage.listRawObjects(`${GENERATIONS_PREFIX}/${gen}/prev`)
|
||||
} catch {
|
||||
paths = []
|
||||
return []
|
||||
}
|
||||
const records: GenerationRecord[] = []
|
||||
for (const p of paths) {
|
||||
const record = (await this.storage.readRawObject(p)) as GenerationRecord | null
|
||||
if (record) records.push(record)
|
||||
}
|
||||
if (records.length > 0) return records
|
||||
// Two-tier: folded generations serve their record-set from the segment.
|
||||
const packed = await this.segments?.readRecords(gen)
|
||||
return packed ? (packed.map((r) => r.record) as GenerationRecord[]) : []
|
||||
return records
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -759,58 +600,6 @@ export class GenerationStore {
|
|||
* @returns The committed generation and its commit timestamp.
|
||||
* @throws GenerationConflictError when the CAS expectation fails.
|
||||
*/
|
||||
/**
|
||||
* The generation fact log, or `null` when the storage layer cannot host one.
|
||||
* Consumers scan committed facts through it (`scanFacts` / `segmentPaths`).
|
||||
*/
|
||||
getFactLog(): FactLog | null {
|
||||
return this.factLog
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Build one commit's AFTER-IMAGE fact by reading canonical
|
||||
* state back for every touched id — under the commit mutex, immediately
|
||||
* after the operations applied, so canonical IS the after-image (and the
|
||||
* reads are page-cache-warm: the operations just wrote these files). An
|
||||
* absent id (both legs null) becomes a body-less TOMBSTONE — the delete
|
||||
* fact needs no body, so removal never requires reading the removed thing.
|
||||
* The fact's blobHashes are extracted from the AFTER records (the content
|
||||
* this generation's state references), unlike the history path's
|
||||
* before-image hashes.
|
||||
*/
|
||||
private async buildCommitFact(args: {
|
||||
generation: number
|
||||
timestamp: number
|
||||
nouns: string[]
|
||||
verbs: string[]
|
||||
meta?: Record<string, unknown>
|
||||
}): Promise<CommitFact> {
|
||||
const ops: FactOp[] = []
|
||||
const afterRecords: GenerationRecord[] = []
|
||||
for (const id of args.nouns) {
|
||||
const after = await this.storage.readNounRaw(id)
|
||||
const absent = after.metadata === null && after.vector === null
|
||||
ops.push({ kind: 'noun', id, record: absent ? null : after })
|
||||
if (!absent) afterRecords.push({ kind: 'noun', metadata: after.metadata, vector: after.vector })
|
||||
}
|
||||
for (const id of args.verbs) {
|
||||
const after = await this.storage.readVerbRaw(id)
|
||||
const absent = after.metadata === null && after.vector === null
|
||||
ops.push({ kind: 'verb', id, record: absent ? null : after })
|
||||
if (!absent) afterRecords.push({ kind: 'verb', metadata: after.metadata, vector: after.vector })
|
||||
}
|
||||
const blobHashes = this.storage.extractBlobHashesFromRecords
|
||||
? this.storage.extractBlobHashesFromRecords(afterRecords)
|
||||
: []
|
||||
return {
|
||||
generation: args.generation,
|
||||
timestamp: args.timestamp,
|
||||
ops,
|
||||
...(args.meta ? { meta: args.meta } : {}),
|
||||
...(blobHashes.length > 0 ? { blobHashes } : {})
|
||||
}
|
||||
}
|
||||
|
||||
async commitTransaction(args: {
|
||||
touched: TouchedIds
|
||||
meta?: Record<string, unknown>
|
||||
|
|
@ -944,24 +733,6 @@ export class GenerationStore {
|
|||
await this.storage.flushWriteBarrier?.()
|
||||
faultPoint('after-execute')
|
||||
|
||||
// Fact log (dual-write): append + fsync this generation's AFTER-IMAGE
|
||||
// fact BEFORE the commit point, inside the same durability window —
|
||||
// so a crash can only leave the log AHEAD (open truncates), never a
|
||||
// committed generation without its fact. A real abort below this point
|
||||
// compensates via dropAbove in the catch. An append failure fails the
|
||||
// write, loudly — a silent fact gap would be a lie a replay discovers.
|
||||
if (this.factLog) {
|
||||
const fact = await this.buildCommitFact({
|
||||
generation: gen,
|
||||
timestamp,
|
||||
nouns,
|
||||
verbs,
|
||||
...(args.meta ? { meta: args.meta } : {})
|
||||
})
|
||||
await this.factLog.append(fact)
|
||||
await this.factLog.sync()
|
||||
}
|
||||
|
||||
// -- 5. Counter + manifest rename (COMMIT POINT) ----------------------
|
||||
await this.persistCounterUnlocked()
|
||||
faultPoint('before-manifest-rename')
|
||||
|
|
@ -984,9 +755,6 @@ export class GenerationStore {
|
|||
timestamp,
|
||||
bytes: delta.bytes ?? 0
|
||||
})
|
||||
if (this.historyBytesTotal !== null) {
|
||||
this.historyBytesTotal += delta.bytes ?? 0
|
||||
}
|
||||
this.extendChains(gen, nouns, verbs)
|
||||
const logEntry: TxLogEntry = { generation: gen, timestamp, ...(args.meta && { meta: args.meta }) }
|
||||
await this.storage.appendTxLogLine(JSON.stringify(logEntry))
|
||||
|
|
@ -1032,13 +800,6 @@ export class GenerationStore {
|
|||
// over-count-safe; the scrub restores exactness
|
||||
}
|
||||
}
|
||||
// Fact-log compensation: a real (non-crash) abort after the fact was
|
||||
// appended must take the fact back out — the generation never
|
||||
// committed. A crash instead reaches open(), whose truncation does the
|
||||
// same reconcile from disk.
|
||||
if (this.factLog && this.factLog.headGeneration() >= gen) {
|
||||
await this.factLog.dropAbove(gen - 1)
|
||||
}
|
||||
// Return the reservation when no concurrent bump consumed a later
|
||||
// number, so a failed transaction leaves generation() unchanged.
|
||||
if (this.counter === gen) this.counter = gen - 1
|
||||
|
|
@ -1230,14 +991,6 @@ export class GenerationStore {
|
|||
this.pendingBuffer.set(gen, { nouns: nounBefore, verbs: verbBefore, timestamp })
|
||||
this.pendingGens.push(gen)
|
||||
this.extendChains(gen, nouns, verbs)
|
||||
// The adopted generation is committed — it gets its fact like any
|
||||
// other (durability rides the group-commit flush, same as the
|
||||
// buffered history).
|
||||
if (this.factLog) {
|
||||
await this.factLog.append(
|
||||
await this.buildCommitFact({ generation: gen, timestamp, nouns, verbs })
|
||||
)
|
||||
}
|
||||
prodLog.warn(
|
||||
`[GenerationStore] Recovered a failed rollback FORWARD: single-op write ` +
|
||||
`committed as generation ${gen} because its canonical undo could not be ` +
|
||||
|
|
@ -1267,17 +1020,6 @@ export class GenerationStore {
|
|||
this.pendingBuffer.set(gen, { nouns: nounBefore, verbs: verbBefore, timestamp })
|
||||
this.pendingGens.push(gen)
|
||||
this.extendChains(gen, nouns, verbs)
|
||||
// Fact log (dual-write): the acked write's AFTER-IMAGE fact, appended
|
||||
// now (read back warm, under the mutex — group-commit means flush-time
|
||||
// canonical only holds the LATEST state, so each generation's after-image
|
||||
// exists only here). Durability rides the group-commit flush, exactly
|
||||
// like the buffered before-image history: a crash before the flush loses
|
||||
// the fact AND the generation together — never a torn state.
|
||||
if (this.factLog) {
|
||||
await this.factLog.append(
|
||||
await this.buildCommitFact({ generation: gen, timestamp, nouns, verbs })
|
||||
)
|
||||
}
|
||||
this.schedulePendingFlush()
|
||||
return { generation: gen, timestamp }
|
||||
})
|
||||
|
|
@ -1399,12 +1141,6 @@ export class GenerationStore {
|
|||
// ONE fsync for the whole window — the durability-batching win.
|
||||
await this.storage.syncRawObjects(stagedPaths)
|
||||
|
||||
// Fact log (dual-write): make the window's buffered facts durable in the
|
||||
// same batch, BEFORE the commit point below — so a crash can only leave
|
||||
// the log AHEAD of the counter (open truncates), never a committed
|
||||
// generation without its durable fact.
|
||||
await this.factLog?.sync()
|
||||
|
||||
// Test-only crash simulation: a throwing injector here leaves the staged
|
||||
// group-commit generation dirs on disk with NO manifest advance — the
|
||||
// exact "crashed mid-flush" state recovery must DROP-WITHOUT-RESTORE
|
||||
|
|
@ -1440,9 +1176,6 @@ export class GenerationStore {
|
|||
timestamp: buf.timestamp,
|
||||
bytes: genBytes.get(gen) ?? 0
|
||||
})
|
||||
if (this.historyBytesTotal !== null) {
|
||||
this.historyBytesTotal += genBytes.get(gen) ?? 0
|
||||
}
|
||||
this.pendingBuffer.delete(gen)
|
||||
}
|
||||
this.pendingGens = []
|
||||
|
|
@ -1875,15 +1608,9 @@ export class GenerationStore {
|
|||
if (pending) {
|
||||
return (kind === 'noun' ? pending.nouns : pending.verbs).get(id) ?? null
|
||||
}
|
||||
const live = (await this.storage.readRawObject(
|
||||
return (await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/prev/${id}.json`
|
||||
)) as GenerationRecord | null
|
||||
if (live) return live
|
||||
// Two-tier: the packed tier serves folded generations (live-tier-wins).
|
||||
if (this.segments?.hasGeneration(gen)) {
|
||||
return (await this.segments.readRecord(gen, kind, id)) as GenerationRecord | null
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -2230,21 +1957,6 @@ export class GenerationStore {
|
|||
`${GENERATIONS_PREFIX}/${gen}/tx.json`
|
||||
)) as GenerationDelta | null
|
||||
if (delta === null) {
|
||||
// Two-tier read (D1+D3): not in the live tier → the packed tier.
|
||||
// Live-tier-wins ordering (a crash mid-fold leaves a duplicate, never
|
||||
// a gap), so the segment lookup runs only after the live miss.
|
||||
const packed = await this.segments?.readDelta(gen)
|
||||
if (packed) {
|
||||
const d = packed.delta as GenerationDelta
|
||||
const entry = {
|
||||
nouns: new Set(d.nouns),
|
||||
verbs: new Set(d.verbs),
|
||||
timestamp: packed.timestamp,
|
||||
bytes: d.bytes ?? 0
|
||||
}
|
||||
this.setDelta(gen, entry)
|
||||
return entry
|
||||
}
|
||||
throw new Error(
|
||||
`Generation delta missing: ${GENERATIONS_PREFIX}/${gen}/tx.json ` +
|
||||
`(store corrupted or records removed outside compactHistory())`
|
||||
|
|
@ -2284,26 +1996,17 @@ export class GenerationStore {
|
|||
/**
|
||||
* @description Total serialized bytes of the ON-DISK generational history —
|
||||
* the sum of every committed generation's recorded `bytes`. Backs the
|
||||
* `maxBytes` and adaptive retention caps. O(1) after the first call: the
|
||||
* total is computed by ONE walk over committed deltas, then maintained
|
||||
* incrementally at every commit (+bytes) and reclaim (−bytes) and dropped on
|
||||
* a wholesale state replacement (restore). Without the running total, the
|
||||
* adaptive auto-compaction on every flush() re-walked the ENTIRE history —
|
||||
* O(committed generations) file reads per flush past the delta-cache bound —
|
||||
* which is how a 70k-generation production brain turned every write into a
|
||||
* full-tail scan (SELF-GENERATIONS-GROWTH). Pending (un-flushed) generations
|
||||
* are excluded (they are not on disk).
|
||||
* `maxBytes` and adaptive retention caps. Reads each committed generation's
|
||||
* delta (cached; a re-read only for cache-evicted ones) — O(committed
|
||||
* generations), bounded by retention itself and invoked only at compaction
|
||||
* time. Pending (un-flushed) generations are excluded (they are not on disk).
|
||||
* @returns The total on-disk history byte count.
|
||||
*/
|
||||
async historyBytes(): Promise<number> {
|
||||
if (this.historyBytesTotal !== null) {
|
||||
return this.historyBytesTotal
|
||||
}
|
||||
let total = 0
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
total += (await this.getDelta(gen)).bytes
|
||||
}
|
||||
this.historyBytesTotal = total
|
||||
return total
|
||||
}
|
||||
|
||||
|
|
@ -2326,94 +2029,6 @@ export class GenerationStore {
|
|||
* @param options - Retention caps (see {@link CompactHistoryOptions}).
|
||||
* @returns Count of removed record-sets and the new horizon.
|
||||
*/
|
||||
/**
|
||||
* @description The REPACKER (D1+D3+repacking): fold cold live-tier
|
||||
* generations into sealed segments — re-representation, never deletion.
|
||||
* Every record and delta stays readable (asOf/chains unchanged); the
|
||||
* per-generation directories are deleted only AFTER their segment is
|
||||
* durable (crash between = duplicate representation, resolved
|
||||
* live-tier-wins by every reader; never a gap). This is the transform that
|
||||
* takes a 70k-file history to tens of segment files, and the ONLY history
|
||||
* transform permitted under the archival profile.
|
||||
*
|
||||
* Folds oldest-first, contiguous from the packed boundary, in batches, and
|
||||
* stops at the live window ({@link GenerationStore.REPACK_LIVE_WINDOW})
|
||||
* or when `timeBudgetMs` is spent — an early stop is a consistent prefix;
|
||||
* the next pass resumes.
|
||||
*/
|
||||
async repackHistory(options?: { timeBudgetMs?: number; batchGenerations?: number }): Promise<{
|
||||
foldedGenerations: number
|
||||
segmentsCreated: number
|
||||
}> {
|
||||
if (!this.segments) return { foldedGenerations: 0, segmentsCreated: 0 }
|
||||
const segments = this.segments
|
||||
return this.withMutex(async () => {
|
||||
const deadline =
|
||||
options?.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined
|
||||
const batchSize = options?.batchGenerations ?? 512
|
||||
const coldCeiling = this.committed - GenerationStore.REPACK_LIVE_WINDOW
|
||||
const packedThrough =
|
||||
segments.segments().length > 0
|
||||
? segments.segments()[segments.segments().length - 1].lastGeneration
|
||||
: 0
|
||||
|
||||
// Cold, unpacked, committed generations — ascending, contiguous scan.
|
||||
const eligible: number[] = []
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
if (gen > coldCeiling) break
|
||||
if (gen <= packedThrough) continue // already packed (dup fold barred)
|
||||
if (this.pendingBuffer.has(gen)) continue // un-flushed = live by definition
|
||||
eligible.push(gen)
|
||||
}
|
||||
|
||||
let folded = 0
|
||||
let segmentsCreated = 0
|
||||
for (let i = 0; i < eligible.length; i += batchSize) {
|
||||
if (deadline !== undefined && Date.now() >= deadline) break
|
||||
const batch = eligible.slice(i, i + batchSize)
|
||||
const foldInput: FoldGeneration[] = []
|
||||
for (const gen of batch) {
|
||||
const delta = (await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/tx.json`
|
||||
)) as GenerationDelta | null
|
||||
if (delta === null) {
|
||||
// Already folded by a prior crashed pass whose dirs were removed,
|
||||
// or damage — getDelta's two-tier read decides which, loudly,
|
||||
// when someone asks. Skip; never fold a generation we cannot read.
|
||||
continue
|
||||
}
|
||||
const records: FoldGeneration['records'] = []
|
||||
for (const [kind, ids] of [
|
||||
['noun', delta.nouns] as const,
|
||||
['verb', delta.verbs] as const
|
||||
]) {
|
||||
for (const id of ids) {
|
||||
const record = await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/prev/${id}.json`
|
||||
)
|
||||
if (record) records.push({ kind, id, record })
|
||||
}
|
||||
}
|
||||
foldInput.push({ generation: gen, timestamp: delta.timestamp, delta, records })
|
||||
}
|
||||
if (foldInput.length === 0) continue
|
||||
await segments.fold(foldInput)
|
||||
segmentsCreated++
|
||||
// Segment + manifest durable → the live copies retire.
|
||||
for (const g of foldInput) {
|
||||
await this.storage.removeRawPrefix(`${GENERATIONS_PREFIX}/${g.generation}`)
|
||||
}
|
||||
folded += foldInput.length
|
||||
}
|
||||
if (folded > 0) {
|
||||
prodLog.info(
|
||||
`[GenerationStore] repacked ${folded} cold generation(s) into ${segmentsCreated} segment(s) — history preserved, file count reduced`
|
||||
)
|
||||
}
|
||||
return { foldedGenerations: folded, segmentsCreated }
|
||||
})
|
||||
}
|
||||
|
||||
async compact(options?: CompactHistoryOptions): Promise<CompactHistoryResult> {
|
||||
return this.withMutex(async () => {
|
||||
const minPinned = this.minPinnedGeneration()
|
||||
|
|
@ -2421,11 +2036,6 @@ export class GenerationStore {
|
|||
const maxAge = options?.maxAge
|
||||
const maxBytes = options?.maxBytes
|
||||
const ageCutoff = maxAge !== undefined ? Date.now() - maxAge : undefined
|
||||
// Bounded maintenance pass (8.9.0): stop reclaiming once the budget is
|
||||
// spent. Safe mid-loop — reclamation is oldest-first, so an early stop
|
||||
// leaves a consistent contiguous prefix and the next pass resumes.
|
||||
const deadline =
|
||||
options?.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined
|
||||
const noCaps =
|
||||
maxGenerations === undefined && maxAge === undefined && maxBytes === undefined
|
||||
|
||||
|
|
@ -2439,7 +2049,6 @@ export class GenerationStore {
|
|||
for (const gen of [...this.committedGensAsc()]) {
|
||||
// Pins are always exempt: never reclaim a generation a live pin needs.
|
||||
if (gen > minPinned) break // committedGensAsc ascending → nothing newer is eligible either
|
||||
if (deadline !== undefined && Date.now() >= deadline) break // budget spent — resume next pass
|
||||
const delta = await this.getDelta(gen)
|
||||
if (!noCaps) {
|
||||
const violatesCount = maxGenerations !== undefined && remainingCount > maxGenerations
|
||||
|
|
@ -2471,9 +2080,6 @@ export class GenerationStore {
|
|||
|
||||
await this.storage.removeRawPrefix(`${GENERATIONS_PREFIX}/${gen}`)
|
||||
this.deltaCache.delete(gen)
|
||||
if (this.historyBytesTotal !== null) {
|
||||
this.historyBytesTotal -= delta.bytes
|
||||
}
|
||||
|
||||
// AFTER the record-set is gone (over-count-only crash ordering):
|
||||
// release its history references and reclaim any blob left with zero
|
||||
|
|
@ -2505,16 +2111,6 @@ export class GenerationStore {
|
|||
// Reclaimed generations leave the per-id chains stale → rebuild on next read.
|
||||
this.invalidateChains()
|
||||
this.horizonGen = Math.max(this.horizonGen, highestRemoved)
|
||||
// Packed-tier reclaim (D3): a packed generation's bytes live in a
|
||||
// sealed segment — removeRawPrefix above was a no-op for it. Drop
|
||||
// WHOLE segments now fully below the horizon; a partially-reclaimed
|
||||
// segment keeps its bytes until the boundary passes it (the frozen
|
||||
// partial-segments-wait rule; logical reclamation above still holds —
|
||||
// the generations left committedRanges and asOf below the horizon
|
||||
// throws regardless).
|
||||
if (this.segments) {
|
||||
await this.segments.dropSegmentsBelow(this.horizonGen + 1)
|
||||
}
|
||||
const manifest: GenerationManifest = {
|
||||
version: 1,
|
||||
generation: this.committed,
|
||||
|
|
@ -2544,9 +2140,6 @@ export class GenerationStore {
|
|||
async reopenAfterRestore(floorGeneration: number): Promise<void> {
|
||||
await this.withMutex(async () => {
|
||||
this.deltaCache.clear()
|
||||
// The running history-byte total describes the REPLACED store — drop it;
|
||||
// the next historyBytes() re-seeds with one walk over the new state.
|
||||
this.historyBytesTotal = null
|
||||
// A wholesale state replacement invalidates any buffered single-op
|
||||
// history — discard the pending tier (its live writes are gone with the
|
||||
// replaced store).
|
||||
|
|
|
|||
|
|
@ -116,16 +116,6 @@ export interface TransactOptions {
|
|||
* record is staged.
|
||||
*/
|
||||
ifAtGeneration?: number
|
||||
/**
|
||||
* Budget (ms) for the atomic apply phase. When omitted, the budget SCALES
|
||||
* with the batch: `max(30 000, opCount × 2 000)` — production imports on
|
||||
* network-attached disks measure ~2 s per operation, so a flat 30 s budget
|
||||
* silently capped honest bulk work at ~15 operations. A tripped budget
|
||||
* rolls the whole batch back and throws a retryable
|
||||
* `TransactionTimeoutError` naming the operation it stopped at, the batch
|
||||
* size, and the elapsed/budget times.
|
||||
*/
|
||||
timeoutMs?: number
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -176,15 +166,6 @@ export interface CompactHistoryOptions {
|
|||
* of each surviving generation's serialized record set (`GenerationDelta.bytes`).
|
||||
*/
|
||||
maxBytes?: number
|
||||
/**
|
||||
* Stop reclaiming after this many milliseconds even if caps are still
|
||||
* exceeded (8.9.0). Compaction is maintenance — a bounded pass keeps
|
||||
* `close()` (and any explicit maintenance window) from stalling on a large
|
||||
* backlog; the next pass resumes where this one stopped (reclamation is
|
||||
* oldest-first, so an early stop is always a consistent prefix). Unset =
|
||||
* run to completion.
|
||||
*/
|
||||
timeBudgetMs?: number
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -202,36 +183,6 @@ export interface CompactHistoryResult {
|
|||
horizon: number
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Result of `brain.historyStats()` — the read-only generational
|
||||
* history footprint, for fleet audits and ops doors. A pool operator runs this
|
||||
* per brain to size retention exposure (how much MVCC history each brain
|
||||
* carries and under which policy) without touching any data.
|
||||
*/
|
||||
export interface HistoryStats {
|
||||
/** Committed generation record-sets currently on disk. */
|
||||
generations: number
|
||||
/** Total on-disk history bytes across those record-sets. */
|
||||
bytes: number
|
||||
/** Oldest committed generation still on disk (null when history is empty). */
|
||||
oldestGeneration: number | null
|
||||
/** Newest committed generation (null when history is empty). */
|
||||
newestGeneration: number | null
|
||||
/** Commit timestamp (ms) of the oldest on-disk generation. */
|
||||
oldestTimestamp: number | null
|
||||
/** Commit timestamp (ms) of the newest on-disk generation. */
|
||||
newestTimestamp: number | null
|
||||
/** Compaction horizon — generations below it were reclaimed. */
|
||||
horizon: number
|
||||
/** The effective retention mode this brain runs under. */
|
||||
retentionMode: 'all' | 'adaptive' | 'explicit'
|
||||
/**
|
||||
* The adaptive byte budget in force (coordinator-driven or the local
|
||||
* free-memory probe); null under 'all' or explicit caps.
|
||||
*/
|
||||
effectiveBudgetBytes: number | null
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// Db surfaces
|
||||
// ============================================================================
|
||||
|
|
@ -472,22 +423,6 @@ export interface GenerationStorage {
|
|||
/** Read all lines of `_system/tx-log.jsonl` (empty array if absent). */
|
||||
readTxLogLines(): Promise<string[]>
|
||||
|
||||
/**
|
||||
* OPTIONAL binary raw-byte primitives — the substrate for the generation
|
||||
* fact log's append-only CRC-framed segments. Feature-detected: a storage
|
||||
* layer that omits them hosts no fact log (dual-write is skipped; readers
|
||||
* fall back to canonical enumeration). Paths are used VERBATIM (no
|
||||
* suffixing). Append durability rides `syncRawObjects` at the commit
|
||||
* barrier, exactly like the staged history files.
|
||||
*/
|
||||
appendRawBytes?(path: string, bytes: Uint8Array): Promise<void>
|
||||
/** Read a raw binary file whole; absent → null; a real fault throws. */
|
||||
readRawBytes?(path: string): Promise<Uint8Array | null>
|
||||
/** Replace a raw binary file atomically (tmp → fsync → rename). */
|
||||
writeRawBytes?(path: string, bytes: Uint8Array): Promise<void>
|
||||
/** Byte size of a raw binary file, or null when absent. */
|
||||
rawByteSize?(path: string): Promise<number | null>
|
||||
|
||||
/**
|
||||
* OPTIONAL temporal-blob contract (implemented by blob-aware storage; the
|
||||
* generation store treats the hashes as opaque strings). Extract the
|
||||
|
|
|
|||
|
|
@ -282,16 +282,6 @@ export class GraphAdjacencyIndex implements GraphIndexProvider {
|
|||
}
|
||||
|
||||
hasMore = result.hasMore
|
||||
if (hasMore && (!result.nextCursor || result.nextCursor === cursor)) {
|
||||
// A stalled cursor with hasMore=true would re-read the same page
|
||||
// forever — a silent full-CPU loop at cold open. Abort loudly; a
|
||||
// graph read failing beats a process that spins without a log line.
|
||||
throw new Error(
|
||||
`GraphAdjacencyIndex: verb walk stalled after ${count} verbs — storage returned ` +
|
||||
`hasMore=true with ${result.nextCursor ? 'a non-advancing' : 'no'} cursor. ` +
|
||||
`Aborting the cold-load; run brain.repairIndex() if this persists.`
|
||||
)
|
||||
}
|
||||
cursor = result.nextCursor
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -1,214 +0,0 @@
|
|||
/**
|
||||
* @module graph/graphAudit
|
||||
* @description Read-only graph-truth audit — the graph sibling of `repairIndex()`'s
|
||||
* diagnosis half. Verifies three layers against each other without mutating anything:
|
||||
*
|
||||
* 1. CANONICAL verb records (the storage walk — the source of truth)
|
||||
* 2. the RELATIONSHIP READ PATH (`related()` with all visibility tiers — exactly
|
||||
* what application reads like a VFS `readdir` consult)
|
||||
* 3. ENTITY ENDPOINTS (does each verb's source/target still exist?)
|
||||
*
|
||||
* and classifies every discrepancy into the three failure families production
|
||||
* incidents have shown:
|
||||
*
|
||||
* - `missingFromReads` — a canonical verb record the read path does NOT return
|
||||
* for its source: PRESENT BUT INVISIBLE (adjacency/membership staleness).
|
||||
* - `danglingEndpoints` — a canonical verb whose endpoint entity is gone:
|
||||
* the SCAR class (write-path loss / partial delete).
|
||||
* - `readOnlyVerbIds` — the read path returns an edge with NO canonical
|
||||
* record: GHOST edges (stale index entries).
|
||||
*
|
||||
* `visibilityHiddenCount` is reported separately: an internal/system edge that is
|
||||
* indexed and present but hidden from DEFAULT reads is working as designed — the
|
||||
* audit reads with all tiers included so design-hiding is never misclassified as
|
||||
* index loss.
|
||||
*
|
||||
* Full counts are always exact; only the example LISTS are capped (`maxExamples`)
|
||||
* — a capped report says so via `truncatedExamples`, never silently.
|
||||
*/
|
||||
|
||||
import { prodLog } from '../utils/logger.js'
|
||||
|
||||
/** One discrepant relationship, identified fully enough to inspect by hand. */
|
||||
export interface GraphAuditDiscrepancy {
|
||||
verbId: string
|
||||
from: string
|
||||
to: string
|
||||
type: string
|
||||
}
|
||||
|
||||
export interface GraphAuditReport {
|
||||
/** True iff every discrepancy count is zero — `related()` returns canonical truth. */
|
||||
coherent: boolean
|
||||
verbsInCanonical: number
|
||||
entitiesInCanonical: number
|
||||
/** Distinct source entities whose read path was actually consulted (coverage honesty). */
|
||||
sourcesChecked: number
|
||||
|
||||
/** PRESENT BUT INVISIBLE: canonical records the read path omits. */
|
||||
missingFromReadsCount: number
|
||||
missingFromReads: GraphAuditDiscrepancy[]
|
||||
|
||||
/** SCAR CLASS: canonical verbs with a missing endpoint entity. */
|
||||
danglingEndpointsCount: number
|
||||
danglingEndpoints: Array<GraphAuditDiscrepancy & { missingEnd: 'from' | 'to' | 'both' }>
|
||||
|
||||
/** GHOST EDGES: read-path verb ids with no canonical record. */
|
||||
readOnlyCount: number
|
||||
readOnlyVerbIds: string[]
|
||||
|
||||
/** Canonical verbs hidden from DEFAULT reads by design (internal/system visibility). */
|
||||
visibilityHiddenCount: number
|
||||
|
||||
/** Example lists above were capped at maxExamples; counts remain exact. */
|
||||
truncatedExamples: boolean
|
||||
durationMs: number
|
||||
}
|
||||
|
||||
/** A canonical verb record, as the audit needs it. */
|
||||
export interface AuditVerbRecord {
|
||||
id: string
|
||||
type: string
|
||||
sourceId: string
|
||||
targetId: string
|
||||
visibility?: string
|
||||
}
|
||||
|
||||
/** The seams the audit runs over — injected so the walk is testable in isolation. */
|
||||
export interface GraphAuditDeps {
|
||||
/** Stream every canonical entity id (id-only; no per-entity reads needed). */
|
||||
eachNounId(consume: (id: string) => void): Promise<void>
|
||||
/** Stream every canonical verb record. */
|
||||
eachVerb(consume: (verb: AuditVerbRecord) => void): Promise<void>
|
||||
/**
|
||||
* The END-TO-END relationship read for one source, ALL visibility tiers
|
||||
* included — must be the same path application reads consult.
|
||||
*/
|
||||
readRelationsFrom(sourceId: string): Promise<Array<{ id: string }>>
|
||||
}
|
||||
|
||||
export interface GraphAuditOptions {
|
||||
/** Cap on entries per example list (counts stay exact). Default 100. */
|
||||
maxExamples?: number
|
||||
}
|
||||
|
||||
export async function runGraphAudit(
|
||||
deps: GraphAuditDeps,
|
||||
options: GraphAuditOptions = {}
|
||||
): Promise<GraphAuditReport> {
|
||||
const maxExamples = options.maxExamples ?? 100
|
||||
const started = Date.now()
|
||||
|
||||
// 1. Canonical entity ids — endpoint existence oracle.
|
||||
const entityIds = new Set<string>()
|
||||
await deps.eachNounId((id) => entityIds.add(id))
|
||||
|
||||
// 2. Canonical verb walk: group by source, check endpoints, note visibility.
|
||||
const canonicalVerbIds = new Set<string>()
|
||||
const bySource = new Map<string, AuditVerbRecord[]>()
|
||||
let verbsInCanonical = 0
|
||||
let visibilityHiddenCount = 0
|
||||
let danglingEndpointsCount = 0
|
||||
const danglingEndpoints: GraphAuditReport['danglingEndpoints'] = []
|
||||
|
||||
await deps.eachVerb((verb) => {
|
||||
verbsInCanonical++
|
||||
canonicalVerbIds.add(verb.id)
|
||||
const list = bySource.get(verb.sourceId)
|
||||
if (list) list.push(verb)
|
||||
else bySource.set(verb.sourceId, [verb])
|
||||
|
||||
if (verb.visibility === 'internal' || verb.visibility === 'system') {
|
||||
visibilityHiddenCount++
|
||||
}
|
||||
|
||||
const fromMissing = !entityIds.has(verb.sourceId)
|
||||
const toMissing = !entityIds.has(verb.targetId)
|
||||
if (fromMissing || toMissing) {
|
||||
danglingEndpointsCount++
|
||||
if (danglingEndpoints.length < maxExamples) {
|
||||
danglingEndpoints.push({
|
||||
verbId: verb.id,
|
||||
from: verb.sourceId,
|
||||
to: verb.targetId,
|
||||
type: verb.type,
|
||||
missingEnd: fromMissing && toMissing ? 'both' : fromMissing ? 'from' : 'to'
|
||||
})
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
// 3. Per-source read-path comparison. A verb must be returned by the read
|
||||
// path of ITS OWN source — the exact consult a readdir/traversal makes.
|
||||
let missingFromReadsCount = 0
|
||||
const missingFromReads: GraphAuditDiscrepancy[] = []
|
||||
let readOnlyCount = 0
|
||||
const readOnlyVerbIds: string[] = []
|
||||
const readOnlySeen = new Set<string>()
|
||||
|
||||
for (const [sourceId, verbs] of bySource) {
|
||||
const readIds = new Set((await deps.readRelationsFrom(sourceId)).map((r) => r.id))
|
||||
|
||||
for (const verb of verbs) {
|
||||
if (!readIds.has(verb.id)) {
|
||||
missingFromReadsCount++
|
||||
if (missingFromReads.length < maxExamples) {
|
||||
missingFromReads.push({
|
||||
verbId: verb.id,
|
||||
from: verb.sourceId,
|
||||
to: verb.targetId,
|
||||
type: verb.type
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (const readId of readIds) {
|
||||
if (!canonicalVerbIds.has(readId) && !readOnlySeen.has(readId)) {
|
||||
readOnlySeen.add(readId)
|
||||
readOnlyCount++
|
||||
if (readOnlyVerbIds.length < maxExamples) {
|
||||
readOnlyVerbIds.push(readId)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const coherent =
|
||||
missingFromReadsCount === 0 && danglingEndpointsCount === 0 && readOnlyCount === 0
|
||||
|
||||
const report: GraphAuditReport = {
|
||||
coherent,
|
||||
verbsInCanonical,
|
||||
entitiesInCanonical: entityIds.size,
|
||||
sourcesChecked: bySource.size,
|
||||
missingFromReadsCount,
|
||||
missingFromReads,
|
||||
danglingEndpointsCount,
|
||||
danglingEndpoints,
|
||||
readOnlyCount,
|
||||
readOnlyVerbIds,
|
||||
visibilityHiddenCount,
|
||||
truncatedExamples:
|
||||
missingFromReadsCount > missingFromReads.length ||
|
||||
danglingEndpointsCount > danglingEndpoints.length ||
|
||||
readOnlyCount > readOnlyVerbIds.length,
|
||||
durationMs: Date.now() - started
|
||||
}
|
||||
|
||||
if (coherent) {
|
||||
prodLog.info(
|
||||
`[GraphAudit] coherent: ${verbsInCanonical} verbs across ${bySource.size} sources — ` +
|
||||
`the read path returns canonical truth (${report.durationMs}ms)`
|
||||
)
|
||||
} else {
|
||||
prodLog.warn(
|
||||
`[GraphAudit] INCOHERENT: ${missingFromReadsCount} present-but-invisible, ` +
|
||||
`${danglingEndpointsCount} dangling-endpoint, ${readOnlyCount} ghost ` +
|
||||
`(of ${verbsInCanonical} canonical verbs, ${bySource.size} sources, ` +
|
||||
`${visibilityHiddenCount} visibility-hidden by design) — ${report.durationMs}ms`
|
||||
)
|
||||
}
|
||||
|
||||
return report
|
||||
}
|
||||
|
|
@ -41,14 +41,6 @@ export interface DeduplicationStats {
|
|||
* - Import-scoped deduplication (no cross-contamination)
|
||||
* - 3-tier strategy (ID → Name → Similarity)
|
||||
* - Uses existing indexes (EntityIdMapper, MetadataIndexManager, TypeAware HNSW)
|
||||
*
|
||||
* Lifecycle: ONE instance per brain, owned by Brainy (getBackgroundDeduplicator)
|
||||
* so the debounce genuinely spans imports and brain.close() cancels pending
|
||||
* work via cancelPending() — this pass merge-DELETES duplicate entities, so it
|
||||
* must never fire against a closed brain. The enableDeduplication gate lives
|
||||
* at the scheduling call site (ImportCoordinator); scheduleDedup itself is
|
||||
* unconditional. The timer is unref'd — a pending pass never holds the
|
||||
* process open.
|
||||
*/
|
||||
export class BackgroundDeduplicator {
|
||||
private brain: Brainy
|
||||
|
|
@ -75,15 +67,12 @@ export class BackgroundDeduplicator {
|
|||
clearTimeout(this.debounceTimer)
|
||||
}
|
||||
|
||||
// Schedule for 5 minutes from now. unref'd: a pending dedup pass must
|
||||
// never hold the process open (exit-hang class) — if the process exits
|
||||
// first, the pass simply never runs; imports are already durable.
|
||||
// Schedule for 5 minutes from now
|
||||
this.debounceTimer = setTimeout(() => {
|
||||
this.runBatchDedup().catch(error => {
|
||||
prodLog.error('[BackgroundDedup] Batch dedup failed:', error)
|
||||
})
|
||||
}, 5 * 60 * 1000)
|
||||
this.debounceTimer.unref?.()
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
|
|||
|
|
@ -13,6 +13,7 @@
|
|||
import { Brainy } from '../brainy.js'
|
||||
import { FormatDetector, SupportedFormat } from './FormatDetector.js'
|
||||
import { ImportHistory, type ImportHistoryEntry } from './ImportHistory.js'
|
||||
import { BackgroundDeduplicator } from './BackgroundDeduplicator.js'
|
||||
import { SmartExcelImporter } from '../importers/SmartExcelImporter.js'
|
||||
import { SmartPDFImporter } from '../importers/SmartPDFImporter.js'
|
||||
import { SmartCSVImporter } from '../importers/SmartCSVImporter.js'
|
||||
|
|
@ -111,12 +112,7 @@ export interface ValidImportOptions {
|
|||
/** Confidence threshold for entities */
|
||||
confidenceThreshold?: number
|
||||
|
||||
/**
|
||||
* Enable entity deduplication (default: true). Gates BOTH passes: the
|
||||
* inline merge during import AND the debounced background pass that runs
|
||||
* ~5 minutes after the last import (which merge-DELETES duplicate entities).
|
||||
* Set false for deployments that must never auto-remove records.
|
||||
*/
|
||||
/** Enable entity deduplication across imports */
|
||||
enableDeduplication?: boolean
|
||||
|
||||
/** Similarity threshold for deduplication (0-1) */
|
||||
|
|
@ -290,6 +286,7 @@ export class ImportCoordinator {
|
|||
private brain: Brainy
|
||||
private detector: FormatDetector
|
||||
private history: ImportHistory
|
||||
private backgroundDedup: BackgroundDeduplicator
|
||||
private excelImporter: SmartExcelImporter
|
||||
private pdfImporter: SmartPDFImporter
|
||||
private csvImporter: SmartCSVImporter
|
||||
|
|
@ -303,6 +300,7 @@ export class ImportCoordinator {
|
|||
this.brain = brain
|
||||
this.detector = new FormatDetector()
|
||||
this.history = new ImportHistory(brain)
|
||||
this.backgroundDedup = new BackgroundDeduplicator(brain)
|
||||
this.excelImporter = new SmartExcelImporter(brain)
|
||||
this.pdfImporter = new SmartPDFImporter(brain)
|
||||
this.csvImporter = new SmartCSVImporter(brain)
|
||||
|
|
@ -1461,16 +1459,9 @@ export class ImportCoordinator {
|
|||
}
|
||||
}
|
||||
|
||||
// Schedule background deduplication (debounced 5 minutes, brain-owned so
|
||||
// close() can cancel it). Honors the same enableDeduplication gate as the
|
||||
// inline pass — false means NO dedup, inline or background.
|
||||
if (
|
||||
trackingContext &&
|
||||
trackingContext.importId &&
|
||||
options.enableDeduplication !== false
|
||||
) {
|
||||
const backgroundDedup = await this.brain.getBackgroundDeduplicator()
|
||||
backgroundDedup.scheduleDedup(trackingContext.importId)
|
||||
// Schedule background deduplication (debounced 5 minutes)
|
||||
if (trackingContext && trackingContext.importId) {
|
||||
this.backgroundDedup.scheduleDedup(trackingContext.importId)
|
||||
}
|
||||
|
||||
return {
|
||||
|
|
|
|||
25
src/index.ts
25
src/index.ts
|
|
@ -28,16 +28,6 @@ export type { FileVersion } from './vfs/types.js'
|
|||
|
||||
// Export diagnostics result type
|
||||
export type { DiagnosticsResult } from './brainy.js'
|
||||
export type {
|
||||
GraphAuditReport,
|
||||
GraphAuditDiscrepancy
|
||||
} from './graph/graphAudit.js'
|
||||
export {
|
||||
checkOsLimits,
|
||||
NOFILE_POOL_FLOOR,
|
||||
MAX_MAP_COUNT_POOL_FLOOR
|
||||
} from './utils/osLimits.js'
|
||||
export type { OsLimitsReport } from './utils/osLimits.js'
|
||||
|
||||
// Export Brainy configuration and types
|
||||
export type {
|
||||
|
|
@ -201,26 +191,11 @@ export type {
|
|||
TxLogEntry,
|
||||
CompactHistoryOptions,
|
||||
CompactHistoryResult,
|
||||
HistoryStats,
|
||||
ChangedIds,
|
||||
DiffResult,
|
||||
HistoryVersion,
|
||||
EntityHistory
|
||||
} from './db/types.js'
|
||||
// The generation fact log — sequential after-image scan surface
|
||||
// (brain.scanFacts / brain.factSegmentPaths) for index heals and replays.
|
||||
export type {
|
||||
CommitFact,
|
||||
FactOp,
|
||||
FactScanBatch,
|
||||
SCANFACTS_FIRST_BATCH_MS,
|
||||
FactScanHandle
|
||||
} from './db/factLog.js'
|
||||
// The generalized family stamp — which source generation a projection
|
||||
// reflects + the surface that verifies it whole; one verifier, both member
|
||||
// modes (enumerated byte-exact / rollup invariants).
|
||||
export { readFamilyStamp, verifyFamilyStamp, ENTITY_TREE_STAMP_PATH } from './db/familyStamp.js'
|
||||
export type { FamilyStamp, StampMembers, StampVerdict } from './db/familyStamp.js'
|
||||
// Optional provider capability for generation-aware native indexes
|
||||
export { isVersionedIndexProvider } from './plugin.js'
|
||||
export type { VersionedIndexProvider } from './plugin.js'
|
||||
|
|
|
|||
|
|
@ -19,29 +19,3 @@ export type { MemoryInfo, CacheAllocationStrategy } from './utils/memoryDetectio
|
|||
// HNSWNounWithMetadata. First-party plugins (Cor) use this to stay in
|
||||
// lockstep with the entity shape contract.
|
||||
export { resolveEntityField, STANDARD_ENTITY_FIELDS } from './coreTypes.js'
|
||||
|
||||
// The generalized family stamp — ONE verifier, both member modes, shared
|
||||
// verbatim with native providers so stamp verification is literally one
|
||||
// function, never two synchronized copies. Providers write the same shape
|
||||
// (their set-swap rewrites stamps, so adoption is migration-free).
|
||||
export {
|
||||
readFamilyStamp,
|
||||
writeFamilyStamp,
|
||||
verifyFamilyStamp,
|
||||
FAMILY_STAMPS_PREFIX,
|
||||
ENTITY_TREE_STAMP_PATH
|
||||
} from './db/familyStamp.js'
|
||||
export type {
|
||||
FamilyStamp,
|
||||
StampMembers,
|
||||
StampVerdict,
|
||||
EnumeratedMember,
|
||||
StampStorage
|
||||
} from './db/familyStamp.js'
|
||||
|
||||
// Generation fact-log types — the scan surface a provider reaches through the
|
||||
// storage capability (`storage.scanFacts` / `storage.factLogHeadGeneration` /
|
||||
// `storage.factSegmentPaths`, wired by the host brain at init). Providers
|
||||
// never construct a FactLog themselves: open() is writer-side (it reconciles
|
||||
// by truncating/rewriting) and there is exactly one writer.
|
||||
export type { CommitFact, FactOp, FactScanBatch, FactScanHandle } from './db/factLog.js'
|
||||
|
|
|
|||
|
|
@ -96,14 +96,6 @@ export class FileSystemStorage extends BaseStorage {
|
|||
private static readonly WRITER_STALE_THRESHOLD_MS = 60_000
|
||||
private writerLockHeartbeat?: NodeJS.Timeout
|
||||
private writerLockInfo?: WriterLockInfo
|
||||
/**
|
||||
* The currently-executing heartbeat refresh, if any. `releaseWriterLock()`
|
||||
* awaits it before unlinking: clearInterval() stops FUTURE ticks but not a
|
||||
* tick already in flight, and a straggler landing after the unlink would
|
||||
* RE-CREATE the lock file — a phantom lock blocking the next writer until
|
||||
* the stale TTL expires (the pool-eviction reopen case).
|
||||
*/
|
||||
private writerHeartbeatInFlight?: Promise<void>
|
||||
|
||||
// Flush-request RPC state. The writer polls `locks/_flush_requests/` for
|
||||
// new `.req` files and emits `.ack` files in `locks/_flush_responses/` after
|
||||
|
|
@ -701,17 +693,6 @@ export class FileSystemStorage extends BaseStorage {
|
|||
*/
|
||||
private static readonly SNAPSHOT_BYTE_COPY_DIRS = new Set<string>(['_id_mapper'])
|
||||
|
||||
/**
|
||||
* Nested path PREFIXES whose files are byte-copied into snapshots, not
|
||||
* hard-linked — for append-in-place files below the top level. The
|
||||
* generation fact log's tail segment is appended in place between rotations;
|
||||
* a hard-linked tail would let post-snapshot appends reach through into the
|
||||
* snapshot. (Sealed segments are immutable and would be link-safe, but the
|
||||
* prefix rule keeps the discipline simple; segments are bounded by the
|
||||
* rotation threshold, so the copy cost is small.)
|
||||
*/
|
||||
private static readonly SNAPSHOT_BYTE_COPY_PREFIXES: string[] = ['_generations/facts/']
|
||||
|
||||
/**
|
||||
* Top-level directories excluded from snapshots: process-local lock state
|
||||
* (writer lock, flush-request RPC files) must never travel with the data, and
|
||||
|
|
@ -889,75 +870,6 @@ export class FileSystemStorage extends BaseStorage {
|
|||
}
|
||||
}
|
||||
|
||||
// ==========================================================================
|
||||
// Binary raw-byte primitives — the substrate for append-only log-structured
|
||||
// files (the generation fact log's CRC-framed segments). Paths are used
|
||||
// VERBATIM (no .gz/.bin suffixing). Append durability rides syncRawObjects
|
||||
// at the commit barrier, like every other staged write.
|
||||
// ==========================================================================
|
||||
|
||||
/**
|
||||
* Append bytes to a raw binary file, creating it (and parent directories)
|
||||
* when absent. NOT fsync'd here — the caller batches durability via
|
||||
* `syncRawObjects` at its commit barrier.
|
||||
*/
|
||||
public async appendRawBytes(rawPath: string, bytes: Uint8Array): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
const fullPath = path.join(this.rootDir, rawPath)
|
||||
await fs.promises.mkdir(path.dirname(fullPath), { recursive: true })
|
||||
await fs.promises.appendFile(fullPath, bytes)
|
||||
}
|
||||
|
||||
/**
|
||||
* Read a raw binary file whole. Absent → `null`; a real IO fault throws —
|
||||
* a present-but-unreadable log segment must never read as "no facts".
|
||||
*/
|
||||
public async readRawBytes(rawPath: string): Promise<Uint8Array | null> {
|
||||
await this.ensureInitialized()
|
||||
try {
|
||||
const buf: Buffer = await fs.promises.readFile(path.join(this.rootDir, rawPath))
|
||||
return new Uint8Array(buf.buffer, buf.byteOffset, buf.byteLength)
|
||||
} catch (error: any) {
|
||||
if (isAbsentError(error)) return null
|
||||
throw error
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Replace a raw binary file atomically: write-new → fsync → rename. The
|
||||
* reconcile primitive (e.g. truncating a fact-log tail back to committed
|
||||
* truth after a crash) — a crash mid-replace leaves either the old file or
|
||||
* the new one, never a mix.
|
||||
*/
|
||||
public async writeRawBytes(rawPath: string, bytes: Uint8Array): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
const fullPath = path.join(this.rootDir, rawPath)
|
||||
await fs.promises.mkdir(path.dirname(fullPath), { recursive: true })
|
||||
const tmpPath = `${fullPath}.tmp.${Date.now()}.${Math.random().toString(36).slice(2)}`
|
||||
const handle = await fs.promises.open(tmpPath, 'w')
|
||||
try {
|
||||
await handle.writeFile(bytes)
|
||||
await handle.sync()
|
||||
} finally {
|
||||
await handle.close()
|
||||
}
|
||||
await fs.promises.rename(tmpPath, fullPath)
|
||||
}
|
||||
|
||||
/**
|
||||
* Byte size of a raw binary file, or `null` when absent.
|
||||
*/
|
||||
public async rawByteSize(rawPath: string): Promise<number | null> {
|
||||
await this.ensureInitialized()
|
||||
try {
|
||||
const stat = await fs.promises.stat(path.join(this.rootDir, rawPath))
|
||||
return stat.size
|
||||
} catch (error: any) {
|
||||
if (isAbsentError(error)) return null
|
||||
throw error
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Snapshot the entire store into `targetPath` as a hard-link farm
|
||||
* (Cassandra-style: instant, space-shared). Safe because every data file
|
||||
|
|
@ -1007,8 +919,7 @@ export class FileSystemStorage extends BaseStorage {
|
|||
// msync/truncate reach through into the snapshot.
|
||||
if (
|
||||
FileSystemStorage.SNAPSHOT_BYTE_COPY_PATHS.has(normalized) ||
|
||||
FileSystemStorage.SNAPSHOT_BYTE_COPY_DIRS.has(normalized.split('/')[0]) ||
|
||||
FileSystemStorage.SNAPSHOT_BYTE_COPY_PREFIXES.some((p) => normalized.startsWith(p))
|
||||
FileSystemStorage.SNAPSHOT_BYTE_COPY_DIRS.has(normalized.split('/')[0])
|
||||
) {
|
||||
await fs.promises.copyFile(sourceFile, targetFile)
|
||||
continue
|
||||
|
|
@ -1772,158 +1683,79 @@ export class FileSystemStorage extends BaseStorage {
|
|||
const lockFile = path.join(this.lockDir, FileSystemStorage.WRITER_LOCK_FILE)
|
||||
const os = await import('node:os')
|
||||
const hostname = os.hostname()
|
||||
const myPid = typeof process !== 'undefined' && process.pid ? process.pid : 0
|
||||
const now = new Date().toISOString()
|
||||
|
||||
// Bounded acquire loop. The CLAIM itself is an atomic create-exclusive
|
||||
// write (O_EXCL) — two processes racing an ABSENT lock can never both
|
||||
// succeed, which closes the read-then-write window where the loser used
|
||||
// to keep running unlocked, silently. An EEXIST loser loops, re-reads,
|
||||
// and handles whatever it finds honestly (fresh foreign lock → loud
|
||||
// throw; stale/forced → verified takeover).
|
||||
const MAX_ATTEMPTS = 3
|
||||
for (let attempt = 0; attempt < MAX_ATTEMPTS; attempt++) {
|
||||
const now = new Date().toISOString()
|
||||
const existing = await this.readWriterLock()
|
||||
|
||||
if (existing) {
|
||||
// Same-process re-open: a second Brainy instance in this Node process
|
||||
// (e.g. test "simulate server restart" patterns, or a consumer that
|
||||
// explicitly re-instantiates without closing first). This isn't the
|
||||
// dangerous cross-process case the lock exists to prevent — the two
|
||||
// instances share a memory space and can't silently diverge from each
|
||||
// other beyond what their callers already see. Warn and take over.
|
||||
if (existing.pid === myPid && existing.hostname === hostname && !options?.force) {
|
||||
console.warn(
|
||||
`[brainy] Re-acquiring writer lock for ${this.rootDir} held by the same process (PID ${existing.pid}). ` +
|
||||
`If you intended to keep the previous Brainy instance alive, this is a bug — close it first.`
|
||||
)
|
||||
const info: WriterLockInfo = {
|
||||
pid: myPid,
|
||||
hostname,
|
||||
startedAt: now,
|
||||
lastHeartbeat: now,
|
||||
version: getBrainyVersion(),
|
||||
rootDir: this.rootDir
|
||||
}
|
||||
await this.writeFileAtomic(lockFile, JSON.stringify(info, null, 2))
|
||||
this.installWriterLock(info)
|
||||
return info
|
||||
}
|
||||
|
||||
const stale = !options?.force && (await this.isWriterLockStale(existing))
|
||||
if (!options?.force && !stale) {
|
||||
const existing = await this.readWriterLock()
|
||||
if (existing && !options?.force) {
|
||||
// Same-process re-open: a second Brainy instance in this Node process
|
||||
// (e.g. test "simulate server restart" patterns, or a consumer that
|
||||
// explicitly re-instantiates without closing first). This isn't the
|
||||
// dangerous cross-process case the lock exists to prevent — the two
|
||||
// instances share a memory space and can't silently diverge from each
|
||||
// other beyond what their callers already see. Warn and take over the lock.
|
||||
if (existing.pid === (typeof process !== 'undefined' ? process.pid : 0) &&
|
||||
existing.hostname === hostname) {
|
||||
console.warn(
|
||||
`[brainy] Re-acquiring writer lock for ${this.rootDir} held by the same process (PID ${existing.pid}). ` +
|
||||
`If you intended to keep the previous Brainy instance alive, this is a bug — close it first.`
|
||||
)
|
||||
} else {
|
||||
const stale = await this.isWriterLockStale(existing)
|
||||
if (!stale) {
|
||||
// Consumer-facing error contract: callers detect this case via
|
||||
// err.code and read the holder's details from err.lockInfo.
|
||||
throw this.writerLockedError(existing)
|
||||
const err = new Error(
|
||||
`Another writer holds this Brainy directory.\n` +
|
||||
` PID: ${existing.pid} on host ${existing.hostname}\n` +
|
||||
` Started: ${existing.startedAt}\n` +
|
||||
` Heartbeat: ${existing.lastHeartbeat}\n` +
|
||||
` Version: ${existing.version}\n` +
|
||||
` Directory: ${this.rootDir}\n\n` +
|
||||
`For diagnostic queries against this live store, use:\n` +
|
||||
` const reader = await Brainy.openReadOnly({ storage: { type: 'filesystem', path: '${this.rootDir}' } })\n\n` +
|
||||
`If you have verified the existing lock is stale (e.g. a crashed writer on a different host that PID liveness cannot reach), pass { force: true }.`
|
||||
) as Error & { code: string; lockInfo: WriterLockInfo }
|
||||
err.code = 'BRAINY_WRITER_LOCKED'
|
||||
err.lockInfo = existing
|
||||
throw err
|
||||
}
|
||||
|
||||
console.warn(
|
||||
options?.force
|
||||
? `[brainy] Force-overwriting writer lock for ${this.rootDir} ` +
|
||||
`(was held by PID ${existing.pid} on ${existing.hostname}).`
|
||||
: `[brainy] Overwriting stale writer lock for ${this.rootDir} ` +
|
||||
`(PID ${existing.pid} on ${existing.hostname} appears dead).`
|
||||
`[brainy] Overwriting stale writer lock for ${this.rootDir} ` +
|
||||
`(PID ${existing.pid} on ${existing.hostname} appears dead).`
|
||||
)
|
||||
// Takeover: verify the file still holds the lock we judged (a live
|
||||
// successor may have claimed meanwhile), then remove it and fall
|
||||
// through to the atomic claim below. A racing claimer who beats us to
|
||||
// the create simply wins — our next loop iteration reads their fresh
|
||||
// lock and throws honestly. (Advisory file locking has no
|
||||
// compare-and-delete; staleness requiring a 60s-old heartbeat keeps
|
||||
// the residual verify-to-unlink window practically unreachable.)
|
||||
const recheck = await this.readWriterLock()
|
||||
if (
|
||||
recheck &&
|
||||
(recheck.pid !== existing.pid ||
|
||||
recheck.startedAt !== existing.startedAt ||
|
||||
recheck.lastHeartbeat !== existing.lastHeartbeat)
|
||||
) {
|
||||
continue // the lock changed hands while we deliberated — re-evaluate
|
||||
}
|
||||
try {
|
||||
await fs.promises.unlink(lockFile)
|
||||
} catch (err: any) {
|
||||
if (err.code !== 'ENOENT') throw err
|
||||
}
|
||||
}
|
||||
|
||||
const info: WriterLockInfo = {
|
||||
pid: myPid,
|
||||
hostname,
|
||||
startedAt: existing && options?.force ? existing.startedAt : now,
|
||||
lastHeartbeat: now,
|
||||
version: getBrainyVersion(),
|
||||
rootDir: this.rootDir
|
||||
}
|
||||
|
||||
// The atomic claim: create-exclusive, so exactly ONE racer wins.
|
||||
try {
|
||||
await fs.promises.writeFile(lockFile, JSON.stringify(info, null, 2), { flag: 'wx' })
|
||||
} catch (err: any) {
|
||||
if (err.code === 'EEXIST') {
|
||||
continue // someone else claimed between our read and create — re-evaluate
|
||||
}
|
||||
throw err
|
||||
}
|
||||
|
||||
this.installWriterLock(info)
|
||||
return info
|
||||
} else if (existing && options?.force) {
|
||||
console.warn(
|
||||
`[brainy] Force-overwriting writer lock for ${this.rootDir} ` +
|
||||
`(was held by PID ${existing.pid} on ${existing.hostname}).`
|
||||
)
|
||||
}
|
||||
|
||||
// Attempts exhausted: something is claiming this directory faster than we
|
||||
// can evaluate it. Read whoever holds it now and fail loudly with their
|
||||
// details rather than degrading into a lockless open.
|
||||
const holder = await this.readWriterLock()
|
||||
if (holder) throw this.writerLockedError(holder)
|
||||
throw new Error(
|
||||
`Failed to acquire the writer lock for ${this.rootDir} after ${MAX_ATTEMPTS} attempts — ` +
|
||||
`the lock file is being contended. Retry, or inspect ${lockFile}.`
|
||||
)
|
||||
}
|
||||
const info: WriterLockInfo = {
|
||||
pid: typeof process !== 'undefined' && process.pid ? process.pid : 0,
|
||||
hostname,
|
||||
startedAt: existing && options?.force ? existing.startedAt : now,
|
||||
lastHeartbeat: now,
|
||||
version: getBrainyVersion(),
|
||||
rootDir: this.rootDir
|
||||
}
|
||||
|
||||
/** Record lock ownership + start the unref'd heartbeat. */
|
||||
private installWriterLock(info: WriterLockInfo): void {
|
||||
await this.writeFileAtomic(lockFile, JSON.stringify(info, null, 2))
|
||||
this.writerLockInfo = info
|
||||
|
||||
// Heartbeat — rewrite lastHeartbeat every WRITER_HEARTBEAT_MS so other
|
||||
// processes can tell a live writer from one that crashed without releasing.
|
||||
this.writerLockHeartbeat = setInterval(() => {
|
||||
const tick = this.refreshWriterLockHeartbeat().catch((err) => {
|
||||
// ENOENT = the lock (or its directory) vanished mid-refresh — the
|
||||
// store was released or removed under us; the next acquire recreates
|
||||
// it. Benign by construction; anything else stays loud.
|
||||
if ((err as NodeJS.ErrnoException)?.code !== 'ENOENT') {
|
||||
console.warn('[brainy] Failed to refresh writer lock heartbeat:', err)
|
||||
}
|
||||
})
|
||||
this.writerHeartbeatInFlight = tick.finally(() => {
|
||||
if (this.writerHeartbeatInFlight === tick) {
|
||||
this.writerHeartbeatInFlight = undefined
|
||||
}
|
||||
this.refreshWriterLockHeartbeat().catch((err) => {
|
||||
console.warn('[brainy] Failed to refresh writer lock heartbeat:', err)
|
||||
})
|
||||
}, FileSystemStorage.WRITER_HEARTBEAT_MS)
|
||||
if (typeof this.writerLockHeartbeat.unref === 'function') {
|
||||
// Don't keep the event loop alive just for the heartbeat.
|
||||
this.writerLockHeartbeat.unref()
|
||||
}
|
||||
}
|
||||
|
||||
/** The consumer-facing BRAINY_WRITER_LOCKED error, holder details attached. */
|
||||
private writerLockedError(existing: WriterLockInfo): Error {
|
||||
const err = new Error(
|
||||
`Another writer holds this Brainy directory.\n` +
|
||||
` PID: ${existing.pid} on host ${existing.hostname}\n` +
|
||||
` Started: ${existing.startedAt}\n` +
|
||||
` Heartbeat: ${existing.lastHeartbeat}\n` +
|
||||
` Version: ${existing.version}\n` +
|
||||
` Directory: ${this.rootDir}\n\n` +
|
||||
`For diagnostic queries against this live store, use:\n` +
|
||||
` const reader = await Brainy.openReadOnly({ storage: { type: 'filesystem', path: '${this.rootDir}' } })\n\n` +
|
||||
`If you have verified the existing lock is stale (e.g. a crashed writer on a different host that PID liveness cannot reach), pass { force: true }.`
|
||||
) as Error & { code: string; lockInfo: WriterLockInfo }
|
||||
err.code = 'BRAINY_WRITER_LOCKED'
|
||||
err.lockInfo = existing
|
||||
return err
|
||||
return info
|
||||
}
|
||||
|
||||
public override async releaseWriterLock(): Promise<void> {
|
||||
|
|
@ -1931,15 +1763,6 @@ export class FileSystemStorage extends BaseStorage {
|
|||
clearInterval(this.writerLockHeartbeat)
|
||||
this.writerLockHeartbeat = undefined
|
||||
}
|
||||
// Drain an in-flight heartbeat tick BEFORE unlinking: clearInterval stops
|
||||
// future ticks only, and a straggler write landing after the unlink would
|
||||
// re-create the lock as a phantom (blocking the next writer until the
|
||||
// stale TTL). After the drain, any refresh is either fully landed (we
|
||||
// unlink its output below) or not started (it sees writerLockInfo
|
||||
// undefined and returns).
|
||||
if (this.writerHeartbeatInFlight) {
|
||||
await this.writerHeartbeatInFlight
|
||||
}
|
||||
if (!this.writerLockInfo) {
|
||||
return
|
||||
}
|
||||
|
|
|
|||
|
|
@ -133,11 +133,6 @@ export class MemoryStorage extends BaseStorage {
|
|||
*/
|
||||
protected async deleteObjectFromPath(path: string): Promise<void> {
|
||||
this.objectStore.delete(path)
|
||||
// Filesystem parity: on disk, objects and raw BYTE files are both just
|
||||
// files — unlink removes whichever exists. Without this, deleteRawObject
|
||||
// on a raw-bytes path (fact-log/generation segments) silently no-ops on
|
||||
// memory storage: the delete "succeeds" and the bytes remain.
|
||||
this.rawBytesStore.delete(path)
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -227,45 +222,6 @@ export class MemoryStorage extends BaseStorage {
|
|||
return [...this.txLogLines]
|
||||
}
|
||||
|
||||
// ===========================================================================
|
||||
// Binary raw-byte primitives — in-memory mirror of the filesystem adapter's
|
||||
// append-only substrate (the generation fact log's segments), so memory
|
||||
// brains dual-write facts too and the compat suite runs on both adapters.
|
||||
// ===========================================================================
|
||||
|
||||
/** Raw binary files, keyed by verbatim path. */
|
||||
private rawBytesStore: Map<string, Uint8Array> = new Map()
|
||||
|
||||
/** Append bytes to a raw binary file, creating it when absent. */
|
||||
public async appendRawBytes(rawPath: string, bytes: Uint8Array): Promise<void> {
|
||||
const existing = this.rawBytesStore.get(rawPath)
|
||||
if (!existing) {
|
||||
this.rawBytesStore.set(rawPath, bytes.slice())
|
||||
return
|
||||
}
|
||||
const merged = new Uint8Array(existing.length + bytes.length)
|
||||
merged.set(existing, 0)
|
||||
merged.set(bytes, existing.length)
|
||||
this.rawBytesStore.set(rawPath, merged)
|
||||
}
|
||||
|
||||
/** Read a raw binary file whole (a copy); absent → null. */
|
||||
public async readRawBytes(rawPath: string): Promise<Uint8Array | null> {
|
||||
const bytes = this.rawBytesStore.get(rawPath)
|
||||
return bytes ? bytes.slice() : null
|
||||
}
|
||||
|
||||
/** Replace a raw binary file (atomic by construction in memory). */
|
||||
public async writeRawBytes(rawPath: string, bytes: Uint8Array): Promise<void> {
|
||||
this.rawBytesStore.set(rawPath, bytes.slice())
|
||||
}
|
||||
|
||||
/** Byte size of a raw binary file, or null when absent. */
|
||||
public async rawByteSize(rawPath: string): Promise<number | null> {
|
||||
const bytes = this.rawBytesStore.get(rawPath)
|
||||
return bytes ? bytes.length : null
|
||||
}
|
||||
|
||||
/**
|
||||
* Serialize the entire in-memory store to a directory in the exact layout
|
||||
* the filesystem adapter uses (uncompressed JSON objects, `_blobs/*.bin`
|
||||
|
|
@ -398,7 +354,6 @@ export class MemoryStorage extends BaseStorage {
|
|||
public async clear(): Promise<void> {
|
||||
this.objectStore.clear()
|
||||
this.blobStore.clear()
|
||||
this.rawBytesStore.clear()
|
||||
this.txLogLines = []
|
||||
this.statistics = null
|
||||
|
||||
|
|
|
|||
|
|
@ -377,14 +377,7 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
id.startsWith('statistics_') ||
|
||||
id === 'statistics' ||
|
||||
id.startsWith('__chunk__') || // Metadata index chunks (roaring bitmap data)
|
||||
id.startsWith('__sparse_index__') || // Metadata sparse indices (zone maps + bloom filters)
|
||||
id.startsWith('__aggregation_') || // Aggregation engine definitions + state (routing is
|
||||
// identical to the unknown-key fallback these keys hit
|
||||
// before being listed here — this only kills the
|
||||
// per-boot "Unknown key format" warning)
|
||||
isSingletonSystemKey(id) // Known singletons (e.g. brainy:entityIdMapper) hit the
|
||||
// same warn-then-route fallback without this — the
|
||||
// routing below already handles them identically
|
||||
id.startsWith('__sparse_index__') // Metadata sparse indices (zone maps + bloom filters)
|
||||
|
||||
if (isSystemKey) {
|
||||
if (isSingletonSystemKey(id)) {
|
||||
|
|
@ -849,64 +842,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
return this.deleteObjectFromPath(path)
|
||||
}
|
||||
|
||||
// ==========================================================================
|
||||
// Fact-scan capability — the seam through which an index provider holding
|
||||
// only `storage` reaches the generation fact log. The HOST brain wires the
|
||||
// source at init (a closure over its live fact log, so restore/reopen stays
|
||||
// transparent); providers call storage.scanFacts?.(...) and fall back to the
|
||||
// enumeration walk when it returns null. Providers never construct a
|
||||
// fact-log reader themselves — the log's open path is writer-side.
|
||||
// ==========================================================================
|
||||
|
||||
/** The host-wired fact-scan source (closures over the live generation state). */
|
||||
private _factScanSource: {
|
||||
factLog: () => import('../db/factLog.js').FactLog | null
|
||||
committedGeneration: () => number
|
||||
} | null = null
|
||||
|
||||
/** HOST-ONLY: wire (or clear) the fact-scan capability's source. */
|
||||
public setFactScanSource(
|
||||
source: {
|
||||
factLog: () => import('../db/factLog.js').FactLog | null
|
||||
committedGeneration: () => number
|
||||
} | null
|
||||
): void {
|
||||
this._factScanSource = source
|
||||
}
|
||||
|
||||
/** Open a scan over committed facts, or `null` when no fact log exists. */
|
||||
public scanFacts(options?: {
|
||||
fromGeneration?: number
|
||||
toGeneration?: number
|
||||
kinds?: Array<'noun' | 'verb'>
|
||||
batchSize?: number
|
||||
}): import('../db/factLog.js').FactScanHandle | null {
|
||||
const log = this._factScanSource?.factLog()
|
||||
return log ? log.scanFacts(options) : null
|
||||
}
|
||||
|
||||
/** The fact log's head generation, or `null` when no fact log exists. */
|
||||
public factLogHeadGeneration(): number | null {
|
||||
const log = this._factScanSource?.factLog()
|
||||
return log ? log.headGeneration() : null
|
||||
}
|
||||
|
||||
/**
|
||||
* The COMMITTED generation watermark (the manifest truth) — the value a
|
||||
* projection's `sourceGeneration` compares against. Exposed here so a
|
||||
* provider never parses the store's private manifest format. `null` when
|
||||
* the capability is unwired (no host, or a bare adapter).
|
||||
*/
|
||||
public committedGeneration(): number | null {
|
||||
return this._factScanSource ? this._factScanSource.committedGeneration() : null
|
||||
}
|
||||
|
||||
/** Sealed fact-segment paths (zero-copy handoff); empty when none. */
|
||||
public factSegmentPaths(options?: { fromGeneration?: number }): string[] {
|
||||
const log = this._factScanSource?.factLog()
|
||||
return log ? log.segmentPaths(options) : []
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Remove the container that held a canonical entity's leg files,
|
||||
* called after both legs are deleted so a delete leaves NOTHING behind — no
|
||||
|
|
@ -1681,7 +1616,7 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
/**
|
||||
* Delete a noun from storage
|
||||
*/
|
||||
public async deleteNoun(id: string, priorMetadata?: NounMetadata | null): Promise<void> {
|
||||
public async deleteNoun(id: string): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
// FULL removal (live-HEAD hygiene): remove BOTH canonical legs AND the
|
||||
|
|
@ -1696,9 +1631,7 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
// success); a REAL fault must surface loudly (loud errors, never quiet
|
||||
// losses) and must never silently skip the count decrement — so this is NO
|
||||
// LONGER wrapped in a blind catch that masked faults as "file didn't exist".
|
||||
// `priorMetadata` (the caller's pre-delete read) keeps the decrement honest
|
||||
// even when the canonical read inside returns null (replace race / ghost).
|
||||
await this.deleteNounMetadata(id, priorMetadata)
|
||||
await this.deleteNounMetadata(id)
|
||||
|
||||
// Remove the now-empty entity container (a no-op for key/prefix stores).
|
||||
await this.removeCanonicalContainer(getNounVectorPath(id))
|
||||
|
|
@ -2068,20 +2001,11 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
// Cursor (8.0): resume token carrying the (shard, nounId) of the last returned
|
||||
// noun — the noun mirror of getVerbsWithPagination. When present it supersedes
|
||||
// `offset` and resumes the shard walk immediately AFTER that position, so a full
|
||||
// walk is O(N) instead of the O(N²) of offset paging. (Previously the cursor was
|
||||
// ignored, which was latent — the only multi-page consumer used a single big
|
||||
// page — until small chunk sizes needed page 2 and an offset-0-on-every-call
|
||||
// walk never terminated.)
|
||||
// walk is O(N) instead of the O(N²) of offset paging. Malformed/foreign tokens
|
||||
// decode to null → offset fallback. (Previously the cursor was ignored, which
|
||||
// was latent — the only multi-page consumer used a single big page — until small
|
||||
// chunk sizes needed page 2 and an offset-0-on-every-call walk never terminated.)
|
||||
const cursor = this.decodeNounWalkCursor(options.cursor)
|
||||
if (options.cursor && cursor === null) {
|
||||
// A supplied-but-undecodable resume token must FAIL, not silently restart
|
||||
// at offset 0 — to a while(hasMore) caller the silent fallback re-serves
|
||||
// page 1 forever: an unbounded CPU loop wearing pagination's clothes.
|
||||
throw BrainyError.storage(
|
||||
`getNouns: invalid pagination cursor '${options.cursor}' — cannot resume this walk. ` +
|
||||
`Restart it without a cursor.`
|
||||
)
|
||||
}
|
||||
const collected: Array<{ noun: HNSWNounWithMetadata; shard: number }> = []
|
||||
|
||||
// Peek one past the window so `hasMore` is decidable. Cursor mode collects one
|
||||
|
|
@ -2230,14 +2154,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
|
||||
const { limit, offset = 0, filter } = options
|
||||
const cursor = this.decodeNounWalkCursor(options.cursor)
|
||||
if (options.cursor && cursor === null) {
|
||||
// Same law as getNouns/getVerbs: an undecodable resume token FAILS
|
||||
// instead of silently restarting the walk at offset 0.
|
||||
throw BrainyError.storage(
|
||||
`getNounIds: invalid pagination cursor '${options.cursor}' — cannot resume this walk. ` +
|
||||
`Restart it without a cursor.`
|
||||
)
|
||||
}
|
||||
const collected: Array<{ id: string; shard: number }> = []
|
||||
const peekCount = cursor ? limit + 1 : offset + limit + 1
|
||||
const startShard = cursor ? cursor.shard : 0
|
||||
|
|
@ -2383,17 +2299,8 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
// (shard, verbId) of the last returned verb. When present it SUPERSEDES `offset`
|
||||
// and resumes the shard walk immediately AFTER that position, so a full walk is
|
||||
// O(N) total instead of the O(N²) of offset paging (which re-scans from shard 0
|
||||
// every page).
|
||||
// every page). Malformed / foreign tokens decode to null → offset fallback.
|
||||
const cursor = this.decodeVerbWalkCursor(options.cursor)
|
||||
if (options.cursor && cursor === null) {
|
||||
// A supplied-but-undecodable resume token must FAIL, not silently restart
|
||||
// at offset 0 — to a while(hasMore) caller the silent fallback re-serves
|
||||
// page 1 forever: an unbounded CPU loop wearing pagination's clothes.
|
||||
throw BrainyError.storage(
|
||||
`getVerbs: invalid pagination cursor '${options.cursor}' — cannot resume this walk. ` +
|
||||
`Restart it without a cursor.`
|
||||
)
|
||||
}
|
||||
|
||||
// Each collected entry remembers its shard so nextCursor can point at the exact
|
||||
// (shard, id) resume position.
|
||||
|
|
@ -3029,15 +2936,14 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
/**
|
||||
* Delete a verb from storage
|
||||
*/
|
||||
public async deleteVerb(id: string, priorMetadata?: VerbMetadata | null): Promise<void> {
|
||||
public async deleteVerb(id: string): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
// FULL removal (see deleteNoun): both canonical legs + the verb container.
|
||||
// The blind catch that masked real faults as "no metadata file" is gone — a
|
||||
// genuine absence is already a no-op downstream; a real fault surfaces loudly.
|
||||
// `priorMetadata` keeps the count decrement honest on a null internal read.
|
||||
await this.deleteVerb_internal(id) // vectors.json leg
|
||||
await this.deleteVerbMetadata(id, priorMetadata) // metadata leg + count decrement
|
||||
await this.deleteVerbMetadata(id) // metadata leg + count decrement
|
||||
await this.removeCanonicalContainer(getVerbVectorPath(id))
|
||||
}
|
||||
/**
|
||||
|
|
@ -3690,32 +3596,19 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
}
|
||||
|
||||
/**
|
||||
* Delete noun metadata from storage (ID-first, O(1) delete).
|
||||
*
|
||||
* @param id - The entity id.
|
||||
* @param priorRecord - OPTIONAL already-known metadata of the entity being
|
||||
* removed (e.g. the pre-delete read `remove()` performs, or a captured
|
||||
* before-image). The count decrement must never REQUIRE re-reading the
|
||||
* thing being removed: when the canonical read here returns `null` (a
|
||||
* replace race, or a partial-delete ghost from an earlier version) the
|
||||
* decrement falls back to this record instead of being silently skipped —
|
||||
* the skip permanently inflated the persisted totals (adds counted, paired
|
||||
* removals not decremented), and `Math.max(totalNounCount, scanned)` made
|
||||
* the inflation unfixable by any disk cleanup.
|
||||
* Delete noun metadata from storage (ID-first, O(1) delete)
|
||||
*/
|
||||
public async deleteNounMetadata(id: string, priorRecord?: NounMetadata | null): Promise<void> {
|
||||
public async deleteNounMetadata(id: string): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
// Direct O(1) delete with ID-first path. Read the canonical record BEFORE
|
||||
// removing it: the per-type and subtype decrements are sourced from the
|
||||
// entity's own metadata (`noun` type, `subtype`, `visibility`) rather than an
|
||||
// id-keyed cache, keeping type-statistics honest across deletes — symmetric
|
||||
// with the increments in `saveNounMetadata_internal()`. A null read falls
|
||||
// back to the caller-provided prior record (see @param priorRecord).
|
||||
// with the increments in `saveNounMetadata_internal()`.
|
||||
const path = getNounMetadataPath(id)
|
||||
const read = await this.readCanonicalObject(path)
|
||||
const record = await this.readCanonicalObject(path)
|
||||
await this.deleteCanonicalObject(path)
|
||||
const record = read ?? priorRecord
|
||||
|
||||
const priorType = record?.noun as NounType | undefined
|
||||
// 8.0 visibility: an internal/system entity was never added to `nounCountsByType`
|
||||
|
|
@ -3889,20 +3782,16 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
/**
|
||||
* Delete verb metadata from storage (ID-first, O(1) delete)
|
||||
*/
|
||||
public async deleteVerbMetadata(id: string, priorRecord?: VerbMetadata | null): Promise<void> {
|
||||
public async deleteVerbMetadata(id: string): Promise<void> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
// Direct O(1) delete with ID-first path. Read the canonical record BEFORE
|
||||
// removing it so every decrement is sourced from the edge's own metadata
|
||||
// (`verb` type, `subtype`, `visibility`) rather than an id-keyed cache —
|
||||
// symmetric with the increments in `saveVerbMetadata_internal()`. A null
|
||||
// read falls back to the caller-provided prior record: the decrement must
|
||||
// never REQUIRE re-reading the thing being removed (a silent skip minted
|
||||
// permanent counter inflation — see deleteNounMetadata).
|
||||
// symmetric with the increments in `saveVerbMetadata_internal()`.
|
||||
const path = getVerbMetadataPath(id)
|
||||
const read = await this.readCanonicalObject(path)
|
||||
const record = await this.readCanonicalObject(path)
|
||||
await this.deleteCanonicalObject(path)
|
||||
const record = read ?? priorRecord
|
||||
|
||||
const priorVerb = record?.verb as VerbType | undefined
|
||||
// Symmetric count decrement (previously OMITTED — verb deletes touched neither the
|
||||
|
|
@ -4339,17 +4228,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
this.nounCountsByType = new Uint32Array(NOUN_TYPE_COUNT)
|
||||
this.verbCountsByType = new Uint32Array(VERB_TYPE_COUNT)
|
||||
|
||||
// The SAME walk also rebuilds the user-facing scalar totals + per-type maps
|
||||
// persisted in counts.json (totalNounCount / totalVerbCount / entityCounts /
|
||||
// verbCounts). Previously only the type-statistics arrays were rebuilt and
|
||||
// the total was computed just to LOG it — so a drifted persisted scalar
|
||||
// (deletes whose decrement was skipped) survived every "rebuild" forever,
|
||||
// and Math.max(totalNounCount, scanned) made the inflation unfixable by any
|
||||
// disk cleanup. This method is now the SANCTIONED RECOUNT: one canonical
|
||||
// walk, every counter rollup rebuilt and persisted from it.
|
||||
const countedNouns = new Map<string, number>()
|
||||
const countedVerbs = new Map<string, number>()
|
||||
|
||||
// Scan noun shards
|
||||
for (let shard = 0; shard < 256; shard++) {
|
||||
const shardHex = shard.toString(16).padStart(2, '0')
|
||||
|
|
@ -4370,7 +4248,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
if (typeIndex >= 0 && typeIndex < NOUN_TYPE_COUNT) {
|
||||
this.nounCountsByType[typeIndex]++
|
||||
}
|
||||
countedNouns.set(metadata.noun, (countedNouns.get(metadata.noun) || 0) + 1)
|
||||
}
|
||||
}
|
||||
} catch (error) {
|
||||
|
|
@ -4402,7 +4279,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
if (typeIndex >= 0 && typeIndex < VERB_TYPE_COUNT) {
|
||||
this.verbCountsByType[typeIndex]++
|
||||
}
|
||||
countedVerbs.set(metadata.verb, (countedVerbs.get(metadata.verb) || 0) + 1)
|
||||
}
|
||||
}
|
||||
} catch (error) {
|
||||
|
|
@ -4419,19 +4295,7 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
|
||||
const totalVerbs = this.verbCountsByType.reduce((sum, count) => sum + count, 0)
|
||||
const totalNouns = this.nounCountsByType.reduce((sum, count) => sum + count, 0)
|
||||
|
||||
// The sanctioned recount half: replace the user-facing scalar totals + the
|
||||
// per-type maps with the walk's truth and PERSIST them (counts.json), so an
|
||||
// inflated persisted counter is actually corrected — not merely out-voted
|
||||
// in memory until the next reopen rehydrates the stale file.
|
||||
this.entityCounts = countedNouns
|
||||
this.verbCounts = countedVerbs
|
||||
this.totalNounCount = totalNouns
|
||||
this.totalVerbCount = totalVerbs
|
||||
this.countCache.clear()
|
||||
await this.persistCounts()
|
||||
|
||||
prodLog.info(`[BaseStorage] Rebuilt counts: ${totalNouns} nouns, ${totalVerbs} verbs (scalar + per-type persisted)`)
|
||||
prodLog.info(`[BaseStorage] Rebuilt counts: ${totalNouns} nouns, ${totalVerbs} verbs`)
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
|
|||
|
|
@ -37,24 +37,6 @@ const DEFAULT_OPTIONS: Required<TransactionOptions> = {
|
|||
maxRollbackRetries: 3
|
||||
}
|
||||
|
||||
/**
|
||||
* The apply-phase budget for a batch of `opCount` operations.
|
||||
*
|
||||
* An explicit override wins untouched. Otherwise the budget SCALES with the
|
||||
* batch: `max(30 000 ms, opCount × 2 000 ms)`. The per-op term is calibrated
|
||||
* from field data — bulk imports on network-attached disks measure ~2 s per
|
||||
* operation (each op pays canonical writes + fsync + index maintenance) — so
|
||||
* a flat 30 s budget silently capped honest work at ~15 operations while
|
||||
* looking generous for small batches. Scaling keeps small transacts
|
||||
* fast-failing and gives bulk ones a budget proportional to the work they
|
||||
* actually asked for; a trip still rolls back atomically and throws a
|
||||
* retryable, fully-labeled TransactionTimeoutError.
|
||||
*/
|
||||
export function transactTimeoutBudget(opCount: number, override?: number): number {
|
||||
if (override !== undefined) return override
|
||||
return Math.max(30_000, opCount * 2_000)
|
||||
}
|
||||
|
||||
/**
|
||||
* Transaction class
|
||||
*/
|
||||
|
|
@ -132,11 +114,7 @@ export class Transaction implements TransactionContext {
|
|||
// into the catch below and rolls back like any other failure — it must
|
||||
// never bypass rollback.
|
||||
if (Date.now() - this.startTime > this.options.timeout) {
|
||||
throw new TransactionTimeoutError(this.options.timeout, i, {
|
||||
elapsedMs: Date.now() - this.startTime,
|
||||
totalOperations: this.operations.length,
|
||||
operationName: this.operations[i]?.name
|
||||
})
|
||||
throw new TransactionTimeoutError(this.options.timeout, i)
|
||||
}
|
||||
|
||||
const operation = this.operations[i]
|
||||
|
|
|
|||
|
|
@ -77,24 +77,11 @@ export class InvalidTransactionStateError extends TransactionError {
|
|||
export class TransactionTimeoutError extends TransactionError {
|
||||
constructor(
|
||||
timeoutMs: number,
|
||||
operationIndex: number,
|
||||
telemetry?: {
|
||||
elapsedMs?: number
|
||||
totalOperations?: number
|
||||
operationName?: string
|
||||
}
|
||||
operationIndex: number
|
||||
) {
|
||||
const progress =
|
||||
telemetry?.totalOperations !== undefined
|
||||
? `${operationIndex}/${telemetry.totalOperations}`
|
||||
: String(operationIndex)
|
||||
const name = telemetry?.operationName ? ` ('${telemetry.operationName}')` : ''
|
||||
const elapsed =
|
||||
telemetry?.elapsedMs !== undefined ? `${telemetry.elapsedMs}ms elapsed, ` : ''
|
||||
super(
|
||||
`Transaction timed out at operation ${progress}${name} — ${elapsed}budget ${timeoutMs}ms. ` +
|
||||
`The batch rolled back atomically; retry with a higher timeoutMs or a smaller batch.`,
|
||||
{ timeoutMs, operationIndex, ...telemetry }
|
||||
`Transaction timed out after ${timeoutMs}ms at operation ${operationIndex}`,
|
||||
{ timeoutMs, operationIndex }
|
||||
)
|
||||
this.name = 'TransactionTimeoutError'
|
||||
}
|
||||
|
|
|
|||
|
|
@ -127,33 +127,22 @@ export class DeleteNounMetadataOperation implements Operation {
|
|||
|
||||
constructor(
|
||||
private readonly storage: StorageAdapter,
|
||||
private readonly id: string,
|
||||
/**
|
||||
* OPTIONAL already-known metadata of the entity being removed (the caller's
|
||||
* pre-delete read). Removal must never REQUIRE re-reading the thing being
|
||||
* removed: if the reads here return null (replace race, or a ghost left by
|
||||
* an earlier version), the count decrement downstream falls back to this
|
||||
* record instead of being silently skipped — the skip minted permanent
|
||||
* counter inflation (adds counted, paired removals not decremented).
|
||||
*/
|
||||
private readonly priorMetadata?: NounMetadata | null
|
||||
private readonly id: string
|
||||
) {}
|
||||
|
||||
async execute(): Promise<RollbackAction> {
|
||||
// Capture the FULL before-image (both legs) so the undo restores the whole
|
||||
// entity — a metadata-only rollback would leave the vector leg unrestored.
|
||||
// A null metadata read falls back to the caller's pre-delete read.
|
||||
const previousNoun = await this.storage.getNoun(this.id)
|
||||
const previousMetadata = (await this.storage.getNounMetadata(this.id)) ?? this.priorMetadata ?? null
|
||||
const previousMetadata = await this.storage.getNounMetadata(this.id)
|
||||
|
||||
if (!previousNoun && !previousMetadata) {
|
||||
// Nothing to delete - no rollback needed
|
||||
return async () => {}
|
||||
}
|
||||
|
||||
// Full removal: both canonical legs + the entity container + count decrement
|
||||
// (the prior record keeps the decrement honest on a null canonical read).
|
||||
await this.storage.deleteNoun(this.id, previousMetadata)
|
||||
// Full removal: both canonical legs + the entity container + count decrement.
|
||||
await this.storage.deleteNoun(this.id)
|
||||
|
||||
// Return rollback action
|
||||
return async () => {
|
||||
|
|
@ -279,9 +268,9 @@ export class DeleteVerbMetadataOperation implements Operation {
|
|||
return async () => {}
|
||||
}
|
||||
|
||||
// Delete verb (metadata + vector). The pre-read rides along so the count
|
||||
// decrement never depends on re-reading the record being removed.
|
||||
await this.storage.deleteVerb(this.id, previousMetadata)
|
||||
// Delete verb (metadata + vector)
|
||||
// Note: StorageAdapter has deleteVerb but not deleteVerbMetadata
|
||||
await this.storage.deleteVerb(this.id)
|
||||
|
||||
// Return rollback action
|
||||
return async () => {
|
||||
|
|
|
|||
|
|
@ -1737,10 +1737,7 @@ export interface BrainyConfig {
|
|||
* Under Model-B EVERY write (`transact()` AND single-op `add`/`update`/
|
||||
* `remove`/`relate`) produces an immutable generation record-set serving
|
||||
* historical reads (`asOf()`, pinned `Db` values). Without compaction those
|
||||
* accumulate, so Brainy **auto-compacts at `close()`** (time-bounded per
|
||||
* pass; 8.9.0 removed compaction from `flush()` — flush is durability work
|
||||
* and never pays maintenance costs). A long-lived writer that never closes
|
||||
* accumulates history until its next explicit `compactHistory()` call.
|
||||
* accumulate, so Brainy **auto-compacts on every `flush()` and `close()`**.
|
||||
* Live `Db` pins are ALWAYS exempt from reclamation, in every mode.
|
||||
*
|
||||
* Modes:
|
||||
|
|
@ -1757,7 +1754,7 @@ export interface BrainyConfig {
|
|||
* the oldest unpinned generations while ANY supplied cap is exceeded
|
||||
* (predictable ops). `maxAge` in ms; `maxBytes` total history bytes.
|
||||
*
|
||||
* `autoCompact: false` disables the automatic close() compaction (manage
|
||||
* `autoCompact: false` disables the automatic flush/close compaction (manage
|
||||
* manually via `brain.compactHistory()`). `budgetBytes` is the settable
|
||||
* adaptive byte budget a coordinator drives (also via `brain.setRetentionBudget()`).
|
||||
* Long-term archives belong in `db.persist(path)` snapshots, which compaction
|
||||
|
|
@ -1775,7 +1772,7 @@ export interface BrainyConfig {
|
|||
maxBytes?: number
|
||||
/** Adaptive byte budget for this brain, driven by a coordinator (e.g. cor). */
|
||||
budgetBytes?: number
|
||||
/** Run compaction automatically at close() (default: true; 8.9.0 — flush() never compacts). */
|
||||
/** Run compaction automatically on flush()/close() (default: true). */
|
||||
autoCompact?: boolean
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -1,43 +0,0 @@
|
|||
/**
|
||||
* @module utils/crc32c
|
||||
* @description CRC-32C (Castagnoli, polynomial 0x1EDC6F41, reflected 0x82F63B78)
|
||||
* — the storage-industry frame checksum (ext4, iSCSI, SCTP, LSM segment files).
|
||||
* Used to frame generation-fact segments: every appended record carries the
|
||||
* CRC-32C of its payload, so a torn tail (crash mid-append) or bit rot is
|
||||
* DETECTED at scan time and never silently read as data.
|
||||
*
|
||||
* Table-driven, dependency-free reference implementation. Native providers may
|
||||
* substitute a hardware-accelerated (SSE4.2 / ARMv8 CRC) implementation — the
|
||||
* polynomial is the contract, byte-identical results required.
|
||||
*/
|
||||
|
||||
/** The 256-entry lookup table for the reflected CRC-32C polynomial. */
|
||||
const TABLE: Uint32Array = (() => {
|
||||
const table = new Uint32Array(256)
|
||||
for (let n = 0; n < 256; n++) {
|
||||
let c = n
|
||||
for (let k = 0; k < 8; k++) {
|
||||
c = c & 1 ? 0x82f63b78 ^ (c >>> 1) : c >>> 1
|
||||
}
|
||||
table[n] = c >>> 0
|
||||
}
|
||||
return table
|
||||
})()
|
||||
|
||||
/**
|
||||
* Compute the CRC-32C checksum of a byte buffer.
|
||||
*
|
||||
* Known-answer vectors (RFC 3720 appendix / the standard test suite):
|
||||
* - ASCII "123456789" → 0xE3069283
|
||||
* - 32 zero bytes → 0x8A9136AA
|
||||
*
|
||||
* @param bytes - The payload to checksum.
|
||||
* @returns The CRC-32C as an unsigned 32-bit integer.
|
||||
*/
|
||||
export function crc32c(bytes: Uint8Array): number {
|
||||
let crc = 0xffffffff
|
||||
for (let i = 0; i < bytes.length; i++) {
|
||||
crc = TABLE[(crc ^ bytes[i]) & 0xff] ^ (crc >>> 8)
|
||||
}
|
||||
return (crc ^ 0xffffffff) >>> 0
|
||||
}
|
||||
|
|
@ -1867,26 +1867,9 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
// not once per AND-clause inside it.
|
||||
const unindexedFields: string[] = []
|
||||
|
||||
for (const [rawField, condition] of Object.entries(filter)) {
|
||||
for (const [field, condition] of Object.entries(filter)) {
|
||||
// Skip logical operators
|
||||
if (rawField === 'allOf' || rawField === 'anyOf' || rawField === 'not') continue
|
||||
|
||||
// Metadata is FLATTENED at index time (metadata.entry.title indexes as
|
||||
// entry.title), so a `metadata.`-prefixed where key is almost always
|
||||
// the caller spelling the STORAGE shape rather than the index shape.
|
||||
// Accept both spellings: when the key as spelled is unindexed but its
|
||||
// stripped spelling is, query the stripped one. A literal nested
|
||||
// custom key named `metadata` still wins when indexed as spelled
|
||||
// (checked first), so that rare shape keeps working.
|
||||
let field = rawField
|
||||
if (
|
||||
rawField.startsWith('metadata.') &&
|
||||
this.columnStore &&
|
||||
!this.columnStore.hasField(rawField) &&
|
||||
this.columnStore.hasField(rawField.slice('metadata.'.length))
|
||||
) {
|
||||
field = rawField.slice('metadata.'.length)
|
||||
}
|
||||
if (field === 'allOf' || field === 'anyOf' || field === 'not') continue
|
||||
|
||||
let fieldResults: string[] = []
|
||||
|
||||
|
|
|
|||
|
|
@ -1,141 +0,0 @@
|
|||
/**
|
||||
* @module utils/osLimits
|
||||
* @description Detect-and-warn for OS resource limits that bite at POOL scale.
|
||||
*
|
||||
* A single brain rarely notices them, but a pool of brains — especially with a
|
||||
* native accelerator memory-mapping many index files per brain — consumes file
|
||||
* descriptors and memory mappings multiplicatively. On stock Linux defaults
|
||||
* (RLIMIT_NOFILE soft 1024, vm.max_map_count 65530) the failure arrives as
|
||||
* EMFILE or a failed mmap deep inside an index open, long after the real cause
|
||||
* (the limit) stopped being visible. This module reads the limits at open and
|
||||
* WARNS ONCE per process with the exact raise commands, so the operator learns
|
||||
* the fix before the incident instead of from it.
|
||||
*
|
||||
* Read-only and Linux-only by construction: both sources are `/proc` files.
|
||||
* On platforms where they are absent the check reports nulls and stays silent —
|
||||
* no limit read means no claim made, never a guessed warning.
|
||||
*/
|
||||
|
||||
import * as fs from 'node:fs'
|
||||
import { prodLog } from './logger.js'
|
||||
|
||||
/** Soft-NOFILE floor below which pool-scale use is at EMFILE risk. */
|
||||
export const NOFILE_POOL_FLOOR = 65536
|
||||
|
||||
/** vm.max_map_count floor below which mmap-heavy native indexes are at risk. */
|
||||
export const MAX_MAP_COUNT_POOL_FLOOR = 262144
|
||||
|
||||
export interface OsLimitsReport {
|
||||
/** RLIMIT_NOFILE soft limit (null when unreadable; Infinity for 'unlimited'). */
|
||||
nofileSoft: number | null
|
||||
/** RLIMIT_NOFILE hard limit (null when unreadable; Infinity for 'unlimited'). */
|
||||
nofileHard: number | null
|
||||
/** vm.max_map_count (null when unreadable). */
|
||||
maxMapCount: number | null
|
||||
/** Human-actionable warnings for limits below the pool floors. Empty = fine. */
|
||||
warnings: string[]
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse the `Max open files` row of a `/proc/<pid>/limits` document into
|
||||
* soft/hard values. Returns nulls when the row is absent or malformed.
|
||||
*/
|
||||
export function parseProcLimits(content: string): { soft: number | null; hard: number | null } {
|
||||
const line = content.split('\n').find((l) => l.startsWith('Max open files'))
|
||||
if (!line) return { soft: null, hard: null }
|
||||
const m = line.match(/^Max open files\s+(\S+)\s+(\S+)/)
|
||||
if (!m) return { soft: null, hard: null }
|
||||
const parse = (v: string): number | null => {
|
||||
if (v === 'unlimited') return Infinity
|
||||
const n = Number.parseInt(v, 10)
|
||||
return Number.isNaN(n) ? null : n
|
||||
}
|
||||
return { soft: parse(m[1]), hard: parse(m[2]) }
|
||||
}
|
||||
|
||||
/**
|
||||
* Assess readable limits against the pool floors. Pure — feed it any values.
|
||||
* A null (unreadable) limit produces NO warning: no measurement, no claim.
|
||||
*/
|
||||
export function assessOsLimits(limits: {
|
||||
nofileSoft: number | null
|
||||
nofileHard: number | null
|
||||
maxMapCount: number | null
|
||||
}): string[] {
|
||||
const warnings: string[] = []
|
||||
|
||||
if (limits.nofileSoft !== null && limits.nofileSoft < NOFILE_POOL_FLOOR) {
|
||||
const hardNote =
|
||||
limits.nofileHard !== null && limits.nofileHard >= NOFILE_POOL_FLOOR
|
||||
? ` (the hard limit ${limits.nofileHard === Infinity ? 'unlimited' : limits.nofileHard} already allows it — raise the soft limit only)`
|
||||
: ''
|
||||
warnings.push(
|
||||
`RLIMIT_NOFILE soft limit is ${limits.nofileSoft} — below the ${NOFILE_POOL_FLOOR} recommended ` +
|
||||
`for pool-scale use (a pool of brains with a native accelerator opens many index files per brain; ` +
|
||||
`the failure mode is EMFILE deep inside an index open). Raise with \`ulimit -n ${NOFILE_POOL_FLOOR}\` ` +
|
||||
`or LimitNOFILE=${NOFILE_POOL_FLOOR} in the service unit${hardNote}.`
|
||||
)
|
||||
}
|
||||
|
||||
if (limits.maxMapCount !== null && limits.maxMapCount < MAX_MAP_COUNT_POOL_FLOOR) {
|
||||
warnings.push(
|
||||
`vm.max_map_count is ${limits.maxMapCount} — below the ${MAX_MAP_COUNT_POOL_FLOOR} recommended ` +
|
||||
`for mmap-heavy native indexes at pool scale (each mapped index segment consumes map entries; ` +
|
||||
`the failure mode is a failed mmap mid-heal). Raise with ` +
|
||||
`\`sysctl -w vm.max_map_count=${MAX_MAP_COUNT_POOL_FLOOR}\` (persist in /etc/sysctl.d/).`
|
||||
)
|
||||
}
|
||||
|
||||
return warnings
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the limits from /proc and assess them. `readFile` is injectable for
|
||||
* tests; absent/unreadable sources yield nulls (and therefore no warnings).
|
||||
*/
|
||||
export async function checkOsLimits(
|
||||
readFile: (path: string) => Promise<string> = async (p) => fs.promises.readFile(p, 'utf-8')
|
||||
): Promise<OsLimitsReport> {
|
||||
let nofileSoft: number | null = null
|
||||
let nofileHard: number | null = null
|
||||
let maxMapCount: number | null = null
|
||||
|
||||
try {
|
||||
const parsed = parseProcLimits(await readFile('/proc/self/limits'))
|
||||
nofileSoft = parsed.soft
|
||||
nofileHard = parsed.hard
|
||||
} catch {
|
||||
// Not Linux (or /proc unavailable) — no measurement, no claim.
|
||||
}
|
||||
|
||||
try {
|
||||
const raw = (await readFile('/proc/sys/vm/max_map_count')).trim()
|
||||
const n = Number.parseInt(raw, 10)
|
||||
maxMapCount = Number.isNaN(n) ? null : n
|
||||
} catch {
|
||||
// Not Linux — same rule.
|
||||
}
|
||||
|
||||
const warnings = assessOsLimits({ nofileSoft, nofileHard, maxMapCount })
|
||||
return { nofileSoft, nofileHard, maxMapCount, warnings }
|
||||
}
|
||||
|
||||
/** Once-per-process latch so a brain pool warns once, not once per brain. */
|
||||
let osLimitsWarned = false
|
||||
|
||||
/**
|
||||
* Run the check and warn (once per process) about limits below the pool
|
||||
* floors. Called from brain open; safe everywhere (silent off-Linux).
|
||||
*/
|
||||
export async function warnOnLowOsLimits(): Promise<void> {
|
||||
if (osLimitsWarned) return
|
||||
osLimitsWarned = true
|
||||
try {
|
||||
const report = await checkOsLimits()
|
||||
for (const warning of report.warnings) {
|
||||
prodLog.warn(`[Brainy] OS limit check: ${warning}`)
|
||||
}
|
||||
} catch {
|
||||
// The check must never affect open — measurement-only.
|
||||
}
|
||||
}
|
||||
|
|
@ -209,10 +209,8 @@ export interface ValidationConfigOptions {
|
|||
}
|
||||
|
||||
/**
|
||||
* Auto-configured limits based on system resources.
|
||||
* Derived from memory (explicit overrides > reserved memory > container limit >
|
||||
* free memory). Query timing is recorded for diagnostics only — it never
|
||||
* changes the cap (see `recordQuery`).
|
||||
* Auto-configured limits based on system resources
|
||||
* These adapt to available memory and observed performance
|
||||
*/
|
||||
export class ValidationConfig {
|
||||
private static instance: ValidationConfig | null = null
|
||||
|
|
@ -325,23 +323,24 @@ export class ValidationConfig {
|
|||
}
|
||||
|
||||
/**
|
||||
* Record query timing for diagnostics. Telemetry ONLY — never mutates the cap.
|
||||
*
|
||||
* `maxLimit` is a MEMORY-protection bound; query duration says nothing about
|
||||
* memory-per-result, so duration must never drive it. An earlier version
|
||||
* "learned" here: while the lifetime-average query time exceeded 1s it shrank
|
||||
* `maxLimit` by 20% per recorded query down to a floor of 1000 — below the
|
||||
* documented `MIN_AUTO_QUERY_LIMIT` (10 000) and with no recovery once slow
|
||||
* samples poisoned the cumulative average. On a production host a burst of
|
||||
* slow aggregate queries silently strangled every consumer's `find()` to
|
||||
* 1000 while the error message blamed "available free memory" — exactly the
|
||||
* silent throttling this module's own contract forbids. The cap now comes
|
||||
* from its construction-time basis (or explicit overrides) alone.
|
||||
* Learn from actual usage to adjust limits
|
||||
*/
|
||||
recordQuery(duration: number, resultCount: number) {
|
||||
void resultCount
|
||||
this.queryCount++
|
||||
this.avgQueryTime = (this.avgQueryTime * (this.queryCount - 1) + duration) / this.queryCount
|
||||
|
||||
// Only auto-adjust if not using explicit overrides
|
||||
if (this.limitBasis !== 'override') {
|
||||
// If queries are consistently fast with large results, increase limits
|
||||
if (this.avgQueryTime < 100 && resultCount > this.maxLimit * 0.8) {
|
||||
this.maxLimit = Math.min(this.maxLimit * 1.5, 100000)
|
||||
}
|
||||
|
||||
// If queries are slow, reduce limits
|
||||
if (this.avgQueryTime > 1000) {
|
||||
this.maxLimit = Math.max(this.maxLimit * 0.8, 1000)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -1954,27 +1954,15 @@ export class VirtualFileSystem implements IVirtualFileSystem {
|
|||
const newParentPath = this.getParentPath(newPath)
|
||||
|
||||
if (oldParentPath !== newParentPath) {
|
||||
// Remove the OLD parent's containment edge(s) — by edge id, resolved from
|
||||
// the graph's own adjacency (the removal law: a removal never requires
|
||||
// reading the thing being removed). This step used to be skipped as "not
|
||||
// critical", which left the moved entity a child of BOTH directories:
|
||||
// readdir(oldDir) kept listing it, re-creating the old path showed the
|
||||
// name twice, and tree-walking consumers saw the file in two places.
|
||||
// Remove from old parent
|
||||
if (oldParentPath) {
|
||||
const oldParentId = await this.pathResolver.resolve(oldParentPath)
|
||||
const staleEdges = await this.brain.related({
|
||||
from: oldParentId,
|
||||
to: entityId,
|
||||
type: VerbType.Contains
|
||||
})
|
||||
for (const edge of staleEdges) {
|
||||
await this.brain.unrelate(edge.id)
|
||||
}
|
||||
// unrelate takes the relation ID, not params - need to find and remove relation
|
||||
// For now, skip unrelate as it's not critical for rename
|
||||
}
|
||||
|
||||
// Add to the new parent. The root ('/') is a REAL parent — skipping it
|
||||
// orphaned a move-to-root out of readdir('/') entirely.
|
||||
if (newParentPath) {
|
||||
// Add to new parent
|
||||
if (newParentPath && newParentPath !== '/') {
|
||||
const newParentId = await this.pathResolver.resolve(newParentPath)
|
||||
await this.brain.relate({
|
||||
from: newParentId,
|
||||
|
|
@ -2015,95 +2003,6 @@ export class VirtualFileSystem implements IVirtualFileSystem {
|
|||
this.triggerWatchers(newPath, 'rename')
|
||||
}
|
||||
|
||||
/**
|
||||
* Reconcile every VFS entity's containment edges against its canonical
|
||||
* `metadata.path` — the path is the truth (maintained by write/rename); the
|
||||
* `Contains` edges are a projection of it. Heals the "cosmetic ghost" class
|
||||
* left by pre-fix renames that added the new parent's edge without removing
|
||||
* the old one (an entity listed in TWO directories; a re-created old path
|
||||
* showing its name twice), plus duplicate edges from the same parent left by
|
||||
* concurrent writers.
|
||||
*
|
||||
* CONSERVATIVE by design: only VFS containment edges (subtype
|
||||
* `'vfs-contains'` or `metadata.isVFS`) are ever touched — a user's own
|
||||
* knowledge-graph `Contains` edge between the same entities is never
|
||||
* removed. An entity whose expected parent path has no entity is logged
|
||||
* loudly and left alone (never orphaned further). Operator-invoked via
|
||||
* `brain.repairIndex()`.
|
||||
*
|
||||
* The canonical pagination walk (not an index query) is deliberate: this is
|
||||
* a repair op — the projections are the thing under suspicion, so the walk
|
||||
* reads the source of truth.
|
||||
*
|
||||
* @returns Counts of stale edges removed and missing expected edges restored.
|
||||
*/
|
||||
async repairContainment(): Promise<{ removed: number; restored: number }> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
// Pass 1: canonical walk → every VFS entity's id + path.
|
||||
const idByPath = new Map<string, string>()
|
||||
const vfsEntities: Array<{ id: string; path: string }> = []
|
||||
let cursor: string | undefined
|
||||
for (;;) {
|
||||
const page = await (this.brain as any).storage.getNounsWithPagination({ limit: 500, cursor })
|
||||
for (const noun of page.items) {
|
||||
const meta = (noun as any).metadata ?? noun
|
||||
const p = meta?.path
|
||||
if (meta?.vfsType && typeof p === 'string') {
|
||||
idByPath.set(p, noun.id)
|
||||
vfsEntities.push({ id: noun.id, path: p })
|
||||
}
|
||||
}
|
||||
if (!page.hasMore) break
|
||||
cursor = page.nextCursor
|
||||
}
|
||||
|
||||
let removed = 0
|
||||
let restored = 0
|
||||
for (const { id, path } of vfsEntities) {
|
||||
if (path === '/') continue // the root has no parent
|
||||
const expectedParentId = idByPath.get(this.getParentPath(path))
|
||||
if (!expectedParentId) {
|
||||
console.warn(
|
||||
`[VFS] repairContainment: no entity found for parent of ${path} — leaving its edges untouched.`
|
||||
)
|
||||
continue
|
||||
}
|
||||
|
||||
const incoming = await this.brain.related({ to: id, type: VerbType.Contains })
|
||||
let expectedSeen = false
|
||||
for (const edge of incoming) {
|
||||
const isVfsEdge = edge.subtype === 'vfs-contains' || (edge.metadata as any)?.isVFS === true
|
||||
if (!isVfsEdge) continue // never touch user knowledge edges
|
||||
if (edge.from === expectedParentId && !expectedSeen) {
|
||||
expectedSeen = true // keep exactly one correct edge
|
||||
continue
|
||||
}
|
||||
// Stale parent (a pre-fix rename ghost) or a duplicate of the correct
|
||||
// edge (concurrent-writer artifact) — remove it, loudly.
|
||||
await this.brain.unrelate(edge.id)
|
||||
removed++
|
||||
console.warn(
|
||||
`[VFS] repairContainment: removed ${edge.from === expectedParentId ? 'duplicate' : 'stale'} ` +
|
||||
`containment edge ${edge.from} -> ${id} (${path})`
|
||||
)
|
||||
}
|
||||
if (!expectedSeen) {
|
||||
await this.brain.relate({
|
||||
from: expectedParentId,
|
||||
to: id,
|
||||
type: VerbType.Contains,
|
||||
subtype: 'vfs-contains',
|
||||
metadata: { isVFS: true }
|
||||
})
|
||||
restored++
|
||||
console.warn(`[VFS] repairContainment: restored missing containment edge for ${path}`)
|
||||
}
|
||||
}
|
||||
|
||||
return { removed, restored }
|
||||
}
|
||||
|
||||
/**
|
||||
* Copy a file or directory to a new path.
|
||||
*
|
||||
|
|
|
|||
|
|
@ -1,169 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/aggregate-reserved-fields
|
||||
* @description One field-resolution law across the whole aggregation + query
|
||||
* surface (SELF-AGGREGATE-DELETE-DRIFT). Laws:
|
||||
* (1) aggregates grouped by a RESERVED field (subtype) decrement on delete —
|
||||
* the delete-side entity view carries every reserved field, so the
|
||||
* decrement finds its group (counts must never drift from ground truth);
|
||||
* (2) same for update: moving an entity between reserved-field groups
|
||||
* decrements the old group and increments the new one (no double-count);
|
||||
* (3) aggregation source.where on a reserved field (subtype) FILTERS instead
|
||||
* of silently matching nothing;
|
||||
* (4) removeMany refuses empty/invalid selectors loudly (bare array, empty
|
||||
* object, ids: []) instead of resolving as a silent no-op;
|
||||
* (5) find() accepts both where spellings: flattened (entry.title) and
|
||||
* storage-shaped (metadata.entry.title) resolve to the same rows.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
|
||||
const stubEmbedding = async (text: string): Promise<number[]> => {
|
||||
const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(hash + i))
|
||||
}
|
||||
|
||||
describe('aggregation + query field-resolution law', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('reserved-field groupBy decrements on delete (the drift bug)', async () => {
|
||||
brain.defineAggregate({
|
||||
name: 'by_subtype',
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['subtype'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
|
||||
const ids: string[] = []
|
||||
for (let i = 0; i < 5; i++) {
|
||||
ids.push(
|
||||
await brain.add({
|
||||
data: `doc-${i}`,
|
||||
type: NounType.Document,
|
||||
subtype: 'note',
|
||||
metadata: { team: 'alpha' }
|
||||
})
|
||||
)
|
||||
}
|
||||
let groups = await brain.queryAggregate('by_subtype')
|
||||
expect(groups).toHaveLength(1)
|
||||
expect(groups[0].groupKey).toEqual({ subtype: 'note' })
|
||||
expect(groups[0].metrics.count).toBe(5)
|
||||
|
||||
await brain.remove(ids[0])
|
||||
await brain.flush()
|
||||
|
||||
groups = await brain.queryAggregate('by_subtype')
|
||||
expect(groups[0].metrics.count).toBe(4)
|
||||
const live = await brain.find({ type: NounType.Document, limit: 100 })
|
||||
expect(groups[0].metrics.count).toBe(live.length)
|
||||
})
|
||||
|
||||
it('reserved-field groupBy moves between groups on update (no double-count)', async () => {
|
||||
brain.defineAggregate({
|
||||
name: 'by_subtype',
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['subtype'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
const id = await brain.add({
|
||||
data: 'doc-move',
|
||||
type: NounType.Document,
|
||||
subtype: 'draft'
|
||||
})
|
||||
await brain.update({ id, subtype: 'published' })
|
||||
|
||||
const groups = await brain.queryAggregate('by_subtype')
|
||||
const byKey = Object.fromEntries(
|
||||
groups.map((g) => [String(g.groupKey.subtype), g.metrics.count])
|
||||
)
|
||||
expect(byKey['published']).toBe(1)
|
||||
// The old group must be gone or zero — never still counting the entity.
|
||||
expect(byKey['draft'] ?? 0).toBe(0)
|
||||
})
|
||||
|
||||
it('source.where on a reserved field filters instead of matching nothing', async () => {
|
||||
brain.defineAggregate({
|
||||
name: 'notes_only',
|
||||
source: { type: NounType.Document, where: { subtype: 'note' } },
|
||||
groupBy: ['team'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'n1',
|
||||
type: NounType.Document,
|
||||
subtype: 'note',
|
||||
metadata: { team: 'alpha' }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'd1',
|
||||
type: NounType.Document,
|
||||
subtype: 'draft',
|
||||
metadata: { team: 'alpha' }
|
||||
})
|
||||
|
||||
const groups = await brain.queryAggregate('notes_only')
|
||||
expect(groups).toHaveLength(1)
|
||||
expect(groups[0].metrics.count).toBe(1) // the note, never the draft
|
||||
})
|
||||
|
||||
it('removeMany refuses empty/invalid selectors loudly', async () => {
|
||||
const id = await brain.add({ data: 'keep-me', type: NounType.Document })
|
||||
|
||||
// Bare array passed positionally — the classic silent no-op.
|
||||
await expect(
|
||||
brain.removeMany([id] as unknown as Parameters<typeof brain.removeMany>[0])
|
||||
).rejects.toThrow(/bare array/)
|
||||
// Empty selector object.
|
||||
await expect(
|
||||
brain.removeMany({} as Parameters<typeof brain.removeMany>[0])
|
||||
).rejects.toThrow(/requires a selector/)
|
||||
// Explicit empty id list.
|
||||
await expect(brain.removeMany({ ids: [] })).rejects.toThrow(/ids: \[\]/)
|
||||
|
||||
// Nothing was deleted by any of the refused calls.
|
||||
expect(await brain.get(id)).toBeTruthy()
|
||||
})
|
||||
|
||||
it('find() accepts both flattened and metadata.-prefixed where spellings', async () => {
|
||||
await brain.add({
|
||||
data: 'nested-doc',
|
||||
type: NounType.Document,
|
||||
metadata: { entry: { title: 'T1' }, classifier: { contextHints: { vfsPath: '/n/a.md' } } }
|
||||
})
|
||||
await brain.flush()
|
||||
|
||||
const flat = await brain.find({
|
||||
type: NounType.Document,
|
||||
where: { 'entry.title': 'T1' },
|
||||
limit: 10
|
||||
})
|
||||
const prefixed = await brain.find({
|
||||
type: NounType.Document,
|
||||
where: { 'metadata.entry.title': 'T1' },
|
||||
limit: 10
|
||||
})
|
||||
const deepPrefixed = await brain.find({
|
||||
type: NounType.Document,
|
||||
where: { 'metadata.classifier.contextHints.vfsPath': '/n/a.md' },
|
||||
limit: 10
|
||||
})
|
||||
expect(flat).toHaveLength(1)
|
||||
expect(prefixed).toHaveLength(1)
|
||||
expect(prefixed[0].id).toBe(flat[0].id)
|
||||
expect(deepPrefixed).toHaveLength(1)
|
||||
})
|
||||
})
|
||||
|
|
@ -1,311 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/aggregation-state-persistence
|
||||
* @description The boot-order contract for the aggregation engine. Five laws:
|
||||
* (1) STATE ADOPTION — a reopen with an unchanged defineAggregate() adopts the
|
||||
* persisted state and performs NO store walk. (The pre-fix behavior: the
|
||||
* synchronous define always beat the async init, flagged a backfill, and
|
||||
* the first query wiped the just-loaded state and re-walked the whole
|
||||
* store — every restart, forever.)
|
||||
* (2) SINGLE-FLIGHT + BATCH — concurrent cold queries across multiple pending
|
||||
* aggregates share exactly ONE store walk; a query never wipes another's
|
||||
* partial progress and M pending aggregates cost one enumeration, not M.
|
||||
* (3) CHANGED DEFINITION — a real definition change still backfills, exactly.
|
||||
* (4) PERSISTED-ONLY DEFINITIONS — an app that does not re-define at boot can
|
||||
* query a persisted aggregate without racing a spurious "not defined".
|
||||
* (5) QUIET KEYS — the engine's persistence keys (__aggregation_*) are
|
||||
* recognized system keys: no "Unknown key format" warning at boot.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy } from '../../src/index.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
import type { AggregateDefinition } from '../../src/types/brainy.types.js'
|
||||
import { prodLog } from '../../src/utils/logger.js'
|
||||
|
||||
const SPENDING: AggregateDefinition = {
|
||||
name: 'spending',
|
||||
source: { type: NounType.Event, where: { domain: 'financial' } },
|
||||
groupBy: ['category'],
|
||||
metrics: {
|
||||
total: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' }
|
||||
}
|
||||
}
|
||||
|
||||
/** Same name, different metrics — a REAL definition change (hash differs). */
|
||||
const SPENDING_CHANGED: AggregateDefinition = {
|
||||
...SPENDING,
|
||||
metrics: {
|
||||
total: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' },
|
||||
average: { op: 'avg', field: 'amount' }
|
||||
}
|
||||
}
|
||||
|
||||
describe('aggregation state persistence — boot-order contract', () => {
|
||||
let dir: string
|
||||
|
||||
beforeEach(() => {
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-agg-persist-'))
|
||||
})
|
||||
|
||||
afterEach(() => {
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
const open = async (): Promise<any> => {
|
||||
const b: any = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
/** Count store walks by intercepting the storage adapter's getNouns. */
|
||||
const countWalks = (brain: any): { count: () => number } => {
|
||||
const storage = brain.storage
|
||||
const orig = storage.getNouns.bind(storage)
|
||||
let calls = 0
|
||||
storage.getNouns = async (opts: unknown) => {
|
||||
calls++
|
||||
return orig(opts)
|
||||
}
|
||||
return { count: () => calls }
|
||||
}
|
||||
|
||||
const seed = async (brain: any): Promise<void> => {
|
||||
for (let i = 0; i < 12; i++) {
|
||||
await brain.add({
|
||||
data: `tx ${i}`,
|
||||
type: NounType.Event,
|
||||
metadata: {
|
||||
domain: 'financial',
|
||||
category: i % 2 === 0 ? 'food' : 'transport',
|
||||
amount: 10 + i
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
it('adopts persisted state on reopen with an unchanged definition — zero walks', async () => {
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
const before = await brain1.queryAggregate('spending')
|
||||
expect(before.length).toBe(2)
|
||||
await brain1.close()
|
||||
|
||||
const brain2 = await open()
|
||||
brain2.defineAggregate(SPENDING) // the standard declarative boot pattern
|
||||
await brain2.getNounCount() // settle init paths before counting walks
|
||||
const walks = countWalks(brain2)
|
||||
|
||||
const after = await brain2.queryAggregate('spending')
|
||||
|
||||
expect(walks.count()).toBe(0)
|
||||
const key = (r: any) => r.groupKey.category
|
||||
expect(
|
||||
after.map((r: any) => [key(r), r.metrics.total, r.metrics.count]).sort()
|
||||
).toEqual(
|
||||
before.map((r: any) => [key(r), r.metrics.total, r.metrics.count]).sort()
|
||||
)
|
||||
await brain2.close()
|
||||
})
|
||||
|
||||
it('adopted state keeps accumulating: post-reopen writes land on top of it', async () => {
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
await brain1.queryAggregate('spending')
|
||||
await brain1.close()
|
||||
|
||||
const brain2 = await open()
|
||||
brain2.defineAggregate(SPENDING)
|
||||
// A write BEFORE the first query: if it lands before state adoption the
|
||||
// engine must choose an exact rescan over adoption — either way the
|
||||
// result must include all 13 entities.
|
||||
await brain2.add({
|
||||
data: 'late tx',
|
||||
type: NounType.Event,
|
||||
metadata: { domain: 'financial', category: 'food', amount: 100 }
|
||||
})
|
||||
const rows = await brain2.queryAggregate('spending')
|
||||
const food = rows.find((r: any) => r.groupKey.category === 'food')
|
||||
expect(food.metrics.count).toBe(7) // 6 seeded + 1 late
|
||||
await brain2.close()
|
||||
})
|
||||
|
||||
it('a changed definition still backfills — exactly once, with correct results', async () => {
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
await brain1.queryAggregate('spending')
|
||||
await brain1.close()
|
||||
|
||||
const brain2 = await open()
|
||||
brain2.defineAggregate(SPENDING_CHANGED)
|
||||
await brain2.getNounCount()
|
||||
const walks = countWalks(brain2)
|
||||
|
||||
const rows = await brain2.queryAggregate('spending')
|
||||
|
||||
expect(walks.count()).toBe(1) // 12 entities = one page = one getNouns call
|
||||
const food = rows.find((r: any) => r.groupKey.category === 'food')
|
||||
expect(food.metrics.count).toBe(6)
|
||||
expect(food.metrics.average).toBeCloseTo(food.metrics.total / 6)
|
||||
await brain2.close()
|
||||
})
|
||||
|
||||
it('concurrent cold queries across two aggregates share exactly ONE walk', async () => {
|
||||
const brain = await open()
|
||||
brain.defineAggregate(SPENDING)
|
||||
brain.defineAggregate({
|
||||
...SPENDING,
|
||||
name: 'by_category_count',
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
await seed(brain)
|
||||
await brain.getNounCount()
|
||||
const walks = countWalks(brain)
|
||||
|
||||
const results = await Promise.all([
|
||||
brain.queryAggregate('spending'),
|
||||
brain.queryAggregate('by_category_count'),
|
||||
brain.queryAggregate('spending'),
|
||||
brain.queryAggregate('by_category_count'),
|
||||
brain.queryAggregate('spending'),
|
||||
brain.queryAggregate('by_category_count')
|
||||
])
|
||||
|
||||
expect(walks.count()).toBe(1) // 12 entities = one page; one walk fills both
|
||||
for (const rows of results) {
|
||||
const total = rows.reduce((s: number, r: any) => s + r.metrics.count, 0)
|
||||
expect(total).toBe(12)
|
||||
}
|
||||
|
||||
// Warm re-query: converged, no further walks.
|
||||
await brain.queryAggregate('spending')
|
||||
expect(walks.count()).toBe(1)
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('persisted-only definitions are queryable without re-defining at boot', async () => {
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
await brain1.queryAggregate('spending')
|
||||
await brain1.close()
|
||||
|
||||
const brain2 = await open()
|
||||
// NO defineAggregate — the app relies on the persisted definition.
|
||||
const walks = countWalks(brain2)
|
||||
const rows = await brain2.queryAggregate('spending')
|
||||
expect(walks.count()).toBe(0) // persisted state adopted here too
|
||||
expect(rows.length).toBe(2)
|
||||
await brain2.close()
|
||||
})
|
||||
|
||||
it('generation-mismatched persisted state is rescanned once, loudly — never adopted', async () => {
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
await brain1.queryAggregate('spending')
|
||||
await brain1.close()
|
||||
|
||||
// Simulate the copied-store incident class: a fact-log truncation (or an
|
||||
// unclean shutdown) leaves the committed watermark different from the
|
||||
// generation the flushed state was stamped with.
|
||||
const tamper: any = await open()
|
||||
const key = '__aggregation_state_spending__'
|
||||
const stored = await tamper.storage.getMetadata(key)
|
||||
expect(typeof stored.sourceGeneration).toBe('number') // the stamp is really persisted
|
||||
await tamper.storage.saveMetadata(key, {
|
||||
...stored,
|
||||
sourceGeneration: stored.sourceGeneration + 5
|
||||
})
|
||||
await tamper.close()
|
||||
|
||||
const warnSpy = vi.spyOn(prodLog, 'warn')
|
||||
const brain2 = await open()
|
||||
brain2.defineAggregate(SPENDING)
|
||||
await brain2.getNounCount()
|
||||
const walks = countWalks(brain2)
|
||||
|
||||
const rows = await brain2.queryAggregate('spending')
|
||||
|
||||
expect(walks.count()).toBe(1) // exactly ONE rescan — no silent adopt, no spin
|
||||
const food = rows.find((r: any) => r.groupKey.category === 'food')
|
||||
expect(food.metrics.count).toBe(6) // rescan produced exact results
|
||||
expect(
|
||||
warnSpy.mock.calls.some(args => String(args[0]).includes('rescanning instead of adopting'))
|
||||
).toBe(true) // and it said so out loud
|
||||
warnSpy.mockRestore()
|
||||
await brain2.close()
|
||||
})
|
||||
|
||||
it('a failing walk is loud, non-destructive, and latched — never a silent retry loop', async () => {
|
||||
// Fresh define + seeded writes: the write hooks have populated LIVE state,
|
||||
// and the first-query rescan is still pending. The incident shape
|
||||
// (wipe-before-scan + no try/catch + per-query re-walk) would have wiped
|
||||
// that live state and silently re-walked on every query.
|
||||
const brain: any = await open()
|
||||
brain.defineAggregate(SPENDING)
|
||||
await seed(brain)
|
||||
expect(brain._aggregationIndex.queryAggregate({ name: 'spending' }).length).toBe(2)
|
||||
|
||||
const storage = brain.storage
|
||||
const origGetNouns = storage.getNouns.bind(storage)
|
||||
let walkAttempts = 0
|
||||
storage.getNouns = async () => {
|
||||
walkAttempts++
|
||||
throw new Error('injected storage failure')
|
||||
}
|
||||
|
||||
// First query: the walk fails LOUDLY with the storage error.
|
||||
await expect(brain.queryAggregate('spending')).rejects.toThrow('injected storage failure')
|
||||
expect(walkAttempts).toBe(1)
|
||||
|
||||
// Live state was NOT destroyed by the failed walk (staging was dropped).
|
||||
expect(brain._aggregationIndex.queryAggregate({ name: 'spending' }).length).toBe(2)
|
||||
|
||||
// Second query inside the cooldown: instant loud failure, NO new walk.
|
||||
await expect(brain.queryAggregate('spending')).rejects.toThrow('failure cooldown')
|
||||
expect(walkAttempts).toBe(1)
|
||||
|
||||
// Heal the storage + expire the cooldown: one fresh walk succeeds exactly.
|
||||
storage.getNouns = origGetNouns
|
||||
brain._aggregationBackfillFailure.at = Date.now() - 60_000
|
||||
const rows = await brain.queryAggregate('spending')
|
||||
const food = rows.find((r: any) => r.groupKey.category === 'food')
|
||||
expect(food.metrics.count).toBe(6)
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('aggregation persistence keys never log "Unknown key format"', async () => {
|
||||
const warnSpy = vi.spyOn(prodLog, 'warn')
|
||||
const brain1 = await open()
|
||||
brain1.defineAggregate(SPENDING)
|
||||
await seed(brain1)
|
||||
await brain1.queryAggregate('spending')
|
||||
await brain1.close()
|
||||
|
||||
const brain2 = await open()
|
||||
brain2.defineAggregate(SPENDING)
|
||||
await brain2.queryAggregate('spending')
|
||||
await brain2.close()
|
||||
|
||||
const offenders = warnSpy.mock.calls
|
||||
.map(args => String(args[0]))
|
||||
.filter(
|
||||
msg =>
|
||||
msg.includes('Unknown key format') &&
|
||||
(msg.includes('__aggregation_') || msg.includes('brainy:entityIdMapper'))
|
||||
)
|
||||
expect(offenders).toEqual([])
|
||||
warnSpy.mockRestore()
|
||||
})
|
||||
})
|
||||
|
|
@ -1,85 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/background-dedup-lifecycle
|
||||
* @description The post-import background deduplication pass (a merge-DELETE
|
||||
* writer) obeys the same contract as the inline pass. Laws:
|
||||
* (1) enableDeduplication:false schedules NO background pass — the brain-owned
|
||||
* deduplicator is never even constructed;
|
||||
* (2) by default the pass IS scheduled, brain-owned, with an unref'd timer
|
||||
* (a pending pass never holds the process open);
|
||||
* (3) repeated imports debounce into ONE pending batch on ONE instance
|
||||
* (per-coordinator instances used to arm one timer per import);
|
||||
* (4) close() cancels pending work — no delete pass can fire after close.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
|
||||
const ROWS = [
|
||||
{ name: 'Alice Zephyr', role: 'engineer' },
|
||||
{ name: 'Bob Quill', role: 'writer' }
|
||||
]
|
||||
|
||||
// Keep imports fast and deterministic — dedup scheduling is what's under test.
|
||||
const FAST = {
|
||||
enableNeuralExtraction: false,
|
||||
enableRelationshipInference: false,
|
||||
enableConceptExtraction: false
|
||||
} as const
|
||||
|
||||
// Deterministic stub embedder (hnsw-rebuild.test.ts pattern) — dedup
|
||||
// scheduling never inspects vector CONTENT, so skip the WASM model load.
|
||||
const stubEmbedding = async (text: string): Promise<number[]> => {
|
||||
const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0)
|
||||
const vector = new Array(384).fill(0).map((_, i) => Math.sin(hash + i))
|
||||
return vector
|
||||
}
|
||||
|
||||
describe('background dedup lifecycle', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('enableDeduplication:false schedules no background pass at all', async () => {
|
||||
await brain.import(ROWS, { ...FAST, enableDeduplication: false })
|
||||
expect((brain as any)._backgroundDedup).toBeUndefined()
|
||||
})
|
||||
|
||||
it('default schedules a brain-owned pass with an unref-ed timer', async () => {
|
||||
await brain.import(ROWS, { ...FAST })
|
||||
const dedup = (brain as any)._backgroundDedup
|
||||
expect(dedup).toBeDefined()
|
||||
expect(dedup.pendingImports.size).toBe(1)
|
||||
const timer = dedup.debounceTimer
|
||||
expect(timer).toBeDefined()
|
||||
// Node timers expose hasRef(); an unref'd timer must not hold the process.
|
||||
expect(typeof timer.hasRef).toBe('function')
|
||||
expect(timer.hasRef()).toBe(false)
|
||||
})
|
||||
|
||||
it('imports debounce into one pending batch on one brain-owned instance', async () => {
|
||||
await brain.import(ROWS, { ...FAST })
|
||||
const first = (brain as any)._backgroundDedup
|
||||
await brain.import([{ name: 'Cara Vex', role: 'analyst' }], { ...FAST })
|
||||
expect((brain as any)._backgroundDedup).toBe(first)
|
||||
expect(first.pendingImports.size).toBe(2)
|
||||
})
|
||||
|
||||
it('close() cancels pending background dedup', async () => {
|
||||
await brain.import(ROWS, { ...FAST })
|
||||
const dedup = (brain as any)._backgroundDedup
|
||||
expect(dedup.debounceTimer).toBeDefined()
|
||||
await brain.close()
|
||||
expect(dedup.debounceTimer).toBeUndefined()
|
||||
expect(dedup.pendingImports.size).toBe(0)
|
||||
})
|
||||
})
|
||||
|
|
@ -1,122 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/counter-recount
|
||||
* @description Counter honesty. Two laws under test:
|
||||
* (1) REMOVAL NEVER REQUIRES RE-READING THE REMOVED RECORD — the count
|
||||
* decrement falls back to the caller's pre-delete read when the canonical
|
||||
* metadata re-read returns null (replace race / ghost), instead of being
|
||||
* silently skipped. The skip minted permanent inflation: adds counted,
|
||||
* paired removals not decremented, and Math.max(totalNounCount, scanned)
|
||||
* pinned the inflated scalar forever.
|
||||
* (2) THE SANCTIONED RECOUNT — repairIndex() unconditionally recomputes and
|
||||
* PERSISTS every counter rollup (scalar totals + per-type maps +
|
||||
* type-statistics) from one canonical walk, so an already-inflated brain
|
||||
* is permanently corrected (survives reopen).
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy } from '../../src/index.js'
|
||||
|
||||
/** Absolute path of one entity's `<id>` directory (or null if not present). */
|
||||
function entityDir(root: string, id: string): string | null {
|
||||
const base = path.join(root, 'entities', 'nouns')
|
||||
if (!fs.existsSync(base)) return null
|
||||
for (const shard of fs.readdirSync(base)) {
|
||||
const candidate = path.join(base, shard, id)
|
||||
if (fs.existsSync(candidate)) return candidate
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
describe('counter honesty — removal without re-reading + the sanctioned recount', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
const open = async () => {
|
||||
const b: any = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true,
|
||||
dimensions: 384
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-recount-'))
|
||||
brain = await open()
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('the counter returns to baseline across add→remove→re-create cycles (no drift)', async () => {
|
||||
const baseline = await brain.storage.getNounCount()
|
||||
for (let i = 0; i < 4; i++) {
|
||||
const id = await brain.add({ data: `cycle ${i}`, type: 'document', metadata: { i } })
|
||||
expect(await brain.storage.getNounCount()).toBe(baseline + 1)
|
||||
await brain.remove(id)
|
||||
expect(await brain.storage.getNounCount()).toBe(baseline)
|
||||
}
|
||||
})
|
||||
|
||||
it('deleteNoun decrements from the provided prior record when the canonical re-read is null', async () => {
|
||||
const id = await brain.add({ data: 'to be ghosted', type: 'document', metadata: { g: 1 } })
|
||||
await brain.flush()
|
||||
const baseline = await brain.storage.getNounCount()
|
||||
|
||||
// Capture the pre-delete read (what remove() holds), then simulate the
|
||||
// replace-race / ghost shape: the metadata leg vanishes before the delete's
|
||||
// internal re-read.
|
||||
const prior = await brain.storage.getNounMetadata(id)
|
||||
expect(prior).not.toBeNull()
|
||||
const eDir = entityDir(dir, id)!
|
||||
for (const f of fs.readdirSync(eDir)) {
|
||||
if (f.startsWith('metadata.json')) fs.rmSync(path.join(eDir, f))
|
||||
}
|
||||
|
||||
// Without the prior record this decrement used to be silently skipped.
|
||||
await brain.storage.deleteNoun(id, prior)
|
||||
expect(await brain.storage.getNounCount()).toBe(baseline - 1)
|
||||
// Full removal still holds: nothing left on disk.
|
||||
expect(entityDir(dir, id)).toBeNull()
|
||||
})
|
||||
|
||||
it('repairIndex() recounts an inflated persisted scalar over CLEAN shelves — and it survives reopen', async () => {
|
||||
for (let i = 0; i < 3; i++) {
|
||||
await brain.add({ data: `real ${i}`, type: 'document', metadata: { i } })
|
||||
}
|
||||
await brain.flush()
|
||||
const honest = await brain.storage.getNounCount()
|
||||
// Pagination's totalCount has its own baseline: it enumerates EVERYTHING
|
||||
// (including the internal VFS root), while the user-facing scalar counts
|
||||
// only public entities — so the two legitimately differ by the internals.
|
||||
const pageHonest = (await brain.storage.getNounsWithPagination({ limit: 1000, offset: 0 })).totalCount
|
||||
|
||||
// Simulate the historical drift: an inflated persisted scalar (deletes whose
|
||||
// decrement was skipped). Persist it so a reopen rehydrates the lie.
|
||||
;(brain.storage as any).totalNounCount = honest + 52
|
||||
await (brain.storage as any).persistCounts()
|
||||
await brain.close()
|
||||
brain = await open()
|
||||
expect(await brain.storage.getNounCount()).toBe(honest + 52) // the lie survived reopen
|
||||
|
||||
// The sanctioned recount — unconditional in repairIndex (no orphans needed).
|
||||
await brain.repairIndex()
|
||||
expect(await brain.storage.getNounCount()).toBe(honest)
|
||||
|
||||
// Permanently: the corrected counter survives another reopen.
|
||||
await brain.close()
|
||||
brain = await open()
|
||||
expect(await brain.storage.getNounCount()).toBe(honest)
|
||||
|
||||
// And the paginated totalCount (the Math.max consumer) is honest too —
|
||||
// back to ITS baseline, no longer pinned high by the inflated scalar.
|
||||
const page = await brain.storage.getNounsWithPagination({ limit: 1000, offset: 0 })
|
||||
expect(page.totalCount).toBe(pageHonest)
|
||||
})
|
||||
})
|
||||
|
|
@ -515,15 +515,7 @@ describe('8.0 Db API — generational MVCC', () => {
|
|||
await brain.transact([{ op: 'update', id: uid('compact-e'), metadata: { v: 4 } }])
|
||||
).release()
|
||||
|
||||
// History record-sets only — the generation FACT LOG also lives under
|
||||
// `_generations/` (at `facts/`) and is deliberately NOT reclaimed by
|
||||
// history compaction (facts are the future canonical, not undo history).
|
||||
const historyRecords = async (): Promise<number> =>
|
||||
(await storage.listRawObjects('_generations')).filter(
|
||||
(p: string) => !p.startsWith('_generations/facts/')
|
||||
).length
|
||||
|
||||
const recordsBefore = await historyRecords()
|
||||
const recordsBefore = (await storage.listRawObjects('_generations')).length
|
||||
expect(recordsBefore).toBeGreaterThan(0)
|
||||
|
||||
// Compact while pinned: record-sets above the pin survive, pinned reads stay correct.
|
||||
|
|
@ -536,7 +528,7 @@ describe('8.0 Db API — generational MVCC', () => {
|
|||
const second = await brain.compactHistory()
|
||||
expect(first.removedGenerations + second.removedGenerations).toBeGreaterThan(0)
|
||||
|
||||
const recordsAfter = await historyRecords()
|
||||
const recordsAfter = (await storage.listRawObjects('_generations')).length
|
||||
expect(recordsAfter).toBeLessThan(recordsBefore)
|
||||
expect(recordsAfter).toBe(0)
|
||||
|
||||
|
|
@ -1280,32 +1272,24 @@ describe('8.0 Db API — generational MVCC', () => {
|
|||
await expect(reopened.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
})
|
||||
|
||||
it('Model-B retention — flush() NEVER compacts (8.9.0); adaptive reclaim runs at close()', async () => {
|
||||
// Default brain → ADAPTIVE retention with a driven byte budget far below
|
||||
// the accumulated history (~13 generations of full-vector before-images).
|
||||
// The 8.9.0 law: flush() is durability-only — it must not reclaim even
|
||||
// when the budget is exceeded (reclaim-on-flush blocked production writes
|
||||
// for 25-191s). Maintenance runs at close(), time-bounded.
|
||||
const { brain, dir } = await openFsBrain()
|
||||
it('Model-B retention — setRetentionBudget drives adaptive reclaim on flush; live data intact', async () => {
|
||||
// Default brain → ADAPTIVE retention. A coordinator (e.g. cor's ResourceManager)
|
||||
// pushes a byte budget via setRetentionBudget(); auto-compaction on flush() reclaims
|
||||
// oldest history down toward it. Each update's before-image carries the full prior
|
||||
// 384-dim vector (~KBs), so ~13 generations far exceed a few-KB budget.
|
||||
const { brain } = await openFsBrain()
|
||||
const a = uid('ret-budget')
|
||||
await brain.add({ id: a, type: NounType.Document, data: 'v0', vector: vec(1), metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 12; v++) await brain.update({ id: a, metadata: { v } })
|
||||
|
||||
brain.setRetentionBudget(6000) // ~6 KB — well below the accumulated history
|
||||
await brain.flush()
|
||||
await brain.flush() // group-commit + adaptive auto-compaction under the budget
|
||||
|
||||
// flush() paid durability only: nothing reclaimed, all history readable.
|
||||
expect(generationStoreOf(brain).horizon()).toBe(0)
|
||||
const probe = await brain.asOf(1) // readable proves nothing was reclaimed…
|
||||
await probe.release() // …and MUST be released: a held pin would (correctly)
|
||||
// protect every newer generation through the close() compaction below.
|
||||
await brain.close() // ← THE auto-compaction site now
|
||||
|
||||
// close() reclaimed under the budget; live record intact; horizon durable.
|
||||
const { brain: reopened } = await openFsBrain(dir)
|
||||
expect(generationStoreOf(reopened).horizon()).toBeGreaterThan(0)
|
||||
await expect(reopened.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
expect((await reopened.get(a))?.metadata?.v).toBe(12)
|
||||
// History was reclaimed (the horizon advanced past the oldest generations)…
|
||||
expect(generationStoreOf(brain).horizon()).toBeGreaterThan(0)
|
||||
await expect(brain.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
// …but the budget reclaims HISTORY only — the live record is untouched.
|
||||
expect((await brain.get(a))?.metadata?.v).toBe(12)
|
||||
})
|
||||
|
||||
// ==========================================================================
|
||||
|
|
|
|||
|
|
@ -1,163 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/entity-tree-stamp
|
||||
* @description The entity tree's FAMILY STAMP: written at flush/close with
|
||||
* `sourceGeneration` (the committed generation the canonical tree reflects)
|
||||
* plus rollup invariants (entity/relationship counts); verified at open by
|
||||
* comparison — coherent on a clean close, LOUD on genuine incoherence
|
||||
* (tampered counters), healed by repairIndex()'s recount + re-stamp. One
|
||||
* verifier reads both member modes.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import {
|
||||
Brainy,
|
||||
readFamilyStamp,
|
||||
verifyFamilyStamp,
|
||||
ENTITY_TREE_STAMP_PATH,
|
||||
type FamilyStamp
|
||||
} from '../../src/index.js'
|
||||
import { prodLog } from '../../src/utils/logger.js'
|
||||
|
||||
describe('entity-tree family stamp', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
const open = async () => {
|
||||
const b: any = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true,
|
||||
dimensions: 384
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-stamp-'))
|
||||
brain = await open()
|
||||
})
|
||||
afterEach(async () => {
|
||||
vi.restoreAllMocks()
|
||||
await brain.close?.().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('flush() writes the stamp: rollup mode, counts + sourceGeneration match live state', async () => {
|
||||
for (let i = 0; i < 3; i++) await brain.add({ data: `s${i}`, type: 'document', metadata: { i } })
|
||||
await brain.flush()
|
||||
|
||||
const stamp = (await readFamilyStamp(brain.storage, ENTITY_TREE_STAMP_PATH)) as FamilyStamp
|
||||
expect(stamp).not.toBeNull()
|
||||
expect(stamp.family).toBe('entity-tree')
|
||||
expect(stamp.members.mode).toBe('rollup')
|
||||
const invariants = (stamp.members as any).invariants
|
||||
expect(invariants.nounCount).toBe(await brain.storage.getNounCount())
|
||||
expect(invariants.verbCount).toBe(await brain.storage.getVerbCount())
|
||||
expect(stamp.sourceGeneration).toBe(brain.generation())
|
||||
expect(stamp.generation).toBeGreaterThanOrEqual(1)
|
||||
})
|
||||
|
||||
it('a cleanly-closed store reopens COHERENT (no incoherence warning)', async () => {
|
||||
await brain.add({ data: 'clean', type: 'document', metadata: {} })
|
||||
await brain.close()
|
||||
|
||||
const warn = vi.spyOn(prodLog, 'warn')
|
||||
brain = await open()
|
||||
const stampWarnings = warn.mock.calls.filter((c) => String(c[0]).includes('entity-tree stamp'))
|
||||
expect(stampWarnings).toEqual([])
|
||||
})
|
||||
|
||||
it('tampered counters surface as INCOHERENT at open; repairIndex() heals + re-stamps', async () => {
|
||||
for (let i = 0; i < 3; i++) await brain.add({ data: `t${i}`, type: 'document', metadata: { i } })
|
||||
await brain.close()
|
||||
|
||||
// Simulate counter drift AFTER the stamp was written: inflate the
|
||||
// persisted scalar the way the historical decrement-skip did.
|
||||
brain = await open()
|
||||
;(brain.storage as any).totalNounCount += 52
|
||||
await (brain.storage as any).persistCounts()
|
||||
await brain.close()
|
||||
// The close boundary re-stamps with the inflated counter — so tamper the
|
||||
// STAMP instead for a deterministic mismatch: stamped counts differ from
|
||||
// the (inflated) live ones at the NEXT open only if the stamp is older.
|
||||
// Rewrite the stamp with the honest counts + current sourceGeneration.
|
||||
const raw = JSON.parse(
|
||||
require('node:zlib')
|
||||
.gunzipSync(fs.readFileSync(path.join(dir, `${ENTITY_TREE_STAMP_PATH}.gz`)))
|
||||
.toString('utf-8')
|
||||
) as FamilyStamp
|
||||
const honest = { ...raw }
|
||||
;(honest.members as any).invariants.nounCount -= 52
|
||||
fs.writeFileSync(
|
||||
path.join(dir, `${ENTITY_TREE_STAMP_PATH}.gz`),
|
||||
require('node:zlib').gzipSync(JSON.stringify(honest))
|
||||
)
|
||||
|
||||
const warn = vi.spyOn(prodLog, 'warn')
|
||||
brain = await open()
|
||||
const incoherent = warn.mock.calls.filter((c) => String(c[0]).includes('INCOHERENT'))
|
||||
expect(incoherent.length).toBeGreaterThanOrEqual(1)
|
||||
expect(String(incoherent[0][0])).toMatch(/nounCount/)
|
||||
|
||||
// The heal: recount from canonical + re-stamp → next open is quiet.
|
||||
await brain.repairIndex()
|
||||
await brain.close()
|
||||
const warn2 = vi.spyOn(prodLog, 'warn')
|
||||
brain = await open()
|
||||
const stillIncoherent = warn2.mock.calls.filter((c) => String(c[0]).includes('INCOHERENT'))
|
||||
expect(stillIncoherent).toEqual([])
|
||||
})
|
||||
|
||||
it('the one verifier handles both member modes', () => {
|
||||
const rollup: FamilyStamp = {
|
||||
family: 'x',
|
||||
generation: 1,
|
||||
committedAt: new Date().toISOString(),
|
||||
sourceGeneration: 5,
|
||||
members: { mode: 'rollup', invariants: { nounCount: 10 } }
|
||||
}
|
||||
expect(verifyFamilyStamp(rollup, 5, { nounCount: 10 })).toEqual({ state: 'coherent' })
|
||||
expect(verifyFamilyStamp(rollup, 5, { nounCount: 11 }).state).toBe('incoherent')
|
||||
expect(verifyFamilyStamp(rollup, 9, { nounCount: 10 })).toEqual({
|
||||
state: 'behind',
|
||||
stampSource: 5,
|
||||
head: 9
|
||||
})
|
||||
expect(verifyFamilyStamp(rollup, 3, { nounCount: 10 }).state).toBe('incoherent') // ahead of head
|
||||
expect(verifyFamilyStamp(null, 5, {})).toEqual({ state: 'absent' })
|
||||
|
||||
const enumerated: FamilyStamp = {
|
||||
family: 'y',
|
||||
generation: 1,
|
||||
committedAt: new Date().toISOString(),
|
||||
sourceGeneration: 2,
|
||||
members: { mode: 'enumerated', files: [{ path: 'a.bin', bytes: 128 }] }
|
||||
}
|
||||
expect(verifyFamilyStamp(enumerated, 2, { 'a.bin': 128 })).toEqual({ state: 'coherent' })
|
||||
expect(verifyFamilyStamp(enumerated, 2, { 'a.bin': 64 }).state).toBe('incoherent')
|
||||
expect(verifyFamilyStamp(enumerated, 2, {}).state).toBe('incoherent') // missing member
|
||||
|
||||
// Rollup invariants may be STRING fingerprints (e.g. a per-tree SHA-256):
|
||||
// strict equality either way; a type mismatch reads as incoherence.
|
||||
const fingerprinted: FamilyStamp = {
|
||||
family: 'z',
|
||||
generation: 1,
|
||||
committedAt: new Date().toISOString(),
|
||||
sourceGeneration: 3,
|
||||
members: { mode: 'rollup', invariants: { treeDigest: 'abc123', rows: 42 } }
|
||||
}
|
||||
expect(verifyFamilyStamp(fingerprinted, 3, { treeDigest: 'abc123', rows: 42 })).toEqual({
|
||||
state: 'coherent'
|
||||
})
|
||||
expect(verifyFamilyStamp(fingerprinted, 3, { treeDigest: 'deadbeef', rows: 42 }).state).toBe(
|
||||
'incoherent'
|
||||
)
|
||||
expect(verifyFamilyStamp(fingerprinted, 3, { treeDigest: 'abc123', rows: '42' }).state).toBe(
|
||||
'incoherent' // type mismatch never passes
|
||||
)
|
||||
})
|
||||
})
|
||||
|
|
@ -1,125 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/fact-log-contracts
|
||||
* @description Pinned durability + stability contracts for the fact log.
|
||||
*
|
||||
* (1) FSYNC-BEFORE-ACK: an acknowledged write's fact survives an abrupt
|
||||
* process end (no flush, no close — reopen from disk).
|
||||
* - transact(): HOLDS TODAY — the fact is fsync'd before transact returns.
|
||||
* - single-op: PINNED AS `it.fails` — today's group-commit batches
|
||||
* DURABILITY (ack precedes the group fsync; a hard kill loses the fact
|
||||
* AND the generation together, coherently — the documented Model-B
|
||||
* contract, fine while the tree is authoritative). The destination
|
||||
* (ack-at-log) requires group commit to become LATENCY batching: the
|
||||
* ack waits for the shared fsync. When that lands, this pin flips red —
|
||||
* remove `.fails` and the contract is permanent. No cliff to discover.
|
||||
*
|
||||
* (2) SCAN STABILITY UNDER ROTATION: a scan handle opened before segment
|
||||
* rotation yields exactly its snapshot — byte-identical facts, no gaps,
|
||||
* no duplicates, and no bleed-in of facts appended after the snapshot.
|
||||
* (The reclaim-during-scan variant lands with fact-log compaction, which
|
||||
* does not exist yet — segments only rotate today, never reclaim.)
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy, type CommitFact } from '../../src/index.js'
|
||||
import { MemoryStorage } from '../../src/storage/adapters/memoryStorage.js'
|
||||
import { FactLog, type FactLogStorage } from '../../src/db/factLog.js'
|
||||
|
||||
describe('fsync-before-ack contract (fact durability at the ack boundary)', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
const open = async () => {
|
||||
const b: any = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true,
|
||||
dimensions: 384
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-factack-'))
|
||||
brain = await open()
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('transact(): the fact is durable the moment the ack returns (kill-after-ack safe)', async () => {
|
||||
const receipt = await brain.transact([
|
||||
{ op: 'add', type: 'document', metadata: { durable: 1 }, data: 'ack-at-commit' }
|
||||
])
|
||||
// Abrupt end: no flush(), no close() — a new instance reads only disk.
|
||||
brain = await open()
|
||||
const facts: CommitFact[] = []
|
||||
for await (const b of brain.scanFacts()!.batches()) facts.push(...b.facts)
|
||||
expect(facts.some((f) => f.generation === receipt.generation)).toBe(true)
|
||||
})
|
||||
|
||||
// PINNED (flips red when group commit becomes latency batching — then
|
||||
// remove `.fails` and the ack-at-log contract is permanent on every path).
|
||||
it.fails('single-op: the fact is durable the moment the ack returns (the ack-at-log target)', async () => {
|
||||
await brain.add({ data: 'acked single-op', type: 'document', metadata: { n: 1 } })
|
||||
const ackedHead = brain.scanFacts()!.headGeneration
|
||||
// Abrupt end immediately after the ack — before any flush window.
|
||||
brain = await open()
|
||||
const facts: CommitFact[] = []
|
||||
for await (const b of brain.scanFacts()!.batches()) facts.push(...b.facts)
|
||||
expect(facts.some((f) => f.generation === ackedHead)).toBe(true)
|
||||
})
|
||||
})
|
||||
|
||||
describe('scan stability under rotation (the snapshot contract)', () => {
|
||||
const UUID = (n: number): string => `00000000-0000-4000-8000-${String(n).padStart(12, '0')}`
|
||||
const fact = (generation: number): CommitFact => ({
|
||||
generation,
|
||||
timestamp: 1_700_000_000_000 + generation,
|
||||
ops: [
|
||||
{
|
||||
kind: 'noun',
|
||||
id: UUID(generation),
|
||||
// Padding makes each frame ~1KB so a small rotateBytes forces rotations.
|
||||
record: { metadata: { noun: 'document', pad: 'x'.repeat(900), g: generation }, vector: null }
|
||||
}
|
||||
]
|
||||
})
|
||||
|
||||
it('a scan opened before rotations yields its exact snapshot — no gaps, dups, or bleed-in', async () => {
|
||||
const mem: any = new MemoryStorage()
|
||||
await mem.init()
|
||||
const log = new FactLog(mem as FactLogStorage, { rotateBytes: 4096 }) // ~4 facts per segment
|
||||
await log.open(0)
|
||||
for (let g = 1; g <= 10; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
|
||||
// Open the snapshot, THEN keep appending — forcing further rotations.
|
||||
const scan = log.scanFacts()
|
||||
expect(scan.headGeneration).toBe(10)
|
||||
for (let g = 11; g <= 25; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
expect(log.headGeneration()).toBe(25)
|
||||
|
||||
const seen: number[] = []
|
||||
for await (const batch of scan.batches()) {
|
||||
for (const f of batch.facts) seen.push(f.generation)
|
||||
}
|
||||
// Exactly the snapshot: 1..10 in order, nothing appended-after bleeds in.
|
||||
expect(seen).toEqual([1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
|
||||
expect(scan.summary().factsYielded).toBe(10)
|
||||
|
||||
// And a fresh scan sees everything, across all rotated segments.
|
||||
const all: number[] = []
|
||||
for await (const batch of log.scanFacts().batches()) {
|
||||
for (const f of batch.facts) all.push(f.generation)
|
||||
}
|
||||
expect(all).toEqual(Array.from({ length: 25 }, (_, i) => i + 1))
|
||||
expect(log.segmentPaths().length).toBeGreaterThanOrEqual(2) // rotations actually happened
|
||||
})
|
||||
})
|
||||
|
|
@ -1,219 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/fact-log-dual-write
|
||||
* @description The generation fact log end-to-end through real commits: every
|
||||
* committed generation (single-op AND transact) appends its AFTER-IMAGE fact
|
||||
* at the commit point; removals append body-less tombstones; an aborted
|
||||
* transaction leaves no fact; facts survive reopen and continue monotonically;
|
||||
* the scan surface (brain.scanFacts) carries the frozen telemetry shape; and
|
||||
* the fact-log namespace is protected against prefix-nuking.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy, ProtectedArtifactError, type CommitFact } from '../../src/index.js'
|
||||
|
||||
async function allFacts(brain: any): Promise<CommitFact[]> {
|
||||
const scan = brain.scanFacts()
|
||||
expect(scan).not.toBeNull()
|
||||
const facts: CommitFact[] = []
|
||||
for await (const batch of scan!.batches()) facts.push(...batch.facts)
|
||||
return facts
|
||||
}
|
||||
|
||||
describe('fact log dual-write (memory adapter)', () => {
|
||||
let brain: any
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
brain = new Brainy({ requireSubtype: false, storage: { type: 'memory' }, silent: true, dimensions: 384 })
|
||||
await brain.init()
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
})
|
||||
|
||||
it('every single-op write appends its after-image fact; a remove appends a tombstone', async () => {
|
||||
const id = await brain.add({ data: 'first', type: 'document', metadata: { rev: 1 } })
|
||||
await brain.update({ id, metadata: { rev: 2 } })
|
||||
await brain.remove(id)
|
||||
|
||||
const facts = await allFacts(brain)
|
||||
// add + update + remove each committed a generation (the remove may span
|
||||
// cascade ops but is ONE generation). Facts are monotonic.
|
||||
const gens = facts.map((f) => f.generation)
|
||||
expect([...gens].sort((a, b) => a - b)).toEqual(gens)
|
||||
expect(facts.length).toBeGreaterThanOrEqual(3)
|
||||
|
||||
// The add fact carries the after-image of the new entity.
|
||||
const addFact = facts.find((f) => f.ops.some((op) => op.id === id && op.record !== null))
|
||||
expect(addFact).toBeDefined()
|
||||
|
||||
// The remove fact carries a body-less tombstone for the id.
|
||||
const removeFact = facts[facts.length - 1]
|
||||
const tombstone = removeFact.ops.find((op) => op.id === id)
|
||||
expect(tombstone).toBeDefined()
|
||||
expect(tombstone!.record).toBeNull()
|
||||
expect(tombstone!.kind).toBe('noun')
|
||||
})
|
||||
|
||||
it('the update fact holds the NEW state (after-image, not before)', async () => {
|
||||
const id = await brain.add({ data: 'versioned', type: 'document', metadata: { v: 'old' } })
|
||||
await brain.update({ id, metadata: { v: 'new' } })
|
||||
|
||||
const facts = await allFacts(brain)
|
||||
const updateFact = facts[facts.length - 1]
|
||||
const op = updateFact.ops.find((o) => o.id === id)!
|
||||
expect(op.record).not.toBeNull()
|
||||
expect((op.record!.metadata as any).v).toBe('new')
|
||||
})
|
||||
|
||||
it('a transact commits ONE fact carrying all its ops, with meta', async () => {
|
||||
const receipt = await brain.transact(
|
||||
[
|
||||
{ op: 'add', type: 'document', metadata: { part: 1 }, data: 'a' },
|
||||
{ op: 'add', type: 'document', metadata: { part: 2 }, data: 'b' }
|
||||
],
|
||||
{ meta: { source: 'batch-import' } }
|
||||
)
|
||||
|
||||
const facts = await allFacts(brain)
|
||||
const txFact = facts.find((f) => f.generation === receipt.generation)
|
||||
expect(txFact).toBeDefined()
|
||||
expect(txFact!.ops.filter((op) => op.kind === 'noun').length).toBeGreaterThanOrEqual(2)
|
||||
expect(txFact!.meta).toEqual({ source: 'batch-import' })
|
||||
})
|
||||
|
||||
it('an aborted transact leaves NO fact (absent = never committed)', async () => {
|
||||
const id = await brain.add({ data: 'cas target', type: 'document', metadata: { n: 1 } })
|
||||
const before = (await allFacts(brain)).length
|
||||
|
||||
await expect(
|
||||
brain.transact([{ op: 'update', id, ifRev: 999, metadata: { n: 2 } }])
|
||||
).rejects.toThrow()
|
||||
|
||||
const after = await allFacts(brain)
|
||||
expect(after.length).toBe(before)
|
||||
})
|
||||
|
||||
it('fact generations line up with the transaction log', async () => {
|
||||
await brain.add({ data: 'x', type: 'document', metadata: {} })
|
||||
await brain.add({ data: 'y', type: 'document', metadata: {} })
|
||||
await brain.flush()
|
||||
|
||||
const facts = await allFacts(brain)
|
||||
const logGens = new Set((await brain.transactionLog()).map((e: any) => e.generation))
|
||||
for (const f of facts) {
|
||||
expect(logGens.has(f.generation)).toBe(true)
|
||||
}
|
||||
})
|
||||
|
||||
it('the storage fact-scan capability serves a provider holding only `storage`', async () => {
|
||||
// An index provider receives `storage` — never the brain — and reaches the
|
||||
// fact log through the host-wired capability (it must never construct its
|
||||
// own fact-log reader: the log's open path is writer-side).
|
||||
const id = await brain.add({ data: 'via storage', type: 'document', metadata: { s: 1 } })
|
||||
await brain.remove(id)
|
||||
|
||||
const storage = brain.storage
|
||||
expect(typeof storage.scanFacts).toBe('function')
|
||||
expect(storage.factLogHeadGeneration()).toBe(brain.scanFacts()!.headGeneration)
|
||||
|
||||
const viaStorage: CommitFact[] = []
|
||||
for await (const b of storage.scanFacts()!.batches()) viaStorage.push(...b.facts)
|
||||
const viaBrain: CommitFact[] = []
|
||||
for await (const b of brain.scanFacts()!.batches()) viaBrain.push(...b.facts)
|
||||
expect(viaStorage.map((f) => f.generation)).toEqual(viaBrain.map((f) => f.generation))
|
||||
expect(storage.factSegmentPaths()).toEqual(brain.factSegmentPaths())
|
||||
})
|
||||
|
||||
it('scan telemetry carries the frozen shape end-to-end', async () => {
|
||||
for (let i = 0; i < 5; i++) await brain.add({ data: `t${i}`, type: 'document', metadata: { i } })
|
||||
|
||||
const scan = brain.scanFacts({ batchSize: 2 })!
|
||||
expect(scan.headGeneration).toBeGreaterThanOrEqual(5)
|
||||
expect(scan.approxFactCount).toBeGreaterThanOrEqual(5)
|
||||
let batches = 0
|
||||
for await (const b of scan.batches()) {
|
||||
batches++
|
||||
expect(b.factCount).toBe(b.facts.length)
|
||||
expect(b.firstGeneration).toBe(b.facts[0].generation)
|
||||
expect(b.lastGeneration).toBe(b.facts[b.facts.length - 1].generation)
|
||||
expect(b.byteSize).toBeGreaterThan(0)
|
||||
expect(typeof b.segmentId).toBe('string')
|
||||
}
|
||||
expect(batches).toBeGreaterThan(1)
|
||||
expect(scan.summary().factsYielded).toBe(scan.approxFactCount)
|
||||
})
|
||||
})
|
||||
|
||||
describe('fact log dual-write (filesystem adapter — durability + protection)', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
const open = async () => {
|
||||
const b: any = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true,
|
||||
dimensions: 384
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-factlog-'))
|
||||
brain = await open()
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('facts survive close + reopen and appends continue monotonically', async () => {
|
||||
const id = await brain.add({ data: 'persist me', type: 'document', metadata: { k: 1 } })
|
||||
await brain.remove(id)
|
||||
await brain.close()
|
||||
|
||||
brain = await open()
|
||||
const facts = await allFacts(brain)
|
||||
expect(facts.length).toBeGreaterThanOrEqual(2)
|
||||
const headBefore = facts[facts.length - 1].generation
|
||||
|
||||
await brain.add({ data: 'after reopen', type: 'document', metadata: { k: 2 } })
|
||||
const facts2 = await allFacts(brain)
|
||||
expect(facts2[facts2.length - 1].generation).toBeGreaterThan(headBefore)
|
||||
})
|
||||
|
||||
it('the fact segments exist on disk under _generations/facts/ with zero-padded names', async () => {
|
||||
await brain.add({ data: 'on disk', type: 'document', metadata: {} })
|
||||
await brain.flush()
|
||||
const factsDir = path.join(dir, '_generations', 'facts')
|
||||
const files = fs.readdirSync(factsDir)
|
||||
// The manifest rides the store's JSON object discipline (gzip on disk).
|
||||
expect(files.some((f) => f.startsWith('manifest.json'))).toBe(true)
|
||||
const segs = files.filter((f) => /^seg-\d{20}\.bfl$/.test(f))
|
||||
expect(segs.length).toBeGreaterThanOrEqual(1)
|
||||
})
|
||||
|
||||
it('the fact-log namespace is PROTECTED: a prefix-nuke is refused', async () => {
|
||||
await brain.add({ data: 'protected', type: 'document', metadata: {} })
|
||||
await expect(brain.storage.removeRawPrefix('_generations/facts')).rejects.toBeInstanceOf(
|
||||
ProtectedArtifactError
|
||||
)
|
||||
// Per-generation history cleanup remains unaffected (no false intersect).
|
||||
await expect(brain.storage.removeRawPrefix('_generations/999999')).resolves.toBeUndefined()
|
||||
})
|
||||
|
||||
it('transact facts are durable-on-return (no flush needed before reopen)', async () => {
|
||||
const receipt = await brain.transact([
|
||||
{ op: 'add', type: 'document', metadata: { durable: true }, data: 'tx' }
|
||||
])
|
||||
// Simulate an abrupt end: no flush(), no close() — reopen from disk.
|
||||
brain = await open()
|
||||
const facts = await allFacts(brain)
|
||||
expect(facts.some((f) => f.generation === receipt.generation)).toBe(true)
|
||||
})
|
||||
})
|
||||
|
|
@ -1,164 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/graph-audit
|
||||
* @description brain.auditGraph() — the read-only graph-truth instrument.
|
||||
* Laws under test:
|
||||
* (1) a healthy brain audits COHERENT: every canonical verb is returned by the
|
||||
* read path of its source, endpoints exist, no ghosts;
|
||||
* (2) design-hidden edges (internal/system visibility) are counted separately
|
||||
* and never misclassified as index loss;
|
||||
* (3) a verb whose endpoint entity was destroyed at the storage layer (the
|
||||
* scar class) is flagged as a dangling endpoint, loudly;
|
||||
* (4) the classification core flags present-but-invisible and ghost edges
|
||||
* exactly (exercised via injected seams — manufacturing a genuinely stale
|
||||
* adjacency index end-to-end would require corrupting internals the
|
||||
* public API rightly refuses to corrupt).
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy } from '../../src/index.js'
|
||||
import { NounType, VerbType } from '../../src/types/graphTypes.js'
|
||||
import { runGraphAudit, type AuditVerbRecord } from '../../src/graph/graphAudit.js'
|
||||
|
||||
describe('brain.auditGraph() — graph-truth audit', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
beforeEach(async () => {
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-graph-audit-'))
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
silent: true
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
async function seedGraph(): Promise<{ ids: string[]; verbIds: string[] }> {
|
||||
const ids: string[] = []
|
||||
for (let i = 0; i < 5; i++) {
|
||||
ids.push(
|
||||
await brain.add({
|
||||
data: `entity ${i}`,
|
||||
type: NounType.Concept,
|
||||
metadata: { n: i }
|
||||
})
|
||||
)
|
||||
}
|
||||
const verbIds: string[] = []
|
||||
verbIds.push(await brain.relate({ from: ids[0], to: ids[1], type: VerbType.RelatedTo }))
|
||||
verbIds.push(await brain.relate({ from: ids[0], to: ids[2], type: VerbType.Contains }))
|
||||
verbIds.push(await brain.relate({ from: ids[1], to: ids[3], type: VerbType.DependsOn }))
|
||||
verbIds.push(await brain.relate({ from: ids[3], to: ids[4], type: VerbType.RelatedTo }))
|
||||
return { ids, verbIds }
|
||||
}
|
||||
|
||||
it('audits a healthy brain as coherent, with exact counts', async () => {
|
||||
const { ids } = await seedGraph()
|
||||
const report = await brain.auditGraph()
|
||||
|
||||
expect(report.coherent).toBe(true)
|
||||
expect(report.verbsInCanonical).toBe(4)
|
||||
expect(report.entitiesInCanonical).toBeGreaterThanOrEqual(ids.length) // VFS root etc. may add system nouns
|
||||
expect(report.sourcesChecked).toBe(3) // ids[0], ids[1], ids[3]
|
||||
expect(report.missingFromReadsCount).toBe(0)
|
||||
expect(report.danglingEndpointsCount).toBe(0)
|
||||
expect(report.readOnlyCount).toBe(0)
|
||||
expect(report.truncatedExamples).toBe(false)
|
||||
})
|
||||
|
||||
it('counts design-hidden edges separately and stays coherent', async () => {
|
||||
const { ids } = await seedGraph()
|
||||
await brain.relate({
|
||||
from: ids[2],
|
||||
to: ids[4],
|
||||
type: VerbType.RelatedTo,
|
||||
visibility: 'internal'
|
||||
})
|
||||
|
||||
const report = await brain.auditGraph()
|
||||
expect(report.coherent).toBe(true) // hidden-by-design is NOT a discrepancy
|
||||
expect(report.verbsInCanonical).toBe(5)
|
||||
expect(report.visibilityHiddenCount).toBeGreaterThanOrEqual(1)
|
||||
})
|
||||
|
||||
it('flags a destroyed endpoint as a dangling verb (the scar class)', async () => {
|
||||
const { ids } = await seedGraph()
|
||||
// Destroy ids[4] at the STORAGE layer (bypassing remove(), which would
|
||||
// also delete its verbs) — the historical partial-delete scar shape.
|
||||
await brain.storage.deleteNoun(ids[4])
|
||||
|
||||
const report = await brain.auditGraph()
|
||||
expect(report.coherent).toBe(false)
|
||||
expect(report.danglingEndpointsCount).toBe(1)
|
||||
expect(report.danglingEndpoints[0].to).toBe(ids[4])
|
||||
expect(report.danglingEndpoints[0].missingEnd).toBe('to')
|
||||
})
|
||||
})
|
||||
|
||||
describe('runGraphAudit classification core (injected seams)', () => {
|
||||
const verb = (id: string, from: string, to: string): AuditVerbRecord => ({
|
||||
id,
|
||||
type: 'relatedTo',
|
||||
sourceId: from,
|
||||
targetId: to
|
||||
})
|
||||
|
||||
const deps = (opts: {
|
||||
nouns: string[]
|
||||
verbs: AuditVerbRecord[]
|
||||
reads: Record<string, string[]> // sourceId -> verb ids the read path returns
|
||||
}) => ({
|
||||
eachNounId: async (consume: (id: string) => void) => {
|
||||
for (const id of opts.nouns) consume(id)
|
||||
},
|
||||
eachVerb: async (consume: (v: AuditVerbRecord) => void) => {
|
||||
for (const v of opts.verbs) consume(v)
|
||||
},
|
||||
readRelationsFrom: async (sourceId: string) =>
|
||||
(opts.reads[sourceId] ?? []).map((id) => ({ id }))
|
||||
})
|
||||
|
||||
it('flags a canonical verb the read path omits — present but invisible', async () => {
|
||||
const report = await runGraphAudit(
|
||||
deps({
|
||||
nouns: ['A', 'B', 'C'],
|
||||
verbs: [verb('v1', 'A', 'B'), verb('v2', 'A', 'C')],
|
||||
reads: { A: ['v1'] } // v2 exists canonically but reads miss it
|
||||
})
|
||||
)
|
||||
expect(report.coherent).toBe(false)
|
||||
expect(report.missingFromReadsCount).toBe(1)
|
||||
expect(report.missingFromReads[0].verbId).toBe('v2')
|
||||
})
|
||||
|
||||
it('flags a read-path edge with no canonical record — a ghost', async () => {
|
||||
const report = await runGraphAudit(
|
||||
deps({
|
||||
nouns: ['A', 'B'],
|
||||
verbs: [verb('v1', 'A', 'B')],
|
||||
reads: { A: ['v1', 'ghost-9'] }
|
||||
})
|
||||
)
|
||||
expect(report.coherent).toBe(false)
|
||||
expect(report.readOnlyCount).toBe(1)
|
||||
expect(report.readOnlyVerbIds).toEqual(['ghost-9'])
|
||||
})
|
||||
|
||||
it('caps example lists but keeps counts exact, and says so', async () => {
|
||||
const verbs = Array.from({ length: 10 }, (_, i) => verb(`v${i}`, 'A', 'B'))
|
||||
const report = await runGraphAudit(
|
||||
deps({ nouns: ['A', 'B'], verbs, reads: { A: [] } }),
|
||||
{ maxExamples: 3 }
|
||||
)
|
||||
expect(report.missingFromReadsCount).toBe(10)
|
||||
expect(report.missingFromReads.length).toBe(3)
|
||||
expect(report.truncatedExamples).toBe(true)
|
||||
})
|
||||
})
|
||||
|
|
@ -1,186 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/history-repacking
|
||||
* @description The D1+D3 two-tier history lifecycle end-to-end on a real
|
||||
* brain. Laws: (1) repacking is RE-REPRESENTATION — after folding, every
|
||||
* asOf() read below the fold boundary answers exactly as before, across a
|
||||
* cold reopen; (2) folded per-generation directories are physically gone
|
||||
* (the file-count cure is real, not cosmetic); (3) repack + reclaim compose:
|
||||
* bounded retention after repacking drops whole segments and asOf below the
|
||||
* horizon throws GenerationCompactedError; (4) repackHistory is explicit
|
||||
* API and time-bounded (spent budget = consistent no-op).
|
||||
*
|
||||
* Uses a tiny REPACK_LIVE_WINDOW override so a small history has a cold
|
||||
* tier at all (the production window is 1024).
|
||||
*/
|
||||
import { describe, it, expect, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as path from 'node:path'
|
||||
import * as os from 'node:os'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
import { GenerationStore } from '../../src/db/generationStore.js'
|
||||
import { GenerationCompactedError } from '../../src/db/errors.js'
|
||||
import { SEGMENTS_PREFIX } from '../../src/db/generationSegments.js'
|
||||
|
||||
const stub = async (text: string): Promise<number[]> => {
|
||||
const h = text.split('').reduce((a, c) => a + c.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(h + i))
|
||||
}
|
||||
|
||||
const openBrain = async (dir: string): Promise<Brainy> => {
|
||||
const brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
embeddingFunction: stub
|
||||
})
|
||||
await brain.init()
|
||||
return brain
|
||||
}
|
||||
|
||||
describe('history repacking — the two-tier lifecycle', () => {
|
||||
const dirs: string[] = []
|
||||
const tempDir = (): string => {
|
||||
const d = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-repack-'))
|
||||
dirs.push(d)
|
||||
return d
|
||||
}
|
||||
const originalWindow = GenerationStore.REPACK_LIVE_WINDOW
|
||||
|
||||
afterEach(() => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = originalWindow
|
||||
for (const d of dirs.splice(0)) {
|
||||
try {
|
||||
fs.rmSync(d, { recursive: true, force: true })
|
||||
} catch {
|
||||
/* best effort */
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
it('repack preserves every historical read across cold reopen; folded dirs are gone', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 3
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
|
||||
const id = await brain.add({
|
||||
data: 'versioned-entity',
|
||||
type: NounType.Document,
|
||||
metadata: { v: 0 }
|
||||
})
|
||||
for (let v = 1; v <= 10; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
|
||||
// Ground truth BEFORE repacking: capture asOf views for early generations.
|
||||
const before: Record<number, number> = {}
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await brain.asOf(g)
|
||||
before[g] = (await db.get(id))?.metadata?.v as number
|
||||
await db.release()
|
||||
}
|
||||
|
||||
const result = await brain.repackHistory()
|
||||
expect(result.foldedGenerations).toBeGreaterThan(0)
|
||||
expect(result.segmentsCreated).toBeGreaterThan(0)
|
||||
|
||||
// The folded per-generation directories are PHYSICALLY gone…
|
||||
const genDirs = fs
|
||||
.readdirSync(path.join(dir, '_generations'), { withFileTypes: true })
|
||||
.filter((e) => e.isDirectory() && /^\d+$/.test(e.name)).length
|
||||
expect(genDirs).toBeLessThanOrEqual(4) // live window (3) + at most the newest
|
||||
// …and the segment tier exists (the filesystem adapter stores objects
|
||||
// gzipped, so the manifest may live at either spelling).
|
||||
const segDir = path.join(dir, SEGMENTS_PREFIX)
|
||||
expect(
|
||||
fs.existsSync(path.join(segDir, 'manifest.json')) ||
|
||||
fs.existsSync(path.join(segDir, 'manifest.json.gz'))
|
||||
).toBe(true)
|
||||
expect(fs.readdirSync(segDir).some((f) => f.endsWith('.bgs'))).toBe(true)
|
||||
|
||||
// Same asOf answers from the packed tier, same process…
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await brain.asOf(g)
|
||||
expect((await db.get(id))?.metadata?.v).toBe(before[g])
|
||||
await db.release()
|
||||
}
|
||||
await brain.close()
|
||||
|
||||
// …and across a COLD REOPEN (manifest discovery, no live dirs to list).
|
||||
const reopened = await openBrain(dir)
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await reopened.asOf(g)
|
||||
expect((await db.get(id))?.metadata?.v).toBe(before[g])
|
||||
await db.release()
|
||||
}
|
||||
expect((await reopened.get(id))?.metadata?.v).toBe(10) // live state untouched
|
||||
await reopened.close()
|
||||
})
|
||||
|
||||
it('repack + bounded reclaim compose: whole segments drop, horizon is loud', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'reclaim-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 8; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
await brain.repackHistory()
|
||||
|
||||
// Reclaim down to the 3 newest generations — packed segments below the
|
||||
// horizon drop whole; asOf below throws loudly.
|
||||
const res = await brain.compactHistory({ maxGenerations: 3 })
|
||||
expect(res.removedGenerations).toBeGreaterThan(0)
|
||||
await expect(brain.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
expect((await brain.get(id))?.metadata?.v).toBe(8)
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('generationDigest: reopen-stable, divergence-sensitive, loud below the horizon', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'digest-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
await brain.repackHistory()
|
||||
|
||||
const gen = brain.generation()
|
||||
const atHead = await brain.generationDigest(gen)
|
||||
const atMid = await brain.generationDigest(3)
|
||||
expect(atHead).toMatch(/^[0-9a-f]{8}$/)
|
||||
expect(atMid).not.toBe(atHead) // more history ⇒ different digest
|
||||
await brain.close()
|
||||
|
||||
// Reopen-stable: same history, same digests (packed prefix stability).
|
||||
const reopened = await openBrain(dir)
|
||||
expect(await reopened.generationDigest(gen)).toBe(atHead)
|
||||
expect(await reopened.generationDigest(3)).toBe(atMid)
|
||||
|
||||
// New history diverges the head digest.
|
||||
await reopened.update({ id, metadata: { v: 7 } })
|
||||
await reopened.flush()
|
||||
expect(await reopened.generationDigest(reopened.generation())).not.toBe(atHead)
|
||||
|
||||
// Below the horizon: LOUD, never a silent pin of reclaimed history.
|
||||
await reopened.compactHistory({ maxGenerations: 2 })
|
||||
await expect(reopened.generationDigest(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
await reopened.close()
|
||||
})
|
||||
|
||||
it('a spent time budget is a consistent no-op; the next pass resumes', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'budget-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
|
||||
const bounded = await brain.repackHistory({ timeBudgetMs: 0 })
|
||||
expect(bounded).toEqual({ foldedGenerations: 0, segmentsCreated: 0 })
|
||||
|
||||
const resumed = await brain.repackHistory()
|
||||
expect(resumed.foldedGenerations).toBeGreaterThan(0)
|
||||
const db = await brain.asOf(3)
|
||||
expect((await db.get(id))?.metadata?.v).toBeDefined()
|
||||
await db.release()
|
||||
await brain.close()
|
||||
})
|
||||
})
|
||||
|
|
@ -1,138 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/lens-consistency
|
||||
* @description The three metadata "lenses" over one corpus must agree with
|
||||
* canonical ground truth id-for-id, warm AND after a cold reopen:
|
||||
* - combined: find({ type: T, where: { subtype: S } })
|
||||
* - subtype-only: find({ where: { subtype: S } })
|
||||
* - type-only: find({ type: T })
|
||||
* Ported from the fresh-brain probe that closed the type+subtype lens-drop
|
||||
* investigation (a restored pre-8.2.2 torn capture had entities visible to the
|
||||
* subtype-only lens but dropped by the combined lens — "0 of 2 migrated, all
|
||||
* gates green"). The corpus is seeded through the REAL write API — never
|
||||
* restored bytes — which is what made the original datapoint decisive. The
|
||||
* invariants: every lens matches an unfiltered canonical scan exactly (no
|
||||
* missing ids, no extras) and combined ⊆ subtype-only always holds.
|
||||
*/
|
||||
import { describe, it, expect, beforeAll, afterAll } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy } from '../../src/index.js'
|
||||
|
||||
/** The corpus: 7 (type, subtype) pairs, uneven counts, incl. the incident's 2-of-a-pair shape. */
|
||||
const CORPUS: Array<{ type: string; subtype: string; count: number }> = [
|
||||
{ type: 'proposition', subtype: 'decision', count: 2 }, // the incident shape: "0 of 2"
|
||||
{ type: 'concept', subtype: 'decision', count: 3 },
|
||||
{ type: 'task', subtype: 'decision', count: 2 },
|
||||
{ type: 'concept', subtype: 'action', count: 4 },
|
||||
{ type: 'message', subtype: 'note', count: 5 },
|
||||
{ type: 'message', subtype: 'ship', count: 3 },
|
||||
{ type: 'document', subtype: 'guide', count: 4 }
|
||||
]
|
||||
|
||||
/** Canonical ground truth: unfiltered enumeration, post-filtered IN THE TEST. */
|
||||
async function groundTruth(
|
||||
brain: any,
|
||||
match: { type?: string; subtype?: string }
|
||||
): Promise<Set<string>> {
|
||||
const ids = new Set<string>()
|
||||
let cursor: string | undefined
|
||||
for (;;) {
|
||||
const page = await brain.storage.getNounsWithPagination({ limit: 500, cursor })
|
||||
for (const noun of page.items) {
|
||||
// Hydrated shape: `type`/`subtype` are TOP-LEVEL; `metadata` holds only
|
||||
// custom user fields (vfsType is one — the VFS plumbing marker).
|
||||
const n = noun as any
|
||||
if (n.metadata?.vfsType) continue // VFS plumbing is not corpus
|
||||
if (!n.type || !n.subtype) continue
|
||||
if (match.type && n.type !== match.type) continue
|
||||
if (match.subtype && n.subtype !== match.subtype) continue
|
||||
ids.add(n.id)
|
||||
}
|
||||
if (!page.hasMore) break
|
||||
cursor = page.nextCursor
|
||||
}
|
||||
return ids
|
||||
}
|
||||
|
||||
const idSet = (results: Array<{ id: string }>): Set<string> => new Set(results.map((r) => r.id))
|
||||
|
||||
/** Every lens vs ground truth, id-for-id, for every pair in the corpus. */
|
||||
async function assertAllLenses(brain: any): Promise<void> {
|
||||
const types = [...new Set(CORPUS.map((c) => c.type))]
|
||||
const subtypes = [...new Set(CORPUS.map((c) => c.subtype))]
|
||||
|
||||
for (const { type, subtype } of CORPUS) {
|
||||
const combined = idSet(await brain.find({ type, where: { subtype }, limit: 1000 }))
|
||||
const subtypeOnly = idSet(await brain.find({ where: { subtype }, limit: 1000 }))
|
||||
const truthPair = await groundTruth(brain, { type, subtype })
|
||||
const truthSubtype = await groundTruth(brain, { subtype })
|
||||
|
||||
expect([...combined].sort()).toEqual([...truthPair].sort()) // no drops, no extras
|
||||
expect([...subtypeOnly].sort()).toEqual([...truthSubtype].sort())
|
||||
for (const id of combined) expect(subtypeOnly.has(id)).toBe(true) // combined ⊆ subtype-only
|
||||
}
|
||||
|
||||
for (const type of types) {
|
||||
const typeOnly = idSet(await brain.find({ type, limit: 1000 }))
|
||||
const truthType = await groundTruth(brain, { type })
|
||||
expect([...typeOnly].sort()).toEqual([...truthType].sort())
|
||||
}
|
||||
|
||||
// Count cross-check against the corpus definition itself.
|
||||
for (const subtype of subtypes) {
|
||||
const expected = CORPUS.filter((c) => c.subtype === subtype).reduce((s, c) => s + c.count, 0)
|
||||
const got = (await brain.find({ where: { subtype }, limit: 1000 })).length
|
||||
expect(got).toBe(expected)
|
||||
}
|
||||
}
|
||||
|
||||
describe('lens consistency — combined vs subtype-only vs canonical ground truth', () => {
|
||||
let dir: string
|
||||
let brain: any
|
||||
|
||||
beforeAll(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
dir = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-lens-'))
|
||||
brain = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir }, silent: true, dimensions: 384 })
|
||||
await brain.init()
|
||||
// Seed through the REAL write API — never restored bytes.
|
||||
let i = 0
|
||||
for (const { type, subtype, count } of CORPUS) {
|
||||
for (let k = 0; k < count; k++) {
|
||||
await brain.add({ data: `${type} ${subtype} ${i++}`, type, subtype, metadata: { k } })
|
||||
}
|
||||
}
|
||||
await brain.flush()
|
||||
})
|
||||
afterAll(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
fs.rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('WARM: all lenses agree with ground truth id-for-id', async () => {
|
||||
await assertAllLenses(brain)
|
||||
})
|
||||
|
||||
it('COLD REOPEN: all lenses still agree after close + reopen from disk', async () => {
|
||||
await brain.close()
|
||||
brain = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir }, silent: true, dimensions: 384 })
|
||||
await brain.init()
|
||||
await assertAllLenses(brain)
|
||||
})
|
||||
|
||||
it('after an update() flips type AND subtype, every lens tracks the move exactly', async () => {
|
||||
// The historical cross-bucket-staleness path: change (concept, action) -> (task, review).
|
||||
const victims = await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1 })
|
||||
expect(victims.length).toBe(1)
|
||||
const id = victims[0].id
|
||||
await brain.update({ id, type: 'task', subtype: 'review' })
|
||||
|
||||
const oldCombined = idSet(await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1000 }))
|
||||
expect(oldCombined.has(id)).toBe(false) // unposted from the old buckets
|
||||
const newCombined = idSet(await brain.find({ type: 'task', where: { subtype: 'review' }, limit: 1000 }))
|
||||
expect(newCombined.has(id)).toBe(true) // posted to the new buckets
|
||||
const subtypeOnly = idSet(await brain.find({ where: { subtype: 'review' }, limit: 1000 }))
|
||||
expect(subtypeOnly.has(id)).toBe(true)
|
||||
})
|
||||
})
|
||||
|
|
@ -110,89 +110,6 @@ describe('Multi-process safety + read-only mode', () => {
|
|||
// Don't track `blocked` for afterEach cleanup since init failed.
|
||||
})
|
||||
|
||||
it('takes over a STALE foreign lock (dead PID + old heartbeat) and claims atomically', async () => {
|
||||
const { mkdirSync, writeFileSync, readFileSync } = await import('node:fs')
|
||||
const { join } = await import('node:path')
|
||||
const os = await import('node:os')
|
||||
mkdirSync(join(dir, 'locks'), { recursive: true })
|
||||
const tenMinutesAgo = new Date(Date.now() - 10 * 60 * 1000).toISOString()
|
||||
writeFileSync(join(dir, 'locks', '_writer.lock'), JSON.stringify({
|
||||
pid: 999999999, // no such process — provably dead
|
||||
hostname: os.hostname(),
|
||||
startedAt: tenMinutesAgo,
|
||||
lastHeartbeat: tenMinutesAgo,
|
||||
version: '8.0.0',
|
||||
rootDir: dir
|
||||
}))
|
||||
|
||||
writer = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir } })
|
||||
await writer.init() // stale takeover must succeed
|
||||
|
||||
const lock = JSON.parse(readFileSync(join(dir, 'locks', '_writer.lock'), 'utf-8'))
|
||||
expect(lock.pid).toBe(process.pid) // the atomic wx claim installed OUR lock
|
||||
})
|
||||
|
||||
it('the writer-locked error carries the machine-readable contract (code + lockInfo)', async () => {
|
||||
const { mkdirSync, writeFileSync } = await import('node:fs')
|
||||
const { join } = await import('node:path')
|
||||
const os = await import('node:os')
|
||||
mkdirSync(join(dir, 'locks'), { recursive: true })
|
||||
const otherPid = (process as any).ppid || 1
|
||||
writeFileSync(join(dir, 'locks', '_writer.lock'), JSON.stringify({
|
||||
pid: otherPid,
|
||||
hostname: os.hostname(),
|
||||
startedAt: new Date().toISOString(),
|
||||
lastHeartbeat: new Date().toISOString(),
|
||||
version: '8.7.0',
|
||||
rootDir: dir
|
||||
}))
|
||||
|
||||
const blocked = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir } })
|
||||
const err: any = await blocked.init().catch((e) => e)
|
||||
expect(err.code).toBe('BRAINY_WRITER_LOCKED')
|
||||
expect(err.lockInfo?.pid).toBe(otherPid)
|
||||
})
|
||||
|
||||
it('release drains an in-flight heartbeat — no phantom lock re-created after unlink', async () => {
|
||||
// The race (8.9.0): clearInterval stops FUTURE heartbeat ticks, but a
|
||||
// tick already in flight could land its lock rewrite AFTER release's
|
||||
// unlink — re-creating the lock as a phantom that blocks the next
|
||||
// writer until the stale TTL. Simulate the in-flight tick explicitly
|
||||
// and prove release waits for it.
|
||||
writer = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir } })
|
||||
await writer.init()
|
||||
const storage: any = (writer as any).storage
|
||||
|
||||
// An in-flight refresh that is ALREADY PAST its ownership guards
|
||||
// (captured the lock info before release ran) and lands its atomic
|
||||
// rewrite slowly — the exact straggler shape; absent the drain it
|
||||
// writes after the unlink.
|
||||
const { join: joinPath } = await import('node:path')
|
||||
const capturedInfo = { ...storage.writerLockInfo }
|
||||
const lockPath = joinPath(dir, 'locks', '_writer.lock')
|
||||
const slowTick = (async () => {
|
||||
await new Promise((r) => setTimeout(r, 100))
|
||||
await storage.writeFileAtomic(
|
||||
lockPath,
|
||||
JSON.stringify({ ...capturedInfo, lastHeartbeat: new Date().toISOString() })
|
||||
)
|
||||
})()
|
||||
storage.writerHeartbeatInFlight = slowTick.catch(() => {})
|
||||
|
||||
await writer.close() // → releaseWriterLock must drain slowTick first
|
||||
await slowTick.catch(() => {}) // both paths fully settled either way
|
||||
writer = null
|
||||
|
||||
const { existsSync } = await import('node:fs')
|
||||
const { join } = await import('node:path')
|
||||
expect(existsSync(join(dir, 'locks', '_writer.lock'))).toBe(false)
|
||||
|
||||
// And the directory is immediately claimable — no stale-TTL wait.
|
||||
const next = new Brainy({ requireSubtype: false, storage: { type: 'filesystem', path: dir } })
|
||||
await expect(next.init()).resolves.toBeUndefined()
|
||||
await next.close()
|
||||
})
|
||||
|
||||
it('allows a second in-process writer with a warning (same PID)', async () => {
|
||||
// Two Brainy instances in the same Node process: not the dangerous
|
||||
// cross-process case. Should succeed (with a console warning).
|
||||
|
|
|
|||
|
|
@ -1,157 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/vfs-rename-containment
|
||||
* @description A cross-directory rename must MOVE the containment edge, not
|
||||
* accumulate one per parent. The pre-fix rename added the new parent's
|
||||
* Contains edge but skipped removing the old one ("not critical") — leaving
|
||||
* the entity a child of BOTH directories: readdir(oldDir) kept listing it,
|
||||
* re-creating the old path showed the name twice (the duplicate-readdir /
|
||||
* "cosmetic ghost" field report), and tree-walking consumers saw the file in
|
||||
* two places. Also covers the repair sweep (repairIndex → vfs.repairContainment)
|
||||
* that heals ghosts left by earlier versions, and move-to-root (whose edge
|
||||
* used to be skipped entirely).
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import { Brainy, VerbType } from '../../src/index.js'
|
||||
|
||||
describe('VFS rename moves the containment edge (no ghost in the old directory)', () => {
|
||||
let brain: any
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
brain = new Brainy({ requireSubtype: false, storage: { type: 'memory' }, silent: true, dimensions: 384 })
|
||||
await brain.init()
|
||||
await brain.vfs.mkdir('/a', { recursive: true })
|
||||
await brain.vfs.mkdir('/b', { recursive: true })
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
})
|
||||
|
||||
it('cross-directory move: old dir stops listing it, new dir lists it exactly once', async () => {
|
||||
await brain.vfs.writeFile('/a/x.txt', 'v1')
|
||||
await brain.vfs.rename('/a/x.txt', '/b/x.txt')
|
||||
|
||||
expect(await brain.vfs.readdir('/a')).toEqual([])
|
||||
expect(await brain.vfs.readdir('/b')).toEqual(['x.txt'])
|
||||
expect(String(await brain.vfs.readFile('/b/x.txt'))).toBe('v1')
|
||||
})
|
||||
|
||||
it('re-creating the old path lists each name exactly once in each directory', async () => {
|
||||
await brain.vfs.writeFile('/a/x.txt', 'moved away')
|
||||
await brain.vfs.rename('/a/x.txt', '/b/x.txt')
|
||||
await brain.vfs.writeFile('/a/x.txt', 'new file at old path')
|
||||
|
||||
const a = (await brain.vfs.readdir('/a')) as string[]
|
||||
const b = (await brain.vfs.readdir('/b')) as string[]
|
||||
expect(a).toEqual(['x.txt']) // exactly once — the pre-fix ghost made this list the moved entity too
|
||||
expect(b).toEqual(['x.txt'])
|
||||
expect(String(await brain.vfs.readFile('/a/x.txt'))).toBe('new file at old path')
|
||||
expect(String(await brain.vfs.readFile('/b/x.txt'))).toBe('moved away')
|
||||
})
|
||||
|
||||
it('repeated moves never accumulate containment edges', async () => {
|
||||
await brain.vfs.writeFile('/a/f.txt', 'wanderer')
|
||||
for (let i = 0; i < 3; i++) {
|
||||
await brain.vfs.rename('/a/f.txt', '/b/f.txt')
|
||||
await brain.vfs.rename('/b/f.txt', '/a/f.txt')
|
||||
}
|
||||
expect(await brain.vfs.readdir('/a')).toEqual(['f.txt'])
|
||||
expect(await brain.vfs.readdir('/b')).toEqual([])
|
||||
|
||||
// Exactly ONE containment edge exists on the entity.
|
||||
const stat = await brain.vfs.stat('/a/f.txt')
|
||||
const edges = await brain.related({ to: stat.entityId, type: VerbType.Contains })
|
||||
expect(edges).toHaveLength(1)
|
||||
})
|
||||
|
||||
it('move to the root gets a containment edge (used to be skipped → orphan)', async () => {
|
||||
await brain.vfs.writeFile('/a/up.txt', 'to the top')
|
||||
await brain.vfs.rename('/a/up.txt', '/up.txt')
|
||||
|
||||
const rootListing = (await brain.vfs.readdir('/')) as string[]
|
||||
expect(rootListing).toContain('up.txt')
|
||||
expect(await brain.vfs.readdir('/a')).toEqual([])
|
||||
expect(String(await brain.vfs.readFile('/up.txt'))).toBe('to the top')
|
||||
})
|
||||
})
|
||||
|
||||
describe('repairIndex() heals pre-fix containment ghosts (vfs.repairContainment)', () => {
|
||||
let brain: any
|
||||
|
||||
beforeEach(async () => {
|
||||
process.env.BRAINY_DETERMINISTIC_EMBEDDINGS = 'true'
|
||||
brain = new Brainy({ requireSubtype: false, storage: { type: 'memory' }, silent: true, dimensions: 384 })
|
||||
await brain.init()
|
||||
await brain.vfs.mkdir('/a', { recursive: true })
|
||||
await brain.vfs.mkdir('/b', { recursive: true })
|
||||
})
|
||||
afterEach(async () => {
|
||||
await brain.close?.().catch(() => {})
|
||||
})
|
||||
|
||||
/** Reproduce the PRE-FIX defect state: file lives at /b/g.txt but a stale
|
||||
* vfs-contains edge from /a lingers (what old renames left behind). */
|
||||
const synthesizeGhost = async () => {
|
||||
await brain.vfs.writeFile('/b/g.txt', 'ghost target')
|
||||
const fileId = (await brain.vfs.stat('/b/g.txt')).entityId
|
||||
const aId = (await brain.vfs.stat('/a')).entityId
|
||||
await brain.relate({
|
||||
from: aId,
|
||||
to: fileId,
|
||||
type: VerbType.Contains,
|
||||
subtype: 'vfs-contains',
|
||||
metadata: { isVFS: true }
|
||||
})
|
||||
return { fileId, aId }
|
||||
}
|
||||
|
||||
it('a stale old-parent edge is removed; listings become honest', async () => {
|
||||
const { fileId } = await synthesizeGhost()
|
||||
// The defect state is visible: /a lists a file whose path says /b.
|
||||
expect((await brain.vfs.readdir('/a')) as string[]).toContain('g.txt')
|
||||
|
||||
await brain.repairIndex()
|
||||
|
||||
expect(await brain.vfs.readdir('/a')).toEqual([])
|
||||
expect(await brain.vfs.readdir('/b')).toEqual(['g.txt'])
|
||||
const edges = await brain.related({ to: fileId, type: VerbType.Contains })
|
||||
expect(edges).toHaveLength(1)
|
||||
})
|
||||
|
||||
it("a user's own knowledge Contains edge onto the file survives the repair", async () => {
|
||||
const { fileId } = await synthesizeGhost()
|
||||
// A knowledge-graph containment from a NON-directory entity (a collection
|
||||
// curating the file) — same verb TYPE, not a vfs-contains edge. The repair
|
||||
// must remove only VFS containment ghosts, never user knowledge edges.
|
||||
// (Edges dedupe by (from,to,type), so the user edge needs its own source.)
|
||||
const collectionId = await brain.add({
|
||||
data: 'reading list',
|
||||
type: 'collection',
|
||||
metadata: { kind: 'curation' }
|
||||
})
|
||||
await brain.relate({ from: collectionId, to: fileId, type: VerbType.Contains, subtype: 'curates' })
|
||||
|
||||
await brain.repairIndex()
|
||||
|
||||
const edges = await brain.related({ to: fileId, type: VerbType.Contains })
|
||||
const subtypes = edges.map((e: any) => e.subtype).sort()
|
||||
// The stale vfs ghost edge is gone; the correct vfs edge + the user's
|
||||
// curation edge remain.
|
||||
expect(subtypes).toEqual(['curates', 'vfs-contains'])
|
||||
expect(edges.find((e: any) => e.subtype === 'curates')?.from).toBe(collectionId)
|
||||
})
|
||||
|
||||
it('a missing expected edge is restored (entity unreachable from its own directory)', async () => {
|
||||
await brain.vfs.writeFile('/b/lost.txt', 'find me')
|
||||
const fileId = (await brain.vfs.stat('/b/lost.txt')).entityId
|
||||
// Simulate total edge loss (an older damage shape).
|
||||
for (const e of await brain.related({ to: fileId, type: VerbType.Contains })) {
|
||||
await brain.unrelate(e.id)
|
||||
}
|
||||
expect((await brain.vfs.readdir('/b')) as string[]).not.toContain('lost.txt')
|
||||
|
||||
await brain.repairIndex()
|
||||
|
||||
expect((await brain.vfs.readdir('/b')) as string[]).toContain('lost.txt')
|
||||
})
|
||||
})
|
||||
|
|
@ -205,10 +205,10 @@ describe('Metadata index cleanup after remove / removeMany', () => {
|
|||
}
|
||||
})
|
||||
|
||||
it('refuses an empty ids array loudly (a silent no-op is not "graceful")', async () => {
|
||||
// 8.8.2: an empty selector used to resolve successfully having deleted
|
||||
// NOTHING — the caller believed the delete happened. Now it throws.
|
||||
await expect(brain.removeMany({ ids: [] })).rejects.toThrow(/ids: \[\]/)
|
||||
it('handles empty ids array gracefully', async () => {
|
||||
const result = await brain.removeMany({ ids: [] })
|
||||
expect(result.successful).toHaveLength(0)
|
||||
expect(result.failed).toHaveLength(0)
|
||||
})
|
||||
|
||||
it('handles large batch (> 1 chunk) without leaving stale index entries', async () => {
|
||||
|
|
|
|||
|
|
@ -10,7 +10,6 @@
|
|||
|
||||
import { describe, it, expect, beforeEach } from 'vitest'
|
||||
import { TransactionManager } from '../../src/transaction/TransactionManager.js'
|
||||
import { transactTimeoutBudget } from '../../src/transaction/Transaction.js'
|
||||
import { TransactionError } from '../../src/transaction/errors.js'
|
||||
|
||||
describe('TransactionManager', () => {
|
||||
|
|
@ -106,17 +105,14 @@ describe('TransactionManager', () => {
|
|||
const result = await manager.executeTransactionWithResult(async (tx) => {
|
||||
tx.addOperation({
|
||||
execute: async () => {
|
||||
await new Promise(resolve => setTimeout(resolve, 25))
|
||||
await new Promise(resolve => setTimeout(resolve, 10))
|
||||
return async () => {}
|
||||
}
|
||||
})
|
||||
return 'done'
|
||||
})
|
||||
|
||||
// Timer coalescing can fire a setTimeout up to a few ms EARLY under
|
||||
// load, so assert well below the sleep — this tests that time is
|
||||
// MEASURED, not the OS timer's precision.
|
||||
expect(result.executionTimeMs).toBeGreaterThanOrEqual(20)
|
||||
expect(result.executionTimeMs).toBeGreaterThanOrEqual(10)
|
||||
})
|
||||
})
|
||||
|
||||
|
|
@ -329,45 +325,4 @@ describe('TransactionManager', () => {
|
|||
expect(stats1).toEqual(stats2) // Same values
|
||||
})
|
||||
})
|
||||
|
||||
describe('Timeout budget + telemetry', () => {
|
||||
it('transactTimeoutBudget: explicit override wins; default scales with batch size', () => {
|
||||
expect(transactTimeoutBudget(1)).toBe(30_000) // small batches keep the 30s floor
|
||||
expect(transactTimeoutBudget(15)).toBe(30_000) // the old flat cap's break-even point
|
||||
expect(transactTimeoutBudget(100)).toBe(200_000) // 100 ops × 2s — bulk gets an honest budget
|
||||
expect(transactTimeoutBudget(1000, 5_000)).toBe(5_000) // caller override is untouched
|
||||
})
|
||||
|
||||
it('a tripped budget rolls back and names the operation, progress, and budget', async () => {
|
||||
const rolledBack: string[] = []
|
||||
|
||||
const failing = manager.executeTransaction(
|
||||
async (tx) => {
|
||||
tx.addOperation({
|
||||
name: 'slow-first-op',
|
||||
execute: async () => {
|
||||
await new Promise((r) => setTimeout(r, 30))
|
||||
return async () => {
|
||||
rolledBack.push('slow-first-op')
|
||||
}
|
||||
}
|
||||
})
|
||||
tx.addOperation({
|
||||
name: 'never-reached',
|
||||
execute: async () => undefined
|
||||
})
|
||||
},
|
||||
{ timeout: 5 } // the first op's 30ms sleep guarantees the pre-op-2 check trips
|
||||
)
|
||||
|
||||
await expect(failing).rejects.toThrow(TransactionError)
|
||||
const err = await failing.catch((e) => e)
|
||||
expect(err.name).toBe('TransactionTimeoutError')
|
||||
expect(err.message).toContain('operation 1/2') // which op, of how many
|
||||
expect(err.message).toContain("('never-reached')") // its name
|
||||
expect(err.message).toContain('budget 5ms') // the budget that tripped
|
||||
expect(err.message).toContain('rolled back') // the retryability statement
|
||||
expect(rolledBack).toEqual(['slow-first-op']) // the applied op was undone
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -675,104 +675,4 @@ describe('AggregationIndex', () => {
|
|||
await reloaded.close()
|
||||
})
|
||||
})
|
||||
|
||||
// ============= Boot-order reconciliation =============
|
||||
//
|
||||
// The production boot pattern: defineAggregate() is synchronous and always
|
||||
// beats the async init() that loads persisted state. The old code flagged a
|
||||
// backfill at define time and init never cleared it — so the loaded state
|
||||
// was wiped and the whole store re-walked on EVERY restart.
|
||||
|
||||
describe('boot-order reconciliation (define-before-init)', () => {
|
||||
const DEF: AggregateDefinition = {
|
||||
name: 'boot_agg',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['category'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
}
|
||||
|
||||
const entity = (id: string, category: string): Record<string, unknown> => ({
|
||||
id,
|
||||
noun: NounType.Event,
|
||||
metadata: { category },
|
||||
createdAt: Date.now(),
|
||||
updatedAt: Date.now()
|
||||
})
|
||||
|
||||
/** Simulate the previous session: define, contribute, flush. */
|
||||
async function seedAndFlush(store: MemoryStorage): Promise<void> {
|
||||
const first = new AggregationIndex(store)
|
||||
await first.init()
|
||||
first.defineAggregate(DEF)
|
||||
first.onEntityAdded('e1', entity('e1', 'food'))
|
||||
first.onEntityAdded('e2', entity('e2', 'food'))
|
||||
first.onEntityAdded('e3', entity('e3', 'transport'))
|
||||
await first.flush()
|
||||
}
|
||||
|
||||
it('adopts persisted state when define beats init with an unchanged definition', async () => {
|
||||
const store = new MemoryStorage()
|
||||
await store.init()
|
||||
await seedAndFlush(store)
|
||||
|
||||
const second = new AggregationIndex(store)
|
||||
second.defineAggregate(DEF) // synchronous define FIRST — the real boot order
|
||||
await second.init()
|
||||
await second.ready()
|
||||
|
||||
expect(second.getPendingBackfills()).toEqual([])
|
||||
const rows = second.queryAggregate({ name: 'boot_agg' })
|
||||
const food = rows.find(r => r.groupKey.category === 'food')!
|
||||
expect(food.metrics.count).toBe(2)
|
||||
})
|
||||
|
||||
it('a write landing before adoption forces an exact rescan instead', async () => {
|
||||
const store = new MemoryStorage()
|
||||
await store.init()
|
||||
await seedAndFlush(store)
|
||||
|
||||
const second = new AggregationIndex(store)
|
||||
second.defineAggregate(DEF)
|
||||
second.onEntityAdded('e4', entity('e4', 'food')) // lands before init settles
|
||||
await second.init()
|
||||
await second.ready()
|
||||
|
||||
// Adoption would lose e4's contribution — the engine must rescan.
|
||||
expect(second.getPendingBackfills()).toEqual(['boot_agg'])
|
||||
})
|
||||
|
||||
it('init never clobbers a changed app definition registered before it', async () => {
|
||||
const store = new MemoryStorage()
|
||||
await store.init()
|
||||
await seedAndFlush(store)
|
||||
|
||||
const CHANGED: AggregateDefinition = {
|
||||
...DEF,
|
||||
metrics: { count: { op: 'count' }, total: { op: 'sum', field: 'amount' } }
|
||||
}
|
||||
const second = new AggregationIndex(store)
|
||||
second.defineAggregate(CHANGED)
|
||||
await second.init()
|
||||
await second.ready()
|
||||
|
||||
const def = second.getDefinitions().find(d => d.name === 'boot_agg')!
|
||||
expect(Object.keys(def.metrics).sort()).toEqual(['count', 'total'])
|
||||
expect(second.getPendingBackfills()).toEqual(['boot_agg'])
|
||||
})
|
||||
|
||||
it('init alone restores persisted definitions with adopted state, no backfill', async () => {
|
||||
const store = new MemoryStorage()
|
||||
await store.init()
|
||||
await seedAndFlush(store)
|
||||
|
||||
const second = new AggregationIndex(store)
|
||||
await second.init()
|
||||
await second.ready()
|
||||
|
||||
expect(second.hasAggregate('boot_agg')).toBe(true)
|
||||
expect(second.getPendingBackfills()).toEqual([])
|
||||
const rows = second.queryAggregate({ name: 'boot_agg' })
|
||||
expect(rows.reduce((s, r) => s + (r.metrics.count as number), 0)).toBe(3)
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -533,10 +533,8 @@ describe('Brainy Batch Operations', () => {
|
|||
expect(result.successful).toHaveLength(0)
|
||||
|
||||
await brain.updateMany({ items: [] })
|
||||
// removeMany is the exception (8.8.2): an empty id list is a refused
|
||||
// selector, not an empty batch — deleting "nothing" silently was the
|
||||
// bug class (a positional/bare-array call looked identical).
|
||||
await expect(brain.removeMany({ ids: [] })).rejects.toThrow(/ids: \[\]/)
|
||||
await brain.removeMany({ ids: [] })
|
||||
// Should not throw
|
||||
})
|
||||
|
||||
it('should validate batch size limits', async () => {
|
||||
|
|
|
|||
|
|
@ -1,237 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/db/fact-log
|
||||
* @description The generation fact log in isolation: wire-format round-trip
|
||||
* (positional msgpack facts, bin16 uuids, body-less tombstones), crc32c
|
||||
* framing with torn-tail detection, open-time truncation to committed truth
|
||||
* (the log can only ever be AHEAD after a crash; open cuts it back), rotation
|
||||
* with a manifest-first flip, exactly-once scans with the frozen telemetry
|
||||
* shape, and the mmap segment handoff excluding the mutable tail.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach } from 'vitest'
|
||||
import { MemoryStorage } from '../../../src/storage/adapters/memoryStorage.js'
|
||||
import {
|
||||
FactLog,
|
||||
FACTS_PREFIX,
|
||||
type CommitFact,
|
||||
type FactLogStorage,
|
||||
storageSupportsFactLog
|
||||
} from '../../../src/db/factLog.js'
|
||||
import { crc32c } from '../../../src/utils/crc32c.js'
|
||||
|
||||
const UUID = (n: number): string =>
|
||||
`00000000-0000-4000-8000-${String(n).padStart(12, '0')}`
|
||||
|
||||
const fact = (generation: number, overrides?: Partial<CommitFact>): CommitFact => ({
|
||||
generation,
|
||||
timestamp: 1_700_000_000_000 + generation,
|
||||
ops: [
|
||||
{
|
||||
kind: 'noun',
|
||||
id: UUID(generation),
|
||||
record: { metadata: { noun: 'document', title: `doc ${generation}` }, vector: { v: [1, 2] } }
|
||||
}
|
||||
],
|
||||
...overrides
|
||||
})
|
||||
|
||||
describe('crc32c known-answer vectors', () => {
|
||||
it('matches the RFC 3720 test vectors', () => {
|
||||
expect(crc32c(new TextEncoder().encode('123456789'))).toBe(0xe3069283)
|
||||
expect(crc32c(new Uint8Array(32))).toBe(0x8a9136aa)
|
||||
})
|
||||
})
|
||||
|
||||
describe('fact log — round-trip, framing, reconcile, rotation, scan', () => {
|
||||
let storage: FactLogStorage
|
||||
let log: FactLog
|
||||
|
||||
beforeEach(async () => {
|
||||
const mem: any = new MemoryStorage()
|
||||
await mem.init()
|
||||
expect(storageSupportsFactLog(mem)).toBe(true)
|
||||
storage = mem
|
||||
log = new FactLog(storage)
|
||||
await log.open(0)
|
||||
})
|
||||
|
||||
it('facts round-trip byte-exactly: ops, tombstones, meta, blobHashes', async () => {
|
||||
await log.append(fact(1))
|
||||
await log.append(
|
||||
fact(2, {
|
||||
ops: [
|
||||
{ kind: 'verb', id: UUID(21), record: { metadata: { verb: 'contains' }, vector: null } },
|
||||
{ kind: 'noun', id: UUID(22), record: null } // TOMBSTONE
|
||||
],
|
||||
meta: { source: 'test' },
|
||||
blobHashes: ['abc123', 'abc123'] // multiset — duplicates preserved
|
||||
})
|
||||
)
|
||||
await log.sync()
|
||||
|
||||
const scan = log.scanFacts()
|
||||
expect(scan.headGeneration).toBe(2)
|
||||
const all: CommitFact[] = []
|
||||
for await (const batch of scan.batches()) all.push(...batch.facts)
|
||||
|
||||
expect(all).toHaveLength(2)
|
||||
expect(all[0].generation).toBe(1)
|
||||
expect(all[0].ops[0].id).toBe(UUID(1))
|
||||
expect(all[0].ops[0].record?.metadata).toEqual({ noun: 'document', title: 'doc 1' })
|
||||
expect(all[1].ops[0].kind).toBe('verb')
|
||||
expect(all[1].ops[1].record).toBeNull() // the tombstone is body-less
|
||||
expect(all[1].meta).toEqual({ source: 'test' })
|
||||
expect(all[1].blobHashes).toEqual(['abc123', 'abc123'])
|
||||
expect(scan.summary().factsYielded).toBe(2)
|
||||
})
|
||||
|
||||
it('appends are monotonic — a replayed/duplicate generation throws', async () => {
|
||||
await log.append(fact(5))
|
||||
await expect(log.append(fact(5))).rejects.toThrow(/non-monotonic/)
|
||||
await expect(log.append(fact(3))).rejects.toThrow(/non-monotonic/)
|
||||
await expect(log.append(fact(6))).resolves.toBeUndefined() // gaps are fine (aborted reservations)
|
||||
})
|
||||
|
||||
it('a torn tail (partial frame) is detected and ignored — intact prefix survives', async () => {
|
||||
await log.append(fact(1))
|
||||
await log.append(fact(2))
|
||||
await log.sync()
|
||||
|
||||
// Simulate a crash mid-append: chop bytes off the tail file.
|
||||
const tailPath = `${FACTS_PREFIX}/seg-${'1'.padStart(20, '0')}.bfl`
|
||||
const bytes = (await storage.readRawBytes(tailPath))!
|
||||
await storage.writeRawBytes(tailPath, bytes.subarray(0, bytes.length - 7))
|
||||
|
||||
const reopened = new FactLog(storage)
|
||||
await reopened.open(2)
|
||||
expect(reopened.headGeneration()).toBe(1) // fact 2's frame was torn → gone
|
||||
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of reopened.scanFacts().batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([1])
|
||||
})
|
||||
|
||||
it('open() truncates facts beyond committed truth (the crash-ahead shape)', async () => {
|
||||
await log.append(fact(1))
|
||||
await log.append(fact(2))
|
||||
await log.append(fact(3))
|
||||
await log.sync()
|
||||
|
||||
// The store's committed generation is 1 — facts 2..3 never committed.
|
||||
const reopened = new FactLog(storage)
|
||||
await reopened.open(1)
|
||||
expect(reopened.headGeneration()).toBe(1)
|
||||
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of reopened.scanFacts().batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([1])
|
||||
|
||||
// And appends continue cleanly from the truncated head.
|
||||
await reopened.append(fact(2))
|
||||
expect(reopened.headGeneration()).toBe(2)
|
||||
})
|
||||
|
||||
it('a cleared store (committed=0) truncates everything', async () => {
|
||||
await log.append(fact(1))
|
||||
await log.append(fact(2))
|
||||
await log.sync()
|
||||
const reopened = new FactLog(storage)
|
||||
await reopened.open(0)
|
||||
expect(reopened.headGeneration()).toBe(0)
|
||||
})
|
||||
|
||||
it('scan honors fromGeneration/toGeneration inclusively and filters kinds', async () => {
|
||||
for (let g = 1; g <= 6; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
|
||||
const scan = log.scanFacts({ fromGeneration: 2, toGeneration: 4 })
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of scan.batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([2, 3, 4])
|
||||
|
||||
const verbsOnly = log.scanFacts({ kinds: ['verb'] })
|
||||
for await (const b of verbsOnly.batches()) {
|
||||
for (const f of b.facts) expect(f.ops.every((op) => op.kind === 'verb')).toBe(true)
|
||||
}
|
||||
})
|
||||
|
||||
it('batch telemetry carries the frozen shape', async () => {
|
||||
for (let g = 1; g <= 5; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
|
||||
const scan = log.scanFacts({ batchSize: 2 })
|
||||
expect(scan.approxFactCount).toBe(5)
|
||||
const batches = []
|
||||
for await (const b of scan.batches()) batches.push(b)
|
||||
expect(batches.length).toBe(3)
|
||||
expect(batches[0]).toMatchObject({ firstGeneration: 1, lastGeneration: 2, factCount: 2 })
|
||||
expect(batches[0].byteSize).toBeGreaterThan(0)
|
||||
expect(typeof batches[0].segmentId).toBe('string')
|
||||
expect(scan.summary()).toEqual({ factsYielded: 5, segmentsRead: 1 })
|
||||
})
|
||||
|
||||
it('survives reopen: head and content come back from disk', async () => {
|
||||
for (let g = 1; g <= 3; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
|
||||
const reopened = new FactLog(storage)
|
||||
await reopened.open(3)
|
||||
expect(reopened.headGeneration()).toBe(3)
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of reopened.scanFacts().batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([1, 2, 3])
|
||||
})
|
||||
|
||||
it('segmentPaths excludes the mutable tail (mmap handoff = sealed only)', async () => {
|
||||
await log.append(fact(1))
|
||||
await log.sync()
|
||||
expect(log.segmentPaths()).toEqual([]) // only a tail exists — nothing sealed
|
||||
})
|
||||
|
||||
describe('scanFacts liveness contract (Stage-2 D1)', () => {
|
||||
it('a wedged store fails LOUDLY within the first-batch bound — never a silent hang', async () => {
|
||||
// Force a sealed segment (tiny rotateBytes) so the scan must READ from
|
||||
// storage, then wedge that read: the exact production shape (a
|
||||
// backlogged brain whose segment read never returned).
|
||||
const mem: any = new MemoryStorage()
|
||||
await mem.init()
|
||||
const wedgeable = new FactLog(mem, { rotateBytes: 1 })
|
||||
await wedgeable.open(0)
|
||||
await wedgeable.append(fact(1))
|
||||
await wedgeable.append(fact(2)) // second append rotates → seg 1 sealed
|
||||
await wedgeable.sync()
|
||||
|
||||
const realRead = mem.readRawBytes.bind(mem)
|
||||
mem.readRawBytes = (p: string) =>
|
||||
p.includes('facts/seg-') ? new Promise(() => {}) : realRead(p) // hangs forever
|
||||
|
||||
const scan = wedgeable.scanFacts({ firstBatchTimeoutMs: 200 })
|
||||
const started = Date.now()
|
||||
await expect(scan.batches().next()).rejects.toThrow(/no first batch within 200ms/)
|
||||
expect(Date.now() - started).toBeLessThan(5_000) // bound held, not a hang
|
||||
})
|
||||
|
||||
it('a healthy scan is unaffected — first batch well inside the bound, all facts delivered', async () => {
|
||||
for (let g = 1; g <= 5; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
const scan = log.scanFacts({ batchSize: 2 })
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of scan.batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([1, 2, 3, 4, 5])
|
||||
expect(scan.summary().factsYielded).toBe(5)
|
||||
})
|
||||
|
||||
it('consumer think-time between pulls never counts against the producer', async () => {
|
||||
for (let g = 1; g <= 4; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
// Bound tighter than the consumer's pause: only the FIRST pull is
|
||||
// raced, so a slow consumer after batch 1 must not trip the deadline.
|
||||
const gen = log.scanFacts({ batchSize: 2, firstBatchTimeoutMs: 150 }).batches()
|
||||
const first = await gen.next()
|
||||
expect(first.done).toBe(false)
|
||||
await new Promise((r) => setTimeout(r, 400)) // dawdle past the bound
|
||||
const second = await gen.next()
|
||||
expect(second.done).toBe(false)
|
||||
expect((await gen.next()).done).toBe(true)
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
@ -1,150 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/db/generation-segments
|
||||
* @description The generation-segment store (Stage-2 D1+D3 file format).
|
||||
* Laws: (1) fold → read round-trips deltas and records byte-faithfully via
|
||||
* sidecar point-reads; (2) the manifest is the ONLY discovery path — reopen
|
||||
* reads one file, never a listing; (3) a lost/corrupt sidecar rebuilds from
|
||||
* its segment loudly, a damaged SEGMENT fails loudly (never silent wrong
|
||||
* data); (4) D3 reclaim drops whole segments only and bumps compactedBelow;
|
||||
* (5) the packed digest is deterministic across reopen; (6) immutability —
|
||||
* fold refuses overlap with sealed ranges.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach } from 'vitest'
|
||||
import { MemoryStorage } from '../../../src/storage/adapters/memoryStorage.js'
|
||||
import {
|
||||
GenerationSegmentStore,
|
||||
SEGMENTS_PREFIX,
|
||||
type FoldGeneration
|
||||
} from '../../../src/db/generationSegments.js'
|
||||
|
||||
const UUID = (n: number): string => `00000000-0000-4000-8000-${String(n).padStart(12, '0')}`
|
||||
|
||||
const gen = (g: number, recordCount = 2): FoldGeneration => ({
|
||||
generation: g,
|
||||
timestamp: 1_700_000_000_000 + g,
|
||||
delta: { generation: g, nouns: [UUID(g)], verbs: [], bytes: 123 + g },
|
||||
records: Array.from({ length: recordCount }, (_, i) => ({
|
||||
kind: (i % 2 === 0 ? 'noun' : 'verb') as 'noun' | 'verb',
|
||||
id: UUID(g * 100 + i),
|
||||
record: { metadata: { noun: 'document', v: g }, vector: { v: [g, i] } }
|
||||
}))
|
||||
})
|
||||
|
||||
describe('db/GenerationSegmentStore — the D1+D3 packed tier', () => {
|
||||
let storage: MemoryStorage
|
||||
let store: GenerationSegmentStore
|
||||
|
||||
beforeEach(async () => {
|
||||
storage = new MemoryStorage()
|
||||
await storage.init()
|
||||
store = new GenerationSegmentStore(storage as any)
|
||||
await store.open()
|
||||
})
|
||||
|
||||
it('fold → read round-trips deltas and records via sidecar point-reads', async () => {
|
||||
const meta = await store.fold([gen(1), gen(2), gen(3)])
|
||||
expect(meta).toMatchObject({ firstGeneration: 1, lastGeneration: 3, frames: 3 })
|
||||
expect(meta.checksum).toBeGreaterThan(0)
|
||||
|
||||
expect(store.hasGeneration(2)).toBe(true)
|
||||
expect(store.hasGeneration(4)).toBe(false)
|
||||
|
||||
const d2 = await store.readDelta(2)
|
||||
expect(d2?.delta).toEqual({ generation: 2, nouns: [UUID(2)], verbs: [], bytes: 125 })
|
||||
expect(d2?.timestamp).toBe(1_700_000_000_002)
|
||||
|
||||
const records = await store.readRecords(3)
|
||||
expect(records).toHaveLength(2)
|
||||
expect(records![0]).toEqual({
|
||||
kind: 'noun',
|
||||
id: UUID(300),
|
||||
record: { metadata: { noun: 'document', v: 3 }, vector: { v: [3, 0] } }
|
||||
})
|
||||
// Point read by id, both kinds.
|
||||
expect(await store.readRecord(3, 'verb', UUID(301))).toEqual({
|
||||
metadata: { noun: 'document', v: 3 },
|
||||
vector: { v: [3, 1] }
|
||||
})
|
||||
expect(await store.readRecord(3, 'noun', UUID(999))).toBeNull()
|
||||
})
|
||||
|
||||
it('reopen discovers everything from the manifest alone — no listing', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await store.fold([gen(3), gen(4)])
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(reopened.segments()).toHaveLength(2)
|
||||
expect(reopened.hasGeneration(4)).toBe(true)
|
||||
expect((await reopened.readDelta(1))?.timestamp).toBe(1_700_000_000_001)
|
||||
})
|
||||
|
||||
it('a lost sidecar rebuilds from its segment; a damaged segment fails LOUDLY', async () => {
|
||||
const meta = await store.fold([gen(1), gen(2)])
|
||||
const idxPath = `${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.idx`
|
||||
await storage.deleteRawObject(idxPath)
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
// Rebuild path: still serves correct data.
|
||||
expect((await reopened.readRecords(2))!).toHaveLength(2)
|
||||
|
||||
// Now damage the SEGMENT itself: flip a payload byte → CRC mismatch, loud.
|
||||
const segPath = `${SEGMENTS_PREFIX}/${meta.file}`
|
||||
const bytes = (await storage.readRawBytes(segPath))!
|
||||
bytes[bytes.length - 3] ^= 0xff
|
||||
await storage.writeRawBytes(segPath, bytes)
|
||||
const damaged = new GenerationSegmentStore(storage as any)
|
||||
await damaged.open()
|
||||
;(damaged as any).sidecars.clear()
|
||||
await storage.deleteRawObject(idxPath) // force the sequential rebuild over damaged bytes
|
||||
await expect(damaged.readRecords(2)).rejects.toThrow(/CRC mismatch|damaged/)
|
||||
})
|
||||
|
||||
it('D3 reclaim drops whole segments only and bumps compactedBelow', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await store.fold([gen(3), gen(4)])
|
||||
await store.fold([gen(5), gen(6)])
|
||||
|
||||
// Horizon mid-segment-2 (below 4): only segment 1 is FULLY below → drops.
|
||||
const r1 = await store.dropSegmentsBelow(4)
|
||||
expect(r1).toEqual({ dropped: 1, compactedBelow: 3 })
|
||||
expect(store.hasGeneration(1)).toBe(false)
|
||||
expect(store.hasGeneration(3)).toBe(true) // partial segment survives whole
|
||||
|
||||
// Bytes actually gone.
|
||||
expect(await storage.readRawBytes(`${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.bgs`)).toBeNull()
|
||||
|
||||
// Horizon past everything: the rest drop; compactedBelow is durable.
|
||||
const r2 = await store.dropSegmentsBelow(7)
|
||||
expect(r2.dropped).toBe(2)
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(reopened.compactedBelow()).toBe(7)
|
||||
expect(reopened.segments()).toHaveLength(0)
|
||||
})
|
||||
|
||||
it('the packed digest is deterministic across reopen and changes with history', async () => {
|
||||
await store.fold([gen(1), gen(2), gen(3)])
|
||||
const atSeal = await store.digestThroughPacked(3)
|
||||
const midSegment = await store.digestThroughPacked(2)
|
||||
expect(atSeal).not.toBeNull()
|
||||
expect(midSegment).not.toBeNull()
|
||||
expect(midSegment).not.toBe(atSeal)
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(await reopened.digestThroughPacked(3)).toBe(atSeal)
|
||||
expect(await reopened.digestThroughPacked(2)).toBe(midSegment)
|
||||
|
||||
await reopened.fold([gen(4)])
|
||||
expect(await reopened.digestThroughPacked(4)).not.toBe(atSeal)
|
||||
})
|
||||
|
||||
it('sealed segments are immutable — fold refuses overlap, requires ascending input', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await expect(store.fold([gen(2), gen(3)])).rejects.toThrow(/overlaps the packed tier/)
|
||||
await expect(store.fold([gen(4), gen(4)])).rejects.toThrow(/strictly ascending/)
|
||||
await expect(store.fold([])).rejects.toThrow(/at least one generation/)
|
||||
})
|
||||
})
|
||||
|
|
@ -249,11 +249,7 @@ describe('db/GenerationStore', () => {
|
|||
store.release(pinned)
|
||||
const result = await store.compact()
|
||||
expect(result.removedGenerations).toBeGreaterThan(0)
|
||||
// History record-sets only — the fact log (at `_generations/facts/`) is
|
||||
// deliberately NOT reclaimed by history compaction.
|
||||
const remaining = (await storage.listRawObjects(GENERATIONS_PREFIX)).filter(
|
||||
(p: string) => !p.startsWith(`${GENERATIONS_PREFIX}/facts/`)
|
||||
)
|
||||
const remaining = await storage.listRawObjects(GENERATIONS_PREFIX)
|
||||
expect(remaining).toEqual([])
|
||||
})
|
||||
|
||||
|
|
@ -488,95 +484,5 @@ describe('db/GenerationStore', () => {
|
|||
expect(result.removedGenerations).toBe(2)
|
||||
store.release(2)
|
||||
})
|
||||
|
||||
it('timeBudgetMs bounds a pass; the next pass resumes the same prefix', async () => {
|
||||
await manyGens(4)
|
||||
// A spent budget (0ms) stops before reclaiming anything — an early stop
|
||||
// is a consistent prefix, never a partial generation.
|
||||
const bounded = await store.compact({ timeBudgetMs: 0 })
|
||||
expect(bounded.removedGenerations).toBe(0)
|
||||
expect(bounded.horizon).toBe(0)
|
||||
// The next (unbounded) pass picks up exactly where the bounded one
|
||||
// stopped and completes the same work.
|
||||
const resumed = await store.compact()
|
||||
expect(resumed.removedGenerations).toBe(4)
|
||||
expect(resumed.horizon).toBe(4)
|
||||
})
|
||||
})
|
||||
|
||||
// ==========================================================================
|
||||
describe('history-bytes running total (the O(1) retention check)', () => {
|
||||
/** A fresh walk with the cache dropped — ground truth for the invariant. */
|
||||
async function groundTruthBytes(): Promise<number> {
|
||||
;(store as any).historyBytesTotal = null
|
||||
return store.historyBytes()
|
||||
}
|
||||
|
||||
it('is seeded once, then maintained through commits WITHOUT re-walks', async () => {
|
||||
await commitWrite(ID_A, 1)
|
||||
await commitWrite(ID_A, 2)
|
||||
const seeded = await store.historyBytes()
|
||||
expect(seeded).toBe(await groundTruthBytes())
|
||||
|
||||
// From here every read must come from the running total, not a walk:
|
||||
// getDelta re-reads are the walk's cost — commits must not trigger any.
|
||||
const getDeltaSpy = vi.spyOn(store as any, 'getDelta')
|
||||
await commitWrite(ID_B, 1)
|
||||
const afterCommit = await store.historyBytes()
|
||||
expect(getDeltaSpy).not.toHaveBeenCalled()
|
||||
getDeltaSpy.mockRestore()
|
||||
expect(afterCommit).toBe(await groundTruthBytes())
|
||||
})
|
||||
|
||||
it('stays exact through single-op group commits and compaction', async () => {
|
||||
await commitWrite(ID_A, 1)
|
||||
await store.historyBytes() // seed
|
||||
// Single-op path: buffered generations flushed as one group commit.
|
||||
await store.commitSingleOp({
|
||||
touched: { nouns: [ID_B] },
|
||||
execute: async () => {
|
||||
await storage.saveNounMetadata(ID_B, metadataFixture(1))
|
||||
}
|
||||
})
|
||||
await store.flushPendingSingleOps()
|
||||
expect(await store.historyBytes()).toBe(await groundTruthBytes())
|
||||
|
||||
await store.historyBytes() // re-seed after ground-truth reset
|
||||
await store.compact({ maxGenerations: 1 })
|
||||
expect(await store.historyBytes()).toBe(await groundTruthBytes())
|
||||
})
|
||||
|
||||
it('historyStats reports counts, bytes, range, and horizon read-only', async () => {
|
||||
await commitWrite(ID_A, 1)
|
||||
await commitWrite(ID_B, 1)
|
||||
const stats = await store.historyStats()
|
||||
expect(stats.generations).toBe(2)
|
||||
expect(stats.bytes).toBe(await store.historyBytes())
|
||||
expect(stats.oldestGeneration).toBe(1)
|
||||
expect(stats.newestGeneration).toBe(2)
|
||||
expect(stats.oldestTimestamp).toBeLessThanOrEqual(stats.newestTimestamp!)
|
||||
expect(stats.horizon).toBe(0)
|
||||
// Read-only: nothing was reclaimed by asking.
|
||||
expect(store.committedGeneration()).toBe(2)
|
||||
|
||||
await store.compact({ maxGenerations: 1 })
|
||||
const after = await store.historyStats()
|
||||
expect(after.generations).toBe(1)
|
||||
expect(after.oldestGeneration).toBe(2)
|
||||
expect(after.horizon).toBe(1)
|
||||
})
|
||||
|
||||
it('empty history reports null range and zero bytes', async () => {
|
||||
const stats = await store.historyStats()
|
||||
expect(stats).toMatchObject({
|
||||
generations: 0,
|
||||
bytes: 0,
|
||||
oldestGeneration: null,
|
||||
newestGeneration: null,
|
||||
oldestTimestamp: null,
|
||||
newestTimestamp: null,
|
||||
horizon: 0
|
||||
})
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -95,15 +95,9 @@ describe('verb cursor pagination (graph-perf #2)', () => {
|
|||
expect(new Set(cursorSeen)).toEqual(new Set(offsetSeen))
|
||||
})
|
||||
|
||||
it('a foreign/malformed cursor FAILS LOUDLY — never a silent restart from page 1', async () => {
|
||||
// The old behavior (decode-null → silent offset-0 fallback) re-served page 1
|
||||
// forever to any while(hasMore) walker: an unbounded CPU loop with no log
|
||||
// line. An undecodable resume token now refuses the walk instead.
|
||||
await expect(
|
||||
storage.getVerbs({ pagination: { limit: 5, cursor: 'not-a-cv1-token' } })
|
||||
).rejects.toThrow('invalid pagination cursor')
|
||||
await expect(
|
||||
storage.getNouns({ pagination: { limit: 5, cursor: 'not-a-cv1-token' } })
|
||||
).rejects.toThrow('invalid pagination cursor')
|
||||
it('a foreign/malformed cursor falls back gracefully (no throw, starts from the beginning)', async () => {
|
||||
const page = await storage.getVerbs({ pagination: { limit: 5, cursor: 'not-a-cv1-token' } })
|
||||
expect(page.items.length).toBe(5)
|
||||
expect(page.hasMore).toBe(true)
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -1,76 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/utils/osLimits
|
||||
* @description OS-limit detection for pool-scale use. Laws:
|
||||
* (1) the /proc/self/limits parser reads soft/hard NOFILE exactly, including
|
||||
* 'unlimited'; (2) assessment warns ONLY below the pool floors and NEVER
|
||||
* on an unreadable (null) limit — no measurement, no claim; (3) the full
|
||||
* check composes both sources and survives unreadable /proc silently.
|
||||
*/
|
||||
import { describe, it, expect } from 'vitest'
|
||||
import {
|
||||
parseProcLimits,
|
||||
assessOsLimits,
|
||||
checkOsLimits,
|
||||
NOFILE_POOL_FLOOR,
|
||||
MAX_MAP_COUNT_POOL_FLOOR
|
||||
} from '../../../src/utils/osLimits.js'
|
||||
|
||||
const SAMPLE_LIMITS = [
|
||||
'Limit Soft Limit Hard Limit Units',
|
||||
'Max cpu time unlimited unlimited seconds',
|
||||
'Max open files 1024 1048576 files',
|
||||
'Max locked memory 8388608 8388608 bytes'
|
||||
].join('\n')
|
||||
|
||||
describe('osLimits — detect + warn at pool scale', () => {
|
||||
it('parses soft/hard NOFILE from /proc/self/limits, including unlimited', () => {
|
||||
expect(parseProcLimits(SAMPLE_LIMITS)).toEqual({ soft: 1024, hard: 1048576 })
|
||||
expect(
|
||||
parseProcLimits('Max open files unlimited unlimited files')
|
||||
).toEqual({ soft: Infinity, hard: Infinity })
|
||||
expect(parseProcLimits('no such row here')).toEqual({ soft: null, hard: null })
|
||||
})
|
||||
|
||||
it('warns below the floors, stays quiet at or above them', () => {
|
||||
const low = assessOsLimits({ nofileSoft: 1024, nofileHard: 1048576, maxMapCount: 65530 })
|
||||
expect(low).toHaveLength(2)
|
||||
expect(low[0]).toContain('RLIMIT_NOFILE soft limit is 1024')
|
||||
expect(low[0]).toContain(`ulimit -n ${NOFILE_POOL_FLOOR}`)
|
||||
expect(low[0]).toContain('raise the soft limit only') // hard already allows it
|
||||
expect(low[1]).toContain('vm.max_map_count is 65530')
|
||||
expect(low[1]).toContain(`vm.max_map_count=${MAX_MAP_COUNT_POOL_FLOOR}`)
|
||||
|
||||
expect(
|
||||
assessOsLimits({
|
||||
nofileSoft: NOFILE_POOL_FLOOR,
|
||||
nofileHard: Infinity,
|
||||
maxMapCount: MAX_MAP_COUNT_POOL_FLOOR
|
||||
})
|
||||
).toEqual([])
|
||||
})
|
||||
|
||||
it('an unreadable limit makes NO claim — nulls never warn', () => {
|
||||
expect(assessOsLimits({ nofileSoft: null, nofileHard: null, maxMapCount: null })).toEqual([])
|
||||
})
|
||||
|
||||
it('checkOsLimits composes both sources and survives unreadable /proc silently', async () => {
|
||||
const report = await checkOsLimits(async (p) => {
|
||||
if (p === '/proc/self/limits') return SAMPLE_LIMITS
|
||||
if (p === '/proc/sys/vm/max_map_count') return '65530\n'
|
||||
throw new Error('unexpected path')
|
||||
})
|
||||
expect(report.nofileSoft).toBe(1024)
|
||||
expect(report.maxMapCount).toBe(65530)
|
||||
expect(report.warnings).toHaveLength(2)
|
||||
|
||||
const offLinux = await checkOsLimits(async () => {
|
||||
throw Object.assign(new Error('ENOENT'), { code: 'ENOENT' })
|
||||
})
|
||||
expect(offLinux).toEqual({
|
||||
nofileSoft: null,
|
||||
nofileHard: null,
|
||||
maxMapCount: null,
|
||||
warnings: []
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
@ -270,24 +270,27 @@ describe('Zero-Config Parameter Validation', () => {
|
|||
expect(config.availableMemory).toBeGreaterThan(0)
|
||||
})
|
||||
|
||||
it('never mutates the cap from query timing (telemetry only)', () => {
|
||||
const initialLimit = getValidationConfig().maxLimit
|
||||
|
||||
// Fast queries with large results: no silent growth.
|
||||
it('should adapt limits based on query performance', () => {
|
||||
const initialConfig = getValidationConfig()
|
||||
const initialLimit = initialConfig.maxLimit
|
||||
|
||||
// Simulate fast queries with large results
|
||||
for (let i = 0; i < 10; i++) {
|
||||
recordQueryPerformance(50, initialLimit * 0.9)
|
||||
}
|
||||
expect(getValidationConfig().maxLimit).toBe(initialLimit)
|
||||
|
||||
// A burst of catastrophically slow queries must not strangle the cap.
|
||||
// The removed "learning" ratchet shrank it 20% per recorded query down
|
||||
// to a floor of 1000 — below the documented MIN_AUTO_QUERY_LIMIT — and
|
||||
// the error message blamed "available free memory" (a production
|
||||
// incident: every find({ limit: 5000 }) failed on an idle 23GB-free box).
|
||||
for (let i = 0; i < 50; i++) {
|
||||
recordQueryPerformance(90_000, 100)
|
||||
|
||||
const updatedConfig = getValidationConfig()
|
||||
// Limit might increase if performance is good
|
||||
expect(updatedConfig.maxLimit).toBeGreaterThanOrEqual(initialLimit)
|
||||
|
||||
// Simulate slow queries
|
||||
for (let i = 0; i < 10; i++) {
|
||||
recordQueryPerformance(2000, 100)
|
||||
}
|
||||
expect(getValidationConfig().maxLimit).toBe(initialLimit)
|
||||
|
||||
const finalConfig = getValidationConfig()
|
||||
// Limit should decrease if performance is poor
|
||||
expect(finalConfig.maxLimit).toBeLessThanOrEqual(updatedConfig.maxLimit)
|
||||
})
|
||||
})
|
||||
})
|
||||
Loading…
Add table
Add a link
Reference in a new issue