1868 lines
93 KiB
Markdown
1868 lines
93 KiB
Markdown
# @soulcraft/brainy — Release Notes for Consumers
|
||
|
||
This file is the **quick reference for downstream sessions** tracking Brainy changes.
|
||
Full auto-generated changelog: `CHANGELOG.md` · Releases: https://github.com/soulcraftlabs/brainy/releases
|
||
|
||
**How to use:** Brainy is the underlying data engine for downstream applications. Read this when:
|
||
- Upgrading `@soulcraft/brainy` in your application
|
||
- Debugging data, query, or storage behaviour
|
||
- A new Brainy feature is available that you want to adopt
|
||
|
||
---
|
||
|
||
## v8.0.0 — release candidate (`8.0.0-rc.8` published 2026-06-30 on npm tag `rc`)
|
||
|
||
> **rc.8 additions (2026-06-30) — additive, no breaking change:**
|
||
> - **No-freeze (online) whole-brain auto-upgrade.** When a derived-index format changes, a large
|
||
> brain upgrades **without blocking** — the native provider rebuilds the new indexes in the
|
||
> background and serves correct reads from the canonical records meanwhile, then atomically swaps;
|
||
> no minutes-long freeze on the first open/query. (Small brains still auto-rebuild inline on open.)
|
||
> New surface: an optional provider `isMigrating()` deference signal, a public `brain.stampBrainFormat()`
|
||
> the provider calls once its swap verifies, and a `@soulcraft/brainy/brain-format` export so the
|
||
> provider shares the `indexEpoch` constant. No effect without a native provider.
|
||
|
||
> **rc.7 additions (2026-06-30) — additive, no breaking change:**
|
||
> - **8.0 cold-graph self-heal.** `find({ connected })` / `neighbors()` / `related()` never
|
||
> silently return `[]` on a cold open where the graph adjacency loaded its membership but not
|
||
> its source→target edges — it self-heals (rebuild from storage) or throws a loud
|
||
> `GraphIndexNotReadyError`, never empty-for-persisted-data. (The 8.0 equivalent of the 7.33.4
|
||
> fix 7.x consumers already have; gates on the native provider's honest `isReady()` edge-readiness
|
||
> signal.)
|
||
> - **Billion-scale RAM — nothing O(N)-resident on the write/time-travel path.** Eliminated the
|
||
> per-id storage caches (per-type counts now sourced from the record), made the
|
||
> committed-generation ledger an interval set, and bounded the per-id time-travel history chains
|
||
> (hot-window + LRU, with lock-light reconstruction so historical reads never stall writers).
|
||
> Resident memory is now independent of entity count.
|
||
> - **Whole-brain version handshake.** A `_system/brain-format.json` marker + `brain.formatInfo()`
|
||
> + a shared, lockstep `indexEpoch`: a future derived-index format change auto-rebuilds the
|
||
> indexes from the canonical records on open (non-destructively) — the foundation for
|
||
> whole-brain auto-upgrade with the native provider. No effect on an unchanged brain.
|
||
|
||
> **rc.6 additions (2026-06-29) — additive, no breaking change:**
|
||
> - **Open-core perf**: HNSW delete is now O(in-degree) (was O(N²) for bulk delete) via a
|
||
> reverse-adjacency index; the negation/absence operators (`ne`/`exists:false`/`missing:true`)
|
||
> are served as a roaring-bitmap difference instead of materializing the whole corpus.
|
||
> - **Native provider contract** (cor lockstep, optional/feature-detected — no effect without a
|
||
> native provider): a cold-open `probeConsistency()` self-heal hook, and a `getIdsForFilter`
|
||
> page bound so the native index can early-stop the unsorted `find({ type, where, limit })` path.
|
||
> - **Test hygiene**: re-homed previously-unrun test suites into CI + a guard so a test file can
|
||
> never silently fall outside every config again. No source/API change from these.
|
||
|
||
|
||
**Affected products:** every consumer — this is a major release with removed
|
||
surfaces, hard renames (no aliases, no deprecation period), and one flipped
|
||
default. This entry **is** the migration guide: the find/replace pairs, the
|
||
sed snippet, and the upgrade checklist below are the complete story — there is
|
||
no separate migration doc. **`8.0.0-rc.6` is live** — install with
|
||
`npm i @soulcraft/brainy@rc` to try it; the `latest` tag stays on 7.x until GA.
|
||
|
||
> **rc.3–rc.5 additions (2026-06-24):**
|
||
> - **Showcase-quality GA hardening (rc.5).** A full readiness audit closed a cold-init
|
||
> version-coupling bug (a matched native provider could be rejected on first open), turned
|
||
> silent storage-read failures into named `BrainyError`s (no more empty-result-on-error),
|
||
> stopped the JS graph LSM from orphaning compacted SSTables, corrected the public docs to
|
||
> the real API, and documented the two headline methods. Plus a dead/deprecated-code sweep.
|
||
> - **BREAKING — 4 deprecated query operators removed.** `is`→`eq`, `isNot`→`ne`,
|
||
> `greaterEqual`→`gte`, `lessEqual`→`lte`. The canonical operators and their clean long-form
|
||
> aliases (`equals`/`notEquals`/`greaterThan`/`greaterThanOrEqual`/`lessThan`/`lessThanOrEqual`)
|
||
> are unchanged — only the four redundant spellings are gone. Find/replace if you used them.
|
||
> - **`find()` search mode** is now the single `searchMode: SearchMode` option (the redundant,
|
||
> silently-ignored `mode` alias and the unwired `explain` flag were removed from `FindParams`).
|
||
> - The phantom index-integrity guard (find() re-validates every result against its predicate)
|
||
> shipped in rc.3; rc.4 added the `accel.isInitialized` gate + #35 at-gen candidate vectors.
|
||
|
||
> **rc.1 additions (this entry predates them — full writeup lands at GA):** entity-id
|
||
> normalization — `brain.newId()` (UUID v7) + v7 default ids, and non-UUID string ids are
|
||
> transparently normalized to a stable UUID v5 (original key preserved under `_originalId`);
|
||
> `brain.neural()` clustering and `Db.search()` removed (use `find({ vector })` /
|
||
> `find()` + aggregation `GROUP BY`); storage config collapsed to one `path` key (the
|
||
> pre-8.0 aliases now throw); and `queryAggregate` no longer hangs/staled min-max after a delete.
|
||
> **Reserved fields in a `metadata` bag now throw by default** — passing a Brainy-reserved key
|
||
> (`confidence`, `weight`, `subtype`, `visibility`, `service`, `createdBy`, `noun`/`verb`, `data`,
|
||
> `createdAt`, `updatedAt`, `_rev`) inside `metadata` on `add()`/`update()`/`relate()`/`updateRelation()`
|
||
> (untyped/JS callers — TypeScript already blocks it) is rejected with an Error naming the correct
|
||
> param, instead of the old silent remap-or-drop. Pass reserved values as their dedicated top-level
|
||
> params. Opt back into the legacy behavior with `new Brainy({ reservedFieldPolicy: 'warn' | 'remap' })`.
|
||
|
||
> **rc.2 additions (2026-06-21):**
|
||
> - **Graph performance + the `brain.graph` namespace.** New `brain.graph.subgraph(seeds, opts)`
|
||
> (bounded multi-hop neighborhood → `{ nodes, edges, truncated }`) and `brain.graph.export(opts)`
|
||
> (stream the whole graph in one O(N+E) pass — the right primitive for visualizing all data
|
||
> instead of paging per node). `related({ node })` returns every edge incident to an entity in
|
||
> **both directions** in one O(degree) call. Under the hood: `related({ from/to })` now stays
|
||
> O(degree) under the default visibility filter (was a full scan), and the verb **and** noun
|
||
> pagination walks are cursor-based, so a full edge/node walk is **O(N), not O(N²)** (a consumer
|
||
> measured a 19k-edge whole-graph read drop from ~27s toward a single scan). These transparently
|
||
> use a native graph engine when present (the `@soulcraft/cor` 3.0 acceleration layer) and fall
|
||
> back to pure-TS adjacency otherwise.
|
||
> - **Additive ergonomics:** `add({ upsert: true })` (create-or-update in one call — merges into an
|
||
> existing id instead of overwriting), `find({ includeVectors: true })`, and `removeMany`'s batch
|
||
> size is now storage-adaptive (was a hardcoded 10).
|
||
> - **Correctness fixes (apply to all consumers, not just graph users):** `related()` results now
|
||
> carry `visibility`; whole-graph/`getNouns` streaming no longer leaks `system`/`internal` entities;
|
||
> and small-page cursor walks over nouns no longer loop. No API change — just correct behavior.
|
||
|
||
> **rc.3 additions (2026-06-23):**
|
||
> - **Graph analytics on the `brain.graph` namespace.** Three intent-level reads that answer
|
||
> whole-graph questions in one call:
|
||
> - **`brain.graph.rank(opts?)`** → `{ id, score }[]` descending — "which entities matter most"
|
||
> (importance / centrality). `topK` to cap.
|
||
> - **`brain.graph.communities(opts?)`** → `{ groups: string[][], count }` — "which things group
|
||
> together". Weakly-connected components by default; `{ directed: true }` returns
|
||
> strongly-connected components.
|
||
> - **`brain.graph.path(from, to, opts?)`** → `{ nodes, relationships, cost } | null` — the best
|
||
> route between two entities. Fewest hops by default; `{ by: 'weight' }` minimizes summed edge
|
||
> weight (the 0–1 connection strength); `direction`, `type`, and `maxDepth` filters apply.
|
||
> These are **intent contracts, not algorithm contracts** — the question is the promise, the
|
||
> algorithm is the engine's choice. They transparently use the native `@soulcraft/cor` 3.0 graph
|
||
> engine when present, and fall back to pure-TS kernels (PageRank, connected components / Tarjan
|
||
> SCC, BFS / Dijkstra) otherwise. Both paths return identical shapes and respect the default
|
||
> visibility filter (opt in with `includeInternal` / `includeSystem`).
|
||
> - **Filtered vector search keeps its recall (`allowedIds` pushdown).** A `find({ query, where })`
|
||
> that combines semantic search with a metadata filter now restricts the vector walk to the
|
||
> matching candidates *inside* the search (walk-all, collect-allowed) instead of filtering the
|
||
> top-k afterward — so a query whose nearest vectors are all filtered out still returns the best
|
||
> matches that DO pass the filter, rather than coming back empty. No API change. With the native
|
||
> `@soulcraft/cor` 3.0 stack the matched universe is forwarded as an opaque roaring buffer with
|
||
> zero id materialization in TypeScript; the pure-JS path restricts the beam walk with a string set.
|
||
> - **Historical (`asOf`) semantic search can skip the rebuild.** A filtered semantic query at a
|
||
> past generation — `db.asOf(g).find({ query, where })` — no longer always rebuilds an ephemeral
|
||
> in-memory vector index over every at-`g` vector (O(n@G)). When a native versioned vector engine
|
||
> is registered and can serve the pinned generation, Brainy resolves the at-`g` filter universe from
|
||
> its record layer (no rebuild) and routes the vector leg to the engine with the matched ids + the
|
||
> generation; without one it falls back to the materialization (unchanged). The `VectorIndexProvider.search`
|
||
> options gained an optional `generation?: bigint` ("omitted = now") to carry this — additive, the
|
||
> built-in index ignores it. No public API change.
|
||
> - **`find({ where, orderBy })` bounds the sort to the page.** A broad filter + `orderBy`
|
||
> returning one page no longer materializes every matching sorted id (it produced the full sorted
|
||
> match set, then sliced) — the page bound (`offset + limit`) is threaded into the column store's
|
||
> top-K sort, so returning 20 rows from a million matches stays a bounded-K heap. No API change;
|
||
> ordering and pagination are unchanged.
|
||
> - **`brain.graph.subgraph()` accepts a query (query→expand).** The seed selector now takes not
|
||
> just entity id(s) but a `find()` result or a `FindParams` query — `brain.graph.subgraph({ where:
|
||
> { team: 'platform' } }, { depth: 1 })` runs the query and expands the neighborhood of every
|
||
> match in one call. With the native engine a metadata-only query's matched universe is handed to
|
||
> the traversal as an opaque set with no id materialization in TypeScript (the query→expand
|
||
> fusion); the pure-JS path materializes the matched ids and expands from them. Existing
|
||
> id-seeded calls are unchanged.
|
||
|
||
### Headline: Database as a Value
|
||
|
||
8.0 replaces Brainy's two overlapping version-control subsystems (copy-on-write
|
||
branching and per-entity versioning, ~5,100 LOC combined) with **one
|
||
mechanism**: generational MVCC over immutable, generation-stamped records,
|
||
exposed through a Datomic-style immutable database value — the **`Db`**.
|
||
|
||
```ts
|
||
const db = brain.now() // pin the current state — O(1), no I/O
|
||
|
||
await brain.transact([
|
||
{ op: 'update', id: invoiceId, metadata: { status: 'paid' } }
|
||
], { meta: { author: 'billing-service', reason: 'PO-7741' } })
|
||
|
||
await db.get(invoiceId) // still 'pending' — pinned, forever
|
||
await brain.get(invoiceId) // 'paid' — live
|
||
await db.release() // unpin when done
|
||
```
|
||
|
||
What you get:
|
||
|
||
- **`brain.now()`** — pins the current generation as an immutable `Db` view.
|
||
True snapshot isolation: the view reads exactly its pinned state no matter
|
||
what commits afterwards, including deletes. Readers never block writers and
|
||
writers never block readers.
|
||
- **`brain.transact(ops, { meta, ifAtGeneration })`** — atomic multi-write
|
||
batches (`add` / `update` / `remove` / `relate` / `unrelate`) committed as
|
||
exactly one generation. Either every operation applies or none do.
|
||
`ifAtGeneration` is whole-store compare-and-swap (`GenerationConflictError`
|
||
on conflict) — the big sibling of the per-entity `ifRev` CAS from 7.31.
|
||
`meta` is reified Datomic-style into an append-only transaction log,
|
||
readable via `brain.transactionLog()`.
|
||
- **`brain.asOf(generation | Date | snapshotPath, { exclusive? })`** — time
|
||
travel with the **full query surface**: `get()`, `find()` in every mode,
|
||
semantic search, graph traversal, cursors, aggregation, all at the pinned past
|
||
state. `exclusive: true` pins the generation immediately before the target
|
||
(strict-before).
|
||
- **`db.with(ops)`** — speculative writes applied in memory on top of a view.
|
||
Nothing touches disk, the generation counter, or the indexes. What-if
|
||
analysis, then `transact()` the same ops for real.
|
||
- **`db.persist(path)`** — instant self-contained snapshots. On filesystem
|
||
storage they are built from hard links (no entity data is copied; later
|
||
writes to the source can never alter the snapshot because rewrites swap
|
||
inodes). `brain.restore(path, { confirm: true })` replaces the whole store
|
||
from one; `Brainy.load(path)` opens one read-only with the full query
|
||
surface.
|
||
- **Range verbs over history** — answer "what happened BETWEEN two points":
|
||
- **`db.since(Db | generation | Date)`** — the entity and relationship ids
|
||
committed transactions touched after an **exclusive** lower bound (now
|
||
accepts a generation or `Date`, not just a prior `Db`).
|
||
- **`brain.diff(a, b)`** — the touched ids CLASSIFIED as `{ added, removed,
|
||
modified }` (split by nouns/verbs) by resolving each at both endpoints; a
|
||
touched-but-reverted id lands in none of the buckets. Endpoints are a
|
||
generation, `Date`, or `Db`, in either order.
|
||
- **`brain.history(id, { from?, to? })`** — every distinct version of ONE
|
||
entity/relationship over a range, oldest first (`value: null` marks a
|
||
removal); each version equals `asOf(version.generation).get(id)`.
|
||
- **`brain.transactionLog({ from?, to?, limit? })`** — the commit log over an
|
||
**inclusive** generation/`Date` window (contrast `since`'s exclusive lower
|
||
bound); `limit` applies last.
|
||
- **Every write is versioned (Model-B).** History granularity is **per-write**:
|
||
EVERY write — `transact()` AND a single-operation `add`/`update`/`remove`/
|
||
`relate` — is its own immutable generation. A `now()` pin always freezes
|
||
against later writes, and `asOf`/`since`/`diff`/`history`/`transactionLog`
|
||
reflect single-ops exactly like transacts. `transact()` groups several
|
||
operations into ONE atomic generation. Single-op history durability is **async
|
||
group-commit** (the live write is acknowledged immediately; its before-image
|
||
is batched to disk in one fsync) — a hard crash can lose only the last
|
||
un-flushed window's *history*, never live data, and a crash mid-flush is
|
||
recovered by drop-without-restore. A freshly-initialized brain is
|
||
`generation() === 0` with an empty `transactionLog()` (init-time
|
||
infrastructure is the un-versioned baseline); the first user write is gen 1.
|
||
- **`new Brainy({ retention })`** — the retention knob governs auto-compaction on
|
||
`flush()`/`close()`: **unset → ADAPTIVE** (disk/RAM-pressure byte budget,
|
||
zero-config; a coordinator can drive it via `brain.setRetentionBudget(bytes)`)
|
||
· **`'all'`** → unbounded · **`{ maxGenerations?, maxAge?, maxBytes? }`** →
|
||
explicit caps (reclaim oldest-unpinned while ANY cap is exceeded). Live pins
|
||
are ALWAYS exempt.
|
||
- **`brain.compactHistory({ maxGenerations?, maxAge?, maxBytes? })`** — manual
|
||
reclaim on the same caps. Compaction never breaks a pinned read. `diff`/`since`
|
||
throw `GenerationCompactedError` below the horizon; `history` truncates to it.
|
||
|
||
The precise guarantees are documented in:
|
||
|
||
- [docs/concepts/consistency-model.md](docs/concepts/consistency-model.md) —
|
||
the guarantees, each proven by a dedicated test in
|
||
`tests/integration/db-mvcc.test.ts`
|
||
- [docs/guides/snapshots-and-time-travel.md](docs/guides/snapshots-and-time-travel.md)
|
||
— the recipes: backup, restore, time-travel debugging, what-if, audit trails
|
||
- [docs/ADR-001-generational-mvcc.md](docs/ADR-001-generational-mvcc.md) — the
|
||
design record: persisted layout, commit protocol, crash recovery, proof table
|
||
|
||
### Portable export & import (PortableGraph v1)
|
||
|
||
A portable, versioned graph format that serializes part or all of a brain to a
|
||
single JSON document and restores it — the cross-environment, cross-version
|
||
(7.x↔8.0), partial-or-whole companion to the native `persist()` snapshot.
|
||
|
||
```ts
|
||
// Export is a method on the immutable Db, so it composes with now()/asOf()/with()
|
||
const graph = await brain.export({ collection: id }, { includeVectors: true })
|
||
await otherBrain.import(graph, { onConflict: 'merge' }) // dedup-by-id
|
||
|
||
;(await brain.asOf(gen)).export(sel) // time-travel export
|
||
brain.now().with(ops).export(sel) // what-if export
|
||
```
|
||
|
||
- **`brain.export(selector?, options?)` / `db.export(...)`** → a versioned
|
||
`PortableGraph` document. Selectors reuse `find()`'s grammar: `{ ids }`,
|
||
`{ collection }` (alias `memberOf`, transitive `Contains`),
|
||
`{ connected: { from, depth } }`, `{ vfsPath }`, predicate
|
||
(`{ type, subtype, where, service }`), or the whole brain (omit) — and they
|
||
compose. Options: `includeVectors`, `includeContent` (VFS file bytes),
|
||
`includeSystem`, `edges: 'induced' | 'incident' | 'none'`.
|
||
- **`brain.import(graph, options?)`** is polymorphic: a `PortableGraph` document is
|
||
restored as **one atomic transaction** (`onConflict: 'merge' | 'replace' | 'skip'`,
|
||
`reembed: 'auto' | 'never'`, `remapIds` for cloning); a file/buffer routes to the
|
||
existing CSV/PDF/Excel/JSON ingestion. No migration for ingestion callers.
|
||
- **`validatePortableGraph(data)`** — dry-run structural/version/endpoint check before import.
|
||
- **Format** (`format:'brainy-portable-graph'`, `formatVersion: 1`) is identical on
|
||
7.x and 8.0; standard fields (`subtype`/`visibility`/`data`/…) are top-level,
|
||
`metadata` is custom-only. Current-state (no generation history — that lives in
|
||
`persist()`). Types exported from the package root.
|
||
- **Naming:** the type is `PortableGraph` (with `PortableGraphEntity` /
|
||
`PortableGraphRelation`) — it is the portable interchange form of a graph, not a
|
||
backup (that role is `persist()`/`load()`). Renamed from the short-lived
|
||
`BackupData`/`'brainy-backup'` (introduced in 7.32.0, never adopted) with no
|
||
compatibility shim.
|
||
- Guide: [docs/guides/export-and-import.md](docs/guides/export-and-import.md).
|
||
|
||
### Aggregation: distinctCount over any value type
|
||
|
||
`distinctCount` now counts distinct values of **any** type (strings, booleans,
|
||
numbers) — previously it silently returned `0` for non-numeric fields, missing its
|
||
primary use (distinct categories / users / tags). `percentile` (with `p`; median =
|
||
`p: 0.5`), `stddev`, and `variance` are exact and delete-safe alongside the
|
||
write-time `sum`/`count`/`avg`/`min`/`max`.
|
||
|
||
### Removed surfaces and their replacements
|
||
|
||
| Removed in 8.0 | Replacement |
|
||
|---|---|
|
||
| `brain.fork(name)` | Speculation: `db.with(ops)` (in-memory). Long-lived writable copy: `brain.restore()` a persisted snapshot into a fresh data directory. |
|
||
| `brain.checkout(branch)` | Open the snapshot you want — `Brainy.load(path)` read-only, or restore into its own directory. No in-place switching: every handle always sees one unambiguous store. |
|
||
| `brain.listBranches()` / `brain.getCurrentBranch()` / `brain.deleteBranch()` | A "branch" is now a name → snapshot-path mapping owned by your application. |
|
||
| `brain.commit({ message })` | `brain.transact(ops, { meta })` — every batch is an atomic, logged, time-travelable commit; audit fields live in the database, not in commit messages. |
|
||
| `brain.getHistory()` / `brain.streamHistory()` | `brain.transactionLog({ limit })` + `db.since(olderDb)`. |
|
||
| `brain.versions.*` (`save` / `list` / `compare` / `restore` / `prune`) | A pinned `Db` or persisted snapshot captures *every* entity at that moment; `asOf()` reads any entity's past state; `restore()` for whole-store rollback. |
|
||
| `brain.data()` (DataAPI: `clear` / `import` / `export` / `getStats`) | `brain.clear()`, `brain.import()`, `db.persist(path)` (the full-fidelity export format), `brain.restore(path, { confirm: true })`, `brain.stats()`. |
|
||
| Cloud + browser storage adapters (`s3` / `gcs` / `r2` / `opfs` storage types, Azure Blob, plus the `storage.branch` option) | `filesystem` and `memory` only. Cloud backup is now an operator concern: `db.persist()` produces a plain directory — sync it with `gsutil` / `aws s3 cp` / `rclone` / `azcopy`. Passing a removed storage type throws at construction with this exact guidance. |
|
||
| Browser support | Node.js 22 LTS or Bun ≥ 1.0 (`engines` enforces `node: 22.x`); Deno works through its Node compatibility layer. |
|
||
| `brain.migrateToDiskAnn()` / `brain.migrateToHnsw()` | Gone — there is one canonical `'vector'` provider slot. A registered native vector provider selects its own operating mode; there is nothing to call. |
|
||
| Distributed-clustering subsystem — `config.distributed`, the `DistributedRole` enum, the 13 `BRAINY_*` cluster env vars (`BRAINY_DISTRIBUTED`, `BRAINY_ROLE`, `BRAINY_HTTP_PORT`, `BRAINY_WS_PORT`, `BRAINY_DNS`, `BRAINY_SERVICE`, `BRAINY_NAMESPACE`, `BRAINY_CONSENSUS`, `BRAINY_COORDINATOR`, `BRAINY_NODES`, `BRAINY_REPLICAS`, `BRAINY_SHARDS`, `BRAINY_TRANSPORT`), the storage `setDistributedComponents` hook, and the `@soulcraft/brainy/config` preset/augmentation registry | Removed. Brainy 8.0 is a single-process library — there is no coordinator, peer discovery, or consensus to operate. Scale via: the optional native provider (`@soulcraft/cor` — on-disk DiskANN to 10B+ vectors on one machine); per-tenant pools (one instance + storage dir per tenant); and horizontal read scaling (many reader processes against one shared store, single writer). The `mode: 'reader' \| 'writer'` multi-process roles are unchanged. |
|
||
| CLI `fork` / `branch` / `checkout` / `migrate` | CLI `snapshot <path>` / `restore <path>` / `history` / `generation`. |
|
||
| `BrainyZeroConfig` type | Removed. 8.0 zero-config is automatic and internal — `new Brainy()` auto-detects storage, derives HNSW quality from `config.vector.recall`, picks `persistMode` from the adapter, and sizes caches to the detected container-memory limit. There is no config-generation type to import. |
|
||
| `brain.isFullyInitialized()` / `brain.awaitBackgroundInit()` | `await brain.ready`. The built-in filesystem/memory adapters finish initialization synchronously inside `init()`, so a single readiness promise is the whole story — these methods were no-ops once cloud adapters were removed. |
|
||
|
||
**Opening a 7.x store auto-migrates it.** 8.0 stores entities at the root; 7.x
|
||
stored them branch-scoped under `branches/<branch>/`. On first open, 8.0 collapses
|
||
the **HEAD branch** (`config.storage.branch`, default `main`) to the 8.0 layout in
|
||
place and rebuilds all derived state — no action required (`autoMigrate` defaults
|
||
to `true`; set it to `false` to make 8.0 refuse a legacy layout with an explicit
|
||
error instead). Two caveats:
|
||
|
||
- **Non-HEAD branches are not imported.** 8.0 has no COW branches; only the head
|
||
branch's data is migrated. If you care about other branches, **export each one
|
||
while still on 7.x** (`fork → export`, or copy its store).
|
||
- **Back up first for rollback.** The migration mutates the directory in place and
|
||
8.0 does not keep the old layout — copy the data directory before the first 8.0
|
||
open if you need to roll back.
|
||
|
||
### Renames — find/replace pairs
|
||
|
||
Hard renames, no aliases. Every pair below is verified against the 8.0 source.
|
||
|
||
| Brainy 7.x | Brainy 8.0 |
|
||
|---|---|
|
||
| `brain.delete(id)` | `brain.remove(id)` — matches the transact op vocabulary (`add` / `update` / `remove` / `relate` / `unrelate`) |
|
||
| `brain.deleteMany(params)` | `brain.removeMany(params)` |
|
||
| `DeleteManyParams` (type) | `RemoveManyParams` |
|
||
| `brain.getRelations(paramsOrId)` | `brain.related(paramsOrId)` — the same name a pinned `Db` exposes (`db.related()`) |
|
||
| `GetRelationsParams` (type) | `RelatedParams` |
|
||
| CLI `delete <id>` | CLI `remove <id>` (JSON output `{ id, removed: true }`, matching `unrelate`) |
|
||
| `HnswProvider` (from `@soulcraft/brainy/plugin`) | `VectorIndexProvider` |
|
||
| `HNSWIndex` (exported class) | `JsHnswVectorIndex` |
|
||
| Provider keys `'hnsw'` and `'diskann'` | `'vector'` (the only key Brainy consults for the vector index) |
|
||
| `config.hnsw = { quantization, vectorStorage }` | `config.vector = { recall?, persistMode? }` (shape change — see below) |
|
||
| `config.vector.quantization` (and 7.x `config.hnsw.quantization`) | **Removed.** The JS vector path is full-precision (exact float32 distances) only — quantization at scale is the native provider's job (e.g. DiskANN PQ). |
|
||
| `hnswPersistMode: 'immediate' \| 'deferred'` (top level) | `config.vector.persistMode` (auto-selected from the storage adapter when omitted: `'immediate'` on filesystem, `'deferred'` on memory) |
|
||
| `config.storage.type: 's3' \| 'gcs' \| 'r2' \| 'opfs'` | `'filesystem'` (or `'memory'` / `'auto'`) |
|
||
| `config.storage.branch` | Removed (no COW branches) |
|
||
| `brain.stats().indexHealth.hnsw` | `brain.stats().indexHealth.vector` |
|
||
| Storage adapter contract `saveHNSWData()` / `getHNSWData()` | `saveVectorIndexData()` / `getVectorIndexData()` |
|
||
| Cache category `'hnsw'` (`CacheProvider` union) | `'vectors'` |
|
||
| `config.history = { retainGenerations, retainMs, autoCompact }` | `config.retention` — `'all'` \| `'adaptive'` \| `{ maxGenerations?, maxAge?, maxBytes?, budgetBytes?, autoCompact? }`; **unset → adaptive** (was keep-100/7-days). The fields are now **caps**, not floors. |
|
||
| `compactHistory({ retainGenerations, retainMs })` | `compactHistory({ maxGenerations, maxAge, maxBytes })` — `retainGenerations`→`maxGenerations`, `retainMs`→`maxAge` (now upper-bound CAPS), plus new `maxBytes`. |
|
||
| `CompactHistoryOptions.retainGenerations` / `.retainMs` | `.maxGenerations` / `.maxAge` (+ `.maxBytes`) |
|
||
|
||
Mechanical renames as one command (run from your repo root, review the diff):
|
||
|
||
```bash
|
||
find . -name '*.ts' -not -path '*/node_modules/*' -print0 | xargs -0 sed -i \
|
||
-e 's/\bHnswProvider\b/VectorIndexProvider/g' \
|
||
-e 's/\bHNSWIndex\b/JsHnswVectorIndex/g' \
|
||
-e 's/\bsaveHNSWData\b/saveVectorIndexData/g' \
|
||
-e 's/\bgetHNSWData\b/getVectorIndexData/g' \
|
||
-e 's/indexHealth\.hnsw\b/indexHealth.vector/g' \
|
||
-e "s/registerProvider('hnsw'/registerProvider('vector'/g" \
|
||
-e "s/registerProvider('diskann'/registerProvider('vector'/g" \
|
||
-e "s/getProvider('hnsw'/getProvider('vector'/g" \
|
||
-e "s/getProvider('diskann'/getProvider('vector'/g" \
|
||
-e 's/\.getRelations(/.related(/g' \
|
||
-e 's/\bGetRelationsParams\b/RelatedParams/g' \
|
||
-e 's/\.deleteMany(/.removeMany(/g' \
|
||
-e 's/\bDeleteManyParams\b/RemoveManyParams/g'
|
||
```
|
||
|
||
`delete()` → `remove()` is deliberately **not** in the snippet: `.delete(` is
|
||
too common (`Map`/`Set`/storage adapters) for a blind sed. Find your
|
||
`brain.delete(...)` call sites and rename them by hand.
|
||
|
||
The config shape changes need a human (they are not 1:1 textual):
|
||
|
||
```ts
|
||
// 7.x
|
||
new Brainy({
|
||
hnswPersistMode: 'deferred',
|
||
hnsw: { quantization: { enabled: true, bits: 8, rerankMultiplier: 3 } }
|
||
})
|
||
|
||
// 8.0 — config.vector is { recall?, persistMode? }
|
||
new Brainy({
|
||
vector: {
|
||
recall: 'balanced', // 'fast' | 'balanced' | 'accurate'
|
||
persistMode: 'deferred' // optional; auto-selected from storage when omitted
|
||
}
|
||
})
|
||
```
|
||
|
||
`config.vector.recall` replaces the algorithm-internal HNSW knobs: `'balanced'`
|
||
(the default) maps to exactly the 7.x default parameters, so an upgrade without
|
||
explicit knobs changes nothing about search behaviour. `config.vector.quantization`
|
||
and `config.hnsw.vectorStorage` are removed: the JS vector path now computes exact
|
||
float32 distances throughout (no rerank/approximate branch), which is what made
|
||
its quantization a memory *increase* — it stored both the full and the quantized
|
||
vectors in RAM. Quantization at scale belongs to the native provider's own PQ.
|
||
The on-disk vector index file names are **unchanged** (e.g. `hnsw-system.json`) —
|
||
existing filesystem stores need no data migration for this rename; only the API
|
||
surface moved.
|
||
|
||
### Behavior changes
|
||
|
||
- **`subtype` is required by default.** 7.30's opt-in strict mode is now the
|
||
default: every `add()` / `addMany()` / `update()` / `relate()` /
|
||
`relateMany()` / `updateRelation()` rejects writes whose type carries no
|
||
non-empty `subtype` (`src/brainy.ts:9759` — `requireSubtype ?? true`).
|
||
Escape hatches: `requireSubtype: false` (last-resort opt-out for legacy
|
||
data) or `requireSubtype: { except: [NounType.Thing, ...] }` (per-type
|
||
allowlist). Per-type rules registered via `brain.requireSubtype(type, opts)`
|
||
still compose. `brain.audit()` reports entries missing a subtype and the new
|
||
`brain.fillSubtypes(rules)` migration helper backfills them — one rule per
|
||
NounType/VerbType (literal default or per-entry function), applied only to
|
||
entries still missing a subtype, returning
|
||
`{ scanned, filled, skipped, errors, byType }`. Idempotent: re-running fills
|
||
nothing, so a crashed run is resumed by running it again.
|
||
- **Verb ids are Brainy-generated UUIDs — by contract.** `relate()` now throws
|
||
a teaching error if a caller passes an `id` field
|
||
(`src/utils/paramValidation.ts:526`). In 7.x a supplied id was silently
|
||
ignored (a generated UUID was used anyway); 8.0 says so out loud, because
|
||
graph-index providers key verb-int interning on the raw UUID bytes.
|
||
- **`update({ vector })` is honored on its own.** 7.x only applied a supplied
|
||
pre-computed vector when `data` was also passed (it was silently dropped
|
||
otherwise). 8.0 applies an explicit `params.vector` with dimension
|
||
validation (`src/brainy.ts:1702-1713`; proven by
|
||
`tests/unit/brainy/update.test.ts:256`).
|
||
- **`brain.clear()` re-resolves all indexes exactly as `init()` does** —
|
||
including plugin-provided vector/metadata/entityIdMapper factories and VFS
|
||
root re-creation. In 7.x, clearing a plugin-accelerated brain could leave
|
||
the metadata index rebuilt without its native id-mapper wiring.
|
||
- **`find({ connected: { …, subtype } })` with `depth > 1` now works.** 7.30
|
||
threw `NOT_SUPPORTED` for multi-hop subtype-filtered traversal; 8.0
|
||
implements the BFS in the JS graph index and routes to a native provider
|
||
when one is registered.
|
||
- **Every write advances the generation clock; only `transact()` writes
|
||
history.** Single-operation writes (`add` / `update` / `remove` / `relate`
|
||
outside `transact()`) bump `brain.generation()` so watermarks and CAS stay
|
||
sound, but they do not stage historical records — they remain visible
|
||
through earlier pins and are not reported by `db.since()`. Writes you want
|
||
to travel back through go through `transact()`. This is the documented
|
||
contract, stated rather than papered over.
|
||
- **Reserved fields have one canonical location — enforced at every layer.**
|
||
The Brainy-owned field names (`RESERVED_ENTITY_FIELDS`:
|
||
`noun`/`subtype`/`createdAt`/`updatedAt`/`confidence`/`weight`/`service`/
|
||
`data`/`createdBy`/`_rev`; verb mirror `RESERVED_RELATION_FIELDS` with
|
||
`verb` for the type key) are now (1) a **compile error** inside any
|
||
`metadata` param — `add`/`update`/`relate`/`updateRelation` and the
|
||
matching `transact()` ops; (2) **normalized at write time** for untyped
|
||
callers — user-settable fields remap to their dedicated top-level param
|
||
(top-level wins; `update({metadata:{confidence}})` no longer silently
|
||
no-ops, closing a 7.x trap), system-managed fields drop with a one-shot
|
||
warning naming the right path; (3) **split at read time** through one
|
||
canonical helper, so `entity.metadata` / `relation.metadata` contain ONLY
|
||
custom fields on every read path — `get`, `find`, `related` (several
|
||
paginated/by-source/by-target paths previously echoed the full stored
|
||
record, including the `verb` type key, inside `metadata`), batch reads,
|
||
and historical `asOf()` reads. `related()` results now also surface
|
||
`confidence`/`updatedAt` top-level, and `updateRelation()` no longer
|
||
erases a relationship's `service`/`createdBy`. See "Reserved fields" in
|
||
`docs/concepts/consistency-model.md`.
|
||
- **Named errors for not-found contract failures.** `update()`, `relate()`,
|
||
`updateRelation()`, `similar({ to: id })`, `transact()` planning, and
|
||
speculative `db.with()` planning now throw `EntityNotFoundError` /
|
||
`RelationNotFoundError` (both exported from the package root, both carrying
|
||
the missing `id` as a field) instead of a generic `Error`. Message texts are
|
||
unchanged, so existing `/not found/` matching keeps working — `instanceof`
|
||
is now the supported way to branch.
|
||
- **Unchanged from 7.31:** per-entity `_rev`, `update({ ifRev })`
|
||
(`RevisionConflictError`), and `add({ ifAbsent })` work exactly as before,
|
||
and `ifRev` is also accepted on `transact()` update operations (a conflict
|
||
rejects the whole batch).
|
||
|
||
### For provider and plugin authors
|
||
|
||
The provider contracts in `@soulcraft/brainy/plugin` (`src/plugin.ts`) changed
|
||
shape:
|
||
|
||
- **`GraphIndexProvider` speaks BigInt at the boundary.** Reads take entity
|
||
ints and return entity/verb ints as `bigint[]`; the coordinator owns all
|
||
UUID ↔ int conversion. `addVerb(verb, sourceInt, targetInt)` receives the
|
||
resolved endpoint ints (also mirrored on the new optional
|
||
`GraphVerb.sourceInt` / `GraphVerb.targetInt` fields) and returns the
|
||
interned verb int. The batch reverse resolver
|
||
`verbIntsToIds(verbInts: bigint[]): Promise<(string | null)[]>` is
|
||
**required** — the provider owns durable verb-int interning; Brainy keeps
|
||
only a bounded in-memory warm cache. Implementations may stay u32 internally
|
||
(`Number(bigint)` narrowing is lossless under the `EntityIdSpaceExceeded`
|
||
guard) but must speak the bigint contract.
|
||
- **`VersionedIndexProvider`** is a new optional 4-method capability —
|
||
`generation()`, `isGenerationVisible(g)`, `pin(g)`, `release(g)`, BigInt
|
||
generations — feature-detected on every registered index provider. Providers
|
||
are *post-commit appliers*: the storage-record commit is the source of
|
||
truth; on open, a provider behind the committed watermark replays the gap or
|
||
requests a rebuild. Explicit pins override any time-based snapshot retention
|
||
the provider has. Speculative `db.with()` overlays never reach providers. A
|
||
provider implementing this serves historical reads with **no rebuild** —
|
||
the open-core materializer is the correctness baseline, the provider is the
|
||
accelerator.
|
||
- **The vector index registers under `'vector'`** and implements
|
||
`VectorIndexProvider`. The `'hnsw'` and `'diskann'` keys are retired and
|
||
never looked up. `CacheProvider`'s category union uses `'vectors'` in place
|
||
of `'hnsw'`.
|
||
|
||
### Scale and cost — how the mechanisms behave
|
||
|
||
Mechanism descriptions, not benchmark numbers (none are published for 8.0 yet):
|
||
|
||
- `brain.now()` pins in O(1) and adds **zero read overhead until history
|
||
actually moves** — while nothing has committed past the pin, every read
|
||
delegates to the live fast paths.
|
||
- A `transact()` commit pays O(ids touched) extra writes (before-images +
|
||
delta + manifest) — never O(store size). Single-operation writes pay an
|
||
in-memory counter bump with coalesced persistence.
|
||
- `db.persist()` on filesystem storage is a hard-link farm: snapshot creation
|
||
copies no entity data and the snapshot shares disk space with the source.
|
||
Cross-device targets fall back to per-file byte copies; persisting an
|
||
in-memory brain serializes a real, durable, loadable store.
|
||
- Historical **index-accelerated** queries (semantic search, traversal,
|
||
cursors, aggregation) on the open-core path pay a one-time O(n at the
|
||
pinned generation) in-memory index materialization per `Db`, cached until
|
||
`release()`. Record-path reads (`get`, metadata `find`, filter `related`)
|
||
at any reachable generation are served directly from the immutable record
|
||
layer. A native `VersionedIndexProvider` serves the same reads rebuild-free.
|
||
|
||
The full cost model is in
|
||
[docs/concepts/consistency-model.md](docs/concepts/consistency-model.md).
|
||
|
||
### Upgrade checklist (from 7.x)
|
||
|
||
1. **While still on 7.x:** back up your data directory (a plain file copy is
|
||
fine). If you used COW branches, materialize each branch you need — 8.0
|
||
does not read branch state. If you ran a cloud storage adapter, export your
|
||
data with 7.x tooling (`(await brain.data()).export()`) — 8.0 opens
|
||
`filesystem` and `memory` stores only.
|
||
2. **Runtime:** Node.js 22 LTS or Bun ≥ 1.0.
|
||
3. **Config:** change removed storage types to `'filesystem'`; delete
|
||
`storage.branch`; move `config.hnsw` / `hnswPersistMode` to
|
||
`config.vector` per the shape example above.
|
||
4. **Renames:** run the sed snippet, then fix the manual shape changes.
|
||
5. **Removed APIs:** replace `fork` / `checkout` / branch methods / `commit` /
|
||
`getHistory` / `streamHistory` / `versions` / `data()` per the table above.
|
||
6. **Subtypes:** if your data predates subtype discipline, start with
|
||
`requireSubtype: false`, run `brain.audit()`, backfill with
|
||
`brain.fillSubtypes(rules)` (one rule per type — a literal default or a
|
||
function deriving the subtype from each entry), re-run `audit()` until
|
||
`total === 0`, then remove the opt-out so the 8.0 default enforcement
|
||
protects you going forward.
|
||
7. **CLI scripts:** `fork` / `branch` / `checkout` / `migrate` →
|
||
`snapshot <path>` / `restore <path>` / `history` / `generation`.
|
||
8. **Plugin authors:** apply the contract renames and the BigInt graph
|
||
contract; optionally implement `VersionedIndexProvider`.
|
||
9. **Verify, then take your first snapshot:** run your test suite, then
|
||
`const db = brain.now(); await db.persist('/backups/post-8.0-upgrade');
|
||
await db.release()`.
|
||
|
||
What did **not** change: the core query surface (`find` / `search` / `get` /
|
||
`add` / `update` / `relate` and friends), the aggregation engine, the VFS,
|
||
neural extraction, and the on-disk entity/relationship layout — 8.0 creates
|
||
its MVCC bookkeeping (`_system/generation.json`, `_system/manifest.json`,
|
||
`_system/tx-log.jsonl`, `_generations/`) alongside the existing files on
|
||
first use.
|
||
|
||
### Native-provider (Cor) compatibility
|
||
|
||
Brainy 8.0 pairs with `@soulcraft/cor` 3.0 — the renamed successor to the
|
||
`@soulcraft/cortex` 2.x native provider — shipping in lockstep. The old
|
||
`@soulcraft/cortex` 2.x cannot accelerate Brainy 8.0: 8.0 never consults the
|
||
`'hnsw'` / `'diskann'` provider keys, and the graph/column provider contracts
|
||
are BigInt at the boundary. Upgrade both packages together (`@soulcraft/brainy@8`
|
||
+ `@soulcraft/cor@3`).
|
||
|
||
---
|
||
|
||
## v7.31.2 — 2026-06-09
|
||
|
||
**Affected products:** none in behaviour; affects anyone reading Brainy's TypeScript
|
||
type definitions or coreTypes for the `hnsw.quantization.bits` option. Drop-in from 7.31.1.
|
||
|
||
### Fix: stale comment claimed SQ4 vector quantization required a paid native provider
|
||
|
||
Brainy ships a pure-JS SQ4 distance implementation (`distanceSQ4Js` in
|
||
`src/utils/vectorQuantization.ts`) that's been part of the open-core path the whole
|
||
time. Two type-definition comments incorrectly stated SQ4 required a native (paid)
|
||
provider:
|
||
|
||
```ts
|
||
// coreTypes.ts:472 + types/brainy.types.ts:1123 — before
|
||
bits?: 8 | 4 // default: 8 (SQ8). SQ4 requires cortex native.
|
||
```
|
||
|
||
The comment was simply wrong — open-core users can use SQ4 today, and the native SIMD
|
||
acceleration is a drop-in via `setSQ4DistanceImplementation`, not a hard requirement.
|
||
Corrected to:
|
||
|
||
```ts
|
||
bits?: 8 | 4 // default: 8 (SQ8). SQ4 has a pure-JS implementation;
|
||
// cortex's distance:sq4 SIMD provider accelerates it.
|
||
```
|
||
|
||
Both locations updated. Pure documentation; no behaviour change. Caught during an
|
||
upstream open-core-boundary audit — thanks to whoever read the types carefully enough
|
||
to spot the misleading comment.
|
||
|
||
### Cortex compatibility
|
||
|
||
Zero changes required. The fix is documentation only; runtime behaviour was already
|
||
correct (`vectorQuantization.ts:412` ships `distanceSQ4Js`).
|
||
|
||
---
|
||
|
||
## v7.31.1 — 2026-06-09
|
||
|
||
**Affected products:** any consumer running `@soulcraft/brainy` against `FileSystemStorage`
|
||
that triggers concurrent flush / compaction (e.g. explicit `brain.flush()` calls overlapping
|
||
with periodic background compaction, or a backup job + a flush job racing). Production-impacting
|
||
when present: `brain.flush()` throws `ENOENT` on rename, downstream jobs that rely on flush
|
||
(GCS / S3 / Azure backups, snapshot exports) fail continuously. Drop-in from 7.31.0.
|
||
|
||
### Fixes a same-key rename race in `saveBinaryBlob`
|
||
|
||
`FileSystemStorage.saveBinaryBlob` used a bare `${filePath}.tmp` suffix for its atomic-write
|
||
temp file. Two concurrent same-key calls computed the **same** temp path; both `writeFile`d,
|
||
the first `rename` succeeded, the second `rename` fired against a missing temp and threw
|
||
`ENOENT`. The throw propagated up through `brain.flush()` and broke any downstream job that
|
||
called it.
|
||
|
||
```
|
||
[job-queue] gcs-backup: failed — ENOENT: no such file or directory,
|
||
rename '/data/brainy-data/.../_column_index/owner/DELETED.bin.tmp'
|
||
-> '/data/brainy-data/.../_column_index/owner/DELETED.bin'
|
||
```
|
||
|
||
**Patch:** unique per-writer temp suffix (matches the pattern at six other atomic-write
|
||
sites in the same file) + ENOENT swallow on rename (defensive: if the temp is gone, the
|
||
work has already landed; `saveBinaryBlob` is idempotent for a given key) + temp cleanup
|
||
on any other rename failure (no orphan `.tmp.*` files).
|
||
|
||
```ts
|
||
// before (7.31.0)
|
||
const tmpPath = filePath + '.tmp'
|
||
await fs.promises.writeFile(tmpPath, data)
|
||
await fs.promises.rename(tmpPath, filePath)
|
||
|
||
// after (7.31.1)
|
||
const tmpPath = `${filePath}.tmp.${process.pid}.${Date.now()}.${Math.random().toString(36).slice(2)}`
|
||
await fs.promises.writeFile(tmpPath, data)
|
||
try {
|
||
await fs.promises.rename(tmpPath, filePath)
|
||
} catch (err) {
|
||
if ((err as NodeJS.ErrnoException).code === 'ENOENT') return
|
||
await fs.promises.unlink(tmpPath).catch(() => {})
|
||
throw err
|
||
}
|
||
```
|
||
|
||
### Scope
|
||
|
||
The fix is one site — `src/storage/adapters/fileSystemStorage.ts:992-999`. Audit results:
|
||
|
||
- **`FileSystemStorage`** — six sibling atomic-write sites already used unique suffixes;
|
||
only `saveBinaryBlob` was the outlier. All six were rechecked. Clean.
|
||
- **`OPFSStorage`** — writes via WritableStream (no tmp+rename). Not affected.
|
||
- **`GCSStorage` / `R2Storage` / `AzureBlobStorage` / `S3CompatibleStorage`** — use
|
||
object-store `PUT` (atomic at the API). Not affected.
|
||
- **`MemoryStorage`** — in-memory. Not affected.
|
||
- **`HistoricalStorageAdapter`** — read-only. Not affected.
|
||
- **COW / versioning / snapshot / HNSW / aggregation** — all delegate to storage adapters
|
||
via `saveBinaryBlob` / `writeObjectToPath`. They get the fix automatically by virtue of
|
||
using the patched primitive.
|
||
|
||
### Beneficiary surfaces (broader than the reported bug)
|
||
|
||
Column-store compaction (`_column_index/<field>/DELETED.bin`, segment files) is what the
|
||
production report named, but `saveBinaryBlob` is also used by HNSW connection persistence
|
||
(`_hnsw_conn/<nodeId>` — `src/hnsw/hnswIndex.ts:252`). If two flush paths ever raced on
|
||
HNSW persistence for the same node, they would have hit the same ENOENT. No production
|
||
reports of HNSW-side failures, but the fix removes the latent race.
|
||
|
||
### Tests
|
||
|
||
- **New `tests/integration/savebinaryblob-concurrent-rename.test.ts`** (4 tests):
|
||
- 20 concurrent `saveBinaryBlob(sameKey, ...)` calls all resolve without throwing
|
||
- No orphan `.tmp.*` siblings remain in the blob directory after concurrent writes
|
||
- Production-shape: two concurrent compactor passes over 10 column-index fields
|
||
(`owner`, `path`, `permissions`, `vfsType`, `modified`, `createdAt`, `accessed`,
|
||
`updatedAt`, `mimeType`, `size`)
|
||
- Single-writer path still produces correct bytes (no regression)
|
||
- The first three tests reproduce the production `ENOENT: rename '...DELETED.bin.tmp' ->
|
||
'...DELETED.bin'` error verbatim on the pre-7.31.1 code path. Verified by stashing the
|
||
fix and watching the tests fail with the exact production error.
|
||
- Existing suites unchanged: 1468/1468 unit. All integration suites pass.
|
||
|
||
### Cortex compatibility
|
||
|
||
Zero changes required. The bug and the fix are pure JS in the filesystem storage adapter;
|
||
the Cortex mmap-filesystem adapter wraps `FileSystemStorage` and inherits the patch
|
||
automatically.
|
||
|
||
### Forward-compat
|
||
|
||
The bare-`.tmp` pattern is gone repo-wide. 8.0's `Db.persist()` will route through the
|
||
same patched primitive; no additional work needed when the immutable Db API ships.
|
||
|
||
---
|
||
|
||
## v7.31.0 — 2026-06-09
|
||
|
||
**Affected products:** consumers that need multi-writer coordination (job
|
||
schedulers, idempotent state machines, distributed locks, optimistic UI updates)
|
||
or idempotent bootstrap of well-known singletons. Additive; drop-in from 7.30.2.
|
||
|
||
### Per-entity `_rev` + `update({ ifRev })` + `add({ ifAbsent })`
|
||
|
||
CouchDB / PouchDB / ETag-style optimistic concurrency. Every entity now carries a
|
||
monotonic `_rev: number` that Brainy auto-bumps on every successful `update()`.
|
||
Pass it back as `ifRev` to make read-modify-write race-safe without an external
|
||
lock service. `ifAbsent` adds by-ID idempotent insert for singletons / config rows
|
||
/ deterministic-ID bootstraps.
|
||
|
||
**What's new:**
|
||
|
||
- **`entity._rev: number`** — initialized to `1` on `add()`, bumped by `1` on every
|
||
successful `update()` that reaches storage. Surfaced on `get()` (both fast metadata-
|
||
only and full-vector paths), `find()`, `search()`, and on the `Result` flatten layer
|
||
for backward compatibility. Pre-7.31.0 entities without `_rev` are read as `1`.
|
||
- **`update({ id, ..., ifRev: number })`** — optimistic-concurrency check. When
|
||
provided, throws `RevisionConflictError` if the persisted `_rev` no longer matches.
|
||
Carries `{ id, expected, actual }` for principled recovery. Omitting `ifRev` keeps
|
||
the prior unconditional-update behavior.
|
||
- **`add({ id, ifAbsent: true })`** — by-ID idempotent insert. Returns the existing
|
||
`id` without writing if the entity already exists. No throw, no overwrite. Ignored
|
||
when `id` is omitted (a fresh UUID can never collide).
|
||
- **`addMany({ items, ifAbsent: true })`** — applies the flag to every item; per-item
|
||
`ifAbsent` overrides the batch flag.
|
||
|
||
### The lock recipe (single-process or distributed)
|
||
|
||
```ts
|
||
import { Brainy, RevisionConflictError } from '@soulcraft/brainy'
|
||
|
||
const LOCK_ID = '...uuid...'
|
||
|
||
await brain.add({
|
||
id: LOCK_ID,
|
||
type: NounType.Document,
|
||
data: { owner: null, expiresAt: 0 },
|
||
ifAbsent: true
|
||
})
|
||
|
||
async function tryAcquireLock(workerId: string, ttlMs: number) {
|
||
const lock = await brain.get(LOCK_ID)
|
||
if (!lock) throw new Error('lock document missing')
|
||
const state = lock.data as { owner: string | null; expiresAt: number }
|
||
if (state.owner && state.expiresAt > Date.now()) return false
|
||
|
||
try {
|
||
await brain.update({
|
||
id: LOCK_ID,
|
||
data: { owner: workerId, expiresAt: Date.now() + ttlMs },
|
||
ifRev: lock._rev
|
||
})
|
||
return true
|
||
} catch (err) {
|
||
if (err instanceof RevisionConflictError) return false
|
||
throw err
|
||
}
|
||
}
|
||
```
|
||
|
||
### Why this is scoped to `_rev` + `ifAbsent`
|
||
|
||
A public `brain.transaction(fn)` wrapper was on the table but was cut. The internal
|
||
`TransactionManager` exposes raw `Operation` classes that take a `StorageAdapter`
|
||
constructor argument; a high-level facade in 7.31.0 would have meant either
|
||
(a) delegating to `brain.add()` / `update()` / `relate()` — each of which commits
|
||
its own internal transaction before the closure returns, so multi-write rollback
|
||
wouldn't actually work (a footgun), or (b) threading `tx?` through every internal
|
||
write path — ~2 days of refactor with regression surface, and the API shape changes
|
||
again in 8.0 anyway.
|
||
|
||
The SDK-scheduler use case that drove the thread (read-then-CAS lock pattern) is
|
||
fully solved by `_rev` + `ifRev`. Multi-write atomicity is the 8.0 `brain.transact()`
|
||
use case, where the Datomic-style immutable Db makes it atomic by construction. We
|
||
chose not to ship a worse version in the interim.
|
||
|
||
### How `_rev` interacts with branches and the snapshot API
|
||
|
||
| System | What it tracks | When it advances |
|
||
|---|---|---|
|
||
| `_rev` (NEW) | Per-entity write counter | Auto, on every successful `update()` |
|
||
| `brain.versions.save()` | Named snapshots per entity | Explicit — you call `save()` |
|
||
| `brain.fork()` / branches | Whole-brain copy-on-write | Explicit — you call `fork(name)` |
|
||
| VFS file versioning | Per-VFS-file snapshots | Same as `brain.versions.save()` |
|
||
|
||
`_rev` is independent. Each branch has its own copy of every entity, so each
|
||
branch has its own `_rev` per entity (same as every other field under COW).
|
||
Snapshots taken via `brain.versions.save()` capture the entity at that moment
|
||
including its `_rev` at that time; the snapshot's own `version: number` is the
|
||
snapshot index, distinct from `_rev`.
|
||
|
||
### Docs
|
||
|
||
- **New `docs/guides/optimistic-concurrency.md`** (public, indexed at
|
||
soulcraft.com/docs after portal deploy) — full reference: the lock pattern,
|
||
read-modify-write retry, idempotent bootstrap, how `_rev` interacts with branches /
|
||
snapshots / VFS, what's coming in 8.0.
|
||
- `docs/api/README.md` `add()` and `update()` entries get the new params + tips
|
||
pointing at the guide.
|
||
|
||
### Tests
|
||
|
||
- **New `tests/integration/rev-and-ifabsent.test.ts`** (18 tests): rev initialization,
|
||
rev bump on update, ifRev pass / fail / omit / legacy-entity-fallback, ifAbsent
|
||
with custom id / no id / batch propagation / per-item override, plus a full
|
||
SDK-scheduler-style two-concurrent-CAS-updaters scenario.
|
||
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-enforcement
|
||
30/30, strict-mode-self-test 13/13, find-limits 9/9. Unit 1468/1468.
|
||
|
||
### Cortex compatibility
|
||
|
||
Zero changes required. `_rev` is a metadata column already supported by
|
||
`NativeColumnStore`; the auto-bump runs in Brainy JS before any storage call.
|
||
Cortex never sees the per-entity counter — it's a single integer field updated
|
||
alongside every other metadata field on the existing write paths.
|
||
|
||
### 8.0 forward-compat
|
||
|
||
`_rev`, `ifRev`, `RevisionConflictError`, and `ifAbsent` all survive the 8.0 Db
|
||
redesign unchanged. 8.0 layers `brain.transact(tx, { ifAtGeneration })` for
|
||
whole-tx CAS on top of the same per-entity mechanism — per-entity for single-record
|
||
patterns, generation-based for "did the world move under me."
|
||
|
||
---
|
||
|
||
## v7.30.2 — 2026-06-08
|
||
|
||
**Affected products:** any consumer with `find({ limit })` call sites that pass values
|
||
≥ ~9000 — a common safety-cap pattern in production code (`limit: 10_000` against
|
||
type-filtered queries that typically return 10–500 entities). Additive; drop-in
|
||
from 7.30.1.
|
||
|
||
### Why
|
||
|
||
Brainy 7.30 introduced a memory-derived cap on `find({ limit })` to prevent OOM
|
||
from runaway queries. The cap was sound in intent but **~4× too conservative in
|
||
calibration**: the formula assumed 100 KB per result while typical entity footprint
|
||
is 7-10 KB (384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
|
||
free-memory box this capped `limit` at 9000, breaking production dashboards where
|
||
cascading 500s from queries with `limit: 10_000` silently degraded the `Promise.all`
|
||
loading the canvas.
|
||
|
||
7.30.2 recalibrates the formula and changes the enforcement from synchronous-throw
|
||
to two-tier (warn-then-throw) so existing pre-7.30 code doesn't break on upgrade.
|
||
|
||
### Recalibrated formula — `25 KB` per result (was `100 KB`)
|
||
|
||
The auto-configured cap now uses a realistic per-result size:
|
||
|
||
| Memory source | 7.30.0–7.30.1 cap | 7.30.2 cap |
|
||
|---|---|---|
|
||
| 4 GB Cloud Run container | 10 000 | 40 000 |
|
||
| 2 GB container | 5 000 | 20 000 |
|
||
| 900 MB free system memory | 9 000 | ~36 000 |
|
||
| `reservedQueryMemory: 1 GB` | 10 000 | 40 000 |
|
||
|
||
Hard ceiling stays at 100 000. Existing `BrainyConfig.maxQueryLimit` /
|
||
`reservedQueryMemory` overrides continue to work unchanged.
|
||
|
||
### Two-tier enforcement — warn, then throw
|
||
|
||
**Below cap (`limit ≤ maxLimit`):** silent pass. No signal, no friction. Unchanged.
|
||
|
||
**Soft tier (`maxLimit < limit ≤ 2 × maxLimit`)** *NEW*: one-time warning per call
|
||
site (dedup keyed on stack location + limit value), query proceeds. Pre-7.30.2 code
|
||
that relied on the cap silently allowing safety-cap limits keeps working — the
|
||
warning teaches the recipe so consumers can fix it intentionally.
|
||
|
||
**Hard tier (`limit > 2 × maxLimit`)** *NEW*: throw with the same message format the
|
||
warning uses. Real OOM territory; the cap stops being a recommendation and becomes
|
||
a guardrail.
|
||
|
||
The 2× soft margin is chosen to absorb existing safety-cap patterns (`limit: 10_000`
|
||
against a 9 K-cap box) without disabling OOM protection. Real OOM territory on a JS
|
||
in-memory brain is hundreds of thousands of results, not 10× the safety cap.
|
||
|
||
### Improved error / warning message (mirrors the 7.30.1 enforcement-error format)
|
||
|
||
Before (7.30 → 7.30.1):
|
||
```
|
||
limit exceeds auto-configured maximum of 9000 (based on available memory)
|
||
```
|
||
|
||
After (7.30.2):
|
||
```
|
||
find({ limit: 10000 }) exceeds the auto-configured query limit of 9000 (basis:
|
||
available free memory). Choose one:
|
||
• Increase the cap: new Brainy({ maxQueryLimit: 10000 })
|
||
• Reserve more memory: new Brainy({ reservedQueryMemory: 256000000 })
|
||
• Paginate: split the query with { limit, offset } pages
|
||
at OrderService.loadDashboard (/app/src/orders/dashboard.ts:142:18)
|
||
Docs: https://soulcraft.com/docs/guides/find-limits
|
||
```
|
||
|
||
Caller location comes from the same `findCallerLocation()` helper introduced in
|
||
7.30.1 (now extracted to `src/utils/callerLocation.ts` so both subtype enforcement
|
||
and limit enforcement share it).
|
||
|
||
### New docs
|
||
|
||
New `docs/guides/subtypes-and-facets.md`-style guide at
|
||
`docs/guides/find-limits.md` (public, indexed) — explains the cap, the four memory
|
||
sources the auto-config considers, the three escape valves (`maxQueryLimit`,
|
||
`reservedQueryMemory`, pagination), and when to use which. Explicit "pagination is
|
||
the future-proof pattern" callout — the cap can get tighter in 8.0 (Datomic-style
|
||
`Db.find()` may make per-call limits stricter to keep snapshot semantics cheap),
|
||
but pagination keeps working unchanged.
|
||
|
||
`docs/api/README.md` `find()` entry gets a one-paragraph `limit` tip + pointer
|
||
to the new guide.
|
||
|
||
### Tests
|
||
|
||
- **New `tests/integration/find-limits.test.ts`** (9 tests): below-cap silent
|
||
pass; soft-tier warns once per call site (dedup verified by exercising same vs.
|
||
different source lines); soft-tier message format (names all three escape
|
||
valves + docs link); soft-tier message includes caller location; hard-tier
|
||
throws; hard-tier message format same as soft-tier; consumer `maxQueryLimit`
|
||
override raises the cap and shifts both tiers accordingly; **the pre-7.30.2
|
||
regression scenario** explicitly covered — `limit: 10_000` against a
|
||
`maxQueryLimit: 9000` cap now warns and passes instead of throwing.
|
||
- Existing unit tests in `tests/unit/utils/memoryLimits.test.ts` updated to
|
||
reflect the recalibrated formula (5 tests: hardcoded values for
|
||
container / reserved / free-memory paths bumped 4×).
|
||
- Existing `paramValidation.test.ts` "should auto-limit based on system memory"
|
||
test extended to cover the new three-tier semantics (below-cap pass /
|
||
soft-tier silent / hard-tier throw).
|
||
- All prior suites still green: subtype-and-facets 26/26, verb-subtype-and-
|
||
enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.
|
||
|
||
### Cortex compatibility
|
||
|
||
**No Cortex changes required for 7.30.2.** Every change is JS-side: formula
|
||
recalibration runs in `ValidationConfig.constructor()`, two-tier enforcement runs
|
||
in `validateFindParams()`, both fire before any storage/index/Cortex call. Cortex
|
||
2.x and the pending Cortex 3.0 work both stay compatible without code changes.
|
||
|
||
The new `docs/guides/find-limits.md` calls out that 8.0's Datomic-style `Db.find()`
|
||
may tighten per-call limits; that's a Brainy 8.0 / Cortex 3.0 coordination point
|
||
whose rollout staging now also covers query limits.
|
||
|
||
### What consumers should do
|
||
|
||
- **If you've been carrying a `limit: 9000` workaround for the 7.30.0–7.30.1
|
||
cap:** drop it on bump to 7.30.2 — existing `limit: 10_000` patterns now
|
||
warn (teaching the recipe) but no longer throw. The warning is
|
||
one-time-per-call-site, so production logs don't spam.
|
||
- **All other consumers:** no action required if no enforcement-error was being
|
||
hit. If you see new `[Brainy]` warnings in logs after upgrade, follow the
|
||
recipe in the warning or read `docs/guides/find-limits.md`.
|
||
- **SDK wrappers:** wrapper-level pools that auto-construct `Brainy` instances
|
||
should consider surfacing `maxQueryLimit` / `reservedQueryMemory` as
|
||
consumer-facing config; Brainy now documents the knob explicitly.
|
||
|
||
---
|
||
|
||
## v7.30.1 — 2026-06-08
|
||
|
||
**Affected products:** anyone running a brain where a platform layer (SDK, framework wrapper)
|
||
has registered `brain.requireSubtype()` rules on common NounTypes, AND anyone preparing for
|
||
the upcoming Brainy 8.0 default-on strict mode. Additive; drop-in from 7.30.0. No behavior
|
||
change for consumers not using strict-mode enforcement.
|
||
|
||
### Why
|
||
|
||
Production incident 2026-06-08: a consumer using SDK 3.20.0 (which registers
|
||
`requireSubtype()` rules on `NounType.{Event, Collection, Message, Contract, Media, Document}`)
|
||
saw their booking flow start returning 500s because `brain.add({ type: NounType.Event, ... })`
|
||
calls in their codebase lacked `subtype`. An audit of Brainy's OWN source revealed 14 HIGH-risk
|
||
internal write paths that also omit subtype — VFS move/copy/symlink edges, aggregation
|
||
materializer, neural extraction, importers, integrations (Sheets/OData), MCP client, CLI.
|
||
Any consumer running `requireSubtype()` rules on those NounTypes was one step away from breaking
|
||
Brainy's own infrastructure paths, not just their own code. 7.30.1 closes both gaps before
|
||
8.0 ships and makes strict mode the default.
|
||
|
||
### New — `brain.audit()` diagnostic
|
||
|
||
Find entities and relationships missing a `subtype` value, grouped by type. The companion to
|
||
`migrateField()` (and to 8.0's `fillSubtypes()`): answers "what would break if I enabled
|
||
strict subtype enforcement?".
|
||
|
||
```typescript
|
||
const report = await brain.audit()
|
||
// {
|
||
// entitiesWithoutSubtype: { event: 24, document: 3 },
|
||
// relationshipsWithoutSubtype: { relatedTo: 1402 },
|
||
// total: 1429,
|
||
// scanned: 8400,
|
||
// recommendation: 'Found 1429 entries without subtype. Migrate via `brain.migrateField()`
|
||
// (7.x) — or wait for `brain.fillSubtypes()` (8.0) which closes the same
|
||
// gap with caller-supplied rules.'
|
||
// }
|
||
```
|
||
|
||
VFS infrastructure entities are excluded by default (they bypass enforcement via
|
||
`metadata.isVFSEntity` markers). Pass `{ includeVFS: true }` to surface them.
|
||
|
||
### New — Improved enforcement error messages
|
||
|
||
The error fired when subtype enforcement rejects a write now includes:
|
||
|
||
1. **The caller's source location** — extracted from the JavaScript stack so you see your own
|
||
call site, not a Brainy internal frame. Eliminates the "grep your repo for `brain.add`" step.
|
||
2. **Specific guidance** — points at the registered vocabulary when one exists; mentions
|
||
brain-wide strict mode and the `except` escape valve when not; otherwise the
|
||
`brain.requireSubtype()` registration recipe.
|
||
3. **A documentation link** — `https://soulcraft.com/docs/guides/subtypes-and-facets#strict-mode`
|
||
for the canonical migration recipe.
|
||
|
||
Before:
|
||
```
|
||
add(): NounType.Event requires subtype but got undefined. Register vocabulary via brain.requireSubtype().
|
||
```
|
||
|
||
After:
|
||
```
|
||
add(): NounType.event requires subtype but got undefined.
|
||
at OrderService.getOrCreateByToken (/app/src/orders/draft.ts:42:23)
|
||
Pass one of: standard, recurring, milestone.
|
||
Migration recipe: https://soulcraft.com/docs/guides/subtypes-and-facets#strict-mode
|
||
```
|
||
|
||
### Internal subtype labels — Brainy's own infrastructure paths
|
||
|
||
Every internal Brainy write path now sets a stable, queryable `subtype`. Consumers don't need
|
||
to do anything for these — they're documented here so you can query Brainy-managed data:
|
||
|
||
| Code path | NounType / VerbType | Subtype label |
|
||
|---|---|---|
|
||
| VFS root directory `/` | `Collection` | `'vfs-root'` |
|
||
| VFS subdirectories | `Collection` | `'vfs-directory'` |
|
||
| VFS files | mime-driven (e.g. `Document`/`Code`/`Image`) | `'vfs-file'` |
|
||
| VFS symlinks | `File` | `'vfs-symlink'` (NEW — distinct from `'vfs-file'`) |
|
||
| VFS Contains edges (create + move/copy/symlink/batch) | `Contains` | `'vfs-contains'` |
|
||
| Aggregation materialized output | `Measurement` | `'materialized-aggregate'` |
|
||
| Import-source provenance entity | `Document` | `'import-source'` |
|
||
| Importer-extracted entities | extractor-driven type | `'imported'` |
|
||
| Importer placeholder targets | `Thing` | `'import-placeholder'` |
|
||
| Neural extraction | extractor-driven type | `'extracted'` |
|
||
| GoogleSheets API entity writes | request-driven | `'imported-from-sheets'` |
|
||
| OData API entity writes | request-driven | `'imported-from-odata'` |
|
||
| MCP message storage | `Message` | `'mcp-message'` |
|
||
| `brainy add` CLI default | user-supplied type | `'cli-add'` |
|
||
| `brainy relate` CLI default | user-supplied verb | `'cli-relate'` |
|
||
|
||
Query examples:
|
||
|
||
```typescript
|
||
// Every VFS-managed file in your brain
|
||
await brain.find({ subtype: 'vfs-file' })
|
||
|
||
// Document breakdown — distinguishes import-source from extracted/imported/user content
|
||
brain.counts.bySubtype(NounType.Document)
|
||
// → { 'import-source': 12, 'imported': 847, 'extracted': 34, 'vfs-file': 102, ... }
|
||
```
|
||
|
||
### Caller-supplied `defaultSubtype` on importers and extraction
|
||
|
||
Importer + extraction paths now accept a caller-supplied `defaultSubtype` config so consumers
|
||
can tag a whole batch with their own provenance label instead of the Brainy default:
|
||
|
||
```typescript
|
||
// SmartImportOrchestrator / ImportCoordinator / NeuralImport all accept this
|
||
await brain.importer.import(file, {
|
||
defaultSubtype: 'customer-upload-2026q2', // your batch label
|
||
// ...
|
||
})
|
||
```
|
||
|
||
Precedence: extractor-set subtype (highest) → caller's `defaultSubtype` → Brainy default
|
||
(`'imported'` for importers, `'extracted'` for extractors).
|
||
|
||
### CLI `--subtype` flag
|
||
|
||
`brainy add` and `brainy relate` gain a `--subtype <value>` (`-s`) flag for use with
|
||
strict-mode brains:
|
||
|
||
```bash
|
||
brainy add "Avery Brooks — runs the AI lab" --type person --subtype employee
|
||
brainy relate alice manages bob --subtype direct
|
||
```
|
||
|
||
When the flag isn't supplied, the CLI uses `'cli-add'` / `'cli-relate'` as defaults so
|
||
ad-hoc CLI usage still works against strict-mode brains.
|
||
|
||
### Cortex compatibility
|
||
|
||
**No Cortex changes required for 7.30.1.** Every change is JS-side: internal subtype labels are
|
||
arbitrary strings stored transparently by Cortex; `brain.audit()` runs purely on JS via existing
|
||
`storage.getNouns()` / `getVerbs()` pagination; the CLI is JS-only; error messages fire in
|
||
Brainy before any native call.
|
||
|
||
**For Cortex 3.0 (forward-looking, not blocking):**
|
||
|
||
- **Native `audit()` proxy.** For billion-scale brains, `audit()` walks every entity (O(N)).
|
||
A native implementation reading from a "null-subtype" bitmap in the column store would be
|
||
O(buckets).
|
||
- **Strict-mode parity test.** Cortex should mirror Brainy's new
|
||
`tests/integration/strict-mode-self-test.test.ts` against their native paths to catch any
|
||
latent bug where native writes bypass JS validation.
|
||
- **Reserved-label awareness (optional).** Brainy's internal labels (`'vfs-*'`,
|
||
`'materialized-aggregate'`, `'imported'`, `'extracted'`, `'mcp-message'`, `'cli-*'`) become
|
||
a documented part of the 8.0 contract; useful for telemetry that surfaces "X% of entities are
|
||
Brainy-managed infrastructure".
|
||
|
||
### Docs
|
||
|
||
Updated `docs/guides/subtypes-and-facets.md` with a new "Strict mode in practice" section
|
||
covering the SDK_CORE_VOCABULARY pattern, a 4-step migration recipe, the Brainy-internal label
|
||
reference table, and an 8.0 forward-look. `docs/api/README.md` documents `brain.audit()` and
|
||
adds a strict-mode tip to the `add()` / `relate()` reference entries.
|
||
|
||
### Tests
|
||
|
||
- **New `tests/integration/strict-mode-self-test.test.ts`** (13 tests): creates a brain under
|
||
a realistic consumer vocabulary shape + brain-wide strict mode, then exercises every
|
||
internal Brainy path (VFS root/mkdir/writeFile/cp/mv/ln, aggregation engine, audit
|
||
diagnostic, error-message UX). Zero rejections expected.
|
||
- Existing 7.30.0 + 7.29.0 integration suites unchanged: 26/26 + 30/30.
|
||
- Unit suite unchanged: 1468/1468.
|
||
|
||
---
|
||
|
||
## v7.30.0 — 2026-06-05
|
||
|
||
**Affected products:** consumers modeling typed relationships with sub-classification
|
||
(direct vs dotted-line management; spouse / sibling / colleague; collaborator vs competitor;
|
||
etc.), and anyone wanting to enforce the pairing of `type` + `subtype` on every write.
|
||
Additive; drop-in from 7.29.x. No deprecations.
|
||
|
||
### Symmetric — `subtype` on relationships (parity with 7.29.0 nouns)
|
||
|
||
`subtype?: string` is now a first-class standard field on every relationship, mirroring the
|
||
noun-side work shipped in 7.29.0. Verbs and nouns are now first-class peers — every API
|
||
available on the noun side has a verb-side mirror.
|
||
|
||
```typescript
|
||
const ceoId = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'Avery' })
|
||
const vpId = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'Jordan' })
|
||
|
||
await brain.relate({
|
||
from: ceoId,
|
||
to: vpId,
|
||
type: VerbType.ReportsTo,
|
||
subtype: 'direct' // sub-classification on the edge
|
||
})
|
||
```
|
||
|
||
**Read & filter:**
|
||
|
||
```typescript
|
||
// Fast-path filter — column-store hit, not metadata fallback
|
||
const direct = await brain.getRelations({
|
||
from: ceoId,
|
||
type: VerbType.ReportsTo,
|
||
subtype: 'direct'
|
||
})
|
||
|
||
// Set membership
|
||
const all = await brain.getRelations({
|
||
from: ceoId,
|
||
type: VerbType.ReportsTo,
|
||
subtype: ['direct', 'dotted-line']
|
||
})
|
||
|
||
// Traversal filter (depth-1 in JS; multi-hop lands on Cortex native)
|
||
const reports = await brain.find({
|
||
connected: { from: ceoId, via: VerbType.ReportsTo, subtype: 'direct', depth: 1 }
|
||
})
|
||
```
|
||
|
||
### New — `updateRelation()` closes a long-standing gap
|
||
|
||
Verbs previously had no update method — the only way to change a relationship was
|
||
delete-then-recreate, which lost the relation id. 7.30 ships `brain.updateRelation()`:
|
||
|
||
```typescript
|
||
await brain.updateRelation({ id: relId, subtype: 'dotted-line' })
|
||
await brain.updateRelation({ id: relId, weight: 0.5, confidence: 0.9 })
|
||
|
||
// Change verb type — re-indexes in graph adjacency, id preserved
|
||
await brain.updateRelation({ id: relId, type: VerbType.WorksWith })
|
||
```
|
||
|
||
### New — O(1) verb subtype counts via the persisted rollup
|
||
|
||
`_system/verb-subtype-statistics.json` mirrors the noun-side rollup shipped in 7.29.0.
|
||
Per-VerbType-per-subtype counts are maintained incrementally and persisted; reads are O(1):
|
||
|
||
```typescript
|
||
brain.counts.byRelationshipSubtype(VerbType.ReportsTo)
|
||
// → { direct: 12, 'dotted-line': 3 }
|
||
|
||
brain.counts.byRelationshipSubtype(VerbType.ReportsTo, 'direct') // O(1) point
|
||
// → 12
|
||
|
||
brain.counts.topRelationshipSubtypes(VerbType.ReportsTo, 3)
|
||
// → [['direct', 12], ['dotted-line', 3]]
|
||
|
||
brain.relationshipSubtypesOf(VerbType.ReportsTo)
|
||
// → ['direct', 'dotted-line']
|
||
```
|
||
|
||
### New — `brain.requireSubtype(type, options)` per-type enforcement
|
||
|
||
Unified API for noun OR verb types — register specific types as requiring a subtype,
|
||
optionally with a fixed vocabulary:
|
||
|
||
```typescript
|
||
brain.requireSubtype(NounType.Person, {
|
||
values: ['employee', 'customer', 'vendor'],
|
||
required: true
|
||
})
|
||
|
||
brain.requireSubtype(VerbType.ReportsTo, {
|
||
values: ['direct', 'dotted-line'],
|
||
required: true
|
||
})
|
||
|
||
// Now this throws — Person requires subtype:
|
||
await brain.add({ type: NounType.Person, data: 'no subtype' })
|
||
|
||
// And this throws — 'matrix' isn't in the registered vocabulary:
|
||
await brain.relate({ from: a, to: b, type: VerbType.ReportsTo, subtype: 'matrix' })
|
||
```
|
||
|
||
### New — Brain-wide strict mode
|
||
|
||
`new Brainy({ requireSubtype: true })` enforces subtype on every public write across the
|
||
whole brain. Composes with per-type rules; per-type rules win when both apply.
|
||
|
||
```typescript
|
||
// Every write must include subtype
|
||
const brain = new Brainy({ requireSubtype: true })
|
||
|
||
// Exempt specific types (e.g. catch-all Thing)
|
||
const brain2 = new Brainy({
|
||
requireSubtype: { except: [NounType.Thing, NounType.Custom] }
|
||
})
|
||
```
|
||
|
||
When strict mode is on:
|
||
- Every `add()` / `addMany()` / `update()` / `relate()` / `relateMany()` / `updateRelation()`
|
||
rejects writes missing a subtype on a non-exempt type.
|
||
- `addMany()` and `relateMany()` validate every item BEFORE any storage write —
|
||
atomic-fail semantics, no partial writes.
|
||
- Brainy's own infrastructure writes (VFS root, directories, files) bypass via the
|
||
`metadata.isVFSEntity: true` marker so existing consumers' VFS usage continues to work.
|
||
|
||
The brain-wide flag becomes the default in 8.0.0; the type-level `subtype` field becomes
|
||
required at the type system level. The full 8.0 contract upgrade is coordinated through the
|
||
internal platform handoff (`CTX-SUBTYPE-8.0-CONTRACT`).
|
||
|
||
### `migrateField()` extended to verbs
|
||
|
||
The migration helper shipped in 7.29.0 now walks verbs too via the new `entityKind` option:
|
||
|
||
```typescript
|
||
// Migrate verb-side metadata.kind → top-level subtype
|
||
await brain.migrateField({
|
||
from: 'metadata.kind',
|
||
to: 'subtype',
|
||
entityKind: 'verb'
|
||
})
|
||
|
||
// Or walk nouns and verbs in one pass
|
||
await brain.migrateField({
|
||
from: 'metadata.kind',
|
||
to: 'subtype',
|
||
entityKind: 'both'
|
||
})
|
||
```
|
||
|
||
Default is `entityKind: 'noun'` (backward-compatible).
|
||
|
||
### Symmetry — noun + verb capability matrix
|
||
|
||
7.30 closes every gap between nouns and verbs. Every capability available on the noun side
|
||
has a verb-side mirror:
|
||
|
||
| Capability | Nouns | Verbs |
|
||
|---|---|---|
|
||
| `subtype` top-level field | ✓ (7.29) | ✓ (7.30 new) |
|
||
| Standard-field set | `STANDARD_ENTITY_FIELDS` | `STANDARD_VERB_FIELDS` (new) |
|
||
| Field resolver helper | `resolveEntityField` | `resolveVerbField` (new) |
|
||
| Statistics rollup | `_system/subtype-statistics.json` | `_system/verb-subtype-statistics.json` (new) |
|
||
| Counts breakdown | `counts.bySubtype` / `topSubtypes` / `subtypesOf` | `counts.byRelationshipSubtype` / `topRelationshipSubtypes` / `relationshipSubtypesOf` (new) |
|
||
| Fast-path filter | `find({type, subtype})` | `getRelations({type, subtype})` + `find({connected, subtype})` (new) |
|
||
| Update method | `update()` | `updateRelation()` (new — closed pre-7.30 gap) |
|
||
| Migration helper | `migrateField()` | `migrateField({entityKind: 'verb'\|'both'})` (new) |
|
||
| Enforcement | `requireSubtype(NounType, ...)` | `requireSubtype(VerbType, ...)` — one unified API |
|
||
| Brain-wide strict mode | `new Brainy({ requireSubtype })` covers both | same |
|
||
|
||
### Cortex compatibility
|
||
|
||
Verb subtype works under Cortex out of the box via auto-field-indexing. Cortex parity items
|
||
for the verb side (native query planner recognition, `verbSubtypeCountsByType` native rollup,
|
||
per-edge subtype on native graph adjacency for multi-hop traversal filtering) ship in the
|
||
next Cortex release. Not a Brainy blocker.
|
||
|
||
### Docs
|
||
|
||
Full guide: `docs/guides/subtypes-and-facets.md` (extended with Layer V + Enforcement
|
||
sections). The verb subtype + enforcement APIs are documented in `docs/api/README.md`.
|
||
|
||
---
|
||
|
||
## v7.29.0 — 2026-06-04
|
||
|
||
**Affected products:** anyone modeling entities with per-product sub-classification — every
|
||
consumer that's been reaching for `metadata.kind`, a custom `metadata.subtype` field, a
|
||
`data.kind` shape inside the payload, or similar ad-hoc conventions. Additive; drop-in from
|
||
7.28.x. No deprecations.
|
||
|
||
### New — `subtype` promoted to a top-level standard field
|
||
|
||
`subtype?: string` is now a first-class standard field on every entity, alongside `type` /
|
||
`confidence` / `weight`. Use it as the platform-standard primitive for sub-classifying entities
|
||
within a NounType (`Person` → `'employee'` / `'customer'`; `Document` → `'invoice'` / `'contract'`;
|
||
`Event` → `'milestone'` / `'meeting'`):
|
||
|
||
```typescript
|
||
await brain.add({
|
||
data: 'Avery Brooks — runs the AI lab',
|
||
type: NounType.Person,
|
||
subtype: 'employee', // top-level write param
|
||
metadata: { department: 'ai-lab' }
|
||
})
|
||
|
||
// Fast-path filter — column-store hit, not metadata fallback:
|
||
const employees = await brain.find({ type: NounType.Person, subtype: 'employee' })
|
||
|
||
// Set membership:
|
||
const internal = await brain.find({
|
||
type: NounType.Person,
|
||
subtype: ['employee', 'contractor']
|
||
})
|
||
```
|
||
|
||
Flat string, no hierarchy — your vocabulary, your choice. Brainy stores and counts, never validates.
|
||
|
||
### New — O(1) subtype counts via the persisted rollup
|
||
|
||
A new `_system/subtype-statistics.json` rollup is maintained incrementally as entities are added,
|
||
updated, and deleted — mirroring the existing `nounCountsByType` machinery with the same self-heal
|
||
behavior on poison detection. Counts are O(1) at billion scale:
|
||
|
||
```typescript
|
||
brain.counts.bySubtype(NounType.Person)
|
||
// → { employee: 12, customer: 847, vendor: 34 }
|
||
|
||
brain.counts.bySubtype(NounType.Person, 'employee') // O(1) point count
|
||
// → 12
|
||
|
||
brain.counts.topSubtypes(NounType.Person, 3)
|
||
// → [['customer', 847], ['employee', 12], ['vendor', 34]]
|
||
|
||
brain.subtypesOf(NounType.Person)
|
||
// → ['customer', 'employee', 'vendor']
|
||
```
|
||
|
||
### New — `brain.trackField()` for other metadata facets
|
||
|
||
For facets that aren't the *primary* sub-classification (`status`, `source`, `role`, `paradigm`),
|
||
`trackField` registers a field for cardinality + per-NounType breakdown stats without promoting
|
||
each one to a top-level slot. Piggybacks on the existing aggregation engine, so backfill-on-define
|
||
applies — registering on a populated brain scans existing entities on the first query:
|
||
|
||
```typescript
|
||
brain.trackField('status', { perType: true })
|
||
|
||
await brain.counts.byField('status')
|
||
// → { todo: 12, doing: 3, done: 47 }
|
||
|
||
await brain.counts.byField('status', { type: NounType.Task })
|
||
// → { todo: 8, doing: 2, done: 30 }
|
||
|
||
// Opt-in vocabulary validation:
|
||
brain.trackField('priority', { values: ['low', 'medium', 'high'] })
|
||
// brain.add({ ..., metadata: { priority: 'urgent' } }) → throws
|
||
```
|
||
|
||
### New — generic `brain.migrateField()` for one-shot rewrites
|
||
|
||
Streams every entity and copies a value from one field path to another. Supports top-level
|
||
standard fields, `metadata.*`, and `data.*` paths; idempotent; with optional `readBoth: true`
|
||
deprecation window that preserves the source field alongside the new one:
|
||
|
||
```typescript
|
||
// Phase 1: dual-populate (legacy readers still work)
|
||
await brain.migrateField({
|
||
from: 'metadata.kind',
|
||
to: 'subtype',
|
||
readBoth: true
|
||
})
|
||
|
||
// ...readers migrate to subtype at their own pace...
|
||
|
||
// Phase 2: clear the source field
|
||
await brain.migrateField({ from: 'metadata.kind', to: 'subtype' })
|
||
// → { scanned: 1500, migrated: 1500, skipped: 0, errors: [] }
|
||
```
|
||
|
||
Supports `batchSize` and `onProgress` for large brains.
|
||
|
||
### Aggregation composition
|
||
|
||
`subtype` is a standard-field group-by dimension out of the box — no `where:` wrapper, no
|
||
metadata-fallback overhead:
|
||
|
||
```typescript
|
||
await brain.find({
|
||
aggregate: {
|
||
groupBy: ['type', 'subtype'],
|
||
metrics: { count: { op: 'COUNT' } }
|
||
}
|
||
})
|
||
// [
|
||
// { groupKey: { type: 'person', subtype: 'customer' }, count: 847 },
|
||
// { groupKey: { type: 'person', subtype: 'employee' }, count: 12 },
|
||
// { groupKey: { type: 'document', subtype: 'invoice' }, count: 2103 },
|
||
// ...
|
||
// ]
|
||
```
|
||
|
||
### Cortex compatibility
|
||
|
||
Subtype works under Cortex out of the box via auto-field-indexing. Cortex will land its own
|
||
native-side query-planner recognition + `subtypeCountsByType` native rollup in the next Cortex
|
||
release for full parity with `type`'s fast path. Not a Brainy blocker — subtype reads/writes
|
||
function correctly with current Cortex.
|
||
|
||
### Docs
|
||
|
||
Full guide: `docs/guides/subtypes-and-facets.md`. Subtype is documented as a core primitive
|
||
throughout `README.md`, `docs/DATA_MODEL.md`, `docs/architecture/finite-type-system.md`,
|
||
`docs/api/README.md`, and `docs/QUERY_OPERATORS.md`.
|
||
|
||
---
|
||
|
||
## v7.24.0 — 2026-05-26
|
||
|
||
**Affected products:** aggregation/reporting users + anyone calling `extractEntities` on
|
||
multi-entity text. Additive; drop-in from 7.23.x.
|
||
|
||
### New — array-unnest `groupBy` (tag frequency / faceted counts)
|
||
|
||
A `groupBy` dimension can now be `{ field, unnest: true }`: the field holds an array and the
|
||
entity contributes once per **distinct** element. Enables tag-frequency / label-count style
|
||
aggregates:
|
||
|
||
```typescript
|
||
brain.defineAggregate({
|
||
name: 'tag_frequency',
|
||
source: { type: NounType.Document },
|
||
groupBy: [{ field: 'tags', unnest: true }],
|
||
metrics: { count: { op: 'count' } }
|
||
})
|
||
// queryAggregate('tag_frequency', { orderBy: 'count', order: 'desc' })
|
||
```
|
||
|
||
Duplicate elements on one entity count once; an entity with an empty/missing array joins no group.
|
||
|
||
### Perf — entity extraction batch-embeds candidates
|
||
|
||
`extractEntities` / `extractConcepts` now embed all unique candidate spans in a single
|
||
`embedBatch` call instead of one `embed()` per candidate (N sequential model calls before). No
|
||
behavior change — and with vectors reliably available, the embedding signal more consistently
|
||
reinforces correct types (e.g. clearer Organization confidence). Falls back to per-candidate
|
||
embedding if a batch call fails.
|
||
|
||
---
|
||
|
||
## v7.23.0 — 2026-05-26
|
||
|
||
**Affected products:** anyone using aggregation (stats/dashboards), graph traversal, or entity
|
||
extraction. Adds two report APIs and fixes three correctness gaps surfaced by BR-ADV-FEATURES-BUN.
|
||
Drop-in upgrade from 7.22.x.
|
||
|
||
### New — `brain.queryAggregate(name, params)`
|
||
|
||
A first-class report API returning the clean `AggregateResult[]` shape
|
||
(`{ groupKey, metrics, count }[]`) directly, instead of the search-`Result` wrapper that
|
||
`find({ aggregate })` returns. Accepts `where` / `having` / `orderBy` / `order` / `limit` / `offset`.
|
||
|
||
### New — HAVING (filter groups by metric value)
|
||
|
||
`find({ aggregate, having: { revenue: { greaterThan: 1000 } } })` (and the same on
|
||
`queryAggregate`) filters groups by their computed metrics — the analytics equivalent of SQL
|
||
`HAVING`, complementing `where` (which filters group keys). Evaluated per group: **O(groups),
|
||
independent of entity count.**
|
||
|
||
### Fix — aggregates now backfill when defined over existing data (R1)
|
||
|
||
Defining an aggregate on a store that already holds matching entities returned `[]` because
|
||
write-time hooks only saw *future* writes. It now backfills from existing entities on first query
|
||
(one-time scan via `getNouns`, then incremental) — storage-agnostic, so it works under durable
|
||
backends (Cortex) where a brain reopens pre-populated. Also: `groupBy:['noun']` now resolves to
|
||
the entity type instead of a single null group. `find({ aggregate })` rows now expose
|
||
`groupKey`/`metrics`/`count` at the top level (previously only under `.metadata`).
|
||
|
||
### Fix — multi-hop `find({ connected: { depth, via } })` (traversal)
|
||
|
||
`depth` and `via` are now honored at **every hop** (previously only the immediate neighbour was
|
||
returned, and verb filtering applied to hop 1 only). The BFS is bounded by `limit`.
|
||
|
||
### Fix — entity extraction type accuracy (R3)
|
||
|
||
`extractEntities` no longer lets a type indicator in one candidate's surrounding text bleed onto
|
||
neighbours (e.g. "Corp" in "Sarah Chen founded Acme Corp" no longer types "Sarah Chen" as
|
||
Organization). Each candidate is typed by its own span; context may only reinforce the same type.
|
||
|
||
---
|
||
|
||
## v7.22.1 — 2026-05-26
|
||
|
||
**Affected products:** Anyone using `extractEntities()` / `extractConcepts()`, native
|
||
aggregation (`find({ aggregate })`), or multi-hop graph traversal
|
||
(`find({ connected: { depth } })`). Three advanced-API correctness fixes — all reproduce on
|
||
Node, so they were **not** Bun-specific despite the original report (BR-ADV-FEATURES-BUN).
|
||
|
||
### Entity/concept extraction no longer returns `[]`
|
||
|
||
`extractEntities()` / `extractConcepts()` returned empty for entity-rich text. The SmartExtractor
|
||
ensemble scored agreeing signals with a weighted **sum** against an absolute 0.60 gate, so a
|
||
confident low-weight signal (e.g. a name pattern at 0.82) lost selection to a mediocre
|
||
high-weight signal that then failed the gate — dropping the whole result. It now selects and
|
||
gates on a normalized weighted **average**. Also: the `confidence` option now actually controls
|
||
the threshold (was a dead hardcoded 0.60), and the embedding-signal timeout was raised
|
||
100ms → 2000ms so the neural signal isn't silently dropped on slower runtimes.
|
||
|
||
*Known limitation:* type accuracy on dense multi-entity sentences is still imperfect — a strong
|
||
indicator in one entity's surrounding text can bleed onto neighbours. Tracked separately.
|
||
|
||
### `find({ aggregate })` exposes `groupKey` / `metrics` / `count`
|
||
|
||
Aggregation always computed correct values, but rows nested them under `.metadata`, so callers
|
||
expecting the documented `AggregateResult` shape saw empty-looking rows. Those three fields are
|
||
now also present at the top level of each result row.
|
||
|
||
### Multi-hop `find({ connected: { depth } })` honours depth
|
||
|
||
Previously returned only the immediate (1-hop) neighbour at any depth because the graph-search
|
||
path ignored `depth` / `via`. It now performs the full depth-aware traversal.
|
||
|
||
---
|
||
|
||
## v7.22.0 — 2026-05-15
|
||
|
||
**Affected products:** All. Fixes a silent-data-loss class affecting `find()` and
|
||
`brain.stats()`, plus Cortex-compatibility regression from 7.21.0. Recommended
|
||
upgrade for everyone on 7.20.x or 7.21.x.
|
||
|
||
### Two correctness fixes
|
||
|
||
#### 1. `find({ where })` no longer silently returns `[]` (BR-FIND-WHERE-ZERO)
|
||
|
||
The 7.20.0 column-store refactor deleted the legacy sparse-index write path
|
||
but left `getStats()` and `getIds()` reading from sparse indices. New
|
||
workspaces had no sparse-index files, so:
|
||
- `brain.stats().entityCount` was `0` regardless of actual entity count
|
||
- `find({ where: { ... } })` returned `[]` regardless of actual matches
|
||
- `rebuildIndexesIfNeeded()` saw `0` entries → fired the `[Brainy] CRITICAL ...
|
||
Second rebuild result: 0 entries` log → application saw no indexed data
|
||
|
||
7.22.0 makes the ColumnStore the single source of truth post-7.20.0:
|
||
- `MetadataIndex.getStats()` reads from `ColumnStore.getIndexedFields()` and
|
||
`idMapper.size`. Entity count is now accurate across writer → close →
|
||
reader-open cycles.
|
||
- `MetadataIndex.getIdsFromChunks()` throws `BrainyError(FIELD_NOT_INDEXED)`
|
||
when neither column store nor legacy sparse index has the field.
|
||
- `MetadataIndex.getIdsForFilter()` catches per-clause, logs
|
||
`[brainy] find() where-clause referenced unindexed field(s): ...`, returns
|
||
`[]`. Production `find()` is now consistent with `brain.explain()`'s
|
||
`path: 'none'` diagnostic.
|
||
|
||
Separately, `BaseStorage.getNounType()` was hardcoded to return `'thing'`
|
||
since a type-cache removal in commit `42ae5be`. Every noun save attributed
|
||
the entity to `'thing'` regardless of its actual type, poisoning
|
||
`_system/type-statistics.json` and `brain.stats().entitiesByType`. 7.22.0
|
||
restores type-correct attribution:
|
||
- New `nounTypeByIdCache: Map<string, NounType>` on `BaseStorage`, populated
|
||
in `saveNounMetadata_internal` and consumed in `saveNoun_internal`.
|
||
- New `flushCounts()` override that persists `nounCountsByType` /
|
||
`verbCountsByType` alongside the per-entity counter. Readers opening the
|
||
same directory after a writer flush now see correct per-type counts.
|
||
|
||
**Self-heal at init.** If `loadTypeStatistics()` detects the poisoned-state
|
||
signature (all nouns attributed to `'thing'` with multiple non-`thing`
|
||
entities on disk), it auto-runs `rebuildTypeCounts()` once and rewrites the
|
||
file. A `[BaseStorage] Detected poisoned type-statistics.json signature ...`
|
||
warning is logged. Operators don't need to take any action — the first 7.22
|
||
open of a directory poisoned by 7.20.0/7.21.0 fixes it.
|
||
|
||
**`brainy inspect repair`** can also be used to manually trigger
|
||
`rebuildTypeCounts()`. It's now public on `BaseStorage`.
|
||
|
||
#### 2. Cortex 2.2.x no longer crashes 7.21.0 boot (BR-DEFENSIVE-INTERFACE)
|
||
|
||
7.21.0 added new storage-adapter methods (`supportsMultiProcessLocking`,
|
||
`acquireWriterLock`, etc.) and called them unconditionally. Cortex 2.2.0's
|
||
mmap storage adapter bundles a pre-7.21 `BaseStorage` and doesn't define
|
||
them, so a consumer's 7.21.0 upgrade attempt threw
|
||
`TypeError: this.storage.supportsMultiProcessLocking is not a function`
|
||
at boot.
|
||
|
||
7.22.0 introduces `hasStorageMethod(name)` on Brainy and guards every new-
|
||
method call site. Older adapters degrade with a one-line warning at init:
|
||
```
|
||
[brainy] Storage adapter `CortexMmapStorage` predates the 7.21
|
||
multi-process methods. Writer locking and the flush-request RPC are
|
||
disabled for this directory. Upgrade the plugin (e.g.
|
||
`@soulcraft/cortex` ≥2.3.0) for full enforcement.
|
||
```
|
||
|
||
### Dead-code cleanup
|
||
|
||
Removed `MetadataIndexManager.dirtyChunks`, `dirtySparseIndices`, and the
|
||
`flushDirtyMetadata()` no-op machinery. These accumulators stopped being
|
||
populated when the sparse-index write path was deleted in 7.20.0; the
|
||
remaining 6 callers wasted async hops on an empty Map iteration.
|
||
|
||
### Upgrade notes
|
||
|
||
- **Behaviour change:** `find({ where: { unindexedField: ... } })` previously
|
||
returned `[]` silently. It now logs a one-line warning and still returns
|
||
`[]`. The diagnostic message names the field — use
|
||
`brain.explain({ where: { ... } })` for full diagnosis. No code change
|
||
needed for callers that already handle empty results.
|
||
- **Behaviour change (good):** `brain.stats()` and `brain.health()` now
|
||
report accurate per-type counts. If you previously relied on the buggy
|
||
`entitiesByType.thing` over-count, update consumers.
|
||
- **No API breaks.** Pure additive on the public surface.
|
||
|
||
### Cortex consumers
|
||
|
||
Cortex 2.2.x continues to work against 7.22.0 thanks to the defensive
|
||
guards, but for full multi-process protection on Cortex-backed stores and
|
||
to make sure Cortex's native MetadataIndex picks up the find-where-zero
|
||
fix, Cortex 2.3.0 is recommended. Tracked as `CX-MULTIPROC-METHOD` in the
|
||
internal platform handoff.
|
||
|
||
---
|
||
|
||
## v7.21.0 — 2026-05-15
|
||
|
||
**Affected products:** All consumers using filesystem storage (production deploys,
|
||
local development).
|
||
|
||
### Multi-process safety + first-class read-only inspector
|
||
|
||
Filesystem storage now enforces **single-writer, many-reader**. Previously, opening
|
||
a second writer process against a live data directory silently produced wrong query
|
||
results — no error, no warning, just empty or stale data. This release replaces
|
||
that footgun with a hard refusal at init time and a dedicated read-only inspection
|
||
mode.
|
||
|
||
#### New: `Brainy.openReadOnly()`
|
||
```ts
|
||
const reader = await Brainy.openReadOnly({
|
||
storage: { type: 'filesystem', rootDirectory: '/data/brain' }
|
||
})
|
||
const bookings = await reader.find({ where: { entityType: 'booking' } })
|
||
```
|
||
- Does NOT acquire the writer lock — coexists with a live writer.
|
||
- Every mutation method throws `Cannot mutate a read-only Brainy instance`.
|
||
- `flush()` and `close()` are safe (no-op + clean shutdown).
|
||
|
||
#### New: writer lock at init
|
||
- `new Brainy({ ... })` (default `mode: 'writer'`) acquires
|
||
`<rootDir>/locks/_writer.lock` containing `{ pid, hostname, startedAt, lastHeartbeat, version }`.
|
||
- A second writer on the same directory **throws** with the PID, hostname,
|
||
heartbeat, and a pointer to `openReadOnly()`.
|
||
- Stale-lock detection: same hostname + dead PID OR heartbeat > 60s ago → overwritten with a warning.
|
||
- Heartbeat: writer rewrites `lastHeartbeat` every 10s (unref'd timer).
|
||
- Override: pass `{ force: true }` to bypass when you've verified the existing lock is stale.
|
||
|
||
#### New: `brain.requestFlush({ timeoutMs })`
|
||
Cross-process RPC for inspectors to force a fresh snapshot:
|
||
- In-process: just calls `flush()`.
|
||
- Out-of-process: writes a request file, polls for ack. Times out gracefully.
|
||
- Watcher runs in every writer instance (filesystem only).
|
||
|
||
#### New: `brain.stats()`
|
||
Operator-facing summary: `entityCount`, `entitiesByType`, `relationCount`,
|
||
`relationsByType`, `fieldRegistry`, `indexHealth`, `storage.backend`, `writerLock`, `version`.
|
||
Designed for `/api/health` endpoints and incident triage.
|
||
|
||
#### New: `brain.explain(findParams)`
|
||
Shows which index path serves each `where` clause: `column-store` | `sparse-chunked` | `none`.
|
||
**This is the answer to "why is `find()` returning empty?"** — `path: "none"` means
|
||
the field has no index entries and the query will silently return `[]`. Includes
|
||
notes on likely causes (writer hasn't flushed; typo; field genuinely absent).
|
||
|
||
#### New: `brain.health()`
|
||
Invariant-check battery: HNSW vs metadata count parity, field registry sanity,
|
||
`_seeded` entity sweep, writer heartbeat freshness. Returns `{ overall: 'pass' | 'warn' | 'fail', checks: [...] }`.
|
||
|
||
#### New: `brainy inspect` CLI (13 subcommands)
|
||
All read-only by default (internally uses `openReadOnly()`):
|
||
```
|
||
brainy inspect stats <path>
|
||
brainy inspect find <path> --type Event --where '{"status":"paid"}' --limit 20
|
||
brainy inspect get <path> <id>
|
||
brainy inspect relations <path> <id> --direction both
|
||
brainy inspect explain <path> --where '{"entityType":"booking"}'
|
||
brainy inspect health <path>
|
||
brainy inspect sample <path> --type Event --n 20
|
||
brainy inspect fields <path>
|
||
brainy inspect dump <path> --type Event > backup.jsonl
|
||
brainy inspect watch <path> --type Event
|
||
brainy inspect backup <path> /backups/brain.tar
|
||
brainy inspect repair <path> # writer mode — stop live writer first
|
||
brainy inspect diff <pathA> <pathB>
|
||
```
|
||
Default behaviour: ask the writer to flush via the RPC first (skip with `--no-fresh`).
|
||
|
||
### Brainy + Cortex
|
||
|
||
Cortex segments (`*.cidx`) are immutable mmap files; `MANIFEST.json` uses atomic
|
||
rename. The single writer lock at `<rootDir>/locks/_writer.lock` covers both Brainy
|
||
and Cortex. Read-only inspectors mmap Cortex segments with zero coordination.
|
||
|
||
### What's NOT enforced yet
|
||
|
||
- **Cloud storage backends** (S3, GCS, R2, Azure) — no multi-process locking.
|
||
A best-effort warning logs in writer mode against non-filesystem backends.
|
||
- **The `find({ where })`-returns-0 root-cause bug** — tracked separately. The
|
||
new `brain.explain()` and `brain.health()` surface it loudly
|
||
(`path: 'none'`, `index-parity warn`) instead of letting it be silent.
|
||
|
||
### Upgrade notes
|
||
|
||
- **Default behaviour change:** opening a Brainy directory in writer mode while
|
||
another writer is live now THROWS where it previously silently produced wrong
|
||
results. If you have scripts that intentionally co-opened a writer directory,
|
||
switch them to `Brainy.openReadOnly()` or pass `{ force: true }`.
|
||
- **Cloud backends unchanged** — no lock acquisition for S3/GCS/R2/Azure.
|
||
- All existing APIs unchanged. Pure addition.
|
||
|
||
### Docs
|
||
|
||
- README "Single-Writer Model" section.
|
||
- `docs/concepts/multi-process.md` — full model + Cortex compatibility.
|
||
- `docs/guides/inspection.md` — operator recipes.
|
||
- JSDoc on every new API.
|
||
|
||
---
|
||
|
||
## v7.19.10 — 2026-02-24
|
||
|
||
**Affected products:** All Bun/ESM consumers
|
||
|
||
### ESM crypto fix in SSTable
|
||
|
||
Replaced `require('crypto')` with `import { createHash } from 'node:crypto'` in the
|
||
SSTable implementation. Fixes a crash in Bun and strict ESM environments where
|
||
CommonJS `require` is unavailable.
|
||
|
||
No API changes — upgrade and redeploy.
|
||
|
||
---
|
||
|
||
## v7.19.2 — 2026-02-18
|
||
|
||
**Affected products:** All
|
||
|
||
### Metadata index cleanup on delete
|
||
|
||
Fixed: metadata indexes were not cleaned up after `delete()` / `deleteMany()`. Stale
|
||
index entries could cause phantom results in metadata-filtered queries after deletion.
|
||
|
||
No API changes. If you were seeing ghost results in filtered queries, this fixes it.
|
||
|
||
---
|
||
|
||
## v7.18.0 — 2026-02-16
|
||
|
||
**Affected products:** Analytics, reporting, and session-summary consumers
|
||
|
||
### Aggregation engine
|
||
|
||
New `brain.aggregate()` API — incremental SUM, COUNT, AVG, MIN, MAX with GROUP BY
|
||
and time window support. Computes over entity collections without loading all records
|
||
into memory.
|
||
|
||
```typescript
|
||
const result = await brain.aggregate({
|
||
collection: 'bookings',
|
||
metrics: [
|
||
{ field: 'revenue', fn: 'SUM' },
|
||
{ field: 'id', fn: 'COUNT' },
|
||
],
|
||
groupBy: 'staffId',
|
||
timeWindow: { field: 'createdAt', from: startOfMonth, to: now },
|
||
})
|
||
```
|
||
|
||
SDK exposure: `sdk.brainy.aggregate()` — available once SDK is updated to pass through.
|
||
|
||
---
|
||
|
||
## v7.17.0 — 2026-02-09
|
||
|
||
**Affected products:** All (schema evolution, data migrations)
|
||
|
||
### Migration system
|
||
|
||
New `brain.migrate()` API with error handling, validation, and enterprise hardening.
|
||
Run schema migrations reliably across Brainy data directories.
|
||
|
||
```typescript
|
||
await brain.migrate({
|
||
version: 3,
|
||
up: async (brain) => {
|
||
// transform entities, rename fields, etc.
|
||
},
|
||
})
|
||
```
|
||
|
||
---
|
||
|
||
## v7.16.0 — 2026-02-09
|
||
|
||
**Affected products:** All
|
||
|
||
### Data/metadata separation enforced + numeric range queries
|
||
|
||
- Entity `data` and `metadata` fields are now strictly separated at the storage layer
|
||
- Numeric range queries now supported in metadata filters: `{ age: { $gte: 18, $lt: 65 } }`
|
||
- Fixes edge cases where mixed data/metadata storage caused inconsistent query results
|
||
|
||
**Breaking for anyone storing numeric values in metadata and relying on range queries:**
|
||
verify your filter syntax matches the new `$gte/$lte/$gt/$lt` operators.
|
||
|
||
---
|
||
|
||
## v7.15.5 — 2026-02-02
|
||
|
||
**Affected products:** Anyone using `@soulcraft/cortex` plugin
|
||
|
||
### Plugin opt-in clarified
|
||
|
||
Cortex and other plugins are opt-in. Pass explicitly:
|
||
```typescript
|
||
new Brainy({ plugins: ['@soulcraft/cortex'] })
|
||
```
|
||
Without `plugins`, no external plugins are loaded regardless of what's installed.
|
||
|
||
---
|
||
|
||
## v7.15.2 — 2026-02-01
|
||
|
||
**Affected products:** All (data safety)
|
||
|
||
### Graph LSM flush on close
|
||
|
||
Fixed: graph LSM-trees were not flushed on `brain.close()`, risking data loss across
|
||
restarts. Graph edges written in the final seconds before shutdown are now guaranteed
|
||
to be persisted.
|
||
|
||
No API changes — upgrade immediately if running Brainy in a long-lived server process.
|