diff --git a/CHANGELOG.md b/CHANGELOG.md index 4cb9a405..04283b67 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,41 +2,6 @@ All notable changes to this project will be documented in this file. See [standard-version](https://github.com/conventional-changelog/standard-version) for commit guidelines. -### [9.0.0](https://source.soulcraft.com/soulcraft/brainy/compare/v8.11.0...v9.0.0) (2026-08-04) - -- docs: 9.0 namespace-migration guide — the simple story + the mechanical sweep checklist, published for humans and tooling alike (61ab9db2) -- fix(release): storefront leg republishes CI's exact forge artifact — byte-identity by construction, verified by cross-registry shasum before the ceremony reports success (d89df2ed) -- docs: v9.0.0 release notes — the field-addressing law migration ledger; retitle the shipped 8.11.0 canonical-enumeration entry (header went stale at its cut) (55a7512c) -- feat(namespace): merge the field-addressing law train — no special names, system.* scalars, nested-bag storage, epoch-3 index keys (19b477ae) -- feat(namespace): NO SPECIAL NAMES + storage fidelity — the ruled completion of the field-addressing law (24bf6cdb) -- feat(namespace): write-door forgery refusal (user metadata keys may never start 'system.') + refusal messages name both spellings in every branch (the non-colliding case marks system. honestly as NOT valid) — cross-engine message pin alignment (48a6130a) -- feat(namespace): conformance green 19/19 — data-aware did-you-mean on unindexed bare addresses, ordering contract on the column top-K path (never drop, nulls last, ties by id), shape-complete addressed reads (entity views AND raw storage shapes, shadow-proof both scopes), per-key source matching for dotted addresses; refusal classes unified under UnresolvableFieldError (8e962dab) -- feat(namespace): aggregation reads under the law + epoch 3 (the key-split rebuild) + THE ARMING COMMIT — the capability constant, the law module, and the typed refusals export from the package root; both engines' conformance suites light on this signal (7492b6cb) -- feat(namespace): egress guard + validation speak the law — whereMatcher's resolver reads system.* from the record and bare names from the metadata bag only (the bare-system switch is dead); validateFindParams refuses cursor/includeRelations/writeOnly typed (accepted-and-ignored dies as a class), validates order, and parses every orderBy address (c2fb28a2) -- fix(namespace): noun-record updates preserve legacy inline HNSW adjacency — the placeholder-adjacency write stamped out pre-codec records' stored connections (crash-window unreachability); codec-era records were never at risk (empty field is the blob marker); pin covers the legacy shape (4679c894) -- feat(namespace): find's own filter builders speak the frozen keys — params.type/subtype/service become system.* index keys at every construction site (three pipelines + the canonical buildMetadataFilter); the where.type→noun alias is dead (bare 'type' belongs to the user now) (7a28a946) -- feat(namespace): the index speaks the frozen keys — record-frame scalars index under literal 'system.' (legacy 'noun' spelling folds into system.type; plumbing never indexed from a record frame), user fields stay bare in every shape; filter + sorted paths route every address through parseFieldAddress; storage fallbacks read the addressed side of the record (11c724bc) -- docs(namespace): the d.ts JSDoc wave — the sealed field-addressing law on the full find + aggregation surface, present-tense, with the refusal semantics and migration note inline (comment-only; verified zero code lines changed) (fcb24ab6) -- test(namespace): unit pins for the pure law — the ruled maps verbatim (incl. the relation mirror, unpinnable via public API), plumbing refusals both kinds, did-you-mean text (5502abcd) -- fix(namespace): the JS sorted fallback honors the ruled ordering contract — nulls last in BOTH directions (was nulls-first on desc) + deterministic id-ascending tie-break (56deb2e8) -- test(namespace)+docs: the cross-engine conformance suite (self-arming — skips until the resolver exports land) + the public field-addressing docs page; sidebar order deconflicted to 7 (d8d0b55f) -- feat(namespace): the one field-addressing law as a single source of truth — parseFieldAddress + the ruled ten-scalar system maps + plumbing invisibility + refusal builders (module only; query surfaces wire in next) (8f9a9989) -- docs: port the 8.10.3 backport-release changelog entry to main (f6b14d21) -- docs: port the 8.10.2 backport-release changelog entry to main — release branches carry the version bump, main carries the durable record (0b059ac5) -- fix: user metadata named 'level' is a real field everywhere — the engine-internal node layer no longer shadows it in sort/filter/aggregation, and the indexing views stop stamping a phantom 0 into its column; index epoch 2 rebuilds existing brains at first open (1a09be06) -- fix: metadata-only update() never rewrites the noun record — the unconditional whole-vector save turned per-entity stat touches into full rewrites+fsync, amplifying read-heavy sweeps into disk saturation on a production deployment (cb717be2) -- fix(release): double the forge-publish poll budget — the runner executes jobs sequentially and the publish run queues behind the ci matrix (64049631) -- Merge branch 'release/8.11.0' (1865f60a) -- Merge branch 'release/8.10.1' (fc9f0d72) -- chore: the forge is the address — retire the archived mirror from every live surface (415e824a) -- Merge remote-tracking branch 'origin/main' (069a8894) -- Merge branch 'release/8.10.0' (d918c060) -- ci: run the pipeline on the forge (9a5a9ccc) -- feat: two-tier history reads + the repacker + generationDigest — D1+D3 wired end-to-end (1201e255) -- feat: generation-segment store — the D1+D3 packed-tier file format (d8acb377) -- feat: scanFacts liveness contract — first batch or loud failure within a documented bound (f8e6da2b) - - ### [8.11.0](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.1...v8.11.0) (2026-07-27) - docs: the last two archived-host links point home (91ef1c8b) @@ -46,18 +11,6 @@ All notable changes to this project will be documented in this file. See [standa - ci: run the pipeline on the forge (999d0ebb) -### [8.10.3](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.2...v8.10.3) (2026-08-03) - -- docs: dedupe the 8.10.2 release-notes entry the cherry doubled onto the branch (8c956608) -- fix: user metadata named 'level' is a real field everywhere — the engine-internal node layer no longer shadows it in sort/filter/aggregation, and the indexing views stop stamping a phantom 0 into its column; index epoch 2 rebuilds existing brains at first open (958a0859) - - -### [8.10.2](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.1...v8.10.2) (2026-07-29) - -- docs: 8.10.2 consumer release notes — update() write granularity, PathResolver idle-log fix, graph-lsm key recognition (a0123b5b) -- fix: metadata-only update() never rewrites the noun record — the unconditional whole-vector save turned per-entity stat touches into full rewrites+fsync, amplifying read-heavy sweeps into disk saturation on a production deployment (5b65eb82) - - ### [8.10.1](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.0...v8.10.1) (2026-07-24) - refactor: remove the orphaned transaction-result type left behind by the dead-path removal (edf123a5) diff --git a/RELEASES.md b/RELEASES.md index ce7b5f99..2e137e4f 100644 --- a/RELEASES.md +++ b/RELEASES.md @@ -31,7 +31,7 @@ is sometimes cited as a 7.x removal — those methods never existed on 7.x; the --- -## v8.11.0 — 2026-07-27 (canonical enumeration mode for export — storage-walked, canon-complete) +## Unreleased (canonical enumeration mode for export — storage-walked, canon-complete) From a fleet data-migration program's requirement for whole-brain exports that are provably canon-complete: `export()`'s default enumeration for a whole-brain/predicate @@ -74,154 +74,6 @@ to the caller today. on CI**, triggered by the release tag, instead of PUTting the tarball from the laptop over WAN — no change to what gets published or how a consumer installs it. -## v9.0.0 — 2026-08-04 (the field-addressing law: your names and system.*, nothing in between) - -**Major.** One law now governs every field name, on every surface: - -> **Data is either in main space — where you can use ANY name — or it is in -> `system.*`.** - -Read `docs/concepts/field-addressing.md` (published on the docs site) for the -full contract; this entry is the migration ledger. - -### Breaking — query surfaces (`where` / `orderBy` / `groupBy` / aggregation) - -- **A bare field name ALWAYS addresses your metadata.** `orderBy: 'createdAt'` - no longer silently means the engine timestamp — it now refuses with a typed - `UnresolvableFieldError` naming both candidates unless you actually have a - user field of that name. Engine scalars are addressed explicitly: - `system.id`, `system.type`, `system.subtype`, `system.createdAt`, - `system.updatedAt`, `system.confidence`, `system.weight`, - `system.visibility`, `system.service`, `system.createdBy` (relations mirror - with `system.verb`/`system.sourceId`/`system.targetId`). - **Sweep list:** `where: { subtype: … }` → `where: { 'system.subtype': … }` · - `orderBy: 'createdAt'` → `'system.createdAt'` · `groupBy: ['noun']` → - `['system.type']` · any bare `visibility`/`service`/`confidence` filter that - meant the engine value → its `system.*` spelling. Every missed site fails - LOUDLY with the correction in the error message — nothing silently changes - meaning without telling you. -- **Unimplemented `find()` options refuse** (`cursor`, `includeRelations`, - `writeOnly` → `UnsupportedFindOptionError`); `order` is validated; - accepted-and-ignored is dead as a class. -- **The ordering contract is pinned cross-engine:** missing/null `orderBy` - values sort LAST in both directions, ties break by id ascending, and rows - are never dropped from an ordered read. - -### Breaking — write surfaces - -- **There are no reserved metadata names anymore.** `metadata: { confidence, - type, id, level, data, content, … }` are ordinary user fields — stored - verbatim, indexed, filterable, sortable, aggregatable, faithful across - restarts, index rebuilds, and `asOf()` time travel. The 8.x - reserved-key-in-bag throw is GONE; code that relied on it (or on the - `'warn'`/`'remap'` lift) must set engine scalars via their dedicated params - (`confidence`, `weight`, `subtype`, `visibility`, …) — the bag never touches - them now. -- **`reservedFieldPolicy` is removed.** Passing it throws at construction with - the migration note. `RESERVED_ENTITY_FIELDS`/`RESERVED_RELATION_FIELDS` - remain exported but now describe the stored record's engine half, not a ban - list; the `NoReservedEntityKeys`/`NoReservedRelationKeys` types are no-op - (deprecated). -- **The one refused spelling:** a metadata key literally starting `system.` - (namespace forgery) — typed error on `add`/`update`/`relate`/`updateRelation`. -- **Name-based index exclusions are gone.** Fields named `content`, `data`, - `id`, `vector`, … in your bag now INDEX like everything else (they were - silently un-indexed before — `where` on them returned `[]` with no error). - Value-shape rules stay, uniform across all names: arrays >10 never become - posting scalars; long values index hashed. -- **Migration transforms receive one normalized view** (engine fields - top-level, your bag nested under `metadata`) regardless of how old the - stored record is, and must return the same shape — a stray non-engine - top-level key refuses with the fix in the message. - -### Storage format (automatic, no action) - -- New/updated records persist as **nested-bag records** (engine fields - top-level, your bag verbatim under `metadata`, sealed by a format stamp) — - the shape that makes collider names lossless. Old flat records stay - readable forever; nothing rewrites your data in place. -- **Index epoch 3:** derived-index keys split the namespaces (bare user keys · - literal `system.` keys; the legacy `noun` column is gone). Every - brain rebuilds its derived indexes from canonical once, at first open — - observable via `getIndexStatus()`, no manual step. Pair this release with - the same-day native-accelerator release (its peer floor rises to `>=9`). -- Raw-record consumers (fact-log scanners, export tooling): read bags through - the exported shape-aware splitters (`splitNounMetadataRecord` / - `splitVerbMetadataRecord`) — they handle both record eras. - -### Fixed in the same train - -- Default visibility exclusion was a silent no-op under the new addressing on - pre-release builds (internal/system-tier rows could leak into default - reads) — now pinned by conformance tests at every lifecycle boundary. -- Per-type count surfaces (`getStats()`, count-by-type) read the new type - column, with a legacy fallback for pre-rebuild reads. -- Aggregation `source.where` evaluated dotted keys as nested paths — dotted - addresses now match per-key, and the internal per-type counts aggregate - rebuilds itself onto the new keys automatically. - -### Conformance - -Both engines ship a shared self-arming conformance suite (the law cases, the -ordering contract, and the reopen-collider fidelity case: every collider name -written as user data, verified verbatim through live reads, reopen, a forced -epoch rebuild, and time travel). Capability signal: -`FIELD_ADDRESSING_CAPABILITY = 'field-addressing/v1'` plus the typed error -classes, exported from the package root. - -## v8.10.3 — 2026-08-03, 8.10-line backport (natural field names stop colliding with engine internals) - -From a production report: sorting by a user metadata field named `level` silently -returned insertion order — the engine's internal HNSW node layer (also called -`level`) shadowed the user's field in every by-name read, and the indexing path -stamped a hardcoded `0` into the same index column (multi-valued poison). `level` -is a perfectly natural field name (game characters, priorities, floors); the -engine was wrong, not the caller. - -- **`level` is user data now, everywhere.** Engine plumbing no longer resolves by - name, never shadows metadata, and never enters the indexed views. `orderBy: - 'level'`, `where: { level: 9 }`, `groupBy: ['level']` all read YOUR field. - Regression pins: `tests/integration/level-field-shadow.test.ts` (the reporting - consumer's exact repro rows). -- **Index epoch 2.** The derived posting set changed, so every existing brain - rebuilds its metadata index from canonical at first open — poisoned columns - heal automatically; no manual step. First open after upgrade pays one rebuild - (observable via `getIndexStatus()`); pair this release with the same-day - native-accelerator release, which makes `level` indexable on the native path. -- **`transact()` metadata-only updates stop rewriting the vector record** — the - v8.10.2 write-granularity law now covers the batch/plan path too (it was - fixed for `update()` but the transact plan builder still staged the - unconditional save). If you batch stat touches through `transact()`, this is - your write-amplification fix. -- (The "coming next" note this entry carried shipped as v9.0.0 — the - field-addressing law above.) - ---- - -## v8.10.2 — 2026-07-29 (metadata-only updates stop rewriting the vector record) - -From a production incident on a large deployment: a read-heavy sweep that bumped -per-entity stats (metadata-only `update()` calls) saturated the disk — 5.8GB written -in 40 minutes — because every `update()` unconditionally re-persisted the WHOLE noun -record, unchanged vector included, fsynced. - -- **`update()` write granularity fixed at the core.** A metadata-only update (no new - `data`, `vector`, or `type`) now writes the metadata leg and index deltas ONLY — - the vector-bearing noun record is never rewritten. Vector-side writes and HNSW - reindexing still happen exactly when the vector side actually changed. Regression - pins: `tests/integration/update-write-granularity.test.ts`. -- **Consumer guidance:** per-entity stat touches are now cheap, but batch them anyway - (one `transact()` instead of N `update()` calls) — granularity fixes the cost per - touch; batching fixes the count. -- Idle VFS `PathResolver` no longer logs `NaN% hit rate` once a minute (stats log - only on new traffic, at debug level). -- Native graph providers' `graph-lsm-*` storage keys are recognized as system - resources — the per-boot `Unknown key format` warning for them is gone. - -Pairs with the native accelerator's same-day patch release; adopt as one bump. - ---- - ## v8.10.1 — 2026-07-24 (the no-hot-retry contract + warm()'s metadata surface under native providers) From a production incident: a native-provider op ground 38-40s inside a transaction, diff --git a/docs/concepts/field-addressing.md b/docs/concepts/field-addressing.md deleted file mode 100644 index c459021b..00000000 --- a/docs/concepts/field-addressing.md +++ /dev/null @@ -1,233 +0,0 @@ ---- -title: Field addressing: your fields and system fields -slug: concepts/field-addressing -public: true -category: concepts -template: concept -order: 7 -description: The one rule for every query-surface field name — a bare name always means your metadata, system. reaches the ten engine scalars explicitly, and anything else refuses by name. -next: - - guides/namespace-migration - - concepts/consistency-model ---- - -# Field addressing: your fields and system fields - -Every query surface in Brainy — `find()`'s `where`, `orderBy`, aggregation -`groupBy`, and aggregation `source.where` — resolves field names by one rule, -with no exceptions: - -> **A bare field name always means your metadata. `system.` reaches an -> engine scalar, and only when you spell it explicitly.** - -```typescript -await brain.find({ orderBy: 'level' }) // reads entity.metadata.level — YOUR field -await brain.find({ orderBy: 'system.createdAt' }) // reads the engine's createdAt scalar -await brain.find({ orderBy: 'metadata.level' }) // identical to bare 'level' — explicit scope -``` - -There is no priority list, no "try the system field, fall back to metadata" -behavior, and no name that resolves differently depending on what else -happens to exist on your entities. A field called `level`, `score`, -`createdAt`, or `type` in your own `metadata` is read as *your* field, every -time, by its bare name. - -## Why this rule exists - -An internal report from a production deployment found that a user metadata -field literally named `level` was being silently shadowed by the engine's -own internal index layer field of the same name — every sort by `level` -returned insertion order, with no error raised. This rule makes that class of -bug structurally impossible: bare names belong to you, unconditionally, and -anything that isn't yours has to be spelled out. - -## The system scalars - -`system.` addresses exactly ten scalars on an entity — no more, no -fewer: - -| System field | What it is | -|---|---| -| `system.id` | The entity's id | -| `system.type` | The entity's `NounType` | -| `system.subtype` | The per-app sub-classification passed to `add()` | -| `system.createdAt` | When the entity was created | -| `system.updatedAt` | When the entity was last written | -| `system.confidence` | The `confidence` param (0–1) | -| `system.weight` | The `weight` param | -| `system.visibility` | `'public'` / `'internal'` (see the visibility tiers in [Consistency Model](./consistency-model.md)) | -| `system.service` | The multi-tenancy `service` tag | -| `system.createdBy` | Who/what created the entity | - -Relationships mirror the same eight shared scalars (`subtype`, `createdAt`, -`updatedAt`, `confidence`, `weight`, `visibility`, `service`, `createdBy`) -plus three of their own: - -| System field (relationship) | What it is | -|---|---| -| `system.verb` | The relationship's `VerbType` | -| `system.sourceId` | The id of the entity the relationship starts from | -| `system.targetId` | The id of the entity the relationship points to | - -Anything not on these two lists is not a system scalar — `system.` for -any other name refuses (see "Refusal semantics" below), even if that name -sounds like it should be engine-owned. - -## Invisible plumbing — never addressable, in either spelling - -Five names are pure engine internals. They are not reachable as a bare name, -and not reachable as `system.` either — they simply have no place on -the query surface: - -- **`vector`** — the stored embedding. It participates in similarity search - (`query`, `near`, vector `find()`), never in `where`/`orderBy`/`groupBy`. -- **`connections`** — graph adjacency. Reached through `connected` and - `brain.related()`, not through field addressing. -- **`level`** — the internal index layer number used by the nearest-neighbor - graph. It is pure index plumbing with no query-surface meaning at all — - which is exactly why a user field of the same name must never be shadowed - by it. `level` as a bare name is always yours; there is no engine-owned - spelling of it to compete with. -- **`data`** — your entity's content payload, not a scalar. It can be a - string, a number, or an arbitrary object, so sorting or filtering it as a - single comparable value would lie about its actual shape. Content is - reached through the content/text-search APIs (`query`, `searchMode: - 'text'`), not through `where`/`orderBy`. -- **`_rev`** — the per-entity revision counter used for optimistic - concurrency (`ifRev`). It is a CAS token, not a queryable dimension. - -`system.level`, `system.vector`, and `system.data` all refuse for the same -reason: they are not in the ten-scalar system map, full stop. - -## `metadata.` — the explicit spelling of "mine" - -Prefix any field with `metadata.` to say the same thing a bare name already -says, spelled out. The two are interchangeable everywhere a field name is -accepted, including `orderBy`: - -```typescript -await brain.find({ where: { 'customer.tier': 'gold' } }) -await brain.find({ where: { 'metadata.customer.tier': 'gold' } }) // identical -await brain.find({ orderBy: 'metadata.score', order: 'desc' }) // identical to orderBy: 'score' -``` - -Reach for the explicit spelling when it reads more clearly next to a -`system.` field in the same query — for example, sorting by your own `score` -while filtering on `system.confidence`. - -## No special names — the write side - -The same law governs writes: - -> **Data is either in main space, where developers can use anything, or it -> is in `system.*`.** - -There are **no reserved metadata names**. A field called `confidence`, -`type`, `id`, `data`, `content`, or anything else inside your `metadata` bag -is an ordinary user field: it is stored verbatim, indexed, filterable, -sortable, aggregatable, and it survives restarts, index rebuilds, and -time-travel (`asOf`) reads exactly as written — even when an engine scalar -shares its spelling. The engine's values are written only through their -dedicated params (`confidence`, `weight`, `subtype`, `visibility`, …) and -read at `system.`; your bag can never touch them and they can never -shadow your bag. - -```typescript -const id = await brain.add({ - data: 'Ada Lovelace', - type: NounType.Person, - confidence: 0.9, // the ENGINE scalar - metadata: { confidence: 'self-rated' } // YOUR field, same spelling — both live -}) - -await brain.find({ where: { confidence: 'self-rated' } }) // finds it (yours) -await brain.find({ where: { 'system.confidence': 0.9 } }) // finds it (engine's) -``` - -The one spelling a write refuses is a metadata key that literally starts -with `system.` — the explicit address namespace cannot be forged as a user -field name. That refusal is typed and names the fix. - -Value **shape** rules still apply uniformly to every name (they are not name -carve-outs): arrays longer than 10 elements are not turned into posting-list -scalars, and very long values are indexed by hash. - -## Refusal semantics - -A name that resolves to neither your metadata nor a system scalar is a typed -refusal, not a silent empty result and not a guess. Refusals name **both** -candidates, so the fix is always in the error text: - -```typescript -await brain.find({ orderBy: 'createdAt' }) -// UnresolvableFieldError: no metadata field 'createdAt' — did you mean -// system.createdAt or metadata.createdAt? -``` - -`UnresolvableFieldError` is exported from the package root: - -```typescript -import { UnresolvableFieldError } from '@soulcraft/brainy' - -try { - await brain.find({ orderBy: 'createdAt' }) -} catch (err) { - if (err instanceof UnresolvableFieldError) { - // err.message names both candidates — usually enough to fix the call site. - } -} -``` - -A handful of `find()` options are not implemented yet: `cursor`, -`includeRelations`, and `writeOnly`. Rather than accepting them and quietly -ignoring the option, `find()` refuses with `UnsupportedFindOptionError` — -also exported from the package root — so a call site can never believe an -unimplemented option took effect when it didn't. - -## The ordering contract - -`orderBy` behaves identically regardless of which engine (the pure-TypeScript -path or a native accelerator) is serving the query: - -- An entity missing the `orderBy` field, or holding `null` on it, sorts - **LAST — in both `asc` and `desc`**. It is never treated as "smaller than - everything" in one direction and "larger than everything" in the other; it - is simply last, either way. -- Rows are **never dropped** from an ordered read because they lack the - field — a missing value changes position, never presence. -- Ties on the `orderBy` field break by **id ascending**, regardless of the - primary sort direction. - -```typescript -// employees: [{ score: 9 }, { score: 5 }, { /* no score field */ }] -await brain.find({ orderBy: 'score', order: 'desc' }) // [9, 5, missing] — missing is last -await brain.find({ orderBy: 'score', order: 'asc' }) // [5, 9, missing] — missing is STILL last -``` - -## Migrating existing call sites - -If you have call sites written before this rule shipped that rely on a bare -system name — `orderBy: 'createdAt'`, `where: { confidence: { greaterThan: -0.8 } }`, and similar — they now refuse instead of silently resolving to the -engine field. The fix is always in the error: swap the bare name for -`system.` (or `metadata.` if you actually meant your own field -of that name, and it happens to share a name with a system scalar): - -```typescript -// Before: bare 'createdAt' silently meant the engine's timestamp. -await brain.find({ orderBy: 'createdAt' }) - -// After: say which one you meant. -await brain.find({ orderBy: 'system.createdAt' }) // the engine timestamp -await brain.find({ orderBy: 'metadata.createdAt' }) // your own field named createdAt, if you have one -``` - -There is no silent migration path by design — every ambiguous call site -surfaces as a refusal naming its own fix, once, the first time it runs -against the new rule. - -## Where to go next - -- [Consistency Model](./consistency-model.md) — visibility tiers, revision - counters, and the rest of the read/write contract this page's - read-time addressing rule. diff --git a/docs/guides/namespace-migration.md b/docs/guides/namespace-migration.md deleted file mode 100644 index fad3c766..00000000 --- a/docs/guides/namespace-migration.md +++ /dev/null @@ -1,99 +0,0 @@ ---- -title: Migrating to 9.0 — your fields and system fields -slug: guides/namespace-migration -public: true -category: guides -template: guide -order: 1 -description: The simple story of the 9.0 field-addressing change and the mechanical checklist for updating your call sites — every miss fails loudly with the fix in the error. -next: - - concepts/field-addressing ---- - -# Migrating to 9.0 — your fields and system fields - -The one-sentence version: **your data's field names are now completely -yours, the engine's own fields all live behind one `system.` prefix, and -nothing in between can silently go wrong anymore.** - -## What changed, simply - -**1. Any field name just works.** Before 9.0 the engine quietly owned -certain names. A field called `level` could be shadowed by the engine's -internal index layer of the same name (sorts silently returned insertion -order); names like `confidence` or `subtype` were rejected inside -`metadata`; names like `content` or `id` were silently never indexed, so -filtering on them returned nothing. All of that is gone. Any name — -`level`, `confidence`, `type`, `id`, `content`, anything — is stored -exactly as written and works with every feature: filtering, sorting, -grouping, aggregation, search, and time-travel reads. - -**2. The engine's fields moved behind `system.`.** The engine still keeps -its own per-record bookkeeping — creation time, type, confidence, and so -on. Those are reached one way only now: spelled out, e.g. -`system.createdAt`, `system.type`. They are just as queryable and sortable -as before. `orderBy: 'createdAt'` means *your* field named `createdAt`; -`orderBy: 'system.createdAt'` means the engine's timestamp. No guessing, -no priority rules. - -**3. Storage keeps the two physically separate.** New records store your -metadata in its own nested compartment, so a user field named -`confidence` and the engine's confidence live side by side, both intact, -through restarts, index rebuilds, and `asOf()` history. Old records stay -readable forever; nothing rewrites your data. - -**4. Mistakes are loud.** An ambiguous or unknown field name is a typed -error naming the fix. Unimplemented options refuse instead of being -ignored. The only forbidden name in your metadata is one literally -starting with `system.`. - -## The mechanical checklist - -Every missed site fails **loudly** with the correction in the error -message — nothing silently changes meaning. Sweep these patterns: - -| Before (8.x) | After (9.0) | -|---|---| -| `orderBy: 'createdAt'` (meaning the engine timestamp) | `orderBy: 'system.createdAt'` | -| `where: { subtype: 'invoice' }` (the engine subtype) | `where: { 'system.subtype': 'invoice' }` | -| `where: { confidence: { greaterThan: 0.8 } }` (the engine scalar) | `where: { 'system.confidence': { greaterThan: 0.8 } }` | -| `groupBy: ['noun']` or `groupBy: ['type']` | `groupBy: ['system.type']` | -| `where: { visibility: 'internal' }` / `{ service: … }` (engine values) | `'system.visibility'` / `'system.service'` | -| `metadata: { confidence: 0.9 }` expecting a throw or a lift to the engine scalar | it is YOUR field now — set the engine scalar via the `confidence` param | -| `new Brainy({ reservedFieldPolicy: … })` | remove the option (it throws with this note) | -| `find({ cursor })` / `includeRelations` / `writeOnly` | refuse with `UnsupportedFindOptionError` — they were silently ignored before | - -If a bare name in a query was genuinely *your* field all along (`orderBy: -'score'`, `where: { status: 'active' }`), **change nothing** — bare names -mean your fields, always. - -## What happens at first open - -Each existing database rebuilds its derived indexes once, automatically, -at the first open on 9.0 (index epoch 3 — the index keys split the two -namespaces). One-time cost, observable via `getIndexStatus()`; no manual -step, and your stored data is not modified. - -## For tooling and raw-record readers - -If you read raw stored records (fact-log scanners, export tooling), use -the exported shape-aware splitters — they handle both record eras: - -```typescript -import { splitNounMetadataRecord } from '@soulcraft/brainy' -const { reserved, custom } = splitNounMetadataRecord(rawRecord) -// reserved = engine fields · custom = the user's bag, ANY names -``` - -Feature detection (never version-sniff): - -```typescript -import * as brainy from '@soulcraft/brainy' -const lawActive = 'FIELD_ADDRESSING_CAPABILITY' in brainy // 'field-addressing/v1' -``` - -## Where to go next - -- [Field addressing](../concepts/field-addressing.md) — the full contract: - the ten system scalars, the relation mirror, refusal semantics, and the - cross-engine ordering guarantees. diff --git a/package-lock.json b/package-lock.json index af338ad8..29be914a 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "@soulcraft/brainy", - "version": "9.0.0", + "version": "8.11.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@soulcraft/brainy", - "version": "9.0.0", + "version": "8.11.0", "license": "MIT", "dependencies": { "@msgpack/msgpack": "^3.1.2", diff --git a/package.json b/package.json index f4458a1d..cfb05486 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@soulcraft/brainy", - "version": "9.0.0", + "version": "8.11.0", "description": "Universal Knowledge Protocol™ - World's first Triple Intelligence database unifying vector, graph, and document search in one API. Stage 3 CANONICAL: 42 nouns × 127 verbs covering 96-97% of all human knowledge.", "main": "dist/index.js", "module": "dist/index.js", diff --git a/scripts/release.sh b/scripts/release.sh index b386e580..c64d6c5e 100755 --- a/scripts/release.sh +++ b/scripts/release.sh @@ -189,7 +189,7 @@ echo -e "${GREEN}✅ Pushed to origin${NC}\n" # the forge/npmjs pair enough to publish the storefront leg. FORGE_NPM_REG="https://source.soulcraft.com/api/packages/soulcraft/npm/" FORGE_POLL_INTERVAL_S=15 -FORGE_POLL_MAX_ATTEMPTS=80 # 80 × 15s = 20 minutes — the runner is sequential; the publish run queues behind ci.yml jobs +FORGE_POLL_MAX_ATTEMPTS=40 # 40 × 15s = 10 minutes echo -e "${BLUE}9️⃣ Waiting for CI to publish v${NEW_VERSION} to the forge registry (home)...${NC}" FORGE_LANDED=false for ((attempt = 1; attempt <= FORGE_POLL_MAX_ATTEMPTS; attempt++)); do @@ -212,28 +212,10 @@ else fi echo -e "${BLUE}9️⃣½ Publishing to npmjs (storefront, dist-tag: ${NPM_TAG})...${NC}" -# BYTE-IDENTITY LAW: the storefront republishes CI's EXACT artifact — download -# the tarball the forge serves and publish that file, never a fresh local pack -# (a local rebuild can differ byte-wise, and the fleet verifies the pair by -# shasum across registries). -STOREFRONT_TMP="$(mktemp -d)" -(cd "$STOREFRONT_TMP" && npm pack "@soulcraft/brainy@${NEW_VERSION}" "--@soulcraft:registry=${FORGE_NPM_REG}" >/dev/null) -FORGE_TARBALL="$(ls "$STOREFRONT_TMP"/soulcraft-brainy-*.tgz)" -echo -e "${BLUE} forge artifact: $(sha256sum "$FORGE_TARBALL" | cut -d' ' -f1)${NC}" -npm publish "$FORGE_TARBALL" --tag "$NPM_TAG" "--@soulcraft:registry=https://registry.npmjs.org/" -rm -rf "$STOREFRONT_TMP" +npm publish --tag "$NPM_TAG" "--@soulcraft:registry=https://registry.npmjs.org/" # Brainy is the only PUBLIC @soulcraft package — verify visibility after every publish. npm access get status @soulcraft/brainy "--@soulcraft:registry=https://registry.npmjs.org/" || true -# Verify the pair is byte-identical by registry-reported shasum — divergence here -# means the storefront leg must be treated as failed, loudly. -FORGE_SHA=$(npm view "@soulcraft/brainy@${NEW_VERSION}" dist.shasum "--@soulcraft:registry=${FORGE_NPM_REG}" 2>/dev/null || echo "forge-unavailable") -NPMJS_SHA=$(npm view "@soulcraft/brainy@${NEW_VERSION}" dist.shasum "--@soulcraft:registry=https://registry.npmjs.org/" 2>/dev/null || echo "npmjs-unavailable") -if [ "$FORGE_SHA" = "$NPMJS_SHA" ]; then - echo -e "${GREEN}✅ Published to npmjs — byte-identical pair (shasum ${NPMJS_SHA})${NC}\n" -else - echo -e "${RED}❌ REGISTRY DIVERGENCE: forge shasum ${FORGE_SHA} != npmjs shasum ${NPMJS_SHA} — investigate before announcing${NC}\n" - exit 1 -fi +echo -e "${GREEN}✅ Published to npmjs${NC}\n" # Step 11: Release object on the forge (presentational — the tag, CHANGELOG, # and RELEASES.md are the record; this just gives the forge UI a release page). diff --git a/src/aggregation/AggregationIndex.ts b/src/aggregation/AggregationIndex.ts index ca44ac8b..f9382218 100644 --- a/src/aggregation/AggregationIndex.ts +++ b/src/aggregation/AggregationIndex.ts @@ -14,22 +14,7 @@ */ import type { StorageAdapter, HNSWNounWithMetadata } from '../coreTypes.js' -import { parseFieldAddress, readEntityFieldAddress } from '../db/fieldAddressing.js' -import type { HNSWNounWithMetadata as AddressedEntity } from '../coreTypes.js' - -/** - * Read a user-supplied field name under the one addressing law (sealed - * 2026-08-03): bare / `metadata.` = the user's metadata field, `system.` = - * the ruled engine scalar, malformed = typed refusal. The aggregation engine - * NEVER resolves names any other way — the pre-law resolver made bare - * `subtype`/`confidence` read engine scalars, silently shadowing user fields. - */ -function readAddressed(e: unknown, name: string): unknown { - return readEntityFieldAddress( - e as AddressedEntity, - parseFieldAddress(name, 'entity') - ) -} +import { resolveEntityField } from '../coreTypes.js' import type { AggregateDefinition, AggregateGroupState, @@ -110,15 +95,11 @@ function matchesSource(entity: Record, source: AggregateDefinit // live in the custom bag, so those filters could never match anything. if (source.where && Object.keys(source.where).length > 0) { const e = entity as unknown as HNSWNounWithMetadata - for (const [key, condition] of Object.entries(source.where)) { - // Evaluate ONE field at a time under a neutral key: the address may be - // dotted ('system.subtype'), and the filter evaluator would otherwise - // walk dots as a nested path instead of treating the key as an address. - const value = readAddressed(e, key) - if (!matchesMetadataFilter({ v: value }, { v: condition } as Record)) { - return false - } + const resolved: Record = {} + for (const key of Object.keys(source.where)) { + resolved[key] = resolveEntityField(e, key) } + if (!matchesMetadataFilter(resolved, source.where)) return false } return true @@ -148,11 +129,11 @@ function computeGroupKeys( for (const dim of groupBy) { if (typeof dim === 'string') { - const val = readAddressed(e, dim) + const val = resolveEntityField(e, dim) const v = val !== undefined && val !== null ? String(val) : '__null__' for (const k of keys) k[dim] = v } else if ('unnest' in dim) { - const val = readAddressed(e, dim.field) + const val = resolveEntityField(e, dim.field) const raw = Array.isArray(val) ? val : val !== undefined && val !== null ? [val] : [] // Distinct elements: an entity with duplicate tags counts once per distinct tag. const elems = Array.from(new Set(raw.map(x => String(x)))) @@ -164,7 +145,7 @@ function computeGroupKeys( keys = next } else { // Time-windowed field - const val = readAddressed(e, dim.field) + const val = resolveEntityField(e, dim.field) const v = typeof val === 'number' ? bucketTimestamp(val, dim.window) : '__null__' for (const k of keys) k[dim.field] = v } @@ -193,7 +174,7 @@ function computeGroupKey( * in metadata are both handled in one place. */ function getNumericField(entity: Record, field: string): number | undefined { - const val = readAddressed(entity as unknown as HNSWNounWithMetadata, field) + const val = resolveEntityField(entity as unknown as HNSWNounWithMetadata, field) if (typeof val === 'number' && !isNaN(val)) return val if (typeof val === 'string') { const num = parseFloat(val) @@ -1009,7 +990,7 @@ export class AggregationIndex { // distinctCount tracks distinct values of ANY type (strings, numbers, booleans), // keyed by their string form — NOT numeric-coerced, since its primary use is // categorical (distinct categories / users / tags), not numeric columns. - const raw = readAddressed(entity as unknown as HNSWNounWithMetadata, metricDef.field!) + const raw = resolveEntityField(entity as unknown as HNSWNounWithMetadata, metricDef.field!) if (raw !== undefined && raw !== null) { if (!state.valueCounts) state.valueCounts = {} const key = String(raw) @@ -1053,7 +1034,7 @@ export class AggregationIndex { state.count = Math.max(0, state.count - 1) state.sum = Math.max(0, state.sum - 1) } else if (metricDef.op === 'distinctCount') { - const raw = readAddressed(entity as unknown as HNSWNounWithMetadata, metricDef.field!) + const raw = resolveEntityField(entity as unknown as HNSWNounWithMetadata, metricDef.field!) if (raw !== undefined && raw !== null && state.valueCounts) { const key = String(raw) const c = state.valueCounts[key] diff --git a/src/brainy.ts b/src/brainy.ts index 600e4474..007cdb32 100644 --- a/src/brainy.ts +++ b/src/brainy.ts @@ -148,9 +148,7 @@ import { import { NounType, VerbType, TypeUtils } from './types/graphTypes.js' import { splitNounMetadataRecord, - splitVerbMetadataRecord, - buildNounMetadataRecord, - buildVerbMetadataRecord + splitVerbMetadataRecord } from './types/reservedFields.js' import { BrainyInterface } from './types/brainyInterface.js' import type { IntegrationHub, IntegrationHubConfig } from './integrations/core/IntegrationHub.js' @@ -748,21 +746,6 @@ export class Brainy implements BrainyInterface { private lazyRebuildPromise: Promise | null = null constructor(config?: BrainyConfig) { - // The reserved-field write policy died with the field-addressing law: - // every metadata name is the user's now (engine scalars write via their - // dedicated params and read at `system.*`), so there is nothing left for - // the policy to govern. A config still passing it refuses loudly rather - // than being silently ignored. - if (config && 'reservedFieldPolicy' in (config as Record)) { - throw new Error( - `reservedFieldPolicy was removed by the field-addressing law: metadata field ` + - `names are never reserved anymore — every name in the metadata bag is the ` + - `user's and works like any other field. Set engine scalars via their ` + - `dedicated params (confidence, weight, subtype, …) and query them as ` + - `system.. Remove the reservedFieldPolicy option.` - ) - } - // Normalize configuration with defaults this.config = this.normalizeConfig(config) @@ -2035,6 +2018,12 @@ export class Brainy implements BrainyInterface { // Zero-config validation (static import for performance) validateAddParams(params) + // Reserved fields arriving via the metadata bag (untyped callers — the + // compile-time guard stops TypeScript callers) are normalized to their + // canonical top-level location BEFORE any enforcement runs, so a + // remapped subtype participates in subtype-pairing enforcement and the + // indexed metadata bag carries only custom fields. + params = this.remapReservedAddMetadata(params) // Tracked-field vocabulary enforcement (Layer 2). Walks both bags so a // tracked field declared at top level (e.g. 'subtype') and one declared in @@ -2106,46 +2095,39 @@ export class Brainy implements BrainyInterface { ) } - // Prepare metadata for storage: a v2 nested-bag record — engine fields - // top-level, the user's bag nested VERBATIM (any name, including engine - // spellings like `confidence` or `type`, is the user's and survives - // faithfully; the field-addressing law). - const storageMetadata = buildNounMetadataRecord( - { - data: params.data, - noun: params.type, - ...(params.subtype !== undefined && { subtype: params.subtype }), - // visibility: stored only when not 'public' (absent === public, keeps records lean) - ...(params.visibility !== undefined && - params.visibility !== 'public' && { visibility: params.visibility }), - service: params.service, - createdAt: Date.now(), - updatedAt: Date.now(), - _rev: 1, - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.weight !== undefined && { weight: params.weight }), - ...(params.createdBy && { createdBy: params.createdBy }) - }, - { - ...params.metadata, - // Preserve the caller's original (non-UUID) id when normalized, so reads - // can surface it. A real UUID passes through with no _originalId. - ...(originalId !== undefined && { [ORIGINAL_ID_KEY]: originalId }) - } - ) + // Prepare metadata for storage + // data is stored opaquely in the 'data' field - NOT spread into top-level metadata. + // Only metadata fields are queryable via find({ where }). + const storageMetadata = { + ...params.metadata, + // Preserve the caller's original (non-UUID) id when normalized, so reads + // can surface it. A real UUID passes through with no _originalId. + ...(originalId !== undefined && { [ORIGINAL_ID_KEY]: originalId }), + data: params.data, + noun: params.type, + ...(params.subtype !== undefined && { subtype: params.subtype }), + // visibility: stored only when not 'public' (absent === public, keeps records lean) + ...(params.visibility !== undefined && + params.visibility !== 'public' && { visibility: params.visibility }), + service: params.service, + createdAt: Date.now(), + updatedAt: Date.now(), + _rev: 1, + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.weight !== undefined && { weight: params.weight }), + ...(params.createdBy && { createdBy: params.createdBy }) + } // Build entity structure for indexing (NEW - with top-level fields) // Optional fields must use conditional spreading to match storageMetadata exactly. // If undefined values are included as explicit keys, extractIndexableFields indexes // them as '__NULL__' entries that removeFromIndex can never clean up (storageMetadata // omits those keys entirely via conditional spreading, so the fields don't match). - // No `level` here: engine plumbing never enters the indexing view — a - // hardcoded level:0 landed in the SAME flattened index column as user - // metadata named `level`, poisoning it multi-valued ([0, real]). const entityForIndexing = { id, vector, connections: new Map(), + level: 0, type: params.type, ...(params.subtype !== undefined && { subtype: params.subtype }), ...(params.visibility !== undefined && @@ -2643,6 +2625,320 @@ export class Brainy implements BrainyInterface { return entity } + /** One-shot registry for reserved-field warnings (per process, per method+field). */ + private static warnedReservedFields = new Set() + + /** + * @description Resolve the human-readable "correct write path" guidance for a + * reserved field on a given write method. Single source of truth shared by the + * `'throw'` (Error message) and `'warn'` (one-shot warning) paths so the two + * never drift. The trio `confidence` / `weight` / `subtype` and the + * add()/relate()-time fields `service` / `createdBy` / `visibility` map to a + * dedicated param; everything else is system-managed. + * @param method - The public write method the bag arrived through. + * @param field - The reserved field name found in the metadata bag. + * @returns Guidance naming the correct way to set the field. + */ + private reservedWritePath( + method: 'add' | 'update' | 'relate' | 'updateRelation', + field: string + ): string { + const typeParam = "the top-level 'type' param" + switch (field) { + case 'noun': + case 'verb': + return typeParam + case 'data': + return "the top-level 'data' param" + case 'confidence': + return "the 'confidence' param" + case 'weight': + return "the 'weight' param" + case 'subtype': + return "the 'subtype' param" + case 'visibility': + return "the 'visibility' param ('public' | 'internal')" + case 'service': + return method === 'add' + ? "the 'service' param of add()" + : method === 'relate' + ? "the 'service' param of relate()" + : 'nothing — service is fixed at create time' + case 'createdBy': + return method === 'add' + ? "the 'createdBy' param of add()" + : 'nothing — createdBy is system-managed' + case 'createdAt': + return 'nothing — creation time is set automatically' + case 'updatedAt': + return 'nothing — set automatically on every write' + case '_rev': + return method === 'update' + ? "the 'ifRev' param for optimistic concurrency" + : 'nothing — revisions are system-managed' + default: + return 'a dedicated top-level param' + } + } + + /** + * @description Enforce {@link BrainyConfig.reservedFieldPolicy} for reserved + * fields found inside a metadata bag. Called by every write-path remap once + * the bag has been split and at least one reserved key is present. + * + * - `'throw'` (default): throw a clear Error naming every offending key and + * its correct write path. The caller never reaches the remap. + * - `'warn'`: emit a ONE-SHOT (per method+field, per process) warning for + * EVERY reserved key found — both the user-mutable fields that are about to + * be remapped and the system-managed fields that are about to be dropped — + * then fall through to the legacy remap. + * - `'remap'`: silent legacy remap, no warning. + * + * @param method - The public write method the bag arrived through. + * @param reserved - The reserved half of the split metadata bag (non-empty). + * @param reservedListName - `'RESERVED_ENTITY_FIELDS'` or + * `'RESERVED_RELATION_FIELDS'` — named in the thrown Error for discoverability. + * @returns `true` when the caller should proceed with the legacy remap + * (`'warn'` / `'remap'`); `'throw'` never returns (it throws first). + * @throws {Error} When the policy is `'throw'` and any reserved key is present. + */ + private enforceReservedPolicy( + method: 'add' | 'update' | 'relate' | 'updateRelation', + reserved: Partial>, + reservedListName: 'RESERVED_ENTITY_FIELDS' | 'RESERVED_RELATION_FIELDS' + ): boolean { + const policy = this.config.reservedFieldPolicy ?? 'throw' + const keys = Object.keys(reserved) + if (keys.length === 0) return true + + if (policy === 'throw') { + const detail = keys + .map((k) => { + const path = this.reservedWritePath(method, k) + // System-managed fields resolve to a "nothing — …" sentinel; phrase + // those as "is system-managed" rather than "pass it as the nothing". + return path.startsWith('nothing') + ? `metadata.${k} is a reserved field (${path.replace(/^nothing\s*—\s*/, '')}) and cannot be set through ${method}()` + : `metadata.${k} is a reserved field — pass it as ${path} to ${method}()` + }) + .join('; ') + throw new Error( + `${detail} (reserved: see ${reservedListName}). ` + + `Set reservedFieldPolicy:'remap' to opt into legacy remapping, ` + + `or reservedFieldPolicy:'warn' to remap with a warning.` + ) + } + + if (policy === 'warn') { + // One-shot warning for EVERY reserved key (today only system-managed ones + // warn — this closes that gap so user-mutable remaps are visible too). + for (const k of keys) { + this.warnReservedRemapped(method, k, this.reservedWritePath(method, k)) + } + } + + // 'warn' and 'remap' both fall through to the legacy remap. + return true + } + + /** + * @description One-shot (per method+field, per process) warning that a + * reserved field arrived inside a metadata bag under the `'warn'` policy. The + * wording is neutral on "remapped vs dropped" — `reservedWritePath()` already + * tells the caller where the value goes (a dedicated param, or "nothing"). + * @param method - The public write method the bag arrived through. + * @param field - The reserved field name found in the bag. + * @param rightPath - Guidance naming the correct write path. + */ + private warnReservedRemapped(method: string, field: string, rightPath: string): void { + const key = `${method}:${field}` + if (Brainy.warnedReservedFields.has(key)) return + Brainy.warnedReservedFields.add(key) + // System-managed fields resolve to a "nothing — …" sentinel; phrase the + // guidance so it reads cleanly in both the remapped and dropped cases. + const guidance = rightPath.startsWith('nothing') + ? `it is ${rightPath.replace(/^nothing\s*—\s*/, '')} and was dropped` + : `set it via ${rightPath} instead` + prodLog.warn( + `[brainy] ${method}(): '${field}' is a reserved field and was found inside the ` + + `metadata bag — ${guidance}. (Legacy remap applied because ` + + `reservedFieldPolicy is 'warn'. This warning is shown once per field per process.)` + ) + } + + /** + * @description Normalize an `add()` params object with respect to + * Brainy-reserved fields arriving inside `metadata` (untyped callers only — + * the compile-time guard on `AddParams.metadata` stops TypeScript callers). + * Governed by {@link BrainyConfig.reservedFieldPolicy} (default `'throw'`): + * `'throw'` rejects the write naming the offending key(s); `'warn'`/`'remap'` + * fall through to the legacy remap, where fields with a dedicated `add()` + * param (`confidence`, `weight`, `subtype`, `visibility`, `service`, + * `createdBy`) are remapped to that param unless the caller also passed it + * explicitly (top-level wins) and system-managed fields (`noun`, `data`, + * `createdAt`, `updatedAt`, `_rev`) are dropped. A remapped `subtype` flows + * through subtype-pairing enforcement exactly like a top-level one. + * @param params - The caller's add params (not mutated). + * @returns Params with reserved fields normalized out of `metadata`. + * @throws {Error} When `reservedFieldPolicy` is `'throw'` and the bag carries a reserved key. + */ + private remapReservedAddMetadata(params: AddParams): AddParams { + const bag = params.metadata as Record | undefined + if (!bag || typeof bag !== 'object') return params + const { reserved, custom } = splitNounMetadataRecord(bag) + if (Object.keys(reserved).length === 0) return params + + // Policy gate: 'throw' (default) throws here; 'warn' warns once per key then + // remaps; 'remap' silently remaps. (Throw never returns.) + this.enforceReservedPolicy('add', reserved, 'RESERVED_ENTITY_FIELDS') + + const createdBy = reserved.createdBy as { augmentation?: unknown; version?: unknown } | undefined + const createdByValid = + typeof createdBy === 'object' && + createdBy !== null && + typeof createdBy.augmentation === 'string' && + typeof createdBy.version === 'string' + + return { + ...params, + metadata: custom as AddParams['metadata'], + ...(params.confidence === undefined && + typeof reserved.confidence === 'number' && { confidence: reserved.confidence }), + ...(params.weight === undefined && + typeof reserved.weight === 'number' && { weight: reserved.weight }), + ...(params.subtype === undefined && + typeof reserved.subtype === 'string' && { subtype: reserved.subtype }), + ...(params.visibility === undefined && + (reserved.visibility === 'public' || reserved.visibility === 'internal') && { + visibility: reserved.visibility as 'public' | 'internal' + }), + ...(params.service === undefined && + typeof reserved.service === 'string' && { service: reserved.service }), + ...(params.createdBy === undefined && + createdByValid && { createdBy: createdBy as { augmentation: string; version: string } }) + } + } + + /** + * @description Normalize an `update()` params object with respect to + * Brainy-reserved fields arriving inside the metadata patch — the `update()` + * mirror of {@link remapReservedAddMetadata}, closing the historical trap + * where `add({metadata:{confidence}})` lifted the field but + * `update({metadata:{confidence}})` silently dropped it (the patch value + * survived the merge and was then clobbered by the preserve-existing + * spread; a production consumer's confidence-evolution writes no-oped until + * read back). Governed by {@link BrainyConfig.reservedFieldPolicy} (default + * `'throw'`): `'throw'` rejects the write; `'warn'`/`'remap'` remap + * user-mutable fields (`confidence`, `weight`, `subtype`) to their dedicated + * param unless the caller also passed it (top-level wins) and drop everything + * else (`noun`, `data`, `createdAt`, `updatedAt`, `service`, `createdBy`, + * `_rev`) as system-managed or fixed at `add()` time. + * @param params - The caller's update params (not mutated). + * @returns Params with reserved fields normalized out of `metadata`. + * @throws {Error} When `reservedFieldPolicy` is `'throw'` and the bag carries a reserved key. + */ + private remapReservedUpdateMetadata(params: UpdateParams): UpdateParams { + const bag = params.metadata as Record | undefined + if (!bag || typeof bag !== 'object') return params + const { reserved, custom } = splitNounMetadataRecord(bag) + if (Object.keys(reserved).length === 0) return params + + // Policy gate: 'throw' (default) throws; 'warn' warns once per key then + // remaps; 'remap' silently remaps. + this.enforceReservedPolicy('update', reserved, 'RESERVED_ENTITY_FIELDS') + + return { + ...params, + metadata: custom as UpdateParams['metadata'], + ...(params.confidence === undefined && + typeof reserved.confidence === 'number' && { confidence: reserved.confidence }), + ...(params.weight === undefined && + typeof reserved.weight === 'number' && { weight: reserved.weight }), + ...(params.subtype === undefined && + typeof reserved.subtype === 'string' && { subtype: reserved.subtype }) + } + } + + /** + * @description Normalize a `relate()` params object with respect to + * Brainy-reserved fields arriving inside `metadata` — the relationship + * mirror of {@link remapReservedAddMetadata}. Governed by + * {@link BrainyConfig.reservedFieldPolicy} (default `'throw'`): `'throw'` + * rejects the write; `'warn'`/`'remap'` remap fields with a dedicated + * `relate()` param (`confidence`, `weight`, `subtype`, `visibility`, + * `service`) to that param (top-level wins) and drop system-managed fields + * (`verb`, `data`, `createdAt`, `updatedAt`, `createdBy`, `_rev`). + * @param params - The caller's relate params (not mutated). + * @returns Params with reserved fields normalized out of `metadata`. + * @throws {Error} When `reservedFieldPolicy` is `'throw'` and the bag carries a reserved key. + */ + private remapReservedRelateMetadata(params: RelateParams): RelateParams { + const bag = params.metadata as Record | undefined + if (!bag || typeof bag !== 'object') return params + const { reserved, custom } = splitVerbMetadataRecord(bag) + if (Object.keys(reserved).length === 0) return params + + // Policy gate: 'throw' (default) throws; 'warn' warns once per key then + // remaps; 'remap' silently remaps. + this.enforceReservedPolicy('relate', reserved, 'RESERVED_RELATION_FIELDS') + + return { + ...params, + metadata: custom as RelateParams['metadata'], + ...(params.confidence === undefined && + typeof reserved.confidence === 'number' && { confidence: reserved.confidence }), + ...(params.weight === undefined && + typeof reserved.weight === 'number' && { weight: reserved.weight }), + ...(params.subtype === undefined && + typeof reserved.subtype === 'string' && { subtype: reserved.subtype }), + ...(params.visibility === undefined && + (reserved.visibility === 'public' || reserved.visibility === 'internal') && { + visibility: reserved.visibility as 'public' | 'internal' + }), + ...(params.service === undefined && + typeof reserved.service === 'string' && { service: reserved.service }) + } + } + + /** + * @description Normalize an `updateRelation()` params object with respect + * to Brainy-reserved fields arriving inside the metadata patch — the + * relationship mirror of {@link remapReservedUpdateMetadata}. Governed by + * {@link BrainyConfig.reservedFieldPolicy} (default `'throw'`): `'throw'` + * rejects the write; `'warn'`/`'remap'` remap user-mutable fields + * (`confidence`, `weight`, `subtype`, `visibility`) to their dedicated param + * (top-level wins) and drop everything else. + * @param params - The caller's update-relation params (not mutated). + * @returns Params with reserved fields normalized out of `metadata`. + * @throws {Error} When `reservedFieldPolicy` is `'throw'` and the bag carries a reserved key. + */ + private remapReservedUpdateRelationMetadata( + params: UpdateRelationParams + ): UpdateRelationParams { + const bag = params.metadata as Record | undefined + if (!bag || typeof bag !== 'object') return params + const { reserved, custom } = splitVerbMetadataRecord(bag) + if (Object.keys(reserved).length === 0) return params + + // Policy gate: 'throw' (default) throws; 'warn' warns once per key then + // remaps; 'remap' silently remaps. + this.enforceReservedPolicy('updateRelation', reserved, 'RESERVED_RELATION_FIELDS') + + return { + ...params, + metadata: custom as UpdateRelationParams['metadata'], + ...(params.confidence === undefined && + typeof reserved.confidence === 'number' && { confidence: reserved.confidence }), + ...(params.weight === undefined && + typeof reserved.weight === 'number' && { weight: reserved.weight }), + ...(params.subtype === undefined && + typeof reserved.subtype === 'string' && { subtype: reserved.subtype }), + ...(params.visibility === undefined && + (reserved.visibility === 'public' || reserved.visibility === 'internal') && { + visibility: reserved.visibility as 'public' | 'internal' + }) + } + } /** * Update an existing entity @@ -2708,6 +3004,12 @@ export class Brainy implements BrainyInterface { // Reserved fields arriving via the metadata patch are remapped to their // canonical top-level location, mirroring add()'s lift. Without this the // patch value survived the merge but was then clobbered by the + // preserve-existing spreads below — a silent no-op consumers could only + // detect by reading values back. User-mutable fields (confidence, + // weight, subtype) remap unless the same field was also passed top-level + // (top-level wins); system-managed fields are dropped with a one-shot + // warning naming the right path. + params = this.remapReservedUpdateMetadata(params) // Tracked-field vocabulary enforcement (Layer 2). Same as add() — the // metadata bag carries fields registered via trackField(), and subtype is @@ -2774,41 +3076,38 @@ export class Brainy implements BrainyInterface { ? { ...existing.metadata, ...params.metadata } : params.metadata || existing.metadata - // Prepare the updated v2 nested-bag record: engine fields top-level, - // the merged user bag nested verbatim (collider names stay the user's). - const updatedMetadata = buildNounMetadataRecord( - { - data: params.data !== undefined ? params.data : existing.data, - noun: params.type || existing.type, - service: existing.service, - createdAt: existing.createdAt, - updatedAt: Date.now(), - _rev: currentRev + 1, - // Update confidence and weight if provided, otherwise preserve existing - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.weight !== undefined && { weight: params.weight }), - ...(params.confidence === undefined && existing.confidence !== undefined && { confidence: existing.confidence }), - ...(params.weight === undefined && existing.weight !== undefined && { weight: existing.weight }), - // Update subtype if provided, otherwise preserve existing - ...(params.subtype !== undefined && { subtype: params.subtype }), - ...(params.subtype === undefined && existing.subtype !== undefined && { subtype: existing.subtype }), - // Visibility: take the new value if provided, else preserve existing. Stored only - // when the effective value is not 'public' (absent === public, keeps records lean). - // A change to 'public' therefore drops the field entirely. - ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { - visibility: params.visibility ?? existing.visibility - }) - }, - newMetadata as Record - ) + // Prepare updated metadata object + // data is stored opaquely in the 'data' field - NOT spread into top-level metadata. + const updatedMetadata = { + ...newMetadata, + data: params.data !== undefined ? params.data : existing.data, + noun: params.type || existing.type, + service: existing.service, + createdAt: existing.createdAt, + updatedAt: Date.now(), + _rev: currentRev + 1, + // Update confidence and weight if provided, otherwise preserve existing + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.weight !== undefined && { weight: params.weight }), + ...(params.confidence === undefined && existing.confidence !== undefined && { confidence: existing.confidence }), + ...(params.weight === undefined && existing.weight !== undefined && { weight: existing.weight }), + // Update subtype if provided, otherwise preserve existing + ...(params.subtype !== undefined && { subtype: params.subtype }), + ...(params.subtype === undefined && existing.subtype !== undefined && { subtype: existing.subtype }), + // Visibility: take the new value if provided, else preserve existing. Stored only + // when the effective value is not 'public' (absent === public, keeps records lean). + // A change to 'public' therefore drops the field entirely. + ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { + visibility: params.visibility ?? existing.visibility + }) + } - // Build entity structure for metadata index (with top-level fields). - // No `level`: engine plumbing never enters the indexing view (it - // poisoned the flattened user `level` column — VENUE-BRAINY-ORDERBY-NOOP). + // Build entity structure for metadata index (with top-level fields) const entityForIndexing = { id: params.id, vector, connections: new Map(), + level: 0, type: params.type || existing.type, subtype: params.subtype !== undefined ? params.subtype : existing.subtype, ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { @@ -2862,23 +3161,18 @@ export class Brainy implements BrainyInterface { new UpdateNounMetadataOperation(this.storage, params.id, updatedMetadata) ) - // Operations 2-4: vector-record write + HNSW reindex — ONLY when the - // vector side actually changed (new data/vector/type). A metadata-only - // update must never rewrite the noun record: the record carries the - // full vector, so an unconditional save turned every metadata touch - // into a whole-vector rewrite + fsync — under a read-heavy consumer - // sweep that bumps per-entity stats, this amplified into disk - // saturation on a production deployment (SELF-ENGINE-RESTART-GRIND, - // 2026-07-29: 5.8GB written in 40min from ~50 recalls/min). + // Operation 2: Update vector data (will use updated type cache) + tx.addOperation( + new SaveNounOperation(this.storage, { + id: params.id, + vector, + connections: new Map(), + level: 0 + }) + ) + + // Operation 3-4: Update HNSW index (remove and re-add if reindexing needed) if (needsReindexing) { - tx.addOperation( - new SaveNounOperation(this.storage, { - id: params.id, - vector, - connections: new Map(), - level: 0 - }) - ) tx.addOperation( new RemoveFromVectorIndexOperation(this.index, params.id, existing.vector) ) @@ -3741,6 +4035,9 @@ export class Brainy implements BrainyInterface { // engine-minted UUID — relation ids are never caller-supplied here.) params = { ...params, from: resolveEntityId(params.from), to: resolveEntityId(params.to) } + // Reserved fields arriving via the metadata bag are normalized to their + // canonical top-level params before enforcement — mirror of add()'s lift. + params = this.remapReservedRelateMetadata(params) // Subtype pairing enforcement (Layer 3 — 7.30.0). Per-type rules registered // via brain.requireSubtype() compose with the brain-wide strict-mode flag. @@ -3792,28 +4089,25 @@ export class Brainy implements BrainyInterface { (v, i) => (v + toEntity.vector[i]) / 2 ) - // Prepare verb metadata: a v2 nested-bag record — engine fields - // top-level, the user's edge bag nested verbatim (any name is the - // user's; the field-addressing law). + // Prepare verb metadata + // User metadata spread FIRST, then system fields ALWAYS win (prevents collision) // One timestamp for both createdAt and updatedAt so a never-updated edge reports a // stable updatedAt (=== createdAt) instead of a fresh Date.now() fabricated per read. const relateTs = Date.now() - const verbMetadata = buildVerbMetadataRecord( - { - verb: params.type, - ...(params.subtype !== undefined && { subtype: params.subtype }), - // visibility: stored only when not 'public' (absent === public, keeps records lean) - ...(params.visibility !== undefined && - params.visibility !== 'public' && { visibility: params.visibility }), - weight: params.weight ?? 1.0, - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.service !== undefined && { service: params.service }), - createdAt: relateTs, - updatedAt: relateTs, - ...(params.data !== undefined && { data: params.data }) - }, - (params.metadata as Record) || {} - ) + const verbMetadata = { + ...(params.metadata || {}), + verb: params.type, + ...(params.subtype !== undefined && { subtype: params.subtype }), + // visibility: stored only when not 'public' (absent === public, keeps records lean) + ...(params.visibility !== undefined && + params.visibility !== 'public' && { visibility: params.visibility }), + weight: params.weight ?? 1.0, + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.service !== undefined && { service: params.service }), + createdAt: relateTs, + updatedAt: relateTs, + ...(params.data !== undefined && { data: params.data }) + } // Save to storage (vector and metadata separately) const verb: GraphVerb = { @@ -4045,6 +4339,9 @@ export class Brainy implements BrainyInterface { validateUpdateRelationParams(params) + // Reserved fields arriving via the metadata patch are remapped to their + // canonical top-level params — mirror of update()'s normalization. + params = this.remapReservedUpdateRelationMetadata(params) const existing = await this.storage.getVerb(params.id) if (!existing) { @@ -4073,36 +4370,32 @@ export class Brainy implements BrainyInterface { ? { ...(existingRec.metadata || {}), ...(params.metadata || {}) } : params.metadata || existingRec.metadata - // Build the updated stored record: v2 nested-bag — engine fields - // top-level, the merged user bag nested verbatim (mirror of update()). - const updatedWeight = params.weight ?? existingRec.weight ?? 1.0 - const updatedData = - params.data !== undefined ? params.data : existingRec.data - const updatedMetadata = buildVerbMetadataRecord( - { - verb: newVerbType, - ...(params.subtype !== undefined - ? { subtype: params.subtype } - : existingRec.subtype !== undefined && { subtype: existingRec.subtype }), - // Visibility: new value if provided, else preserve existing; stored only when the - // effective value is not 'public' (a change to 'public' drops the field). - ...(((params.visibility ?? existingRec.visibility) ?? 'public') !== 'public' && { - visibility: params.visibility ?? existingRec.visibility - }), - weight: updatedWeight, - ...(params.confidence !== undefined - ? { confidence: params.confidence } - : existingRec.confidence !== undefined && { confidence: existingRec.confidence }), - // service/createdBy are fixed at relate() time — always carried forward - // (omitting them here silently erased them on every updateRelation()). - ...(existingRec.service !== undefined && { service: existingRec.service }), - ...(existingRec.createdBy !== undefined && { createdBy: existingRec.createdBy }), - createdAt: existingRec.createdAt, - updatedAt: Date.now(), - ...(updatedData !== undefined && { data: updatedData }) - }, - newMetadata as Record - ) + // Build updated stored metadata. System fields ALWAYS win — same shape as relate(). + const updatedMetadata = { + ...newMetadata, + verb: newVerbType, + ...(params.subtype !== undefined + ? { subtype: params.subtype } + : existingRec.subtype !== undefined && { subtype: existingRec.subtype }), + // Visibility: new value if provided, else preserve existing; stored only when the + // effective value is not 'public' (a change to 'public' drops the field). + ...(((params.visibility ?? existingRec.visibility) ?? 'public') !== 'public' && { + visibility: params.visibility ?? existingRec.visibility + }), + weight: params.weight ?? existingRec.weight ?? 1.0, + ...(params.confidence !== undefined + ? { confidence: params.confidence } + : existingRec.confidence !== undefined && { confidence: existingRec.confidence }), + // service/createdBy are fixed at relate() time — always carried forward + // (omitting them here silently erased them on every updateRelation()). + ...(existingRec.service !== undefined && { service: existingRec.service }), + ...(existingRec.createdBy !== undefined && { createdBy: existingRec.createdBy }), + createdAt: existingRec.createdAt, + updatedAt: Date.now(), + ...(params.data !== undefined + ? { data: params.data } + : existingRec.data !== undefined && { data: existingRec.data }) + } // Build the verb view used by the graph index — top-level fields mirror relate()'s. const verbForIndex: GraphVerb = { @@ -4118,9 +4411,9 @@ export class Brainy implements BrainyInterface { ...(((params.visibility ?? existingRec.visibility) ?? 'public') !== 'public' && { visibility: params.visibility ?? existingRec.visibility }), - weight: updatedWeight, + weight: updatedMetadata.weight, metadata: newMetadata, - data: updatedData, + data: updatedMetadata.data, createdAt: existingRec.createdAt } @@ -5318,7 +5611,7 @@ export class Brainy implements BrainyInterface { this._aggregationIndex!.defineAggregate({ name: aggregateName, source: {}, - groupBy: perType ? [name, 'system.type'] : [name], + groupBy: perType ? [name, 'noun'] : [name], metrics: { count: { op: 'count' } } }) } @@ -5329,12 +5622,7 @@ export class Brainy implements BrainyInterface { * and `counts.byField()` agree on the convention. */ private fieldCountsAggregateName(name: string): string { - // v2 suffix: the per-type dimension moved from the legacy 'noun' alias to - // 'system.type' under the addressing law — a NEW name makes the ensure - // block re-define and BACKFILL from canonical instead of silently serving - // the old-dim definition (whose 'noun' key now reads user metadata and - // would drift). The v1 rows are derived state, superseded not lost. - return `__fieldCounts_v2__${name}` + return `__fieldCounts__${name}` } /** @@ -5726,12 +6014,8 @@ export class Brainy implements BrainyInterface { ): Promise> { const excluded = this.excludedVisibilityTiers(params) if (!excluded) return new Set() - // 'system.visibility' — the engine scalar's frozen address. A bare - // 'visibility' key would address the USER's metadata bag under the - // field-addressing law and silently hide nothing (VFS/system entities - // would leak into every default read). const ids = await this.metadataIndex.getIdsForFilter({ - 'system.visibility': excluded.length === 1 ? excluded[0] : { oneOf: excluded } + visibility: excluded.length === 1 ? excluded[0] : { oneOf: excluded } }) return new Set(ids) } @@ -5856,18 +6140,20 @@ export class Brainy implements BrainyInterface { // Build filter for metadata index let filter: any = {} if (params.where) { - // Where keys pass through UNTOUCHED — the addressing law parses - // them at the index boundary. The old where.type→noun alias is - // dead: bare 'type' is the user's own field now. Object.assign(filter, params.where) + // Alias: where.type → where.noun (storage field name for entity type) + if ('type' in filter && !('noun' in filter)) { + filter.noun = filter.type + delete filter.type + } } - if (params.service) filter['system.service'] = params.service + if (params.service) filter.service = params.service // Subtype (top-level standard field — fast path, not metadata fallback). // Must be assigned BEFORE the type-array expansion below so the spread // into each anyOf branch carries it through. if (params.subtype !== undefined) { - filter['system.subtype'] = Array.isArray(params.subtype) + filter.subtype = Array.isArray(params.subtype) ? { oneOf: params.subtype } : params.subtype } @@ -5875,11 +6161,11 @@ export class Brainy implements BrainyInterface { if (params.type) { const types = Array.isArray(params.type) ? params.type : [params.type] if (types.length === 1) { - filter['system.type'] = types[0] + filter.noun = types[0] } else { filter = { anyOf: types.map(type => ({ - 'system.type': type, + noun: type, ...filter })) } @@ -8061,27 +8347,6 @@ export class Brainy implements BrainyInterface { return this.generationStore.compact(options) } - /** - * @description Repack cold generation history into sealed segments — - * re-representation, never deletion: every record and delta stays readable - * (`asOf()` unchanged); the physical file count drops by orders of - * magnitude. Runs automatically (time-bounded) at `close()`; call this for - * explicit maintenance windows on long-lived writers. The ONLY history - * transform permitted under the archival profile (`retention: 'all'`). - * @param options - `timeBudgetMs` bounds the pass (early stop = consistent - * prefix, next pass resumes); `batchGenerations` sizes each fold. - * @returns Folded generation count and segments created. - */ - async repackHistory(options?: { - timeBudgetMs?: number - batchGenerations?: number - }): Promise<{ foldedGenerations: number; segmentsCreated: number }> { - this.assertWritable('repackHistory') - await this.ensureInitialized() - await this.generationStore.flushPendingSingleOps() - return this.generationStore.repackHistory(options) - } - /** * @description Read-only generational-history footprint for fleet audits: * generation count, total on-disk bytes, generation/timestamp range, the @@ -8109,24 +8374,6 @@ export class Brainy implements BrainyInterface { } } - /** - * @description A deterministic content digest of the generation log through - * `g` (D8 — gate-to-generation provenance): identical history produces the - * identical digest on any machine; divergence produces a different one. - * Release gates and suite verdicts pin `{generation, digest}` and verify - * both at execution time instead of pinning a git commit. O(segments + - * live-tier window), never O(all generations). Throws `RangeError` out of - * range and `GenerationCompactedError` below the horizon — a gate can - * never silently pin reclaimed history. - * @example - * const gate = { generation: brain.generation(), digest: await brain.generationDigest(brain.generation()) } - */ - async generationDigest(g: number): Promise { - await this.ensureInitialized() - await this.generationStore.flushPendingSingleOps() - return this.generationStore.generationDigest(g) - } - /** * @description Drive the adaptive retention byte budget at runtime — the * settable input a machine-level coordinator (e.g. cor's `ResourceManager`, @@ -8985,7 +9232,10 @@ export class Brainy implements BrainyInterface { ): Promise { const { op: _discriminator, ...rawParams } = op validateAddParams(rawParams as AddParams) - const params = rawParams as AddParams + // Same reserved-field normalization as add() — the metadata bag is + // cleaned BEFORE enforcement so a remapped subtype participates in + // subtype-pairing enforcement and only custom fields reach the index. + const params = this.remapReservedAddMetadata(rawParams as AddParams) this.enforceTrackedFieldValues(params.metadata as Record | undefined, 'metadata') this.enforceTrackedFieldValues({ subtype: params.subtype } as Record, 'top-level') this.enforceSubtypeOnAdd('add', params.type, params.subtype, params.metadata) @@ -9060,36 +9310,30 @@ export class Brainy implements BrainyInterface { plan.createdNouns.add(id) const now = Date.now() - // v2 nested-bag record — mirror of add(): engine fields top-level, the - // user's bag nested verbatim (collider names stay the user's). - const storageMetadata = buildNounMetadataRecord( - { - data: params.data, - noun: params.type, - ...(params.subtype !== undefined && { subtype: params.subtype }), - // visibility: stored only when not 'public' (absent === public, keeps records lean) - ...(params.visibility !== undefined && - params.visibility !== 'public' && { visibility: params.visibility }), - service: params.service, - createdAt: now, - updatedAt: now, - _rev: 1, - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.weight !== undefined && { weight: params.weight }), - ...(params.createdBy && { createdBy: params.createdBy }) - }, - { - ...params.metadata, - // Preserve the caller's original (non-UUID) id when normalized — mirror - // of add(). A real UUID passes through with no _originalId. - ...(originalId !== undefined && { [ORIGINAL_ID_KEY]: originalId }) - } - ) + const storageMetadata = { + ...params.metadata, + // Preserve the caller's original (non-UUID) id when normalized — mirror + // of add(). A real UUID passes through with no _originalId. + ...(originalId !== undefined && { [ORIGINAL_ID_KEY]: originalId }), + data: params.data, + noun: params.type, + ...(params.subtype !== undefined && { subtype: params.subtype }), + // visibility: stored only when not 'public' (absent === public, keeps records lean) + ...(params.visibility !== undefined && + params.visibility !== 'public' && { visibility: params.visibility }), + service: params.service, + createdAt: now, + updatedAt: now, + _rev: 1, + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.weight !== undefined && { weight: params.weight }), + ...(params.createdBy && { createdBy: params.createdBy }) + } const entityForIndexing = { id, vector, connections: new Map(), - // no `level` — plumbing never enters the indexing view + level: 0, type: params.type, ...(params.subtype !== undefined && { subtype: params.subtype }), ...(params.visibility !== undefined && @@ -9147,7 +9391,10 @@ export class Brainy implements BrainyInterface { ): Promise { const { op: _discriminator, ...rawParams } = op validateUpdateParams(rawParams as UpdateParams) - const params = rawParams as UpdateParams + // Same reserved-field normalization as update() — user-mutable fields + // remap to their dedicated param (top-level wins), system-managed fields + // drop with a one-shot warning. + const params = this.remapReservedUpdateMetadata(rawParams as UpdateParams) // Id normalization (8.0) — mirror of update(): a natural key resolves to the // canonical UUID add() stored. A real UUID passes through. params.id = resolveEntityId(params.id) @@ -9199,33 +9446,29 @@ export class Brainy implements BrainyInterface { ? { ...existing.metadata, ...params.metadata } : params.metadata || existing.metadata const now = Date.now() - // v2 nested-bag record — mirror of update(): engine fields top-level, - // the merged user bag nested verbatim. - const updatedMetadata = buildNounMetadataRecord( - { - data: params.data !== undefined ? params.data : existing.data, - noun: params.type || existing.type, - service: existing.service, - createdAt: existing.createdAt, - updatedAt: now, - _rev: currentRev + 1, - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.weight !== undefined && { weight: params.weight }), - ...(params.confidence === undefined && - existing.confidence !== undefined && { confidence: existing.confidence }), - ...(params.weight === undefined && - existing.weight !== undefined && { weight: existing.weight }), - ...(params.subtype !== undefined && { subtype: params.subtype }), - ...(params.subtype === undefined && - existing.subtype !== undefined && { subtype: existing.subtype }), - // Visibility: new value if provided, else preserve existing; stored only when the - // effective value is not 'public' (a change to 'public' drops the field). - ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { - visibility: params.visibility ?? existing.visibility - }) - }, - newMetadata as Record - ) + const updatedMetadata = { + ...newMetadata, + data: params.data !== undefined ? params.data : existing.data, + noun: params.type || existing.type, + service: existing.service, + createdAt: existing.createdAt, + updatedAt: now, + _rev: currentRev + 1, + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.weight !== undefined && { weight: params.weight }), + ...(params.confidence === undefined && + existing.confidence !== undefined && { confidence: existing.confidence }), + ...(params.weight === undefined && + existing.weight !== undefined && { weight: existing.weight }), + ...(params.subtype !== undefined && { subtype: params.subtype }), + ...(params.subtype === undefined && + existing.subtype !== undefined && { subtype: existing.subtype }), + // Visibility: new value if provided, else preserve existing; stored only when the + // effective value is not 'public' (a change to 'public' drops the field). + ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { + visibility: params.visibility ?? existing.visibility + }) + } // Register for the authoritative under-mutex CAS re-verify + rev re-stamp // (see PlannedTransact.casUpdates). The staged UpdateNounMetadataOperation @@ -9241,7 +9484,7 @@ export class Brainy implements BrainyInterface { id: params.id, vector, connections: new Map(), - // no `level` — plumbing never enters the indexing view + level: 0, type: params.type || existing.type, subtype: params.subtype !== undefined ? params.subtype : existing.subtype, ...(((params.visibility ?? existing.visibility) ?? 'public') !== 'public' && { @@ -9269,22 +9512,16 @@ export class Brainy implements BrainyInterface { } plan.operations.push( - new UpdateNounMetadataOperation(this.storage, params.id, updatedMetadata) + new UpdateNounMetadataOperation(this.storage, params.id, updatedMetadata), + new SaveNounOperation(this.storage, { + id: params.id, + vector, + connections: new Map(), + level: 0 + }) ) - // Noun-record write + HNSW reindex ONLY when the vector side actually - // changed — the same write-granularity law as update(): a metadata-only - // patch must never rewrite the whole vector record. This plan path is the - // one transact() updates ride, so an unconditional save here would - // re-open the read-sweep disk-saturation amplifier for exactly the - // consumers batching their stat touches through transact(). if (needsReindexing) { plan.operations.push( - new SaveNounOperation(this.storage, { - id: params.id, - vector, - connections: new Map(), - level: 0 - }), new RemoveFromVectorIndexOperation(this.index, params.id, existing.vector), new AddToVectorIndexOperation(this.index, params.id, vector) ) @@ -9446,7 +9683,8 @@ export class Brainy implements BrainyInterface { ): Promise { const { op: _discriminator, ...rawParams } = op validateRelateParams(rawParams as RelateParams) - const params = rawParams as RelateParams + // Same reserved-field normalization as relate(). + const params = this.remapReservedRelateMetadata(rawParams as RelateParams) // Id normalization (8.0) — mirror of relate(): resolve BOTH endpoints to the // canonical UUID add() stored, so a relate op may reference either side by // natural key. Real UUIDs pass through. (Relationship ids are engine-minted.) @@ -9496,23 +9734,19 @@ export class Brainy implements BrainyInterface { const id = uuidv4() const relationVector = fromEntity.vector.map((v, i) => (v + toEntity.vector[i]) / 2) const now = Date.now() - // v2 nested-bag record — mirror of relate(): engine fields top-level, - // the user's edge bag nested verbatim. - const verbMetadata = buildVerbMetadataRecord( - { - verb: params.type, - ...(params.subtype !== undefined && { subtype: params.subtype }), - // visibility: stored only when not 'public' (absent === public, keeps records lean) - ...(params.visibility !== undefined && - params.visibility !== 'public' && { visibility: params.visibility }), - weight: params.weight ?? 1.0, - ...(params.confidence !== undefined && { confidence: params.confidence }), - ...(params.service !== undefined && { service: params.service }), - createdAt: now, - ...(params.data !== undefined && { data: params.data }) - }, - (params.metadata as Record) || {} - ) + const verbMetadata = { + ...(params.metadata || {}), + verb: params.type, + ...(params.subtype !== undefined && { subtype: params.subtype }), + // visibility: stored only when not 'public' (absent === public, keeps records lean) + ...(params.visibility !== undefined && + params.visibility !== 'public' && { visibility: params.visibility }), + weight: params.weight ?? 1.0, + ...(params.confidence !== undefined && { confidence: params.confidence }), + ...(params.service !== undefined && { service: params.service }), + createdAt: now, + ...(params.data !== undefined && { data: params.data }) + } const verb: GraphVerb = { id, vector: relationVector, @@ -11101,26 +11335,27 @@ export class Brainy implements BrainyInterface { if (params.where || params.subtype || params.service) { let filter: any = {} if (params.where) { - // Where keys pass through UNTOUCHED — the one addressing law - // parses them at the index boundary (bare = user metadata, - // system.* = engine scalars). The old where.type→noun alias is - // dead: a bare 'type' is the user's own field now. Object.assign(filter, params.where) + // Alias: where.type → where.noun (storage field name for entity type) + if ('type' in filter && !('noun' in filter)) { + filter.noun = filter.type + delete filter.type + } } - if (params.service) filter['system.service'] = params.service + if (params.service) filter.service = params.service if (params.subtype !== undefined) { - filter['system.subtype'] = Array.isArray(params.subtype) + filter.subtype = Array.isArray(params.subtype) ? { oneOf: params.subtype } : params.subtype } if (params.type) { const types = Array.isArray(params.type) ? params.type : [params.type] if (types.length === 1) { - filter['system.type'] = types[0] + filter.noun = types[0] } else { const baseFilter = { ...filter } filter = { - anyOf: types.map(type => ({ 'system.type': type, ...baseFilter })) + anyOf: types.map(type => ({ noun: type, ...baseFilter })) } } } @@ -11170,24 +11405,27 @@ export class Brainy implements BrainyInterface { // Use MetadataIndexManager for efficient filtered streaming let filterObj: any = {} if (filter.where) { - // Where keys pass through — the addressing law parses them at - // the index boundary; the type→noun alias is dead. Object.assign(filterObj, filter.where) + // Alias: where.type → where.noun (storage field name for entity type) + if ('type' in filterObj && !('noun' in filterObj)) { + filterObj.noun = filterObj.type + delete filterObj.type + } } - if (filter.service) filterObj['system.service'] = filter.service + if (filter.service) filterObj.service = filter.service if (filter.subtype !== undefined) { - filterObj['system.subtype'] = Array.isArray(filter.subtype) + filterObj.subtype = Array.isArray(filter.subtype) ? { oneOf: filter.subtype } : filter.subtype } if (filter.type) { const types = Array.isArray(filter.type) ? filter.type : [filter.type] if (types.length === 1) { - filterObj['system.type'] = types[0] + filterObj.noun = types[0] } else { const baseFilterObj = { ...filterObj } filterObj = { - anyOf: types.map(type => ({ 'system.type': type, ...baseFilterObj })) + anyOf: types.map(type => ({ noun: type, ...baseFilterObj })) } } } @@ -11567,7 +11805,7 @@ export class Brainy implements BrainyInterface { // don't have the tracked field at all (e.g. the VFS root) bucket under // '__null__' and would otherwise pollute the count map. if (value === undefined || value === null || value === '__null__') continue - if (options?.type !== undefined && row.groupKey?.['system.type'] !== options.type) continue + if (options?.type !== undefined && row.groupKey?.['noun'] !== options.type) continue const key = String(value) result[key] = (result[key] || 0) + (typeof row.metrics?.count === 'number' ? row.metrics.count : row.count) } @@ -13314,12 +13552,14 @@ export class Brainy implements BrainyInterface { } let filter: any = {} if (params.where) { - // Where keys pass through UNTOUCHED — the one addressing law parses - // them at the index boundary (bare = user metadata, system.* = engine - // scalars, typed refusal otherwise). The old type→noun alias is dead. Object.assign(filter, params.where) + // Alias: where.type → where.noun (storage field name for entity type) + if ('type' in filter && !('noun' in filter)) { + filter.noun = filter.type + delete filter.type + } } - if (params.service) filter['system.service'] = params.service + if (params.service) filter.service = params.service if (params.excludeVFS === true) { filter.vfsType = { exists: false } filter.isVFSEntity = { ne: true } @@ -13327,14 +13567,14 @@ export class Brainy implements BrainyInterface { // Subtype (top-level standard field — fast path). Assigned BEFORE the type-array // expansion below so the spread into each anyOf branch carries it through. if (params.subtype !== undefined) { - filter['system.subtype'] = Array.isArray(params.subtype) ? { oneOf: params.subtype } : params.subtype + filter.subtype = Array.isArray(params.subtype) ? { oneOf: params.subtype } : params.subtype } if (params.type) { const types = Array.isArray(params.type) ? params.type : [params.type] if (types.length === 1) { - filter['system.type'] = types[0] + filter.noun = types[0] } else { - filter = { anyOf: types.map((type) => ({ 'system.type': type, ...filter })) } + filter = { anyOf: types.map((type) => ({ noun: type, ...filter })) } } } return filter @@ -14825,7 +15065,12 @@ export class Brainy implements BrainyInterface { requireSubtype: config?.requireSubtype ?? true, // Multi-process safety mode: config?.mode ?? 'writer', - force: config?.force ?? false + force: config?.force ?? false, + // Reserved-field-in-metadata-bag policy (8.0 — no silent failures). + // Default 'throw': an untyped caller that smuggles a reserved key past + // the compile guard gets a loud Error naming the correct write path. + // 'warn' = remap + one-shot warning per key; 'remap' = legacy silent remap. + reservedFieldPolicy: config?.reservedFieldPolicy ?? 'throw' } } @@ -16222,21 +16467,11 @@ export class Brainy implements BrainyInterface { await this.generationStore.flushPendingSingleOps() } - // Phase 0b: REPACK cold history into sealed segments (D1+D3 — - // re-representation, never deletion; the only history transform under the - // archival profile), then auto-compact per config.retention. Repack runs - // FIRST so bounded-retention reclaim can drop whole segments. Both are - // time-bounded maintenance passes (8.9.0 law: flush() never pays these); - // both are housekeeping — failures warn, never fail a clean shutdown. - if (!this.isReadOnly && this.generationStore) { - try { - await this.generationStore.repackHistory({ timeBudgetMs: 5_000 }) - } catch (error) { - console.warn( - `History repacking failed (non-fatal): ${error instanceof Error ? error.message : String(error)}` - ) - } - } + // Phase 0b: Auto-compact generational history per config.retention (default + // on) BEFORE the generation store closes below. This is THE auto-compaction + // site (8.9.0 — flush() never compacts): time-bounded per pass, respects + // live Db pins and an explicit autoCompact: false; no-op on read-only + // instances. await this.autoCompactHistory() // Phase 1: Flush ALL components in parallel to persist buffered data diff --git a/src/coreTypes.ts b/src/coreTypes.ts index 90fc4462..e0248d17 100644 --- a/src/coreTypes.ts +++ b/src/coreTypes.ts @@ -284,12 +284,7 @@ export const STANDARD_ENTITY_FIELDS: ReadonlySet = new Set([ 'id', 'vector', 'connections', - // 'level' is deliberately ABSENT: it is HNSW plumbing, not an entity field. - // Listing it here made every by-name read of a user metadata field called - // `level` resolve to the engine's internal node layer instead — a silent - // shadow that broke sort/filter/aggregation on a perfectly natural field - // name (VENUE-BRAINY-ORDERBY-NOOP). Engine plumbing is invisible to the - // query surface; a bare `level` reads `entity.metadata.level`. + 'level', 'type', 'subtype', 'visibility', diff --git a/src/db/db.ts b/src/db/db.ts index 68428a7c..c5cbad8b 100644 --- a/src/db/db.ts +++ b/src/db/db.ts @@ -59,6 +59,10 @@ import type { import type { StorageAdapter } from '../coreTypes.js' import { exportGraph } from './portableGraph.js' import type { ExportSelector, ExportOptions, PortableGraph } from './portableGraph.js' +import { + splitNounMetadataRecord, + splitVerbMetadataRecord +} from '../types/reservedFields.js' import { v4 as uuidv4 } from '../universal/uuid.js' import { coerceNewEntityId, resolveEntityId, ORIGINAL_ID_KEY } from '../utils/idNormalization.js' import { EntityNotFoundError } from '../errors/notFound.js' @@ -701,15 +705,23 @@ export class Db { for (const op of ops) { switch (op.op) { case 'add': { - // Field-addressing law: the metadata bag is the user's, VERBATIM — - // no reserved-name lift, no drops. Engine scalars come ONLY from - // their dedicated op fields; a bag field named `confidence` is an - // ordinary user field, exactly as on the committed write path. - const custom = { ...(op.metadata as Record | undefined) } - const confidence = op.confidence - const weight = op.weight - const subtype = op.subtype - const service = op.service + // Reserved-field normalization — mirror of the brain.transact() + // write path: user-settable fields lift to their dedicated field + // (top-level wins), system-managed fields drop, and the entity's + // metadata bag carries ONLY custom fields. Speculative views skip + // the one-shot warnings — committing the same ops through + // `brain.transact()` warns on the real write path. + const { reserved, custom } = splitNounMetadataRecord( + op.metadata as Record | undefined + ) + const confidence = + op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined) + const weight = + op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined) + const subtype = + op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined) + const service = + op.service ?? (typeof reserved.service === 'string' ? reserved.service : undefined) // Id normalization (8.0) — mirror of the committed transact() add // path: a natural key coerces to a STABLE UUID (v5), preserving the @@ -747,12 +759,16 @@ export class Db { `with(): entity ${updateId} not found at generation ${this.gen}` ) } - // Field-addressing law — mirror of the add case: the patch bag is - // the user's verbatim; engine scalars only from dedicated op fields. - const custom = { ...(op.metadata as Record | undefined) } - const confidence = op.confidence - const weight = op.weight - const subtype = op.subtype + // Same reserved-field normalization as the committed update path. + const { reserved, custom } = splitNounMetadataRecord( + op.metadata as Record | undefined + ) + const confidence = + op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined) + const weight = + op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined) + const subtype = + op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined) const mergedMetadata = op.merge !== false ? ({ ...(base.metadata as object), ...custom } as T) @@ -814,14 +830,19 @@ export class Db { } if (duplicate) break - // Field-addressing law — relationship mirror of the add case: the - // edge bag is the user's verbatim; engine scalars only from - // dedicated op fields. - const custom = { ...(op.metadata as Record | undefined) } - const confidence = op.confidence - const weight = op.weight - const subtype = op.subtype - const service = op.service + // Reserved-field normalization — relationship mirror of the add + // op above (and of the committed relate() path). + const { reserved, custom } = splitVerbMetadataRecord( + op.metadata as Record | undefined + ) + const confidence = + op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined) + const weight = + op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined) + const subtype = + op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined) + const service = + op.service ?? (typeof reserved.service === 'string' ? reserved.service : undefined) const id = uuidv4() overlay.verbs.set(id, { diff --git a/src/db/factLog.ts b/src/db/factLog.ts index 94e79700..4c5e95fd 100644 --- a/src/db/factLog.ts +++ b/src/db/factLog.ts @@ -102,26 +102,12 @@ export interface FactScanBatch { segmentId: string } -/** - * Liveness bound on a scan's FIRST batch (Stage-2 co-freeze, D1 contract): - * `batches()` must yield its first batch — or fail loudly — within this many - * ms of the first pull. A backlogged or damaged store may be SLOW, but it may - * never be SILENT: a consumer awaiting the first batch is otherwise - * indistinguishable from a wedge (the exact failure shape a production heal - * hit against a generations-backlogged brain). - */ -export const SCANFACTS_FIRST_BATCH_MS = 10_000 - /** The telemetry a scan OPEN returns (frozen shape). */ export interface FactScanHandle { headGeneration: number segmentCount: number approxFactCount: number - /** - * Ordered batches; a detected gap aborts LOUDLY, never a silent skip. - * Liveness contract: the FIRST batch resolves or rejects within - * {@link SCANFACTS_FIRST_BATCH_MS} of the first pull — never a silent hang. - */ + /** Ordered batches; a detected gap aborts LOUDLY, never a silent skip. */ batches: () => AsyncGenerator /** Close telemetry — the invariant cross-check, valid after iteration ends. */ summary: () => { factsYielded: number; segmentsRead: number } @@ -454,8 +440,6 @@ export class FactLog { toGeneration?: number kinds?: Array<'noun' | 'verb'> batchSize?: number - /** Test override for the first-batch liveness bound (default {@link SCANFACTS_FIRST_BATCH_MS}). */ - firstBatchTimeoutMs?: number }): FactScanHandle { const from = options?.fromGeneration ?? 1 const to = options?.toGeneration ?? this.head @@ -530,43 +514,11 @@ export class FactLog { } } - // Liveness wrapper: the FIRST pull races the contract deadline. Only the - // first — the bound is time-to-first-batch (proof the producer is alive), - // not per-batch pacing; and it runs only while a pull is actually pending, - // so consumer think-time between pulls never counts against the producer. - const firstBatchTimeoutMs = options?.firstBatchTimeoutMs ?? SCANFACTS_FIRST_BATCH_MS - async function* batchesWithLiveness(this: void): AsyncGenerator { - const inner = batches() - let timer: NodeJS.Timeout | undefined - try { - const deadline = new Promise((_, reject) => { - timer = setTimeout( - () => - reject( - new Error( - `fact log: scanFacts produced no first batch within ${firstBatchTimeoutMs}ms ` + - `(liveness contract) — the store is wedged or unreadably slow; aborting scan LOUDLY ` + - `instead of hanging the consumer.` - ) - ), - firstBatchTimeoutMs - ) - timer.unref?.() - }) - const first = await Promise.race([inner.next(), deadline]) - if (first.done) return - yield first.value - } finally { - clearTimeout(timer) - } - yield* inner - } - return { headGeneration: this.head, segmentCount: segments.length + (tailSnapshot.length > 0 ? 1 : 0), approxFactCount, - batches: batchesWithLiveness, + batches, summary: () => ({ factsYielded, segmentsRead }) } } diff --git a/src/db/fieldAddressing.ts b/src/db/fieldAddressing.ts deleted file mode 100644 index 21689319..00000000 --- a/src/db/fieldAddressing.ts +++ /dev/null @@ -1,323 +0,0 @@ -/** - * @module db/fieldAddressing - * @description The one field-addressing law for every query surface (find()'s - * `where` / `orderBy` / `groupBy`, aggregation `source.where`), ruled - * 2026-08-03 after a production incident in which a user metadata field - * named `level` was silently shadowed by the engine's internal HNSW node - * layer (VENUE-BRAINY-ORDERBY-NOOP — thread id kept verbatim as the audit - * key; it names no product): - * - * 1. A BARE field name addresses the user's metadata field. Always. - * No priority resolution, no fallback chain — `orderBy: 'level'` - * reads `entity.metadata.level`, full stop. - * 2. `system.` addresses an engine scalar, reachable ONLY with the - * explicit prefix. The entity map is exactly ten scalars; the relation - * map mirrors it with `verb`/`sourceId`/`targetId` as the structural - * members. - * 3. Engine plumbing (`vector`, `connections`, `level`, `data`, `_rev`) is - * INVISIBLE to the query surface in either spelling — `system.level` - * refuses; bare `level` is the user's field. - * 4. `metadata.` is the explicit spelling of the bare form — - * identical semantics on every path. - * 5. Anything unresolvable refuses with a TYPED error naming both - * candidate spellings — an accepted name either works or refuses; - * there is no third state. - * - * This module is the SINGLE source of truth for the law: parsing, the maps, - * and the refusal builders live here so the JS engine, the provider seams, - * and the cross-engine conformance suite can never drift on the contract. - */ - -import type { HNSWNounWithMetadata, HNSWVerbWithMetadata } from '../coreTypes.js' - -/** - * @description The entity-side `system.*` map — EXACTLY the ten engine - * scalars David ruled queryable (2026-08-03). Adding a name here is a - * cross-engine contract change: the native accelerator's conformance suite - * pins this list verbatim, so any edit must ship as a paired release. - */ -export const SYSTEM_ENTITY_SCALARS: ReadonlySet = new Set([ - 'id', - 'type', - 'subtype', - 'createdAt', - 'updatedAt', - 'confidence', - 'weight', - 'visibility', - 'service', - 'createdBy' -]) - -/** - * @description The relation-side `system.*` map — the verb mirror of - * {@link SYSTEM_ENTITY_SCALARS}: `verb`, `sourceId`, `targetId` are the - * structural members beside the eight shared scalars. Same one law, same - * pairing rule for edits. - */ -export const SYSTEM_RELATION_SCALARS: ReadonlySet = new Set([ - 'verb', - 'sourceId', - 'targetId', - 'subtype', - 'createdAt', - 'updatedAt', - 'confidence', - 'weight', - 'visibility', - 'service', - 'createdBy' -]) - -/** - * @description Engine plumbing — never addressable from the query surface in - * ANY spelling. `level` is the HNSW node layer (the incident field: listing - * it as resolvable shadowed real user data); `data` is the payload container, - * not a scalar — content is reached through the content/text-search APIs, - * and addressing it as a sortable field would lie about its shape. - */ -export const PLUMBING_FIELDS: ReadonlySet = new Set([ - 'vector', - 'connections', - 'level', - 'data', - '_rev' -]) - -/** @description Which record kind a field address is being resolved against. */ -export type FieldAddressKind = 'entity' | 'relation' - -/** - * @description A parsed, law-valid field address. `scope` says which side of - * the record the name lives on; `field` is the unprefixed name to read. - */ -export interface FieldAddress { - /** 'metadata' = the user's field (bare or `metadata.`-prefixed); 'system' = an engine scalar. */ - scope: 'metadata' | 'system' - /** The field name with any scope prefix removed. */ - field: string - /** The exact spelling the caller used — preserved for error text and telemetry. */ - raw: string -} - -/** - * Parse a query-surface field name under the one law. Pure and data-blind: - * this validates the ADDRESS (spelling + map membership), not whether any - * row actually carries the field — data-aware refusals (the did-you-mean - * for a bare system-scalar name no row carries) belong to the query layer, - * which calls {@link buildUnresolvableMessage} with index knowledge. - * - * @param raw - The field name as the caller wrote it (`level`, - * `metadata.level`, `system.createdAt`, …) - * @param kind - Entity or relation resolution (selects the system map) - * @returns The parsed {@link FieldAddress} - * @throws {InvalidFieldAddressError} for a `system.*` name outside the ruled - * map (including every plumbing field) or a malformed spelling — the error - * text enumerates the valid system scalars so the fix is in the message. - * - * @example - * parseFieldAddress('level', 'entity') // { scope: 'metadata', field: 'level' } - * parseFieldAddress('metadata.level', 'entity') // { scope: 'metadata', field: 'level' } - * parseFieldAddress('system.createdAt', 'entity') // { scope: 'system', field: 'createdAt' } - * parseFieldAddress('system.level', 'entity') // throws — plumbing is invisible - */ -export function parseFieldAddress( - raw: string, - kind: FieldAddressKind -): FieldAddress { - const systemMap = - kind === 'entity' ? SYSTEM_ENTITY_SCALARS : SYSTEM_RELATION_SCALARS - - if (raw.startsWith('system.')) { - const field = raw.slice('system.'.length) - if (!systemMap.has(field)) { - throw new InvalidFieldAddressError(raw, kind, systemMap) - } - return { scope: 'system', field, raw } - } - - if (raw.startsWith('metadata.')) { - const field = raw.slice('metadata.'.length) - if (field.length === 0) { - throw new InvalidFieldAddressError(raw, kind, systemMap) - } - return { scope: 'metadata', field, raw } - } - - if (raw.length === 0) { - throw new InvalidFieldAddressError(raw, kind, systemMap) - } - - // Bare name = the user's metadata field. Always. Even when the same name - // exists in the system map — `confidence` as a bare name is the user's - // metadata field named confidence; the engine scalar is system.confidence. - return { scope: 'metadata', field: raw, raw } -} - -/** - * Read the addressed value off an entity. The ONLY sanctioned way a query - * surface turns a {@link FieldAddress} into a value — direct property reads - * against records re-create the shadow class this module exists to kill. - * - * @returns The value, or `undefined` when the record does not carry it - * (missing values sort LAST in both directions per the ordering contract — - * they are never grounds for dropping a row). - */ -export function readEntityFieldAddress( - entity: HNSWNounWithMetadata, - address: FieldAddress -): unknown { - const rec = entity as unknown as Record - const bag = - rec.metadata && typeof rec.metadata === 'object' - ? (rec.metadata as Record) - : null - - if (address.scope === 'system') { - // System scalars live at the record's top level, NEVER in the user's - // bag — a user field named `confidence` must be unreachable from - // system.confidence (and vice versa). Entity views carry the scalars - // top-level directly; record-derived views spell the type `noun`. - const top = rec[address.field] - if (top !== undefined) return top - if (address.field === 'type') return rec.noun - return undefined - } - - // User scope: the bag IS the user's namespace, authoritative — EVERY name - // reads from it, engine spellings included (`bag.confidence` is the user's - // confidence field under the field-addressing law). - if (bag) return bag[address.field] - - // No bag at all: a LEGACY flat record (pre-nested-bag storage). Its keys - // matching system/plumbing names are the ENGINE's — the pre-law write door - // refused user colliders — so a bare system name reads as ABSENT rather - // than resurrecting the shadow this module exists to kill. Same for the - // legacy 'noun' spelling. - if ( - SYSTEM_ENTITY_SCALARS.has(address.field) || - PLUMBING_FIELDS.has(address.field) || - address.field === 'noun' - ) { - return undefined - } - return rec[address.field] -} - -/** - * Relation twin of {@link readEntityFieldAddress}. The stored flat record - * keys the relation type under `verb`; public Relation shapes may carry it - * as `type` — both spellings of the record are read, the ADDRESS is always - * `system.verb`. - */ -export function readRelationFieldAddress( - verb: HNSWVerbWithMetadata, - address: FieldAddress -): unknown { - if (address.scope === 'system') { - const rec = verb as unknown as Record - if (address.field === 'verb') return rec.verb ?? rec.type - return rec[address.field] - } - return verb.metadata?.[address.field] -} - -/** - * Build the ruled did-you-mean refusal text for a bare name that resolved to - * metadata but is UNKNOWN to the index — the data-aware half of the law, - * called by the query layer once it has consulted the known-field set: - * - * "no metadata field 'createdAt' — did you mean system.createdAt or - * metadata.createdAt?" - * - * When the bare name is NOT a system scalar the system candidate is omitted - * (there is only one thing the caller could have meant; the refusal exists - * because refusing beats silently sorting nothing). - */ -export function buildUnresolvableMessage( - raw: string, - kind: FieldAddressKind -): string { - const systemMap = - kind === 'entity' ? SYSTEM_ENTITY_SCALARS : SYSTEM_RELATION_SCALARS - if (systemMap.has(raw)) { - return ( - `no metadata field '${raw}' — did you mean system.${raw} or metadata.${raw}? ` + - `(bare names always address your metadata; engine fields need the system. prefix)` - ) - } - return ( - `no metadata field '${raw}' on this store — nothing carries it, so an ordered or ` + - `filtered read against it cannot mean anything. Spell it metadata.${raw} once the ` + - `field exists, or check the field name (system.${raw} is NOT valid — '${raw}' is ` + - `not one of the engine's system scalars).` - ) -} - -/** - * @description Refusal for a syntactically valid address that resolves to - * NOTHING — a bare name no user field carries. Carries the did-you-mean - * (both candidate spellings when the name collides with a system scalar) so - * the fix ships inside the error. Thrown by the query layer with index - * knowledge, never by the pure parser. - */ -export class UnresolvableFieldError extends Error { - public readonly raw: string - public readonly kind: FieldAddressKind - - constructor(raw: string, kind: FieldAddressKind, messageOverride?: string) { - super(messageOverride ?? buildUnresolvableMessage(raw, kind)) - this.name = 'UnresolvableFieldError' - this.raw = raw - this.kind = kind - } -} - -/** - * @description Refusal for a malformed or out-of-map field ADDRESS — - * `system.` (including all plumbing), an empty - * name, or a bare `metadata.` prefix. The message carries the full valid - * system map so the fix never needs a docs lookup. - */ -export class InvalidFieldAddressError extends UnresolvableFieldError { - constructor(raw: string, kind: FieldAddressKind, systemMap: ReadonlySet) { - const valid = [...systemMap].map((f) => `system.${f}`).join(', ') - super( - raw, - kind, - `'${raw}' is not an addressable ${kind} field. Bare names address your own ` + - `metadata fields; engine fields are exactly: ${valid}. Engine plumbing ` + - `(vector, connections, level, data, _rev) is not part of the query surface.` - ) - this.name = 'InvalidFieldAddressError' - } -} - - -/** - * @description Refusal for a find() option that is accepted by the type - * surface but NOT implemented — an accepted option must work or refuse; - * accepted-and-ignored died as a class (sealed 2026-08-03). Names the - * option and the honest state so nobody discovers a no-op by measurement. - */ -export class UnsupportedFindOptionError extends Error { - public readonly option: string - - constructor(option: string) { - super( - `find() option '${option}' is not implemented — it used to be silently ` + - `ignored, which read as working. Remove it from the call (or track the ` + - `feature request); it will be honored or refused, never swallowed.` - ) - this.name = 'UnsupportedFindOptionError' - this.option = option - } -} - -/** - * @description The capability signal both engines' conformance suites arm on - * (never a version guess): its presence at the package root means the one - * field-addressing law is LIVE on every query surface — bare = user metadata, - * `system.*` = the ruled scalars, plumbing invisible, refusals typed. - */ -export const FIELD_ADDRESSING_CAPABILITY = 'field-addressing/v1' diff --git a/src/db/generationSegments.ts b/src/db/generationSegments.ts deleted file mode 100644 index 0c14b60c..00000000 --- a/src/db/generationSegments.ts +++ /dev/null @@ -1,459 +0,0 @@ -/** - * @module db/generationSegments - * @description The generation-segment store — Stage-2 D1+D3+repacking's file - * format (co-frozen 2026-07-19; design: the d1-d3-repacking spec). - * - * Packs CONSECUTIVE cold generations' record-sets (before-images + delta) - * into append-once segment files with derived sidecar indexes, so history - * scales in SEGMENTS (tens) instead of FILES-PER-GENERATION (hundreds of - * thousands), and cold-open reads ONE manifest instead of listing the - * backlog. Layout under `_generations/segments/`: - * - * - `seg-.bgs` — magic "BGS1", then one frame per - * generation: `u32 payloadLen | u32 crc32c | msgpack payload`. Payload is - * POSITIONAL: `[generation, timestamp, delta, records[], flags]` with - * records `[kindByte, id, record]`. `flags` reserves encoding evolution - * (bit 0 = compressed payload — v1 always 0; a future writer upgrade, - * never a format break). Sealed segments are IMMUTABLE — the fact log's - * own law, generalized. - * - `seg-.idx` — DERIVED sidecar (msgpack): per-generation frame - * offsets (point reads = one ranged read, never a listing) + per-id - * generation postings (per-id chain rebuilds read only what they need). - * Corrupt/missing → rebuilt from its segment in one sequential read, - * loudly. - * - `manifest.json` — the segment catalogue + `compactedBelow` (D3's - * horizon marker). Cold-open reads THIS; the packed backlog is never - * listed. - * - * D3 semantics carried here: bounded-retention reclaim drops WHOLE segments - * at boundaries (O(1) per segment, no rewrite); under the archival profile - * (`retention: 'all'`) nothing here is ever dropped — folding is the only - * transform (re-representation, never deletion). - */ - -import { encode as msgpackEncode, decode as msgpackDecode } from '@msgpack/msgpack' -import { crc32c } from '../utils/crc32c.js' -import type { FactLogStorage } from './factLog.js' -import { prodLog } from '../utils/logger.js' - -/** Directory for segment files + manifest, under the generations prefix. */ -export const SEGMENTS_PREFIX = '_generations/segments' - -/** Target sealed-segment size (co-freeze proposal; tunable on evidence). */ -export const SEGMENT_TARGET_BYTES = 64 * 1024 * 1024 - -const MAGIC = new TextEncoder().encode('BGS1') -const FRAME_PREFIX_BYTES = 8 // u32 payloadLen + u32 crc32c -const MANIFEST_PATH = `${SEGMENTS_PREFIX}/manifest.json` - -/** One generation's fold input — exactly what the live tier holds for it. */ -export interface FoldGeneration { - generation: number - timestamp: number - /** The tx.json delta object, carried verbatim. */ - delta: unknown - /** The before-image record-set (empty for record-less generations). */ - records: Array<{ kind: 'noun' | 'verb'; id: string; record: unknown }> -} - -/** Manifest entry for one sealed segment. */ -export interface SegmentMeta { - file: string - firstGeneration: number - lastGeneration: number - frames: number - bytes: number - /** crc32c of the full segment byte stream — the digest chain's link. */ - checksum: number -} - -interface SegmentManifest { - version: 1 - compactedBelow: number - segments: SegmentMeta[] -} - -interface SidecarIndex { - version: 1 - /** [generation, frameOffset, frameLen] ascending by generation. */ - generations: Array<[number, number, number]> - /** `${kindByte}:${id}` → ascending generations holding a record for it. */ - ids: Record -} - -const segmentFileName = (firstGeneration: number): string => - `seg-${String(firstGeneration).padStart(20, '0')}.bgs` -const sidecarFileName = (firstGeneration: number): string => - `seg-${String(firstGeneration).padStart(20, '0')}.idx` - -/** - * The generation-segment store. Owns the packed tier ONLY — the live - * per-generation tier and the routing between tiers belong to - * `GenerationStore`. All mutating entry points here are called under the - * generation store's commit mutex. - */ -export class GenerationSegmentStore { - private readonly storage: FactLogStorage - private manifest: SegmentManifest = { version: 1, compactedBelow: 0, segments: [] } - /** Sidecar cache — segments are immutable, so entries never invalidate. */ - private readonly sidecars = new Map() - - constructor(storage: FactLogStorage) { - this.storage = storage - } - - /** Load the manifest (ONE read — never a directory listing). */ - async open(): Promise { - const raw = (await this.storage.readRawObject(MANIFEST_PATH)) as SegmentManifest | null - if (raw) { - if (raw.version !== 1) { - throw new Error( - `[GenerationSegments] manifest version ${String(raw.version)} is newer than this ` + - `engine understands — refusing to serve partial history. Upgrade the engine.` - ) - } - this.manifest = raw - } - } - - /** The packed tier's catalogue (ascending, immutable snapshot). */ - segments(): readonly SegmentMeta[] { - return this.manifest.segments - } - - /** D3's horizon marker: generations below this were reclaimed (bounded profiles only). */ - compactedBelow(): number { - return this.manifest.compactedBelow - } - - /** The covering sealed segment for `gen`, or null if it lives outside the packed tier. */ - private coveringSegment(gen: number): SegmentMeta | null { - // Manifest is ascending and ranges never overlap — binary search. - const segs = this.manifest.segments - let lo = 0 - let hi = segs.length - 1 - while (lo <= hi) { - const mid = (lo + hi) >> 1 - const s = segs[mid] - if (gen < s.firstGeneration) hi = mid - 1 - else if (gen > s.lastGeneration) lo = mid + 1 - else return s - } - return null - } - - /** True when `gen` is packed (readable from this tier). */ - hasGeneration(gen: number): boolean { - return this.coveringSegment(gen) !== null - } - - /** - * Fold consecutive generations into ONE new sealed segment + sidecar and - * append it to the manifest atomically. Caller guarantees: `gens` is - * ascending, contiguous with the packed tier (first = last packed + 1 when - * segments exist), and already durable in the live tier. Crash between the - * segment write and the caller's live-tier delete leaves a DUPLICATE - * representation — resolved live-tier-wins by the reader; never a gap. - */ - async fold(gens: FoldGeneration[]): Promise { - if (gens.length === 0) { - throw new Error('[GenerationSegments] fold() requires at least one generation') - } - for (let i = 1; i < gens.length; i++) { - if (gens[i].generation <= gens[i - 1].generation) { - throw new Error('[GenerationSegments] fold() input must be strictly ascending') - } - } - const last = this.manifest.segments[this.manifest.segments.length - 1] - if (last && gens[0].generation <= last.lastGeneration) { - throw new Error( - `[GenerationSegments] fold() overlaps the packed tier: ${gens[0].generation} ≤ ` + - `sealed ${last.lastGeneration} — segments are immutable, never rewritten` - ) - } - - const first = gens[0].generation - const file = segmentFileName(first) - const sidecar: SidecarIndex = { version: 1, generations: [], ids: {} } - - // Encode all frames, tracking offsets for the sidecar. - const parts: Uint8Array[] = [MAGIC] - let offset = MAGIC.length - for (const g of gens) { - const payload = msgpackEncode([ - g.generation, - g.timestamp, - g.delta, - g.records.map((r) => [r.kind === 'noun' ? 0 : 1, r.id, r.record]), - 0 // flags: v1 = uncompressed - ]) - const frame = new Uint8Array(FRAME_PREFIX_BYTES + payload.length) - const view = new DataView(frame.buffer) - view.setUint32(0, payload.length, true) - view.setUint32(4, crc32c(payload), true) - frame.set(payload, FRAME_PREFIX_BYTES) - sidecar.generations.push([g.generation, offset, frame.length]) - for (const r of g.records) { - const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}` - ;(sidecar.ids[key] ??= []).push(g.generation) - } - parts.push(frame) - offset += frame.length - } - const total = parts.reduce((n, p) => n + p.length, 0) - const bytes = new Uint8Array(total) - let at = 0 - for (const p of parts) { - bytes.set(p, at) - at += p.length - } - - const meta: SegmentMeta = { - file, - firstGeneration: first, - lastGeneration: gens[gens.length - 1].generation, - frames: gens.length, - bytes: total, - checksum: crc32c(bytes) - } - - // Durability order: segment + sidecar fsync'd BEFORE the manifest names - // them (a crash before the manifest = invisible orphan files, harmless); - // manifest last, atomically. - const segPath = `${SEGMENTS_PREFIX}/${file}` - const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(first)}` - await this.storage.writeRawBytes(segPath, bytes) - await this.storage.writeRawBytes(idxPath, msgpackEncode(sidecar)) - await this.storage.syncRawObjects([segPath, idxPath]) - const next: SegmentManifest = { - ...this.manifest, - segments: [...this.manifest.segments, meta] - } - await this.storage.writeRawObject(MANIFEST_PATH, next) - await this.storage.syncRawObjects([MANIFEST_PATH]) - this.manifest = next - this.sidecars.set(file, sidecar) - return meta - } - - /** Load (or rebuild, loudly) a segment's sidecar. */ - private async sidecarFor(meta: SegmentMeta): Promise { - const cached = this.sidecars.get(meta.file) - if (cached) return cached - const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(meta.firstGeneration)}` - const raw = await this.storage.readRawBytes(idxPath) - if (raw) { - try { - const idx = msgpackDecode(raw) as SidecarIndex - if (idx.version === 1) { - this.sidecars.set(meta.file, idx) - return idx - } - } catch { - // fall through to rebuild - } - } - // Sidecars are DERIVED: rebuild from the segment, loudly — never serve - // wrong offsets silently. - prodLog.warn( - `[GenerationSegments] sidecar for ${meta.file} missing or unreadable — rebuilding from the segment` - ) - const rebuilt = await this.rebuildSidecar(meta) - await this.storage.writeRawBytes(idxPath, msgpackEncode(rebuilt)) - this.sidecars.set(meta.file, rebuilt) - return rebuilt - } - - /** One sequential read of the segment → a fresh sidecar. Verifies every frame CRC. */ - private async rebuildSidecar(meta: SegmentMeta): Promise { - const frames = await this.readAllFrames(meta) - const idx: SidecarIndex = { version: 1, generations: [], ids: {} } - for (const f of frames) { - idx.generations.push([f.generation, f.offset, f.frameLen]) - for (const r of f.records) { - const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}` - ;(idx.ids[key] ??= []).push(f.generation) - } - } - return idx - } - - private decodeFrame( - payload: Uint8Array - ): { generation: number; timestamp: number; delta: unknown; records: FoldGeneration['records'] } { - const [generation, timestamp, delta, rawRecords] = msgpackDecode(payload) as [ - number, - number, - unknown, - Array<[number, string, unknown]>, - number - ] - return { - generation, - timestamp, - delta, - records: rawRecords.map(([kindByte, id, record]) => ({ - kind: kindByte === 0 ? ('noun' as const) : ('verb' as const), - id, - record - })) - } - } - - private async readAllFrames(meta: SegmentMeta): Promise< - Array & { offset: number; frameLen: number }> - > { - const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`) - if (!bytes) { - throw new Error( - `[GenerationSegments] sealed segment ${meta.file} is MISSING — packed history is damaged; ` + - `refusing to continue silently` - ) - } - const out: Array & { offset: number; frameLen: number }> = [] - let at = MAGIC.length - const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength) - while (at + FRAME_PREFIX_BYTES <= bytes.length) { - const payloadLen = view.getUint32(at, true) - const crc = view.getUint32(at + 4, true) - const payload = bytes.subarray(at + FRAME_PREFIX_BYTES, at + FRAME_PREFIX_BYTES + payloadLen) - if (payload.length !== payloadLen || crc32c(payload) !== crc) { - throw new Error( - `[GenerationSegments] frame CRC mismatch in ${meta.file} at offset ${at} — ` + - `packed history is damaged; refusing to serve it` - ) - } - out.push({ ...this.decodeFrame(payload), offset: at, frameLen: FRAME_PREFIX_BYTES + payloadLen }) - at += FRAME_PREFIX_BYTES + payloadLen - } - return out - } - - /** Read one packed generation's frame via its sidecar offset (one ranged read). */ - private async readFrame( - gen: number - ): Promise | null> { - const meta = this.coveringSegment(gen) - if (!meta) return null - const idx = await this.sidecarFor(meta) - // generations ascending → binary search. - const gens = idx.generations - let lo = 0 - let hi = gens.length - 1 - while (lo <= hi) { - const mid = (lo + hi) >> 1 - if (gens[mid][0] < gen) lo = mid + 1 - else if (gens[mid][0] > gen) hi = mid - 1 - else { - const [, offset, frameLen] = gens[mid] - const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`) - if (!bytes) { - throw new Error(`[GenerationSegments] sealed segment ${meta.file} is MISSING`) - } - const frame = bytes.subarray(offset, offset + frameLen) - const view = new DataView(frame.buffer, frame.byteOffset, frame.byteLength) - const payloadLen = view.getUint32(0, true) - const crc = view.getUint32(4, true) - const payload = frame.subarray(FRAME_PREFIX_BYTES, FRAME_PREFIX_BYTES + payloadLen) - if (payload.length !== payloadLen || crc32c(payload) !== crc) { - throw new Error( - `[GenerationSegments] frame CRC mismatch for generation ${gen} in ${meta.file} — ` + - `packed history is damaged; refusing to serve it` - ) - } - return this.decodeFrame(payload) - } - } - // In the covering range but not present: the packed tier is dense by - // construction (fold packs every generation it is handed, including - // record-less ones) — absence inside a sealed range is damage. - throw new Error( - `[GenerationSegments] generation ${gen} is inside sealed segment ${meta.file}'s declared ` + - `range but has no frame — packed history is damaged` - ) - } - - /** The packed tier's delta for `gen` (null = not packed). */ - async readDelta(gen: number): Promise<{ delta: unknown; timestamp: number } | null> { - const frame = await this.readFrame(gen) - return frame ? { delta: frame.delta, timestamp: frame.timestamp } : null - } - - /** The packed tier's full record-set for `gen` (null = not packed). */ - async readRecords(gen: number): Promise { - const frame = await this.readFrame(gen) - return frame ? frame.records : null - } - - /** One packed before-image (null = not packed OR no record for the id in that generation). */ - async readRecord(gen: number, kind: 'noun' | 'verb', id: string): Promise { - const frame = await this.readFrame(gen) - if (!frame) return null - const hit = frame.records.find((r) => r.kind === kind && r.id === id) - return hit ? hit.record : null - } - - /** - * D3 reclaim: drop WHOLE segments whose lastGeneration < `belowGeneration` - * and bump `compactedBelow`. Partial segments are never dropped — the - * boundary waits. NEVER called under the archival profile (the caller - * enforces retention semantics; this method only executes boundary drops). - */ - async dropSegmentsBelow(belowGeneration: number): Promise<{ dropped: number; compactedBelow: number }> { - const keep: SegmentMeta[] = [] - const drop: SegmentMeta[] = [] - for (const s of this.manifest.segments) { - ;(s.lastGeneration < belowGeneration ? drop : keep).push(s) - } - if (drop.length === 0) { - return { dropped: 0, compactedBelow: this.manifest.compactedBelow } - } - const compactedBelow = Math.max( - this.manifest.compactedBelow, - drop[drop.length - 1].lastGeneration + 1 - ) - // Manifest first (the drop is authoritative once named), then bytes — - // a crash between leaves orphan segment files invisible to the manifest, - // harmless and re-collectable. - const next: SegmentManifest = { ...this.manifest, compactedBelow, segments: keep } - await this.storage.writeRawObject(MANIFEST_PATH, next) - await this.storage.syncRawObjects([MANIFEST_PATH]) - this.manifest = next - for (const s of drop) { - await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${s.file}`) - await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${sidecarFileName(s.firstGeneration)}`) - this.sidecars.delete(s.file) - } - return { dropped: drop.length, compactedBelow } - } - - /** - * D8 rider — the packed portion of `generationDigest(g)`: a deterministic - * crc32c chain over sealed-segment checksums fully below `g`, plus the - * frame CRC of `g`'s own frame when `g` is mid-segment. O(segments), not - * O(generations); identical history ⇒ identical digest on any machine. - * The live-tier portion is composed by the caller. - */ - async digestThroughPacked(g: number): Promise { - let digest = 0 - let covered = false - for (const s of this.manifest.segments) { - if (s.lastGeneration <= g) { - digest = crc32c(new TextEncoder().encode(`${digest}:${s.checksum}`)) - if (s.lastGeneration === g) covered = true - } else if (s.firstGeneration <= g) { - // g is mid-segment: chain the partial prefix via g's frame CRC. - const frame = await this.readFrame(g) - if (frame === null) return null - const idx = await this.sidecarFor(s) - const upTo = idx.generations.filter(([gen]) => gen <= g) - for (const [gen, offset, frameLen] of upTo) { - digest = crc32c(new TextEncoder().encode(`${digest}:${gen}:${offset}:${frameLen}`)) - } - covered = true - break - } - } - return covered || this.manifest.segments.length > 0 ? digest : null - } -} diff --git a/src/db/generationStore.ts b/src/db/generationStore.ts index aede17a4..4e3738d9 100644 --- a/src/db/generationStore.ts +++ b/src/db/generationStore.ts @@ -46,8 +46,6 @@ import type { TxLogEntry } from './types.js' import { FactLog, storageSupportsFactLog, type CommitFact, type FactOp } from './factLog.js' -import { GenerationSegmentStore, type FoldGeneration } from './generationSegments.js' -import { crc32c } from '../utils/crc32c.js' /** * The byte-identical before-images of every id a commit touches, read UNDER @@ -268,21 +266,6 @@ export class GenerationStore { */ private historyBytesTotal: number | null = null - /** - * The packed tier (D1+D3): sealed segments holding folded cold - * generations. Null until {@link open} wires it (and on storage adapters - * without raw-byte primitives — the live tier then carries everything, - * exactly as before the packed tier existed). - */ - private segments: GenerationSegmentStore | null = null - - /** - * Live-tier window: generations newer than `committed - REPACK_LIVE_WINDOW` - * are never folded — the hot tail stays in the per-generation layout the - * write path owns. Matches the resident chain window's scale. - */ - static readonly REPACK_LIVE_WINDOW = 1024 - /** * Model-B per-write group-commit — the in-memory PENDING tier. * @@ -450,33 +433,6 @@ export class GenerationStore { this.factLog = null } - // PACKED TIER (D1+D3): same capability gate as the fact log. Opening - // reads ONE manifest — never a listing of the packed backlog — and seeds - // committedRanges with the sealed ranges so packed generations resolve - // exactly like live ones. - if (storageSupportsFactLog(this.storage)) { - this.segments = new GenerationSegmentStore(this.storage) - await this.segments.open() - const packedRanges = this.segments - .segments() - .map((s): [number, number] => [s.firstGeneration, Math.min(s.lastGeneration, this.committed)]) - .filter(([lo, hi]) => lo <= hi) - if (packedRanges.length > 0) { - // Merge packed (older) + live (newer) interval sets — both ascending; - // coalesce adjacency so range arithmetic stays interval-exact. - const merged: Array<[number, number]> = [] - for (const r of [...packedRanges, ...this.committedRanges].sort((a, b) => a[0] - b[0])) { - const last = merged[merged.length - 1] - if (last && r[0] <= last[1] + 1) last[1] = Math.max(last[1], r[1]) - else merged.push([r[0], r[1]]) - } - this.committedRanges = merged - } - this.horizonGen = Math.max(this.horizonGen, this.segments.compactedBelow() - 1) - } else { - this.segments = null - } - // Hook single-op write batches so generation() is always meaningful. // Suppressed while a transact batch executes (the batch is ONE generation). if (!options?.readOnly) { @@ -544,51 +500,6 @@ export class GenerationStore { * deltas (cache-bounded reads). * @returns Counts, bytes, generation range, and the compaction horizon. */ - /** - * @description D8 (gate-to-generation provenance): a deterministic content - * digest of the generation log THROUGH `g` — identical history ⇒ identical - * digest on any machine; any divergence (different records, different - * order, reclaimed range) ⇒ different digest. Composed from the packed - * tier's sealed-segment checksum chain (O(segments)) plus the live tier's - * per-generation delta digests (O(live window at most)). Release gates pin - * {generation, digest} and verify both at execution time. - * @param g - The generation to digest through (≤ committed). - * @returns A hex digest string, stable across reopen and repacking states - * ONLY for fully-packed prefixes — repacking changes representation, so - * the composed digest is defined over CONTENT: live-tier gens hash their - * delta + record ids, packed gens hash via frame CRCs. A gate should pin - * after a repack pass for long-term stability, or re-pin on repack. - */ - async generationDigest(g: number): Promise { - if (!Number.isInteger(g) || g < 1 || g > this.committed) { - throw new RangeError( - `generationDigest(): generation ${g} is out of range [1, ${this.committed}]` - ) - } - if (g <= this.horizonGen) { - throw new GenerationCompactedError(g, this.horizonGen) - } - let digest = 0 - const enc = new TextEncoder() - if (this.segments) { - const packed = await this.segments.digestThroughPacked(g) - if (packed !== null) digest = packed - } - // Live-tier composition: every committed gen ≤ g not covered by a sealed - // segment hashes its delta content in ascending order. - for (const gen of this.committedGensAsc()) { - if (gen > g) break - if (this.segments?.hasGeneration(gen)) continue - const delta = await this.getDelta(gen) - digest = crc32c( - enc.encode( - `${digest}:${gen}:${delta.timestamp}:${[...delta.nouns].sort().join(',')}:${[...delta.verbs].sort().join(',')}` - ) - ) - } - return digest.toString(16).padStart(8, '0') - } - async historyStats(): Promise<{ generations: number bytes: number @@ -627,17 +538,14 @@ export class GenerationStore { try { paths = await this.storage.listRawObjects(`${GENERATIONS_PREFIX}/${gen}/prev`) } catch { - paths = [] + return [] } const records: GenerationRecord[] = [] for (const p of paths) { const record = (await this.storage.readRawObject(p)) as GenerationRecord | null if (record) records.push(record) } - if (records.length > 0) return records - // Two-tier: folded generations serve their record-set from the segment. - const packed = await this.segments?.readRecords(gen) - return packed ? (packed.map((r) => r.record) as GenerationRecord[]) : [] + return records } /** @@ -1875,15 +1783,9 @@ export class GenerationStore { if (pending) { return (kind === 'noun' ? pending.nouns : pending.verbs).get(id) ?? null } - const live = (await this.storage.readRawObject( + return (await this.storage.readRawObject( `${GENERATIONS_PREFIX}/${gen}/prev/${id}.json` )) as GenerationRecord | null - if (live) return live - // Two-tier: the packed tier serves folded generations (live-tier-wins). - if (this.segments?.hasGeneration(gen)) { - return (await this.segments.readRecord(gen, kind, id)) as GenerationRecord | null - } - return null } /** @@ -2230,21 +2132,6 @@ export class GenerationStore { `${GENERATIONS_PREFIX}/${gen}/tx.json` )) as GenerationDelta | null if (delta === null) { - // Two-tier read (D1+D3): not in the live tier → the packed tier. - // Live-tier-wins ordering (a crash mid-fold leaves a duplicate, never - // a gap), so the segment lookup runs only after the live miss. - const packed = await this.segments?.readDelta(gen) - if (packed) { - const d = packed.delta as GenerationDelta - const entry = { - nouns: new Set(d.nouns), - verbs: new Set(d.verbs), - timestamp: packed.timestamp, - bytes: d.bytes ?? 0 - } - this.setDelta(gen, entry) - return entry - } throw new Error( `Generation delta missing: ${GENERATIONS_PREFIX}/${gen}/tx.json ` + `(store corrupted or records removed outside compactHistory())` @@ -2326,94 +2213,6 @@ export class GenerationStore { * @param options - Retention caps (see {@link CompactHistoryOptions}). * @returns Count of removed record-sets and the new horizon. */ - /** - * @description The REPACKER (D1+D3+repacking): fold cold live-tier - * generations into sealed segments — re-representation, never deletion. - * Every record and delta stays readable (asOf/chains unchanged); the - * per-generation directories are deleted only AFTER their segment is - * durable (crash between = duplicate representation, resolved - * live-tier-wins by every reader; never a gap). This is the transform that - * takes a 70k-file history to tens of segment files, and the ONLY history - * transform permitted under the archival profile. - * - * Folds oldest-first, contiguous from the packed boundary, in batches, and - * stops at the live window ({@link GenerationStore.REPACK_LIVE_WINDOW}) - * or when `timeBudgetMs` is spent — an early stop is a consistent prefix; - * the next pass resumes. - */ - async repackHistory(options?: { timeBudgetMs?: number; batchGenerations?: number }): Promise<{ - foldedGenerations: number - segmentsCreated: number - }> { - if (!this.segments) return { foldedGenerations: 0, segmentsCreated: 0 } - const segments = this.segments - return this.withMutex(async () => { - const deadline = - options?.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined - const batchSize = options?.batchGenerations ?? 512 - const coldCeiling = this.committed - GenerationStore.REPACK_LIVE_WINDOW - const packedThrough = - segments.segments().length > 0 - ? segments.segments()[segments.segments().length - 1].lastGeneration - : 0 - - // Cold, unpacked, committed generations — ascending, contiguous scan. - const eligible: number[] = [] - for (const gen of this.committedGensAsc()) { - if (gen > coldCeiling) break - if (gen <= packedThrough) continue // already packed (dup fold barred) - if (this.pendingBuffer.has(gen)) continue // un-flushed = live by definition - eligible.push(gen) - } - - let folded = 0 - let segmentsCreated = 0 - for (let i = 0; i < eligible.length; i += batchSize) { - if (deadline !== undefined && Date.now() >= deadline) break - const batch = eligible.slice(i, i + batchSize) - const foldInput: FoldGeneration[] = [] - for (const gen of batch) { - const delta = (await this.storage.readRawObject( - `${GENERATIONS_PREFIX}/${gen}/tx.json` - )) as GenerationDelta | null - if (delta === null) { - // Already folded by a prior crashed pass whose dirs were removed, - // or damage — getDelta's two-tier read decides which, loudly, - // when someone asks. Skip; never fold a generation we cannot read. - continue - } - const records: FoldGeneration['records'] = [] - for (const [kind, ids] of [ - ['noun', delta.nouns] as const, - ['verb', delta.verbs] as const - ]) { - for (const id of ids) { - const record = await this.storage.readRawObject( - `${GENERATIONS_PREFIX}/${gen}/prev/${id}.json` - ) - if (record) records.push({ kind, id, record }) - } - } - foldInput.push({ generation: gen, timestamp: delta.timestamp, delta, records }) - } - if (foldInput.length === 0) continue - await segments.fold(foldInput) - segmentsCreated++ - // Segment + manifest durable → the live copies retire. - for (const g of foldInput) { - await this.storage.removeRawPrefix(`${GENERATIONS_PREFIX}/${g.generation}`) - } - folded += foldInput.length - } - if (folded > 0) { - prodLog.info( - `[GenerationStore] repacked ${folded} cold generation(s) into ${segmentsCreated} segment(s) — history preserved, file count reduced` - ) - } - return { foldedGenerations: folded, segmentsCreated } - }) - } - async compact(options?: CompactHistoryOptions): Promise { return this.withMutex(async () => { const minPinned = this.minPinnedGeneration() @@ -2505,16 +2304,6 @@ export class GenerationStore { // Reclaimed generations leave the per-id chains stale → rebuild on next read. this.invalidateChains() this.horizonGen = Math.max(this.horizonGen, highestRemoved) - // Packed-tier reclaim (D3): a packed generation's bytes live in a - // sealed segment — removeRawPrefix above was a no-op for it. Drop - // WHOLE segments now fully below the horizon; a partially-reclaimed - // segment keeps its bytes until the boundary passes it (the frozen - // partial-segments-wait rule; logical reclamation above still holds — - // the generations left committedRanges and asOf below the horizon - // throws regardless). - if (this.segments) { - await this.segments.dropSegmentsBelow(this.horizonGen + 1) - } const manifest: GenerationManifest = { version: 1, generation: this.committed, diff --git a/src/db/whereMatcher.ts b/src/db/whereMatcher.ts index 8dab02fd..c5469209 100644 --- a/src/db/whereMatcher.ts +++ b/src/db/whereMatcher.ts @@ -61,44 +61,41 @@ export class UnsupportedWhereOperatorError extends Error { * @returns The field's value, or `undefined` when absent. */ export function resolveEntityField(entity: Entity, field: string): unknown { - // THE ONE ADDRESSING LAW (sealed 2026-08-03): `system.` reads the - // entity scalar; bare and `metadata.`-prefixed names read the user's - // metadata bag (dotted paths traverse INSIDE the bag). The old bare-name - // switch over system fields is dead — bare `createdAt` is the user's own - // field now; the engine scalar is `system.createdAt`. Plumbing (vector, - // connections, level, data, _rev) is invisible: no spelling reaches it. - if (field.startsWith('system.')) { - switch (field.slice('system.'.length)) { - case 'type': - return entity.type - case 'subtype': - return entity.subtype - case 'id': - return entity.id - case 'createdAt': - return entity.createdAt - case 'updatedAt': - return entity.updatedAt - case 'service': - return entity.service - case 'createdBy': - return entity.createdBy - case 'confidence': - return entity.confidence - case 'weight': - return entity.weight - case 'visibility': - return (entity as unknown as Record).visibility - } - // Out-of-map system spelling: parse refuses these upstream with a typed - // error; reaching here (internal callers only) reads as absent. - return undefined + switch (field) { + case 'noun': + case 'type': + return entity.type + case 'subtype': + return entity.subtype + case 'id': + return entity.id + case 'createdAt': + return entity.createdAt + case 'updatedAt': + return entity.updatedAt + case 'service': + return entity.service + case 'createdBy': + return entity.createdBy + case 'confidence': + return entity.confidence + case 'weight': + return entity.weight + case '_rev': + return entity._rev + case 'data': + return entity.data } - const path = field.startsWith('metadata.') ? field.slice('metadata.'.length) : field - const bag = (entity.metadata ?? {}) as Record - if (!path.includes('.')) return bag[path] - return resolvePath(bag, path) + if (field.includes('.')) { + // Dotted path: resolve against the whole entity first (`metadata.x`), + // then against the metadata bag (`address.city` on nested metadata). + const fromEntity = resolvePath(entity as unknown as Record, field) + if (fromEntity !== undefined) return fromEntity + return resolvePath((entity.metadata ?? {}) as Record, field) + } + + return ((entity.metadata ?? {}) as Record)[field] } /** Walk a dotted path through nested plain objects. */ diff --git a/src/import/ImportCoordinator.ts b/src/import/ImportCoordinator.ts index dd145045..1e1316b7 100644 --- a/src/import/ImportCoordinator.ts +++ b/src/import/ImportCoordinator.ts @@ -22,6 +22,7 @@ import { SmartYAMLImporter } from '../importers/SmartYAMLImporter.js' import { SmartDOCXImporter } from '../importers/SmartDOCXImporter.js' import { VFSStructureGenerator } from '../importers/VFSStructureGenerator.js' import { NounType, VerbType } from '../types/graphTypes.js' +import { splitNounMetadataRecord, splitVerbMetadataRecord } from '../types/reservedFields.js' import { v4 as uuidv4 } from '../universal/uuid.js' import * as fs from 'fs' import * as path from 'path' @@ -870,18 +871,35 @@ export class ImportCoordinator { } /** - * Normalize an extractor/consumer metadata bag for spreading — the - * field-addressing law: the bag is the user's, VERBATIM. No name is - * reserved anymore ('confidence', 'subtype', 'type', … in a source bag - * import as ordinary user fields); the old reserved-key strip was data - * loss under the law and is gone. A forged 'system.'-prefixed key still - * refuses loudly at the write door (`rejectForgedSystemKeys`). + * Strip Brainy-reserved entity keys out of an extractor-supplied metadata bag. + * + * Extractors (and consumer `customMetadata`) can carry reserved keys + * (`confidence`, `subtype`, `weight`, …) inside `metadata`. Brainy 8.0's + * default `reservedFieldPolicy` is `'throw'`, so spreading such a bag into + * `add({ metadata })` would reject the whole import. The import pipeline owns + * the correct write path: user-mutable reserved values are passed as dedicated + * `AddParams` params (see the call sites), so here we simply drop the reserved + * half of the bag and keep only the custom fields that belong in `metadata`. + * * @param bag - The extractor/consumer metadata bag (may be undefined). - * @returns The bag itself, or `{}` for non-object inputs. + * @returns The custom-only metadata (reserved keys removed). */ - private bagVerbatim(bag: Record | undefined | null): Record { + private stripReservedFromBag(bag: Record | undefined | null): Record { if (!bag || typeof bag !== 'object') return {} - return bag + return splitNounMetadataRecord(bag).custom + } + + /** + * Relationship mirror of {@link stripReservedFromBag} — strips reserved verb + * keys (`verb`, `confidence`, `weight`, `subtype`, …) out of an edge metadata + * bag so it carries only custom fields. Reserved values that have a dedicated + * `RelateParams` param are passed there by the call site instead. + * @param bag - The extractor/consumer edge metadata bag (may be undefined). + * @returns The custom-only edge metadata (reserved keys removed). + */ + private stripReservedFromRelationBag(bag: Record | undefined | null): Record { + if (!bag || typeof bag !== 'object') return {} + return splitVerbMetadataRecord(bag).custom } /** @@ -999,7 +1017,7 @@ export class ImportCoordinator { importedAt: trackingContext.importedAt, importFormat: trackingContext.importFormat, importSource: trackingContext.importSource, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromBag(trackingContext.customMetadata) }) } }) @@ -1027,11 +1045,13 @@ export class ImportCoordinator { data: entity.description || entity.name, type: entity.type, subtype: entity.subtype ?? options.defaultSubtype ?? 'imported', - // Engine confidence rides its dedicated param; the bag below is - // the user's verbatim (no name is reserved — field-addressing law). + // `confidence` is a reserved field — pass it as the dedicated param, + // never inside the metadata bag (8.0 reservedFieldPolicy defaults to 'throw'). confidence: entity.confidence, metadata: { - ...this.bagVerbatim(entity.metadata), + // Extractor/consumer bags may smuggle reserved keys — strip them so + // the bag carries only custom fields. + ...this.stripReservedFromBag(entity.metadata), name: entity.name, vfsPath: vfsFile?.path, importedFrom: 'import-coordinator', @@ -1044,7 +1064,7 @@ export class ImportCoordinator { importSource: trackingContext.importSource, sourceRow: row.rowNumber, sourceSheet: row.sheet, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromBag(trackingContext.customMetadata) }) } } @@ -1125,7 +1145,7 @@ export class ImportCoordinator { importIds: [trackingContext.importId], projectId: trackingContext.projectId, importFormat: trackingContext.importFormat, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromRelationBag(trackingContext.customMetadata) }) } } @@ -1160,7 +1180,7 @@ export class ImportCoordinator { confidence: entity.confidence, metadata: { // Strip any reserved keys an extractor smuggled into the bag. - ...this.bagVerbatim(entity.metadata), + ...this.stripReservedFromBag(entity.metadata), name: entity.name, vfsPath: vfsFile?.path, importedFrom: 'import-coordinator', @@ -1174,7 +1194,7 @@ export class ImportCoordinator { importSource: trackingContext.importSource, sourceRow: row.rowNumber, sourceSheet: row.sheet, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromBag(trackingContext.customMetadata) }) } }) @@ -1214,7 +1234,7 @@ export class ImportCoordinator { importIds: [trackingContext.importId], projectId: trackingContext.projectId, importFormat: trackingContext.importFormat, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromRelationBag(trackingContext.customMetadata) }) } }) @@ -1269,7 +1289,7 @@ export class ImportCoordinator { projectId: trackingContext.projectId, importedAt: trackingContext.importedAt, importFormat: trackingContext.importFormat, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromBag(trackingContext.customMetadata) }) } }) @@ -1299,7 +1319,7 @@ export class ImportCoordinator { projectId: trackingContext.projectId, importedAt: trackingContext.importedAt, importFormat: trackingContext.importFormat, - ...this.bagVerbatim(trackingContext.customMetadata) + ...this.stripReservedFromRelationBag(trackingContext.customMetadata) }) } }) @@ -1402,7 +1422,7 @@ export class ImportCoordinator { ...(typeof (rel as any).confidence === 'number' && { confidence: (rel as any).confidence }), ...(typeof (rel as any).weight === 'number' && { weight: (rel as any).weight }), metadata: { - ...this.bagVerbatim(rel.metadata), + ...this.stripReservedFromRelationBag(rel.metadata), relationshipType: 'semantic', // Distinguish from VFS/provenance inferredType: verbType !== rel.type, // Track if type was enhanced originalType: rel.type diff --git a/src/index.ts b/src/index.ts index 3186a6a7..10922adb 100644 --- a/src/index.ts +++ b/src/index.ts @@ -89,12 +89,7 @@ export { RESERVED_ENTITY_FIELDS, RESERVED_RELATION_FIELDS, splitNounMetadataRecord, - splitVerbMetadataRecord, - buildNounMetadataRecord, - buildVerbMetadataRecord, - isNestedBagRecord, - METADATA_RECORD_FORMAT_KEY, - NESTED_BAG_FORMAT + splitVerbMetadataRecord } from './types/reservedFields.js' export type { ReservedEntityField, @@ -111,25 +106,6 @@ export type { // Export Aggregation Engine export { AggregationIndex, AggregateMaterializer, bucketTimestamp, parseBucketRange } from './aggregation/index.js' -// THE ONE FIELD-ADDRESSING LAW (sealed 2026-08-03) — the arming surface both -// engines' conformance suites detect: bare names = user metadata, system.* = -// the ten ruled scalars, plumbing invisible, refusals typed with the fix in -// the message. See docs/concepts/field-addressing.md. -export { - FIELD_ADDRESSING_CAPABILITY, - SYSTEM_ENTITY_SCALARS, - SYSTEM_RELATION_SCALARS, - PLUMBING_FIELDS, - parseFieldAddress, - readEntityFieldAddress, - readRelationFieldAddress, - buildUnresolvableMessage, - InvalidFieldAddressError, - UnresolvableFieldError, - UnsupportedFindOptionError -} from './db/fieldAddressing.js' -export type { FieldAddress, FieldAddressKind } from './db/fieldAddressing.js' - // Export Neural Import (AI data understanding) export { NeuralImport } from './neural/neuralImport.js' export type { @@ -247,7 +223,6 @@ export type { CommitFact, FactOp, FactScanBatch, - SCANFACTS_FIRST_BATCH_MS, FactScanHandle } from './db/factLog.js' // The generalized family stamp — which source generation a projection diff --git a/src/migration/MigrationRunner.ts b/src/migration/MigrationRunner.ts index 6a8a34bd..d2e251a2 100644 --- a/src/migration/MigrationRunner.ts +++ b/src/migration/MigrationRunner.ts @@ -9,67 +9,6 @@ import type { BaseStorage } from '../storage/baseStorage.js' import type { NounMetadata, VerbMetadata } from '../coreTypes.js' import type { Migration, MigrationState, MigrationPreview, MigrationResult, MigrateOptions, MigrationError } from './types.js' import { MIGRATIONS } from './migrations.js' -import { - splitNounMetadataRecord, - splitVerbMetadataRecord, - buildNounMetadataRecord, - buildVerbMetadataRecord, - RESERVED_ENTITY_FIELDS, - RESERVED_RELATION_FIELDS -} from '../types/reservedFields.js' - -const RESERVED_NOUN_SET: ReadonlySet = new Set(RESERVED_ENTITY_FIELDS) -const RESERVED_VERB_SET: ReadonlySet = new Set(RESERVED_RELATION_FIELDS) - -/** - * Normalize a stored record (either era: legacy flat OR v2 nested-bag) into - * THE transform view — the one shape every migration transform receives: - * engine fields top-level, the user's metadata bag nested under `metadata`. - * Transforms never see the storage era; a migration written today works on - * a brain of any age. - */ -function toTransformView( - record: Record, - kind: 'noun' | 'verb' -): Record { - const { reserved, custom } = - kind === 'noun' ? splitNounMetadataRecord(record) : splitVerbMetadataRecord(record) - return { ...reserved, metadata: { ...custom } } -} - -/** - * Convert a transform's returned view back into a stamped v2 stored record. - * LOUD CONTRACT: user fields belong inside `.metadata` — a stray top-level - * key that is not an engine field is a migration bug under the - * field-addressing law (pre-law transforms wrote user fields flat), and it - * refuses with the fix in the message rather than silently dropping or - * silently storing it as an engine key. - */ -function fromTransformView( - view: Record, - kind: 'noun' | 'verb' -): Record { - const reservedSet = kind === 'noun' ? RESERVED_NOUN_SET : RESERVED_VERB_SET - const engine: Record = {} - for (const [key, value] of Object.entries(view)) { - if (key === 'metadata') continue - if (!reservedSet.has(key)) { - throw new Error( - `migration transform returned a top-level key '${key}' that is not an ` + - `engine field — under the field-addressing law user fields live inside ` + - `.metadata (return { ...view, metadata: { ...view.metadata, ${key}: … } }).` - ) - } - engine[key] = value - } - const bag = - view.metadata && typeof view.metadata === 'object' && !Array.isArray(view.metadata) - ? (view.metadata as Record) - : {} - return kind === 'noun' - ? buildNounMetadataRecord(engine, bag) - : buildVerbMetadataRecord(engine, bag) -} const MIGRATION_STATE_KEY = '__migration_state__' const PREVIEW_SAMPLE_SIZE = 5 @@ -186,16 +125,14 @@ export class MigrationRunner { const entityMeta = metadataBatch.get(entity.id) if (!entityMeta) continue - // Transforms see THE view (engine fields + nested user bag), - // never the raw storage era. - const view = toTransformView(entityMeta as Record, 'noun') - const result = this.applyTransforms(view, nounMigrations) + const metadata = entityMeta as Record + const result = this.applyTransforms(metadata, nounMigrations) if (result !== null) { affectedEntities++ if (sampleChanges.length < PREVIEW_SAMPLE_SIZE) { sampleChanges.push({ id: entity.id, - before: view, + before: { ...metadata }, after: result }) } @@ -220,14 +157,14 @@ export class MigrationRunner { const verbMeta = await this.storage.getVerbMetadata(verb.id) if (!verbMeta) continue - const view = toTransformView(verbMeta as Record, 'verb') - const result = this.applyTransforms(view, verbMigrations) + const metadata = verbMeta as Record + const result = this.applyTransforms(metadata, verbMigrations) if (result !== null) { affectedEntities++ if (sampleChanges.length < PREVIEW_SAMPLE_SIZE) { sampleChanges.push({ id: verb.id, - before: view, + before: { ...metadata }, after: result }) } @@ -352,16 +289,9 @@ export class MigrationRunner { if (!entityMeta) continue try { - const transformed = migration.transform( - toTransformView(entityMeta as Record, 'noun') - ) + const transformed = migration.transform(entityMeta as Record) if (transformed !== null) { - // Re-stamp as a v2 record (also upgrades legacy records touched - // by a migration onto the nested-bag shape). - await this.storage.saveNounMetadata( - entity.id, - fromTransformView(transformed, 'noun') as NounMetadata - ) + await this.storage.saveNounMetadata(entity.id, transformed as NounMetadata) modified++ } } catch (err) { @@ -427,14 +357,9 @@ export class MigrationRunner { if (!metadata) continue try { - const transformed = migration.transform( - toTransformView(metadata as Record, 'verb') - ) + const transformed = migration.transform(metadata as Record) if (transformed !== null) { - await this.storage.saveVerbMetadata( - verb.id, - fromTransformView(transformed, 'verb') as VerbMetadata - ) + await this.storage.saveVerbMetadata(verb.id, transformed as VerbMetadata) modified++ } } catch (err) { diff --git a/src/migration/types.ts b/src/migration/types.ts index a63e40b9..2dcc1d1a 100644 --- a/src/migration/types.ts +++ b/src/migration/types.ts @@ -14,19 +14,7 @@ export interface Migration { description: string /** Which entity types this migration applies to */ applies: 'nouns' | 'verbs' | 'both' - /** - * Return the transformed record view, or null if no change needed. - * - * THE VIEW CONTRACT (field-addressing law): the transform receives ONE - * normalized shape regardless of how old the stored record is — engine - * fields top-level (`noun`/`verb`, `subtype`, `confidence`, `weight`, - * timestamps, `_rev`, …) and the USER's metadata bag nested under - * `metadata` (where every name is the user's, engine spellings included). - * Return the same shape: user-field changes go inside `.metadata`; a - * stray non-engine top-level key in the returned object refuses loudly - * (it is the pre-law flat habit, and silently guessing its namespace - * would corrupt data). - */ + /** Return transformed metadata, or null if no change needed */ transform: (metadata: Record) => Record | null } diff --git a/src/neural/neuralImport.ts b/src/neural/neuralImport.ts index ed240a3f..c8d19eb5 100644 --- a/src/neural/neuralImport.ts +++ b/src/neural/neuralImport.ts @@ -7,6 +7,7 @@ import { Brainy } from '../brainy.js' import { NounType, VerbType } from '../types/graphTypes.js' +import { splitNounMetadataRecord, splitVerbMetadataRecord } from '../types/reservedFields.js' import * as fs from '../universal/fs.js' import * as path from '../universal/path.js' // @ts-ignore @@ -802,14 +803,12 @@ export class NeuralImport { data: this.extractMainText(entity.originalData), type: entity.nounType as NounType, subtype: entity.subtype ?? options.defaultSubtype ?? 'extracted', - // Engine confidence rides its dedicated param; the source object - // imports as the user's bag VERBATIM — no name is reserved - // (field-addressing law). + // `confidence` is a reserved field — dedicated param, not metadata + // (8.0 reservedFieldPolicy defaults to 'throw'). confidence: entity.confidence, metadata: { - ...(typeof entity.originalData === 'object' && entity.originalData !== null - ? entity.originalData - : {}), + // Strip any reserved keys the source data smuggled into the bag. + ...splitNounMetadataRecord(entity.originalData).custom, id: entity.suggestedId } }) @@ -823,13 +822,11 @@ export class NeuralImport { type: relationship.verbType as VerbType, subtype: relationship.subtype ?? options.defaultSubtype ?? 'extracted', weight: relationship.weight, - confidence: relationship.confidence, // engine confidence — dedicated param + confidence: relationship.confidence, // reserved field — dedicated param, not metadata metadata: { context: relationship.context, - // The edge bag imports verbatim — no name is reserved. - ...(typeof relationship.metadata === 'object' && relationship.metadata !== null - ? relationship.metadata - : {}) + // Strip any reserved keys smuggled into the edge metadata bag. + ...splitVerbMetadataRecord(relationship.metadata).custom } }) } diff --git a/src/storage/adapters/memoryStorage.ts b/src/storage/adapters/memoryStorage.ts index 1b1f412e..bab9d4d9 100644 --- a/src/storage/adapters/memoryStorage.ts +++ b/src/storage/adapters/memoryStorage.ts @@ -133,11 +133,6 @@ export class MemoryStorage extends BaseStorage { */ protected async deleteObjectFromPath(path: string): Promise { this.objectStore.delete(path) - // Filesystem parity: on disk, objects and raw BYTE files are both just - // files — unlink removes whichever exists. Without this, deleteRawObject - // on a raw-bytes path (fact-log/generation segments) silently no-ops on - // memory storage: the delete "succeeds" and the bytes remain. - this.rawBytesStore.delete(path) } /** diff --git a/src/storage/baseStorage.ts b/src/storage/baseStorage.ts index b78b4a49..6daf09c0 100644 --- a/src/storage/baseStorage.ts +++ b/src/storage/baseStorage.ts @@ -36,8 +36,7 @@ import { BrainyError, ProtectedArtifactError, DerivedArtifactMissingError } from import { MetadataWriteBuffer } from '../utils/metadataWriteBuffer.js' import { splitNounMetadataRecord, - splitVerbMetadataRecord, - isNestedBagRecord + splitVerbMetadataRecord } from '../types/reservedFields.js' /** @@ -383,10 +382,6 @@ export abstract class BaseStorage extends BaseStorageAdapter { // identical to the unknown-key fallback these keys hit // before being listed here — this only kills the // per-boot "Unknown key format" warning) - id.startsWith('graph-lsm-') || // Graph-LSM store manifests written through storage by - // an active native graph provider — same - // warn-then-route fallback as above; listing the family - // silences the per-boot warning on provider-backed brains isSingletonSystemKey(id) // Known singletons (e.g. brainy:entityIdMapper) hit the // same warn-then-route fallback without this — the // routing below already handles them identically @@ -1014,14 +1009,8 @@ export abstract class BaseStorage extends BaseStorageAdapter { const hashes: string[] = [] for (const record of records) { if (record.kind !== 'noun') continue - // The VFS blob pointer (`storage: {type:'blob', hash}`) is a USER-bag - // field: in a v2 nested-bag record it lives inside `metadata`, in a - // legacy flat record it sits at the top level — read shape-aware. - const raw = record.metadata as Record | null - const bag = isNestedBagRecord(raw) - ? (raw!.metadata as Record) - : raw - const storage = (bag as { storage?: { type?: string; hash?: unknown } } | null)?.storage + const storage = (record.metadata as { storage?: { type?: string; hash?: unknown } } | null) + ?.storage if (storage?.type === 'blob' && typeof storage.hash === 'string') { hashes.push(storage.hash) } diff --git a/src/storage/brainFormat.ts b/src/storage/brainFormat.ts index 6ef913d4..2e6488e9 100644 --- a/src/storage/brainFormat.ts +++ b/src/storage/brainFormat.ts @@ -69,15 +69,7 @@ export const BRAIN_FORMAT_PATH = '_system/brain-format.json' * (the 8.0 GA baseline). An on-disk `indexEpoch` that differs from this — or an * absent marker — triggers a full derived-index rebuild on open. */ -// Epoch 3 (2026-08-03, the namespace-law pair): the index key format split -// the two namespaces — user fields keep bare flattened keys, the ten system -// scalars moved to literal 'system.' keys (the legacy 'noun' column -// spelling died with them). Every brain rebuilds its derived indexes from -// canonical at first open onto the frozen keys. -// Epoch 2 (2026-08-03, same day, the interim pair): user metadata fields -// named `level` became indexable on both engines; poisoned multi-valued -// `level` columns healed through the rebuild. -export const EXPECTED_INDEX_EPOCH = 3 +export const EXPECTED_INDEX_EPOCH = 1 /** * @description The data-layer format string this build writes and runs as. diff --git a/src/transaction/operations/StorageOperations.ts b/src/transaction/operations/StorageOperations.ts index 9858219b..316f1ac0 100644 --- a/src/transaction/operations/StorageOperations.ts +++ b/src/transaction/operations/StorageOperations.ts @@ -77,27 +77,8 @@ export class SaveNounOperation implements Operation { ? null : await this.storage.getNoun(this.noun.id) - // PRESERVE stored graph state on updates. Callers stage this op with - // placeholder adjacency ({connections: empty, level: 0}) because the - // vector index owns those values and persists them at flush. Codec-era - // records (2.4.0+) carry an empty connections field by design (adjacency - // lives in a separate compressed blob — the placeholder is harmless), but - // LEGACY pre-codec records store adjacency INLINE: writing the - // placeholder over one stamped out its stored connections, leaving a - // crash window (until the next flush) where a reload found the node - // unreachable. Stale adjacency in that window is tolerable — HNSW - // self-corrects at the reindex flush; EMPTY adjacency is silent recall - // loss. The read above is already paid for rollback; preservation is free. - const toSave: HNSWNoun = - previousNoun && this.noun.connections.size === 0 - ? { - ...this.noun, - connections: previousNoun.connections || this.noun.connections, - level: previousNoun.level ?? this.noun.level - } - : this.noun - - await this.storage.saveNoun(toSave) + // Save new noun + await this.storage.saveNoun(this.noun) // Return rollback action return async () => { diff --git a/src/types/brainy.types.ts b/src/types/brainy.types.ts index 6c1f0ffd..b5f286ed 100644 --- a/src/types/brainy.types.ts +++ b/src/types/brainy.types.ts @@ -320,18 +320,15 @@ export interface AddParams { */ visibility?: 'public' | 'internal' /** - * Structured queryable fields — indexed by MetadataIndex, used in `where` - * filters, `orderBy`, and aggregation. + * Structured queryable fields — indexed by MetadataIndex, used in `where` filters. * - * THE FIELD-ADDRESSING LAW: every name here is YOURS. There are no - * reserved metadata names — `confidence`, `type`, `id`, `level`, `data`, - * `content`, … are ordinary user fields that index, filter, sort, and - * aggregate like any other, and survive faithfully across restarts and - * rebuilds. Engine scalars are set only via their dedicated params - * (`confidence`, `weight`, `subtype`, …) and are queried explicitly as - * `system.` (`where: { 'system.confidence': … }`). The ONE illegal - * spelling is a key starting `'system.'` — the engine's explicit address - * namespace cannot be forged; such a write refuses with a typed error. + * Reserved entity fields (`RESERVED_ENTITY_FIELDS` — `noun`, `subtype`, `visibility`, + * `createdAt`, `updatedAt`, `confidence`, `weight`, `service`, `data`, `createdBy`, + * `_rev`) may NOT appear here — they have dedicated top-level params and the type makes + * a literal reserved key a compile error. Untyped (JavaScript) callers that pass one + * anyway are normalized at write time: user-settable fields remap to their top-level + * param (top-level wins when both are supplied), system-managed fields are dropped with + * a one-shot warning. */ metadata?: EntityMetadataInput /** Custom entity ID. When omitted, a time-ordered UUID v7 is generated; a supplied natural-key string is normalized to a stable UUID v5. */ @@ -389,11 +386,12 @@ export interface UpdateParams { */ visibility?: EntityVisibility /** - * Metadata fields to merge (or replace when `merge: false`). Every name is - * the user's (the field-addressing law) — a patch field named `confidence` - * updates YOUR field of that name, never the engine scalar (use the - * dedicated `confidence` param for that). Keys spelled `'system.…'` refuse - * with a typed error (namespace forgery). + * Metadata fields to merge (or replace when `merge: false`). Reserved entity + * fields (`RESERVED_ENTITY_FIELDS`) may NOT appear here — `confidence` / + * `weight` / `subtype` / `visibility` have dedicated params on this call, and the rest + * are system-managed. A literal reserved key is a compile error; untyped callers + * are normalized at write time (remap user-settable, drop system-managed + * with a one-shot warning). */ metadata?: EntityMetadataPatch merge?: boolean // Merge or replace metadata (default: true) @@ -446,11 +444,11 @@ export interface RelateParams { /** Content for the relationship (optional — overrides auto-computed vector) */ data?: any /** - * Structured queryable fields on the edge. Every name is the user's (the - * field-addressing law) — `verb`, `confidence`, `weight`, … in this bag are - * ordinary user fields; engine scalars ride their dedicated params and are - * addressed as `system.`. Keys spelled `'system.…'` refuse with a - * typed error (namespace forgery). + * Structured queryable fields on the edge. Reserved relationship fields + * (`RESERVED_RELATION_FIELDS` — `verb`, `subtype`, `visibility`, `createdAt`, + * `updatedAt`, `confidence`, `weight`, `service`, `data`, `createdBy`, `_rev`) may NOT + * appear here — they have dedicated params. A literal reserved key is a + * compile error; untyped callers are normalized at write time. */ metadata?: RelationMetadataInput /** Create reverse edge too (default: false) */ @@ -480,9 +478,10 @@ export interface UpdateRelationParams { confidence?: number // New confidence (0-1) data?: any // New content /** - * Metadata fields to merge (or replace when `merge: false`). Every name is - * the user's (the field-addressing law); engine scalars ride their - * dedicated params. Keys spelled `'system.…'` refuse with a typed error. + * Metadata fields to merge (or replace when `merge: false`). Reserved + * relationship fields (`RESERVED_RELATION_FIELDS`) may NOT appear here — + * a literal reserved key is a compile error; untyped callers are + * normalized at write time. */ metadata?: RelationMetadataPatch merge?: boolean // Merge or replace metadata @@ -499,43 +498,6 @@ export interface UpdateRelationParams { * - **Graph:** `connected` for relationship traversal (via GraphAdjacencyIndex) * * See also: [Query Operators](../../docs/QUERY_OPERATORS.md) for all `where` operators. - * - * @remarks - * **Field-addressing law.** Governs every query-surface field name — `where` - * and `orderBy` on this interface, plus `AggregateSource.where` and - * `AggregateDefinition.groupBy` in the aggregation engine: - * - * 1. A bare name (e.g. `'level'`, `'rank'`, `'score'`) always means the - * caller's own metadata field — it reads `entity.metadata.`. There - * is no fallback to an engine-internal field of the same name and no - * priority resolution between the two; metadata wins unconditionally. - * 2. `system.` reaches an engine scalar, explicitly, and only for - * these ten: `id`, `type`, `subtype`, `createdAt`, `updatedAt`, - * `confidence`, `weight`, `visibility`, `service`, `createdBy`. - * 3. `vector`, `connections`, `level` (the engine-internal node field — a - * different thing from a user metadata field also named `level`), - * `data`, and `_rev` are invisible plumbing: neither spelling can - * address them from a query surface. - * 4. `metadata.` is the explicit spelling of the bare form and means - * exactly the same thing as rule 1. - * 5. A name that matches none of the above — most often a bare name that - * collides with one of the ten system-scalar names in rule 2 — REFUSES - * with a typed {@link UnresolvableFieldError} naming both candidates, - * e.g. `no metadata field 'createdAt' — did you mean system.createdAt or - * metadata.createdAt?`. The same loud-refusal principle covers whole - * options: the previously accepted-and-silently-ignored `cursor`, - * `includeRelations`, and `writeOnly` now throw - * {@link UnsupportedFindOptionError} instead of doing nothing. - * 6. **Ordering contract** (identical on the pure-JS engine and the native - * accelerator): rows missing or `null` on the `orderBy` field sort LAST - * in BOTH `asc` and `desc` order and are never dropped from the result; - * ties break by `id` ascending. - * - * Migration note: a call site written against the old rule — e.g. - * `orderBy: 'createdAt'` or `where: { visibility: 'internal' }` meaning the - * engine scalar — now refuses instead of silently reading the wrong field. - * The thrown error names the exact fix (`system.createdAt`). A loud - * refusal with the fix in hand beats a silent behavior flip. */ export interface FindParams { // Vector Intelligence @@ -554,18 +516,7 @@ export interface FindParams { * `{ exists: true }`, `{ missing: true }`) use `where: { subtype: { …operators… } }`. */ subtype?: string | string[] - /** - * Metadata filters using BFO operators (e.g., `{ year: { greaterThan: 2020 } }`). - * Field names follow the field-addressing law — see the `@remarks` on - * {@link FindParams}: a bare key is always the caller's metadata field; - * an engine scalar needs the explicit `system.` form. - * - * @example - * ```typescript - * await brain.find({ where: { level: { greaterThan: 5 } } }) // metadata.level - * await brain.find({ where: { 'system.visibility': 'internal' } }) // engine scalar - * ``` - */ + /** Metadata filters using BFO operators (e.g., `{ year: { greaterThan: 2020 } }`) */ where?: Partial // Visibility @@ -597,49 +548,13 @@ export interface FindParams { // Control options limit?: number // Max results (default: 10) offset?: number // Skip N results - /** - * @deprecated Not implemented. Passing `cursor` throws - * {@link UnsupportedFindOptionError} — it used to be accepted and - * silently ignored, which masked that no cursor pagination ever ran. Use - * `offset` / `limit` until cursor pagination ships. - */ cursor?: string // Cursor-based pagination // Sorting - /** - * Field to sort by. Follows the field-addressing law (see the `@remarks` - * on {@link FindParams}): a bare name (`'level'`, `'rank'`, `'score'`, …) - * always sorts by that metadata field; the ten engine scalars sort only - * via the explicit `system.` form (e.g. `'system.createdAt'`); a - * name that resolves to neither throws {@link UnresolvableFieldError} - * naming the fix. - * - * Ordering contract (identical on the pure-JS engine and the native - * accelerator): rows missing or `null` on this field sort LAST in BOTH - * `asc` and `desc` order and are never dropped from the result; ties - * break by `id` ascending. - * - * @example - * ```typescript - * await brain.find({ orderBy: 'level', order: 'desc' }) // metadata.level - * await brain.find({ orderBy: 'system.createdAt', order: 'desc' }) // engine scalar - * ``` - */ - orderBy?: string - /** - * Sort direction: `'asc'` (default) or `'desc'`. Per the ordering - * contract on `orderBy`, rows missing/`null` on the sorted field sort - * LAST in both directions — `order` never moves them to the front. - */ + orderBy?: string // Field to sort by (e.g., 'createdAt', 'title', 'metadata.priority') order?: 'asc' | 'desc' // Sort direction: 'asc' (default) or 'desc' // Advanced options - /** - * @deprecated Not implemented. Passing `includeRelations` throws - * {@link UnsupportedFindOptionError} — it used to be accepted and - * silently ignored, so no relationships were ever attached. Fetch - * relationships separately via `brain.related()`. - */ includeRelations?: boolean // Include entity relationships excludeVFS?: boolean // Exclude VFS entities from results (default: false - VFS included) service?: string // Multi-tenancy filter @@ -672,11 +587,6 @@ export interface FindParams { } // Performance options - /** - * @deprecated Not implemented. Passing `writeOnly` throws - * {@link UnsupportedFindOptionError} — it used to be accepted and - * silently ignored, so validation was never actually skipped. - */ writeOnly?: boolean // Skip validation for high-speed ingestion // Aggregation @@ -1426,10 +1336,7 @@ export type GroupByDimension = export interface AggregateSource { /** Filter by entity type(s) */ type?: NounType | NounType[] - /** - * Metadata filter — same syntax and field-addressing law as find()'s - * `where` (see the `@remarks` on {@link FindParams}). - */ + /** Metadata filter (same syntax as find({ where })) */ where?: Record /** Multi-tenancy service filter */ service?: string @@ -1443,11 +1350,7 @@ export interface AggregateDefinition { name: string /** Which entities contribute to this aggregate */ source: AggregateSource - /** - * Dimensions to group by — field names follow the same field-addressing - * law as find()'s `where` / `orderBy` (see the `@remarks` on - * {@link FindParams}). - */ + /** Dimensions to group by */ groupBy: GroupByDimension[] /** Named metrics to compute */ metrics: Record @@ -1506,25 +1409,16 @@ export interface AggregateGroupState { export interface AggregateQueryParams { /** Name of the aggregate to query */ name: string - /** - * Filter aggregate groups by their key values — same field-addressing - * law as find() (see the `@remarks` on {@link FindParams}). - */ + /** Filter aggregate groups by their key values */ where?: Record /** * Filter groups by their computed METRIC values (SQL HAVING). Same BFO operators as * `where`, but applied to the derived metric results plus `count`, e.g. * `{ revenue: { greaterThan: 1000 } }`. Evaluated per group (O(groups), independent of - * entity count), before sort/pagination. Metric names and `count` are looked up - * directly, not field-addressed; a group-KEY field used here follows the same - * field-addressing law as find() (see the `@remarks` on {@link FindParams}). + * entity count), before sort/pagination. */ having?: Record - /** - * Sort by metric name (a key from `metrics`, looked up directly) or by a - * group key field — a group key field follows the same field-addressing - * law as find()'s `orderBy` (see the `@remarks` on {@link FindParams}). - */ + /** Sort by metric name or group key field */ orderBy?: string /** Sort direction */ order?: 'asc' | 'desc' @@ -2028,6 +1922,32 @@ export interface BrainyConfig { */ force?: boolean + /** + * How write paths react when an untyped (JavaScript) caller smuggles a + * Brainy-reserved field (`RESERVED_ENTITY_FIELDS` / `RESERVED_RELATION_FIELDS` + * — `confidence`, `weight`, `subtype`, `visibility`, `service`, `createdBy`, + * `noun`/`verb`, `data`, `createdAt`, `updatedAt`, `_rev`) **inside the + * `metadata` bag** of `add()` / `update()` / `relate()` / `updateRelation()` + * (and their `transact()` / `with()` mirrors). TypeScript callers can't write + * these shapes at all — the compile-time guard on the metadata param types + * (`NoReservedEntityKeys` / `NoReservedRelationKeys`) rejects a literal + * reserved key — so this policy only governs untyped callers that slip one + * past the compiler. + * + * - `'throw'` (**default, 8.0**): a reserved key in the bag throws a clear + * `Error` naming the offending key(s) and the correct write path. No silent + * remap, no data loss, no surprise. This is the 8.0 "no silent failures" + * contract. + * - `'warn'`: legacy remapping with a loud, one-shot (per key, per process) + * warning for EVERY reserved key found — user-mutable fields are remapped to + * their dedicated top-level param (top-level wins when both are supplied), + * system-managed fields are dropped. Use while migrating untyped call sites. + * - `'remap'`: the pre-8.0 silent remapping, no warning. Last-resort + * compatibility hatch for code that intentionally relies on the bag path. + * + * @default 'throw' + */ + reservedFieldPolicy?: 'throw' | 'warn' | 'remap' } // ============= Neural API Types ============= diff --git a/src/types/reservedFields.ts b/src/types/reservedFields.ts index 15b585c5..a0606e1d 100644 --- a/src/types/reservedFields.ts +++ b/src/types/reservedFields.ts @@ -1,54 +1,35 @@ /** * @module types/reservedFields - * @description The stored-record layout contract — ONE place that defines - * which keys of a persisted metadata record belong to the ENGINE (top-level - * entity/relationship fields) and how the USER's metadata bag is kept apart - * from them, faithfully, across flush / reopen / rebuild / time travel. + * @description The canonical reserved-field contract — ONE place that defines + * which keys belong to Brainy (top-level entity/relationship fields) and may + * therefore never live inside a `metadata` bag. * - * THE FIELD-ADDRESSING LAW (ruled 2026-08-03, VENUE-BRAINY-ORDERBY-NOOP): - * data is either in main space — where developers can use ANY name, and it - * all works with every database function — or it is in `system.*`. There are - * NO reserved user-facing metadata names anymore: `confidence`, `type`, - * `level`, `data`, `id`, `content` … inside a metadata bag are ordinary user - * fields. The only refused write is a user metadata key literally starting - * with `'system.'` (namespace forgery — see `rejectForgedSystemKeys`). + * Three layers enforce the contract, all driven by the constants below: * - * That law makes name-based storage discrimination unsound for NEW records - * (a user field named `confidence` may now legally sit beside the engine's - * confidence scalar), so persisted metadata records carry the user bag - * NESTED, shape-discriminated by a format stamp: + * 1. **Compile time** — `AddParams.metadata`, `UpdateParams.metadata`, + * `RelateParams.metadata` and `UpdateRelationParams.metadata` are typed so + * a literal reserved key is a TypeScript error (see + * {@link EntityMetadataInput} / {@link RelationMetadataInput}). + * 2. **Write time** — for untyped (JavaScript) callers that smuggle a + * reserved key past the compiler anyway, every write path normalizes the + * bag: user-mutable fields are remapped to their dedicated top-level + * param (top-level wins when both are supplied) and system-managed fields + * are dropped with a one-shot warning naming the correct write path. + * 3. **Read time** — every read path splits the stored flat record through + * {@link splitNounMetadataRecord} / {@link splitVerbMetadataRecord}, so a + * reserved field is surfaced ONLY at top level and `entity.metadata` / + * `relation.metadata` contain ONLY custom fields, always — live reads, + * batch reads, and historical (`asOf`) reads alike. * - * - **v2 (nested-bag)** — `{ …engine fields…, [METADATA_RECORD_FORMAT_KEY]: - * NESTED_BAG_FORMAT, metadata: { …user bag, verbatim… } }`. Built ONLY by - * {@link buildNounMetadataRecord} / {@link buildVerbMetadataRecord}; the - * engine half and the user bag can never collide because they never share - * a level. - * - **legacy (flat)** — engine fields and user fields mixed at one level, - * discriminated BY NAME through the RESERVED_* lists. Sound for legacy - * records precisely because the pre-law write door REFUSED user metadata - * carrying those names — a flat key matching a reserved name IS the - * engine's value in any record the old door admitted. - * - * {@link splitNounMetadataRecord} / {@link splitVerbMetadataRecord} read - * BOTH shapes (stamp first, name split as the legacy fallback) and are the - * single read-side choke point for live, batch, AND historical (`asOf`) - * reads — the generation store snapshots whole records, so time travel - * rides the same split. - * - * The RESERVED_* lists therefore no longer describe a user-facing ban — they - * describe the ENGINE HALF of the stored record layout (and drive the legacy - * split). The write-door remap machinery and the compile-time metadata key - * bans that used to enforce the old contract are gone. + * Documented for consumers in `docs/concepts/consistency-model.md` + * ("Reserved fields"). */ /** - * @description Entity (noun) field names owned by the ENGINE in a stored - * metadata record. In v2 (nested-bag) records these are the legal TOP-LEVEL - * keys beside the nested `metadata` bag; in legacy flat records they drive - * the by-name split. They are NOT a user-facing ban list: since the - * field-addressing law, a user metadata field may carry any of these names - * and remains the user's — it lives inside the nested bag, never at the - * record's top level. + * @description Entity (noun) field names reserved by Brainy. These keys are + * stored in the flat per-entity metadata record alongside custom fields, but + * they belong to Brainy: every read path extracts them to top-level + * `Entity` fields, and no write path accepts them inside `metadata`. * * | Key | Canonical write path | * |-----|----------------------| @@ -138,54 +119,68 @@ export type ReservedRelationField = (typeof RESERVED_RELATION_FIELDS)[number] type IsAny = 0 extends 1 & T ? true : false /** - * @deprecated The compile-time reserved-key ban died with the - * field-addressing law: every name is legal user metadata now. Kept as an - * empty (no-op) guard so external type references keep compiling; it bans - * nothing. + * @description Compile-time tripwire: marks every reserved entity key as + * `never` so an object literal carrying one fails to type-check. Keys that + * `T` itself declares (including via an index signature, where + * `keyof T = string`) are exempted — a consumer who *explicitly* types a + * reserved key into their metadata shape keeps a working (if unwise) type, + * and index-signature metadata types remain assignable. */ -export type NoReservedEntityKeys = unknown +export type NoReservedEntityKeys = { + readonly [K in ReservedEntityField as K extends keyof T ? never : K]?: never +} /** - * @deprecated Relationship mirror of {@link NoReservedEntityKeys} — no-op - * for the same reason. + * @description Relationship mirror of {@link NoReservedEntityKeys}. */ -export type NoReservedRelationKeys = unknown +export type NoReservedRelationKeys = { + readonly [K in ReservedRelationField as K extends keyof T ? never : K]?: never +} + +/** + * @description The metadata bag shape for untyped brains (`T = any`): an + * open index signature (any custom key, any value — exactly the pre-8.0 + * latitude) intersected with the reserved-key guard, whose declared + * `?: never` properties take precedence over the index signature so a + * literal reserved key is still a compile error. + */ +type OpenBag = { [key: string]: any } & Guard /** * @description The type of `AddParams.metadata`: the consumer's metadata - * shape `T`, open. Under the field-addressing law EVERY key is a legal user - * field (engine scalars are written only via their dedicated params and read - * at `system.*`), so no name is banned at compile time. The one illegal - * spelling — a key starting `'system.'` — cannot be expressed as a mapped - * type ban and is refused at runtime (`rejectForgedSystemKeys`). + * shape `T` with reserved entity keys forbidden at compile time. For untyped + * brains (`T = any`) the bag stays open ({@link OpenBag}), so arbitrary + * custom fields remain legal while literal reserved keys still error. */ export type EntityMetadataInput = IsAny extends true - ? { [key: string]: any } - : T + ? OpenBag> + : T & NoReservedEntityKeys /** * @description The type of `UpdateParams.metadata`: a partial patch of the - * consumer's metadata shape. Same openness as {@link EntityMetadataInput}. + * consumer's metadata shape with reserved entity keys forbidden at compile + * time. Same `T = any` handling as {@link EntityMetadataInput}. */ export type EntityMetadataPatch = IsAny extends true - ? { [key: string]: any } - : Partial + ? OpenBag> + : Partial & NoReservedEntityKeys /** * @description The type of `RelateParams.metadata`: the consumer's edge - * metadata shape, open — the relation mirror of {@link EntityMetadataInput}. + * metadata shape with reserved relationship keys forbidden at compile time. */ export type RelationMetadataInput = IsAny extends true - ? { [key: string]: any } - : T + ? OpenBag> + : T & NoReservedRelationKeys /** * @description The type of `UpdateRelationParams.metadata`: a partial patch - * of the consumer's edge metadata shape, open. + * of the consumer's edge metadata shape with reserved relationship keys + * forbidden at compile time. */ export type RelationMetadataPatch = IsAny extends true - ? { [key: string]: any } - : Partial + ? OpenBag> + : Partial & NoReservedRelationKeys /** * @description Result of splitting a stored flat metadata record into its @@ -201,103 +196,6 @@ export interface SplitMetadataRecord { const RESERVED_ENTITY_SET: ReadonlySet = new Set(RESERVED_ENTITY_FIELDS) const RESERVED_RELATION_SET: ReadonlySet = new Set(RESERVED_RELATION_FIELDS) -/** - * @description The format-stamp key of a persisted metadata record. Its - * presence with the exact value {@link NESTED_BAG_FORMAT} marks a v2 - * (nested-bag) record; its absence marks a legacy flat record. The stamp is - * what makes the shape check collision-proof against legacy user data: a - * pre-law record COULD carry a user field named `metadata` (the name was - * never reserved), but it cannot also carry this engine-written stamp. - */ -export const METADATA_RECORD_FORMAT_KEY = '_fmt' - -/** - * @description The nested-bag record format stamp (v2, the field-addressing - * law's storage shape, 2026-08-03): engine fields at top level, the user's - * metadata bag NESTED verbatim under `metadata`. Cross-engine: the native - * provider discriminates record shapes by the same stamp. - */ -export const NESTED_BAG_FORMAT = 2 - -/** - * @description `true` when a persisted record carries the v2 nested-bag - * stamp (and a structurally valid nested bag). - */ -export function isNestedBagRecord( - record: Record | null | undefined -): boolean { - return ( - record !== null && - record !== undefined && - typeof record === 'object' && - record[METADATA_RECORD_FORMAT_KEY] === NESTED_BAG_FORMAT && - typeof record.metadata === 'object' && - record.metadata !== null && - !Array.isArray(record.metadata) - ) -} - -/** - * @description Build a v2 (nested-bag) entity metadata record — THE only - * sanctioned way to construct a persisted noun metadata record. The engine - * half goes top-level; the user bag nests verbatim under `metadata`; the - * format stamp seals the shape. Because the two halves never share a level, - * a user field named `confidence` (or any other engine spelling) survives - * flush / reopen / rebuild / time travel exactly as written. - * @param engineFields - The engine-owned half (keys from - * {@link RESERVED_ENTITY_FIELDS} — `noun`, timestamps, `_rev`, …). - * @param userBag - The consumer's metadata bag, stored verbatim. - * @returns The stamped v2 record. - */ -export function buildNounMetadataRecord( - engineFields: Partial>, - userBag: Record | undefined -): Record { - return { - ...engineFields, - [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, - metadata: { ...(userBag ?? {}) } - } -} - -/** - * @description Build a v2 (nested-bag) relationship metadata record — the - * verb mirror of {@link buildNounMetadataRecord}. - * @param engineFields - The engine-owned half (keys from - * {@link RESERVED_RELATION_FIELDS} — `verb`, `weight`, timestamps, …). - * @param userBag - The consumer's edge metadata bag, stored verbatim. - * @returns The stamped v2 record. - */ -export function buildVerbMetadataRecord( - engineFields: Partial>, - userBag: Record | undefined -): Record { - return { - ...engineFields, - [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, - metadata: { ...(userBag ?? {}) } - } -} - -/** - * @description Shape-first split of a v2 record: the engine half is the top - * level filtered through the reserved list (belt — the builders only ever - * write reserved names there), the user bag is `record.metadata` verbatim. - */ -function splitNestedRecord( - record: Record, - reservedSet: ReadonlySet -): SplitMetadataRecord { - const reserved: Record = {} - for (const [key, value] of Object.entries(record)) { - if (reservedSet.has(key)) reserved[key] = value - } - return { - reserved: reserved as Partial>, - custom: { ...(record.metadata as Record) } - } -} - /** * @description Shared splitter — partitions a record's keys against a * reserved-name set. `null`/`undefined` records split to two empty objects. @@ -324,45 +222,33 @@ function splitRecord( } /** - * @description Split a stored entity (noun) metadata record into engine - * fields and the user's metadata bag — THE canonical read-side split, shape - * aware. v2 (nested-bag) records split by SHAPE: engine half top-level, bag - * = `record.metadata` verbatim (user collider names survive faithfully). - * Legacy flat records split BY NAME through the reserved list — sound for - * them because the pre-law write door refused user metadata carrying those - * names. Every entity read path (live `get()`, batch reads, paginated - * listings, and historical `asOf()` materialization — the generation store - * snapshots whole records) goes through this function, so the two shapes - * can never drift between read paths. - * @param record - The stored metadata record (either shape). - * @returns `reserved` (engine-owned fields) and `custom` (the consumer's metadata bag). + * @description Split a stored entity (noun) flat metadata record into + * reserved fields and custom metadata — THE canonical read-side split. Every + * entity read path (live `get()`, batch reads, paginated listings, and + * historical `asOf()` materialization) goes through this function, so the + * reserved list can never drift between read paths. + * @param record - The stored flat metadata record. + * @returns `reserved` (Brainy-owned fields) and `custom` (the consumer's metadata bag). * @example * const { reserved, custom } = splitNounMetadataRecord(stored) * // reserved.noun → entity.type, reserved.confidence → entity.confidence, … - * // custom → entity.metadata (the user's fields only, always — ANY names) + * // custom → entity.metadata (custom fields only, always) */ export function splitNounMetadataRecord( record: Record | null | undefined ): SplitMetadataRecord { - if (isNestedBagRecord(record)) { - return splitNestedRecord(record as Record, RESERVED_ENTITY_SET) - } return splitRecord(record, RESERVED_ENTITY_SET) } /** - * @description Split a stored relationship (verb) metadata record into - * engine fields and the user's edge metadata bag — the verb mirror of - * {@link splitNounMetadataRecord}, shape aware, used by every relationship - * read path. - * @param record - The stored metadata record (either shape). - * @returns `reserved` (engine-owned fields) and `custom` (the consumer's metadata bag). + * @description Split a stored relationship (verb) flat metadata record into + * reserved fields and custom metadata — the verb mirror of + * {@link splitNounMetadataRecord}, used by every relationship read path. + * @param record - The stored flat metadata record. + * @returns `reserved` (Brainy-owned fields) and `custom` (the consumer's metadata bag). */ export function splitVerbMetadataRecord( record: Record | null | undefined ): SplitMetadataRecord { - if (isNestedBagRecord(record)) { - return splitNestedRecord(record as Record, RESERVED_RELATION_SET) - } return splitRecord(record, RESERVED_RELATION_SET) } diff --git a/src/utils/metadataIndex.ts b/src/utils/metadataIndex.ts index f010560d..fdb17c22 100644 --- a/src/utils/metadataIndex.ts +++ b/src/utils/metadataIndex.ts @@ -5,7 +5,6 @@ */ import { StorageAdapter, resolveEntityField, NounMetadata, VerbMetadata } from '../coreTypes.js' -import { SYSTEM_ENTITY_SCALARS, parseFieldAddress, UnresolvableFieldError } from '../db/fieldAddressing.js' import { ColumnStore } from '../indexes/columnStore/ColumnStore.js' import type { MetadataIndexProvider } from '../plugin.js' import { MetadataIndexCache, MetadataIndexCacheConfig } from './metadataIndexCache.js' @@ -44,8 +43,8 @@ import { BrainyError } from '../errors/brainyError.js' * bucketed field is added (e.g. a compressed float), add it here too. */ const BUCKETED_INDEX_FIELDS: ReadonlySet = new Set([ - 'system.createdAt', - 'system.updatedAt' + 'createdAt', + 'updatedAt' ]) export interface MetadataIndexEntry { @@ -73,11 +72,8 @@ export interface MetadataIndexConfig { maxIndexSize?: number // Max number of entries per field value (default: 10000) rebuildThreshold?: number // Rebuild if index is this % stale (default: 0.1) autoOptimize?: boolean // Auto-cleanup unused entries (default: true) - // NOTE: the name-based indexedFields/excludeFields knobs died with the - // field-addressing law ("no special names"): EVERY user field indexes, - // whatever its name. Bulk-payload protection is value-SHAPE based and - // uniform across all names (large arrays never become posting scalars; - // long values index hashed) — shape is not a name carve-out. + indexedFields?: string[] // Only index these fields (default: all) + excludeFields?: string[] // Never index these fields } export interface MetadataIndexOptions { @@ -188,12 +184,31 @@ export class MetadataIndexManager implements MetadataIndexProvider { this.config = { maxIndexSize: config.maxIndexSize ?? 10000, rebuildThreshold: config.rebuildThreshold ?? 0.1, - autoOptimize: config.autoOptimize ?? true - // No name-based exclude/allow lists — the field-addressing law: every - // user field indexes, whatever its name ('content', 'data', 'id', - // 'vector', … included). Bulk payloads are kept out by uniform value- - // SHAPE rules in extractIndexableFields (arrays >10 never become - // posting scalars; >100-char values index hashed), never by name. + autoOptimize: config.autoOptimize ?? true, + indexedFields: config.indexedFields ?? [], + excludeFields: config.excludeFields ?? [ + // ONLY exclude truly un-indexable fields (binary data, large content) + // Timestamps are NOW indexed with automatic bucketing (prevents pollution) + + // Vectors and embeddings (binary data, already have HNSW indexes) + 'embedding', + 'vector', + 'embeddings', + 'vectors', + + // Large content fields (too large for metadata indexing) + 'content', + 'data', + 'originalData', + '_data', + + // Primary keys (use direct lookups instead) + 'id' + + // NOTE: 'accessed', 'modified', 'createdAt', etc. are NO LONGER excluded! + // They are now indexed with automatic 1-minute bucketing to prevent file pollution + // This enables range queries like: modified > yesterday + ] } // Initialize metadata cache with similar config to search cache @@ -285,7 +300,7 @@ export class MetadataIndexManager implements MetadataIndexProvider { } // Warm the cache with common fields (lazy loading optimization) - // This loads the type column ('system.type') needed for type counts + // This loads the 'noun' sparse index which is needed for type counts await this.warmCache() // Load type counts AFTER warmCache (sparse index is now cached) @@ -334,9 +349,8 @@ export class MetadataIndexManager implements MetadataIndexProvider { * Target: >80% cache hit rate for typical workloads */ async warmCache(): Promise { - // Common columns used in most queries — the frozen system keys, plus - // legacy spellings for a pre-epoch-3 brain read before its rebuild runs. - const commonFields = ['system.type', 'system.service', 'system.createdAt', 'noun'] + // Common fields used in most queries + const commonFields = ['noun', 'type', 'service', 'createdAt'] prodLog.debug(`🔥 Warming metadata cache with common fields: ${commonFields.join(', ')}`) @@ -522,11 +536,9 @@ export class MetadataIndexManager implements MetadataIndexProvider { } /** - * Lazy load entity counts from the type column (O(n) where n = number of - * types). The frozen key is 'system.type' (epoch 3); the legacy 'noun' - * column is read as a fallback for a pre-epoch-3 brain observed before its - * rebuild has run (e.g. a reader-mode open against an old writer). + * Lazy load entity counts from the 'noun' field sparse index (O(n) where n = number of types) * FIX: Previously read from stats.nounCount which was SERVICE-keyed, not TYPE-keyed + * Now computes counts from the sparse index which has the correct type information */ private async lazyLoadCounts(): Promise { try { @@ -536,31 +548,23 @@ export class MetadataIndexManager implements MetadataIndexProvider { this.entityCountsByTypeFixed.fill(0) this.verbCountsByTypeFixed.fill(0) - // PRIMARY (8.0+): rehydrate per-type counts from the column store's - // type column — the authoritative on-disk source after a cold reopen. - // Frozen key first ('system.type', epoch 3), legacy 'noun' as the - // pre-rebuild fallback. + // PRIMARY (8.0+): rehydrate per-type counts from the column store's 'noun' + // field — the authoritative on-disk source after a cold reopen. // // The chunked sparse-index WRITE path was removed in 7.20.0 (commit - // 11be039): new workspaces persist the type column ONLY to the column - // store, never to a sparse-index blob. So the legacy sparse path below - // finds nothing and leaves every count at 0 — which is exactly why - // counts.byType/byTypeEnum/topTypes/allNounTypeCounts all read empty + // 11be039): new workspaces persist the 'noun' field ONLY to the column + // store, never to a `__sparse_index__noun` blob. So the legacy sparse + // path below finds nothing and leaves every count at 0 — which is exactly + // why counts.byType/byTypeEnum/topTypes/allNounTypeCounts all read empty // after close()+reopen while find()/getNounCount() (different sources) // stay correct. The column store's per-value cardinality matches the warm // `updateTypeFieldAffinity` counts EXACTLY because both are driven from the // same `addToIndex` field set, in lockstep, with no visibility gate on // either — so this rehydration reproduces the warm values precisely. - const indexedCols = this.columnStore ? this.columnStore.getIndexedFields() : [] - const typeCol = indexedCols.includes('system.type') - ? 'system.type' - : indexedCols.includes('noun') - ? 'noun' - : null - if (this.columnStore && typeCol) { - const nounValues = await this.columnStore.getFilterValues(typeCol) + if (this.columnStore && this.columnStore.getIndexedFields().includes('noun')) { + const nounValues = await this.columnStore.getFilterValues('noun') for (const value of nounValues) { - const bitmap = await this.columnStore.filter(typeCol, value) + const bitmap = await this.columnStore.filter('noun', value) if (bitmap.size > 0) { // Use the stored value directly as the key (the legacy sparse path // did the same): it is already the normalized type string that @@ -575,17 +579,16 @@ export class MetadataIndexManager implements MetadataIndexProvider { } // LEGACY FALLBACK (pre-7.20.0 workspaces still on the chunked sparse index). - const sparseCol = (await this.loadSparseIndex('system.type')) ? 'system.type' : 'noun' - const nounSparseIndex = await this.loadSparseIndex(sparseCol) + const nounSparseIndex = await this.loadSparseIndex('noun') if (!nounSparseIndex) { - // No column-store type column and no sparse index yet — counts will be + // No column-store 'noun' field and no sparse index yet — counts will be // populated as entities are added. return } // Iterate through all chunks and sum up bitmap sizes by type for (const chunkId of nounSparseIndex.getAllChunkIds()) { - const chunk = await this.chunkManager.loadChunk(sparseCol, chunkId) + const chunk = await this.chunkManager.loadChunk('noun', chunkId) if (chunk) { for (const [type, bitmap] of chunk.entries) { const currentCount = this.totalEntitiesByType.get(type) || 0 @@ -1175,86 +1178,76 @@ export class MetadataIndexManager implements MetadataIndexProvider { return `__HASH_${Math.abs(hash).toString(36)}` } + /** + * Check if field should be indexed + */ + private shouldIndexField(field: string): boolean { + if (this.config.excludeFields.includes(field)) return false + if (this.config.indexedFields.length > 0) { + return this.config.indexedFields.includes(field) + } + return true + } + /** * Extract indexable field-value pairs from entity or metadata * - * Handles BOTH entity structure (with top-level fields) AND record shapes - * - Record-frame system scalars index under literal 'system.' keys - * - The user's metadata bag indexes under bare keys — EVERY name (the - * field-addressing law: no special names; 'level', 'data', 'id', - * 'content', 'vector' in a bag are ordinary user fields) - * - Record-frame plumbing (vector, connections, level, data, _rev, id) - * never indexes — that is namespace routing, not a name carve-out - * - Value-SHAPE rules apply uniformly to all names: arrays >10 never - * become posting scalars; purely numeric key names (array indices) - * skip; >100-char values index hashed (normalizeValue) + * Now handles BOTH entity structure (with top-level fields) AND plain metadata + * - Extracts from top-level fields (confidence, weight, timestamps, type, service, etc.) + * - Also extracts from nested metadata field (custom user fields) + * - Skips HNSW-specific fields (vector, connections, level, id) + * - Maps 'type' → 'noun' for backward compatibility with existing indexes + * + * BUG FIX: Exclude vector embeddings and large arrays from indexing + * BUG FIX: Also exclude purely numeric field names (array indices) + * - Vector fields (384+ dimensions) were creating 825K chunk files for 1,144 entities + * - Arrays converted to objects with numeric keys were still being indexed */ private extractIndexableFields(data: any): Array<{ field: string, value: any }> { const fields: Array<{ field: string, value: any }> = [] - // RECORD-FRAME-ONLY plumbing guard: on an entity/stored-record frame - // these keys are the engine's structural payloads (the 384-dim vector, - // embeddings, the adjacency list, the identity field) and never index. - // This set is NEVER applied inside the user's metadata bag — under the - // field-addressing law every user name indexes; a real vector-sized - // value in a bag is kept out by the uniform array-size shape guard, not - // by its name. - const RECORD_PLUMBING = new Set(['vector', 'embedding', 'embeddings', 'connections', 'id']) + // Fields that should NEVER be indexed: bulk structural payloads that would + // blow up the index (the 384-dim vector, embeddings, the adjacency list). + // These are also caught by the array-size guard below, but naming them is + // belt-and-suspenders. NOTE: `level` was previously here (an HNSW node's + // layer) but it never actually reaches this path — every caller passes a + // metadata bag or Entity record, neither of which carries the node's + // `level` — so its only effect was to silently drop a legitimate USER + // metadata field named `level` (log level, skill level, access level…), + // making `where: { level: … }` return nothing. Removed. (`id` stays: it is + // the reserved entity-identity field, resolved specially by find().) + const NEVER_INDEX = new Set(['vector', 'embedding', 'embeddings', 'connections', 'id']) - // THE FROZEN INDEX KEY FORMAT (cross-engine, sealed 2026-08-03; the native - // accelerator keys identically — epoch 3 rebuilds every brain onto it): - // user fields index under BARE keys exactly as the caller wrote them; - // the ten system scalars index under literal 'system.' keys — the - // key IS the query address, so the two namespaces can never collide - // inside the index again. - // Frame kinds: 'entity-record' = entityForIndexing shape / v2 nested-bag - // stored record (user fields nested under `metadata`; stray top-level - // keys are DROPPED, not guessed); 'flat-record' = the LEGACY stored - // metadata-record shape (user fields flat beside the engine's — sound to - // split by name because the pre-law write door refused user metadata - // carrying engine names, so a flat key matching a system name IS the - // system value); 'user' = inside the metadata bag, where EVERY key is - // the user's and indexes bare — collider names included. - type Frame = 'entity-record' | 'flat-record' | 'user' - const extract = (obj: any, prefix = '', frame: Frame = 'entity-record'): void => { + const extract = (obj: any, prefix = ''): void => { for (const [key, value] of Object.entries(obj)) { - let fullKey = prefix ? `${prefix}.${key}` : key + const fullKey = prefix ? `${prefix}.${key}` : key - if (!prefix && frame !== 'user') { - if (key === 'metadata' && typeof value === 'object' && value !== null && !Array.isArray(value)) { - extract(value, '', 'user') // the user's namespace: bare keys - continue - } - if (key === 'type' || key === 'noun') { - fullKey = 'system.type' // legacy 'noun' spelling folds into the frozen key - } else if (SYSTEM_ENTITY_SCALARS.has(key) && key !== 'id') { - fullKey = `system.${key}` - } else if ( - key === 'data' || key === '_rev' || key === 'level' || key === '_fmt' || - RECORD_PLUMBING.has(key) - ) { - continue // plumbing / identity / format stamp — never indexed from a record frame - } else if (frame === 'entity-record') { - continue // stray entity-frame key: dropped, not guessed - } - // flat-record fallthrough: a non-system, non-plumbing key IS a user - // field (flat beside the engine's, legacy shape) — indexes bare. - } - // User frame: NO name-based skips — every user field indexes, whatever - // its name (the field-addressing law). Only the uniform value-shape - // guards below apply. + // Skip fields in never-index list (CRITICAL: prevents vector indexing bug + HNSW fields) + if (!prefix && NEVER_INDEX.has(key)) continue // Skip purely numeric field names (array indices converted to object keys) // Legitimate field names should never be purely numeric // This catches vectors stored as objects: {0: 0.1, 1: 0.2, ...} if (/^\d+$/.test(key)) continue + // Skip fields based on user configuration + if (!this.shouldIndexField(fullKey)) continue + + // Special handling for metadata field at top level + // Flatten metadata fields to top-level (no prefix) for cleaner queries + // Standard fields are already at top-level, custom fields go in metadata + // By flattening here, queries can use { category: 'B' } instead of { 'metadata.category': 'B' } + if (key === 'metadata' && !prefix && typeof value === 'object' && !Array.isArray(value)) { + extract(value, '') // Flatten to top-level, no prefix + continue + } + // Skip large arrays (> 10 elements) - likely vectors or bulk data if (Array.isArray(value) && value.length > 10) continue if (value && typeof value === 'object' && !Array.isArray(value)) { - // Recurse into nested objects (but not arrays), keeping the frame - extract(value, fullKey, frame) + // Recurse into nested objects (but not arrays) + extract(value, fullKey) } else if (Array.isArray(value) && value.length <= 10) { // Small arrays: index as multi-value field (all with same field name) // Example: tags: ["javascript", "node"] → field="tags", value="javascript" + field="tags", value="node" @@ -1265,21 +1258,16 @@ export class MetadataIndexManager implements MetadataIndexProvider { } } } else { - // Primitive value: index it under the frozen key computed above. - // (The legacy 'type'→'noun' remap is gone — 'noun' columns die at - // the epoch-3 rebuild; system.type is the one spelling.) - fields.push({ field: fullKey, value }) + // Primitive value: index it + // Map 'type' → 'noun' for backward compatibility + const indexField = (!prefix && key === 'type') ? 'noun' : fullKey + fields.push({ field: indexField, value }) } } } if (data && typeof data === 'object') { - // Shape detection for the top frame: an object carrying a nested - // `metadata` bag is the entityForIndexing shape; anything else is the - // flat stored-record shape (user fields flat beside reserved ones). - const entityShaped = - 'metadata' in data && typeof data.metadata === 'object' && data.metadata !== null - extract(data, '', entityShaped ? 'entity-record' : 'flat-record') + extract(data) } // Extract words for hybrid text search @@ -1481,11 +1469,10 @@ export class MetadataIndexManager implements MetadataIndexProvider { prodLog.debug(`Entity ${id} has ${wordFields.length} indexed words (large document)`) } - // Sort fields to process the type column first for type-field affinity - // tracking ('system.type' is the frozen key; 'noun' died at epoch 3). + // Sort fields to process 'noun' field first for type-field affinity tracking fields.sort((a, b) => { - if (a.field === 'system.type') return -1 - if (b.field === 'system.type') return 1 + if (a.field === 'noun') return -1 + if (b.field === 'noun') return 1 return 0 }) @@ -1924,15 +1911,22 @@ export class MetadataIndexManager implements MetadataIndexProvider { // Skip logical operators if (rawField === 'allOf' || rawField === 'anyOf' || rawField === 'not') continue - // THE ONE ADDRESSING LAW (sealed 2026-08-03): every filter key routes - // through parseFieldAddress — bare and 'metadata.'-prefixed spellings - // address the user's fields (indexed under BARE keys), 'system.' - // addresses the ten engine scalars (indexed under their literal - // 'system.' keys). A malformed address (system., - // plumbing in the system spelling) throws typed BEFORE any index read — - // an accepted name either works or refuses. - const address = parseFieldAddress(rawField, 'entity') - const field = address.scope === 'system' ? `system.${address.field}` : address.field + // Metadata is FLATTENED at index time (metadata.entry.title indexes as + // entry.title), so a `metadata.`-prefixed where key is almost always + // the caller spelling the STORAGE shape rather than the index shape. + // Accept both spellings: when the key as spelled is unindexed but its + // stripped spelling is, query the stripped one. A literal nested + // custom key named `metadata` still wins when indexed as spelled + // (checked first), so that rare shape keeps working. + let field = rawField + if ( + rawField.startsWith('metadata.') && + this.columnStore && + !this.columnStore.hasField(rawField) && + this.columnStore.hasField(rawField.slice('metadata.'.length)) + ) { + field = rawField.slice('metadata.'.length) + } let fieldResults: string[] = [] @@ -2213,30 +2207,9 @@ export class MetadataIndexManager implements MetadataIndexProvider { order: 'asc' | 'desc' = 'asc', topK?: number ): Promise { - // THE ONE ADDRESSING LAW — the orderBy address routes through the same - // parse the filter path uses (the historical asymmetry where the filter - // path understood 'metadata.' but the sorted path never did is dead). - // Bare / 'metadata.' → the user's bare index key; 'system.' → the - // literal frozen key; malformed addresses throw typed before any read. - const orderAddress = parseFieldAddress(orderBy, 'entity') - const orderKey = - orderAddress.scope === 'system' ? `system.${orderAddress.field}` : orderAddress.field - - // DATA-AWARE REFUSAL (the did-you-mean): a bare address no user field - // carries cannot mean anything as a sort key — and when the name collides - // with a system scalar the caller almost certainly meant system.. - // Refusing loudly with both candidates beats silently sorting nothing. - if ( - orderAddress.scope === 'metadata' && - !(this.columnStore && this.columnStore.hasField(orderKey)) && - !(await this.loadSparseIndex(orderKey)) - ) { - throw new UnresolvableFieldError(orderAddress.raw, 'entity') - } - // Column store path: O(K log S) sort via k-way merge across segments. // No per-entity storage reads, no precision loss from bucketing. - if (this.columnStore && this.columnStore.hasField(orderKey)) { + if (this.columnStore && this.columnStore.hasField(orderBy)) { // Get filtered IDs from existing roaring bitmap path const hasFilter = filter && Object.keys(filter).length > 0 const filteredIds = hasFilter ? await this.getIdsForFilter(filter) : [] @@ -2256,41 +2229,20 @@ export class MetadataIndexManager implements MetadataIndexProvider { // log K) heap, not a full sort materialization. const k = topK !== undefined ? Math.min(topK, filteredIds.length) : filteredIds.length sortedIntIds = await this.columnStore.filteredSortTopK( - filterBitmap, orderKey, order, k + filterBitmap, orderBy, order, k ) } else { // Unfiltered sort — column store handles the full entity set efficiently sortedIntIds = await this.columnStore.sortTopK( - orderKey, order, topK !== undefined ? Math.min(topK, this.idMapper.size) : this.idMapper.size + orderBy, order, topK !== undefined ? Math.min(topK, this.idMapper.size) : this.idMapper.size ) } // Convert int IDs back to UUIDs. Number() narrowing is lossless — the // shipped EntityIdSpaceExceeded guard caps the JS mapper at u32. - const sortedUuids = sortedIntIds + return sortedIntIds .map(intId => this.idMapper.getUuid(Number(intId))) .filter((uuid): uuid is string => uuid !== undefined) - - // ORDERING CONTRACT (cross-engine, sealed): rows missing the field are - // NEVER dropped — they sort LAST in both directions — and ties break by - // id ascending. The column only contains rows that HAVE the field, so - // (1) re-sort the page deterministically (value, then id) with K cheap - // value reads, and (2) append the filtered rows the column omitted, - // id-ascending, filling any remaining page budget. - const page = await Promise.all( - sortedUuids.map(async id => ({ id, value: await this.getFieldValueForEntity(id, orderKey) })) - ) - page.sort((a, b) => this.compareAddressedValues(a.value, b.value, a.id, b.id, order)) - let result = page.map(p => p.id) - - if (hasFilter) { - const present = new Set(sortedUuids) - if (topK === undefined || result.length < topK) { - const missing = filteredIds.filter(id => !present.has(id)).sort() - result = result.concat(missing) - } - } - return topK !== undefined ? result.slice(0, topK) : result } // Fallback: sparse index path (for fields not yet in column store). @@ -2303,11 +2255,27 @@ export class MetadataIndexManager implements MetadataIndexProvider { const idValuePairs: Array<{ id: string, value: any }> = [] for (const id of filteredIds) { - const value = await this.getFieldValueForEntity(id, orderKey) + const value = await this.getFieldValueForEntity(id, orderBy) idValuePairs.push({ id, value }) } - idValuePairs.sort((a, b) => this.compareAddressedValues(a.value, b.value, a.id, b.id, order)) + idValuePairs.sort((a, b) => { + if (a.value == null && b.value == null) return 0 + if (a.value == null) return order === 'asc' ? 1 : -1 + if (b.value == null) return order === 'asc' ? -1 : 1 + if (a.value === b.value) return 0 + // Numbers compare numerically; everything else by code-point (UTF-8 byte) order. + // This makes the JS fallback sort match cor's native column store exactly + // (numeric i64/f64 vs code-point strings) and stay deterministic across + // environments, unlike the `<` operator's UTF-16 ordering for strings. + let comparison: number + if (typeof a.value === 'number' && typeof b.value === 'number') { + comparison = a.value < b.value ? -1 : 1 + } else { + comparison = compareCodePoints(String(a.value), String(b.value)) + } + return order === 'asc' ? comparison : -comparison + }) const sorted = idValuePairs.map(p => p.id) return topK !== undefined ? sorted.slice(0, topK) : sorted @@ -2341,51 +2309,11 @@ export class MetadataIndexManager implements MetadataIndexProvider { * * @public (called from brainy.ts for sorted queries) */ - /** - * The cross-engine ordering contract in one comparator (sealed 2026-08-03): - * missing/null values sort LAST in BOTH directions — the direction flip - * never moves them to the front — and ties break by id ascending, so an - * ordered read is deterministic and identical on both engines. Numbers - * compare numerically; everything else by code-point (UTF-8 byte) order, - * matching the native column store exactly. - */ - private compareAddressedValues( - aVal: any, - bVal: any, - aId: string, - bId: string, - order: 'asc' | 'desc' - ): number { - const aNull = aVal == null - const bNull = bVal == null - if (aNull || bNull) { - if (aNull && bNull) return aId < bId ? -1 : aId > bId ? 1 : 0 - return aNull ? 1 : -1 - } - let comparison = 0 - if (aVal !== bVal) { - if (typeof aVal === 'number' && typeof bVal === 'number') { - comparison = aVal < bVal ? -1 : 1 - } else { - comparison = compareCodePoints(String(aVal), String(bVal)) - } - } - if (comparison === 0) return aId < bId ? -1 : aId > bId ? 1 : 0 - return order === 'asc' ? comparison : -comparison - } - async getFieldValueForEntity(entityId: string, field: string): Promise { - // `field` arrives as a FROZEN INDEX KEY (bare = user metadata; - // 'system.' = engine scalar). Storage fallbacks read the matching - // side of the record — a system key reads the record scalar, a bare key - // reads the user's metadata bag; the two can never shadow each other. - const systemInner = field.startsWith('system.') ? field.slice('system.'.length) : null - - // Path 1: Bucketed fields need the actual (un-bucketed) value from storage. + // Path 1: Bucketed fields need the actual value from storage. if (BUCKETED_INDEX_FIELDS.has(field)) { const noun = await this.storage.getNoun(entityId) - if (!noun) return undefined - return (noun as unknown as Record)[systemInner as string] + return noun ? resolveEntityField(noun, field) : undefined } // Path 3 precondition: entity must be in the id mapper for bitmap lookup. @@ -2402,11 +2330,7 @@ export class MetadataIndexManager implements MetadataIndexProvider { // yet indexed. resolveEntityField handles the shape contract. if (!sparseIndex) { const noun = await this.storage.getNoun(entityId) - if (!noun) return undefined - if (systemInner !== null) { - return (noun as unknown as Record)[systemInner] - } - return (noun as { metadata?: Record }).metadata?.[field] + return noun ? resolveEntityField(noun, field) : undefined } // Path 3: Search sparse index chunks for this entity's value. @@ -2833,17 +2757,6 @@ export class MetadataIndexManager implements MetadataIndexProvider { // VFS Statistics Methods (uses existing Roaring bitmap infrastructure) // ============================================================================ - /** - * Read the type column's bitmap for one type value — frozen key first - * ('system.type', epoch 3), legacy 'noun' as the pre-rebuild fallback. - */ - private async getTypeBitmap(type: string): Promise { - return ( - (await this.getBitmapFromChunks('system.type', type)) ?? - (await this.getBitmapFromChunks('noun', type)) - ) - } - /** * Get VFS entity count for a specific type using Roaring bitmap intersection * Uses hardware-accelerated SIMD operations (AVX2/SSE4.2) @@ -2852,7 +2765,7 @@ export class MetadataIndexManager implements MetadataIndexProvider { */ async getVFSEntityCountByType(type: string): Promise { const vfsBitmap = await this.getBitmapFromChunks('isVFSEntity', true) - const typeBitmap = await this.getTypeBitmap(type) + const typeBitmap = await this.getBitmapFromChunks('noun', type) if (!vfsBitmap || !typeBitmap) return 0 @@ -2875,7 +2788,7 @@ export class MetadataIndexManager implements MetadataIndexProvider { // Iterate through all known types and compute VFS count via intersection for (const type of this.totalEntitiesByType.keys()) { - const typeBitmap = await this.getTypeBitmap(type) + const typeBitmap = await this.getBitmapFromChunks('noun', type) if (typeBitmap) { const intersection = RoaringBitmap32.and(vfsBitmap, typeBitmap) if (intersection.size > 0) { @@ -3469,21 +3382,18 @@ export class MetadataIndexManager implements MetadataIndexProvider { * Tracks which fields commonly appear with which entity types */ private updateTypeFieldAffinity(entityId: string, field: string, value: any, operation: 'add' | 'remove', metadata?: any): void { - // Only track affinity for user fields (plus the type column itself, - // which drives detection). Engine columns carry the literal 'system.' - // prefix under the frozen key format. - if (field.startsWith('system.') && field !== 'system.type') return + // Only track affinity for non-system fields (but allow 'noun' for type detection) + if (this.config.excludeFields.includes(field) && field !== 'noun') return - // For the type column ('system.type'), the value IS the entity type + // For the 'noun' field, the value IS the entity type let entityType: string | null = null - if (field === 'system.type') { + if (field === 'noun') { // This is the type definition itself entityType = this.normalizeValue(value, field) // Pass field for bucketing! - } else if (metadata && (metadata.noun ?? metadata.type)) { - // Extract entity type from the source shape: stored records carry it - // under 'noun', entity-for-indexing views under 'type'. - entityType = this.normalizeValue(metadata.noun ?? metadata.type, 'system.type') + } else if (metadata && metadata.noun) { + // Extract entity type from metadata + entityType = this.normalizeValue(metadata.noun, 'noun') } else { // No type information available, skip affinity tracking return @@ -3506,9 +3416,8 @@ export class MetadataIndexManager implements MetadataIndexProvider { const currentCount = typeFields.get(field) || 0 typeFields.set(field, currentCount + 1) - // Update total entities of this type (only count once per entity — - // the type column appears exactly once per entity) - if (field === 'system.type') { + // Update total entities of this type (only count once per entity) + if (field === 'noun') { const newCount = this.totalEntitiesByType.get(entityType)! + 1 this.totalEntitiesByType.set(entityType, newCount) @@ -3531,7 +3440,7 @@ export class MetadataIndexManager implements MetadataIndexProvider { } // Update total entities of this type - if (field === 'system.type') { + if (field === 'noun') { const total = this.totalEntitiesByType.get(entityType)! if (total > 1) { const newCount = total - 1 diff --git a/src/utils/paramValidation.ts b/src/utils/paramValidation.ts index 359413d7..ca439524 100644 --- a/src/utils/paramValidation.ts +++ b/src/utils/paramValidation.ts @@ -17,7 +17,6 @@ import { findCallerLocation } from './callerLocation.js' // fallback branches that no supported runtime can reach. import * as os from 'node:os' import * as fs from 'node:fs' -import { parseFieldAddress, UnsupportedFindOptionError } from '../db/fieldAddressing.js' const getSystemMemory = (): number => { if (os) { @@ -467,31 +466,9 @@ export function validateFindParams(params: FindParams): void { throw new Error('cannot specify both query and vector - they are mutually exclusive') } - // ACCEPTED-AND-IGNORED DIED AS A CLASS (sealed 2026-08-03): options the - // engine does not implement REFUSE with a typed error instead of silently - // doing nothing — a production consumer discovered a no-op by measurement - // once; never again. - if (params.cursor !== undefined) { - throw new UnsupportedFindOptionError('cursor') - } - if ((params as Record).includeRelations !== undefined) { - throw new UnsupportedFindOptionError('includeRelations') - } - if ((params as Record).writeOnly !== undefined) { - throw new UnsupportedFindOptionError('writeOnly') - } - - // THE ONE ADDRESSING LAW: the orderBy address must PARSE (bare/metadata. = - // user field, system. = the ruled map, anything else refuses typed - // with the valid map in the message) and order must be a real direction. - if (params.orderBy !== undefined) { - if (typeof params.orderBy !== 'string') { - throw new Error('orderBy must be a string field address') - } - parseFieldAddress(params.orderBy, 'entity') // throws InvalidFieldAddressError on a bad address - } - if (params.order !== undefined && params.order !== 'asc' && params.order !== 'desc') { - throw new Error(`order must be 'asc' or 'desc', got '${String(params.order)}'`) + // Universal truth: can't use both cursor and offset pagination + if (params.cursor !== undefined && params.offset !== undefined) { + throw new Error('cannot use both cursor and offset pagination simultaneously') } // Auto-limit query length based on memory @@ -518,28 +495,7 @@ export function validateFindParams(params: FindParams): void { /** * Validate add parameters */ - -/** - * The namespace cannot be forged: a USER metadata key literally spelled - * 'system.' would collide with the engine's explicit address - * namespace at read time — refuse it at the write door, loudly, with the - * fix in the message (sealed 2026-08-03). - */ -function rejectForgedSystemKeys(metadata: Record | undefined, site: string): void { - if (!metadata) return - for (const key of Object.keys(metadata)) { - if (key.startsWith('system.')) { - throw new Error( - `${site}: metadata key '${key}' is not allowed — the 'system.' prefix is the ` + - `engine's explicit address namespace and cannot be used as a user field name. ` + - `Rename the field (e.g. '${key.slice('system.'.length)}').` - ) - } - } -} - export function validateAddParams(params: AddParams): void { - rejectForgedSystemKeys(params.metadata as Record | undefined, 'add()') // Universal truth: must have data or vector if (!params.data && !params.vector) { throw new Error( @@ -580,7 +536,6 @@ export function validateAddParams(params: AddParams): void { * Validate update parameters */ export function validateUpdateParams(params: UpdateParams): void { - rejectForgedSystemKeys(params.metadata as Record | undefined, 'update()') // Universal truth: must have an ID if (!params.id) { throw new Error('id is required for update') @@ -618,7 +573,6 @@ export function validateUpdateParams(params: UpdateParams): void { * Validate relate parameters */ export function validateRelateParams(params: RelateParams): void { - rejectForgedSystemKeys(params.metadata as Record | undefined, 'relate()') // 8.0 verb-id contract (L.7): verb ids are UUIDs, generated by brainy. // RelateParams has no `id` field — an untyped caller passing one would // previously have it silently ignored (a generated UUID was used instead). @@ -667,7 +621,6 @@ export function validateRelateParams(params: RelateParams): void { * accepts type/subtype/weight/confidence/data/metadata changes. */ export function validateUpdateRelationParams(params: UpdateRelationParams): void { - rejectForgedSystemKeys(params.metadata as Record | undefined, 'updateRelation()') if (!params.id) { throw new Error('id is required for updateRelation') } diff --git a/src/vfs/PathResolver.ts b/src/vfs/PathResolver.ts index 502c95f0..e496c834 100644 --- a/src/vfs/PathResolver.ts +++ b/src/vfs/PathResolver.ts @@ -57,7 +57,6 @@ export class PathResolver { // Statistics private cacheHits = 0 private cacheMisses = 0 - private lastLoggedLookups = 0 // last total the maintenance tick logged stats at private metadataIndexHits = 0 private metadataIndexMisses = 0 private graphTraversalFallbacks = 0 @@ -520,14 +519,10 @@ export class PathResolver { } } - // Log cache statistics only when there is new traffic to report — an - // idle resolver stays silent. 0/0 lookups previously rendered - // "NaN% hit rate" (and the %1000 gate passes at zero), which spammed - // production journals once a minute on every idle VFS. - const totalLookups = this.cacheHits + this.cacheMisses - if (totalLookups > 0 && totalLookups !== this.lastLoggedLookups && totalLookups % 1000 === 0) { - this.lastLoggedLookups = totalLookups - prodLog.debug(`[PathResolver] Cache stats: ${Math.round((this.cacheHits / totalLookups) * 100)}% hit rate, ${this.pathCache.size} entries, ${this.hotPaths.size} hot paths`) + // Log cache statistics (in production, send to monitoring) + const hitRate = this.cacheHits / (this.cacheHits + this.cacheMisses) + if ((this.cacheHits + this.cacheMisses) % 1000 === 0) { + console.log(`[PathResolver] Cache stats: ${Math.round(hitRate * 100)}% hit rate, ${this.pathCache.size} entries, ${this.hotPaths.size} hot paths`) } }, 60000) // Every minute // Cache maintenance must never keep the host process alive. diff --git a/tests/conformance/collider-fidelity.test.ts b/tests/conformance/collider-fidelity.test.ts deleted file mode 100644 index 61c9413d..00000000 --- a/tests/conformance/collider-fidelity.test.ts +++ /dev/null @@ -1,307 +0,0 @@ -/** - * @module tests/conformance/collider-fidelity - * @description THE REOPEN-COLLIDER CONFORMANCE CASE (required cross-engine - * before any RC counts as gates-green — ruled 2026-08-03). The - * field-addressing law's fidelity half: user metadata may carry ANY name — - * including every engine spelling (`confidence`, `type`, `id`, `createdAt`, - * …) and every plumbing name (`level`, `data`, `vector`, `_rev`) — and the - * value survives, verbatim and reachable, across the FULL lifecycle: live - * reads, where/orderBy, flush, close+reopen, a forced epoch rebuild, and - * time travel. The engine scalars stay separately reachable at `system.*` - * the whole way. No halfway states. - * - * Self-arming like the namespace-law suite: skips loudly until the arming - * exports are present, so the suite can sit on a branch ahead of the build. - */ -import { describe, it, expect, beforeAll, afterAll } from 'vitest' -import { mkdtempSync, rmSync } from 'node:fs' -import { tmpdir } from 'node:os' -import { join } from 'node:path' -import * as brainyExports from '../../src/index.js' -import { Brainy, NounType, VerbType } from '../../src/index.js' -import { - BRAIN_FORMAT_PATH, - EXPECTED_INDEX_EPOCH -} from '../../src/storage/brainFormat.js' - -const ARMED = 'UnresolvableFieldError' in brainyExports -const suite = ARMED ? describe : describe.skip -if (!ARMED) { - // eslint-disable-next-line no-console - console.warn( - '[collider-fidelity] SKIPPING: package root does not export the ' + - 'field-addressing law surface yet (UnresolvableFieldError absent).' - ) -} - -/** Every entity system scalar name written as a USER metadata field, with - * unmistakable user values, plus the plumbing names and naturals. */ -const COLLIDER_BAG = { - // the ten entity system scalars, as user fields - id: 'user-id', - type: 'user-type', - subtype: 'user-subtype', - createdAt: 'user-createdAt', - updatedAt: 'user-updatedAt', - confidence: 'user-confidence', - weight: 'user-weight', - visibility: 'user-visibility', - service: 'user-service', - createdBy: 'user-createdBy', - // plumbing names, as user fields - level: 7, - data: 'user-data', - vector: 'user-vector', - _rev: 'user-rev', - // naturals previously silently un-indexed by name - content: 'user-content', - // a plain control field - plain: 'control' -} as const - - -suite('collider fidelity — the reopen-collider case (both suites, ruled)', () => { - let dir: string - let brain: Brainy - let colliderId: string - - const open = async (): Promise => { - const b = new Brainy({ - storage: { type: 'filesystem', path: dir }, - requireSubtype: false - }) - await b.init() - return b - } - - /** The full read battery — run at every lifecycle boundary. */ - const verifyColliderTruth = async (label: string): Promise => { - // 1. get(): the bag comes back verbatim; engine scalars stay engine. - const entity = await brain.get(colliderId) - expect(entity, `${label}: entity readable`).toBeTruthy() - for (const [k, v] of Object.entries(COLLIDER_BAG)) { - expect( - (entity!.metadata as Record)[k], - `${label}: bag.${k} verbatim` - ).toEqual(v) - } - expect(entity!.type, `${label}: engine type intact`).toBe(NounType.Document) - expect(entity!.confidence, `${label}: engine confidence intact`).toBe(0.25) - - // 2. where on collider names (bare = the user's field, always). - for (const [k, v] of [ - ['confidence', 'user-confidence'], - ['type', 'user-type'], - ['id', 'user-id'], - ['content', 'user-content'], - ['data', 'user-data'], - ['level', 7] - ] as const) { - const rows = await brain.find({ where: { [k]: v }, limit: 10 }) - expect( - rows.map((r) => r.id), - `${label}: where {${k}} finds the collider row` - ).toContain(colliderId) - } - - // 3. system.* keeps reading the ENGINE values. - const byEngine = await brain.find({ - where: { 'system.confidence': 0.25 }, - limit: 10 - }) - expect( - byEngine.map((r) => r.id), - `${label}: system.confidence reads the engine scalar` - ).toContain(colliderId) - const byUserSpelledSystem = await brain.find({ - where: { 'system.confidence': 'user-confidence' }, - limit: 10 - }) - expect( - byUserSpelledSystem.map((r) => r.id), - `${label}: the user's value is NOT reachable via system.*` - ).not.toContain(colliderId) - - // 4. orderBy a collider name orders by the USER values. - const ordered = await brain.find({ - type: NounType.Document, - orderBy: 'level', - order: 'desc', - limit: 10 - }) - expect(ordered.length, `${label}: ordered read complete`).toBe(3) - expect( - (ordered[0].metadata as Record).plain, - `${label}: user level orders desc (7 first)` - ).toBe('control') - } - - beforeAll(async () => { - dir = mkdtempSync(join(tmpdir(), 'brainy-collider-')) - brain = await open() - - colliderId = await brain.add({ - data: 'the collider probe document', - type: NounType.Document, - confidence: 0.25, - metadata: { ...COLLIDER_BAG } - }) - // two ordering companions with smaller user `level`s - await brain.add({ - data: 'ordering companion low', - type: NounType.Document, - metadata: { level: 3, plain: 'low' } - }) - await brain.add({ - data: 'ordering companion mid', - type: NounType.Document, - metadata: { level: 5, plain: 'mid' } - }) - }, 120000) - - afterAll(async () => { - await brain.close().catch(() => {}) - rmSync(dir, { recursive: true, force: true }) - }) - - it('LIVE: colliders are the user’s, verbatim and fully queryable', async () => { - await verifyColliderTruth('live') - }) - - it('REOPEN: the restart boundary loses nothing', async () => { - await brain.flush() - await brain.close() - brain = await open() - await verifyColliderTruth('reopen') - }) - - it('REBUILD: a forced epoch rebuild re-indexes the colliders from canonical', async () => { - await brain.close() - // Simulate epoch drift: a missing marker forces the full derived-index - // rebuild at open — the exact path every pre-law brain takes once. - rmSync(join(dir, BRAIN_FORMAT_PATH), { force: true }) - brain = await open() - await verifyColliderTruth('rebuild') - // And the rebuild re-stamps the current epoch. - const marker = await ( - brain as unknown as { - storage: { readRawObject(p: string): Promise<{ indexEpoch?: number } | null> } - } - ).storage.readRawObject(BRAIN_FORMAT_PATH) - expect(marker?.indexEpoch).toBe(EXPECTED_INDEX_EPOCH) - }) - - it('TIME TRAVEL: asOf reads historical collider values faithfully', async () => { - const gen = brain.generation() - await brain.update({ id: colliderId, metadata: { confidence: 'user-confidence-v2' } }) - const now = await brain.get(colliderId) - expect((now!.metadata as Record).confidence).toBe('user-confidence-v2') - - const past = await brain.asOf(gen) - try { - const then = await past.get(colliderId) - expect( - (then!.metadata as Record).confidence, - 'asOf reads the pre-update USER value' - ).toBe('user-confidence') - } finally { - await past.release() - } - // engine scalar untouched throughout - expect(now!.confidence).toBe(0.25) - }) - - it('RELATION MIRROR: edge collider bags survive write → read → reopen', async () => { - const a = await brain.add({ data: 'edge endpoint a', type: NounType.Person, metadata: { plain: 'a' } }) - const b = await brain.add({ data: 'edge endpoint b', type: NounType.Person, metadata: { plain: 'b' } }) - const edgeBag = { - verb: 'user-verb', - confidence: 'user-edge-confidence', - weight: 'user-edge-weight', - subtype: 'user-edge-subtype', - createdAt: 'user-edge-createdAt', - service: 'user-edge-service' - } - const relId = await brain.relate({ - from: a, - to: b, - type: VerbType.RelatedTo, - confidence: 0.5, - metadata: { ...edgeBag } - }) - - const check = async (label: string): Promise => { - const rels = await brain.related({ from: a, type: VerbType.RelatedTo }) - const rel = rels.find((r) => r.id === relId) - expect(rel, `${label}: relation readable`).toBeTruthy() - for (const [k, v] of Object.entries(edgeBag)) { - expect( - (rel!.metadata as Record)[k], - `${label}: edge bag.${k} verbatim` - ).toEqual(v) - } - expect(rel!.confidence, `${label}: engine edge confidence intact`).toBe(0.5) - expect(rel!.type, `${label}: engine verb intact`).toBe(VerbType.RelatedTo) - } - - await check('live') - await brain.flush() - await brain.close() - brain = await open() - await check('reopen') - }) - - it('FORGERY: user metadata keys spelled system.* refuse at every write door', async () => { - await expect( - brain.add({ data: 'forged', type: NounType.Document, metadata: { 'system.confidence': 1 } }) - ).rejects.toThrow(/system\./) - await expect( - brain.update({ id: colliderId, metadata: { 'system.type': 'x' } }) - ).rejects.toThrow(/system\./) - const a = await brain.add({ data: 'forgery endpoint a', type: NounType.Person, metadata: {} }) - const b = await brain.add({ data: 'forgery endpoint b', type: NounType.Person, metadata: {} }) - await expect( - brain.relate({ from: a, to: b, type: VerbType.RelatedTo, metadata: { 'system.verb': 'x' } }) - ).rejects.toThrow(/system\./) - }) - - it('CONFIG: the dead reservedFieldPolicy option refuses loudly, never ignored', () => { - expect( - () => new Brainy({ storage: { type: 'memory' }, reservedFieldPolicy: 'throw' } as never) - ).toThrow(/field-addressing law/) - }) - - it('LEGACY: a pre-law flat record still reads with engine fields top-level', async () => { - const storage = ( - brain as unknown as { - storage: { - saveNoun(n: unknown): Promise - saveNounMetadata(id: string, m: Record): Promise - } - } - ).storage - const legacyId = '00000000-0000-4000-8000-00000000f1a7' - await storage.saveNoun({ id: legacyId, vector: new Array(384).fill(0.01), connections: new Map(), level: 0 }) - // Legacy FLAT shape: engine + user keys mixed at one level, NO _fmt stamp. - // Sound to split by name — the pre-law door refused user colliders. - await storage.saveNounMetadata(legacyId, { - noun: NounType.Document, - confidence: 0.75, - createdAt: 1700000000000, - updatedAt: 1700000000000, - _rev: 1, - legacyField: 'legacy-value' - }) - const entity = await brain.get(legacyId) - expect(entity).toBeTruthy() - expect(entity!.confidence, 'legacy flat confidence = engine').toBe(0.75) - expect( - (entity!.metadata as Record).legacyField, - 'legacy custom field = user bag' - ).toBe('legacy-value') - expect( - (entity!.metadata as Record).confidence, - 'legacy flat engine key never leaks into the bag' - ).toBeUndefined() - }) -}) diff --git a/tests/conformance/namespace-law.test.ts b/tests/conformance/namespace-law.test.ts deleted file mode 100644 index 0231685a..00000000 --- a/tests/conformance/namespace-law.test.ts +++ /dev/null @@ -1,486 +0,0 @@ -/** - * @module tests/conformance/namespace-law - * @description Conformance suite for the ruled field-addressing contract - * announced in RELEASES.md ("Coming next... one field-addressing law — bare - * names = user metadata, `system.` for engine fields, typed refusals - * for unresolvable names"). This suite is the drift-proof shared by this - * engine and its native accelerator: both must satisfy every test here - * bit-for-bit, because they implement the SAME contract independently. - * - * The rule, in full: - * 1. A bare field name in `where` / `orderBy` / `groupBy` / aggregation - * `source.where` ALWAYS means the caller's own `metadata` field. No - * priority resolution, no engine fallback — ever. - * 2. `system.` reaches an engine scalar, and ONLY an engine scalar, - * and ONLY when spelled explicitly. The addressable entity map is exactly - * ten names: id, type, subtype, createdAt, updatedAt, confidence, weight, - * visibility, service, createdBy. The relationship map is system.verb, - * system.sourceId, system.targetId, plus the eight scalars shared with - * entities. - * 3. Some names are invisible plumbing and are never addressable in either - * spelling: vector, connections, level, data, _rev. `system.level`, - * `system.vector`, and `system.data` all refuse — they are not in the - * system map. Bare `level` is a perfectly ordinary user field. - * 4. `metadata.` is the explicit-user-scope spelling: identical - * semantics to the bare spelling, valid everywhere the bare spelling is. - * 5. Anything that resolves to neither a user field nor a system scalar is a - * typed refusal naming both candidates (`UnresolvableFieldError`). - * Unimplemented `find()` options (`cursor`, `includeRelations`, - * `writeOnly`) refuse with `UnsupportedFindOptionError` instead of being - * silently accepted and ignored. - * 6. Ordering is identical on both engines: rows missing/null on the - * `orderBy` field sort LAST in BOTH directions and are never dropped; - * ties break by id ascending. - * - * The motivating incident (told generically — see CLAUDE.md naming rule): an - * internal report from a production deployment showed a user metadata field - * literally named `level` silently shadowed by the engine's internal HNSW - * node layer, breaking sort order with zero errors raised. This contract - * makes that class of bug impossible, and testable forever. - * - * SELF-SKIP: the resolver this suite pins is being built in a parallel - * session and has not landed on every branch yet. Rather than going red on - * a branch that simply hasn't caught up, the suite detects whether the - * contract is live by the one thing any conformant implementation must - * export — `UnresolvableFieldError` from the package root — and skips - * loudly (never silently) until it does. This is the house pattern: a - * sibling engine's gate once went red because a test armed before its - * feature existed. - */ -import { describe, it, expect, beforeEach, afterEach } from 'vitest' -import { Brainy } from '../../src/brainy.js' -import { NounType } from '../../src/types/graphTypes.js' -import * as brainyExports from '../../src/index.js' - -const stubEmbedding = async (text: string): Promise => { - const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0) - return new Array(384).fill(0).map((_, i) => Math.sin(hash + i)) -} - -// Detected purely by the exported error-class NAME — never by reaching into -// implementation internals. Both engines building this contract must export -// it from the package root, so this is a legitimate, implementation-agnostic -// readiness probe. -const lawActive = 'UnresolvableFieldError' in brainyExports -const UnresolvableFieldError = (brainyExports as Record).UnresolvableFieldError as new ( - ...args: any[] -) => Error -const UnsupportedFindOptionError = (brainyExports as Record) - .UnsupportedFindOptionError as new (...args: any[]) => Error - -// Always runs, regardless of lawActive — the loud signal that the rest of -// this file was skipped, and why. -it('namespace law armed?', () => { - if (!lawActive) { - console.warn( - '[conformance] namespace-law suite SKIPPED — UnresolvableFieldError not exported yet; arms when the resolver lands' - ) - } - expect(true).toBe(true) -}) - -/** - * Awaits `promise`, asserting it rejects with an instance of `ErrorClass` - * whose `.message` contains every string in `mustContain`. Fails loudly if - * the promise resolves instead of rejecting. - */ -async function expectRefusal( - promise: Promise, - ErrorClass: new (...args: any[]) => Error, - ...mustContain: string[] -): Promise { - let threw = false - try { - await promise - } catch (err) { - threw = true - expect(err).toBeInstanceOf(ErrorClass) - for (const fragment of mustContain) { - expect((err as Error).message).toContain(fragment) - } - } - expect(threw).toBe(true) -} - -describe.skipIf(!lawActive)('namespace law — bare/system/metadata field addressing', () => { - let brain: Brainy - - beforeEach(async () => { - brain = new Brainy({ - requireSubtype: false, - storage: { type: 'memory' as const }, - embeddingFunction: stubEmbedding - }) - await brain.init() - }) - - afterEach(async () => { - await brain.close() - }) - - /** The star case from the motivating incident: metadata.level 3/9/6. */ - async function addLevelRows(): Promise { - const ids: string[] = [] - for (const level of [3, 9, 6]) { - ids.push( - await brain.add({ - data: `probe level ${level}`, - type: NounType.Person, - subtype: 'ns-law-level', - metadata: { name: `p-${level}`, level } - }) - ) - } - return ids - } - - // ------------------------------------------------------------------- - // Rule 1 — bare field name = the user's metadata field, always. - // ------------------------------------------------------------------- - - it("bare orderBy 'level' reads user metadata, desc and asc (the star case)", async () => { - await addLevelRows() - - const desc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-level', - orderBy: 'level', - order: 'desc', - limit: 100 - }) - expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3]) - - const asc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-level', - orderBy: 'level', - order: 'asc', - limit: 100 - }) - expect(asc.map((r: any) => r.metadata?.level)).toEqual([3, 6, 9]) - }) - - it("bare where { level: N } matches the user's field", async () => { - const ids = await addLevelRows() - const hit = await brain.find({ type: NounType.Person, subtype: 'ns-law-level', where: { level: 9 } }) - expect(hit).toHaveLength(1) - expect(hit[0].id).toBe(ids[1]) - expect(hit[0].metadata?.level).toBe(9) - }) - - // ------------------------------------------------------------------- - // Rule 4 — metadata. is the explicit-user-scope spelling, - // identical semantics to bare, valid on every path including orderBy. - // ------------------------------------------------------------------- - - it("'metadata.level' resolves identically to bare 'level'", async () => { - await addLevelRows() - const desc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-level', - orderBy: 'metadata.level', - order: 'desc', - limit: 100 - }) - expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3]) - }) - - // ------------------------------------------------------------------- - // Rule 2 — system. reaches an engine scalar explicitly. - // ------------------------------------------------------------------- - - it('system.createdAt sorts by entity age', async () => { - const ids: string[] = [] - for (const name of ['first', 'second', 'third']) { - ids.push( - await brain.add({ - data: `aged ${name}`, - type: NounType.Person, - subtype: 'ns-law-aged', - metadata: { name } - }) - ) - // Guarantee distinct createdAt timestamps between adds. - await new Promise((resolve) => setTimeout(resolve, 5)) - } - - const asc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-aged', - orderBy: 'system.createdAt', - order: 'asc', - limit: 100 - }) - expect(asc.map((r: any) => r.id)).toEqual(ids) - - const desc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-aged', - orderBy: 'system.createdAt', - order: 'desc', - limit: 100 - }) - expect(desc.map((r: any) => r.id)).toEqual([...ids].reverse()) - }) - - it('where on system.confidence filters by the engine scalar', async () => { - const highId = await brain.add({ - data: 'high confidence row', - type: NounType.Person, - subtype: 'ns-law-confidence', - confidence: 0.95, - metadata: { name: 'hi' } - }) - await brain.add({ - data: 'low confidence row', - type: NounType.Person, - subtype: 'ns-law-confidence', - confidence: 0.4, - metadata: { name: 'lo' } - }) - - const hit = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-confidence', - where: { 'system.confidence': 0.95 } - }) - expect(hit).toHaveLength(1) - expect(hit[0].id).toBe(highId) - }) - - it('groupBy on system.subtype groups by the engine scalar, not user metadata', async () => { - await brain.add({ data: 'i1', type: NounType.Document, subtype: 'invoice' }) - await brain.add({ data: 'i2', type: NounType.Document, subtype: 'invoice' }) - await brain.add({ data: 'r1', type: NounType.Document, subtype: 'receipt' }) - - brain.defineAggregate({ - name: 'ns_law_by_subtype_system', - source: { type: NounType.Document }, - groupBy: ['system.subtype'], - metrics: { count: { op: 'count' } } - }) - - const groups = await brain.queryAggregate('ns_law_by_subtype_system') - const invoiceGroup = groups.find((g) => Object.values(g.groupKey).includes('invoice')) - const receiptGroup = groups.find((g) => Object.values(g.groupKey).includes('receipt')) - expect(invoiceGroup?.metrics.count).toBe(2) - expect(receiptGroup?.metrics.count).toBe(1) - }) - - // ------------------------------------------------------------------- - // Rule 1 (groupBy face) — bare groupBy dimensions read user metadata, - // never the engine's own notion of the same-sounding name. - // ------------------------------------------------------------------- - - it('groupBy on a bare user metadata field groups by that field', async () => { - await brain.add({ - data: 'd1', - type: NounType.Document, - subtype: 'ns-law-group-bare', - metadata: { team: 'alpha' } - }) - await brain.add({ - data: 'd2', - type: NounType.Document, - subtype: 'ns-law-group-bare', - metadata: { team: 'alpha' } - }) - await brain.add({ - data: 'd3', - type: NounType.Document, - subtype: 'ns-law-group-bare', - metadata: { team: 'beta' } - }) - - brain.defineAggregate({ - name: 'ns_law_by_team_bare', - // system.subtype — bare 'subtype' would address user metadata under the - // law (the exact migration every fleet consumer's aggregates make). - source: { type: NounType.Document, where: { 'system.subtype': 'ns-law-group-bare' } }, - groupBy: ['team'], - metrics: { count: { op: 'count' } } - }) - - const groups = await brain.queryAggregate('ns_law_by_team_bare') - const alphaGroup = groups.find((g) => Object.values(g.groupKey).includes('alpha')) - const betaGroup = groups.find((g) => Object.values(g.groupKey).includes('beta')) - expect(alphaGroup?.metrics.count).toBe(2) - expect(betaGroup?.metrics.count).toBe(1) - }) - - it('where on a bare user metadata field filters normally (score, not a system name)', async () => { - await brain.add({ - data: 'high score', - type: NounType.Person, - subtype: 'ns-law-score', - metadata: { score: 42 } - }) - await brain.add({ - data: 'low score', - type: NounType.Person, - subtype: 'ns-law-score', - metadata: { score: 7 } - }) - - const hit = await brain.find({ type: NounType.Person, subtype: 'ns-law-score', where: { score: 42 } }) - expect(hit).toHaveLength(1) - expect(hit[0].metadata?.score).toBe(42) - }) - - // ------------------------------------------------------------------- - // Rule 5 — typed refusals, naming both candidates. - // ------------------------------------------------------------------- - - it("bare orderBy 'createdAt' refuses when no such metadata field exists — names both candidates", async () => { - await brain.add({ - data: 'no metadata.createdAt here', - type: NounType.Person, - subtype: 'ns-law-refuse-createdAt', - metadata: { name: 'x' } - }) - - await expectRefusal( - brain.find({ - type: NounType.Person, - subtype: 'ns-law-refuse-createdAt', - orderBy: 'createdAt', - limit: 10 - }), - UnresolvableFieldError, - 'system.createdAt', - 'metadata.createdAt' - ) - }) - - // ------------------------------------------------------------------- - // Rule 3 — invisible plumbing refuses in either spelling; system. - // for a name that isn't in the ten-scalar map is unresolvable. - // ------------------------------------------------------------------- - - it('system.level refuses — level is invisible plumbing, never a system scalar', async () => { - await brain.add({ - data: 'has a level metadata field', - type: NounType.Person, - metadata: { level: 5 } - }) - await expectRefusal(brain.find({ orderBy: 'system.level', limit: 10 }), UnresolvableFieldError) - }) - - it('system.vector refuses — vector is invisible plumbing, never a system scalar', async () => { - await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } }) - await expectRefusal(brain.find({ orderBy: 'system.vector', limit: 10 }), UnresolvableFieldError) - }) - - it('system.data refuses — data is a payload container, never a system scalar', async () => { - await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } }) - await expectRefusal(brain.find({ orderBy: 'system.data', limit: 10 }), UnresolvableFieldError) - }) - - // ------------------------------------------------------------------- - // Rule 6 — the ordering contract. - // ------------------------------------------------------------------- - - async function addOrderingProbeRows(): Promise<{ ranked: string[]; missing: string }> { - const low = await brain.add({ - data: 'low score', - type: NounType.Person, - subtype: 'ns-law-ordering', - metadata: { score: 5 } - }) - const high = await brain.add({ - data: 'high score', - type: NounType.Person, - subtype: 'ns-law-ordering', - metadata: { score: 9 } - }) - const missing = await brain.add({ - data: 'no score field at all', - type: NounType.Person, - subtype: 'ns-law-ordering', - metadata: { name: 'no-score' } - }) - return { ranked: [low, high], missing } - } - - it('a row missing the orderBy field sorts LAST in desc — and is never dropped', async () => { - const { ranked, missing } = await addOrderingProbeRows() - const desc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-ordering', - orderBy: 'score', - order: 'desc', - limit: 100 - }) - expect(desc).toHaveLength(3) - expect(desc.map((r: any) => r.id)).toEqual([ranked[1], ranked[0], missing]) - }) - - it('a row missing the orderBy field sorts LAST in asc too — and is never dropped', async () => { - const { ranked, missing } = await addOrderingProbeRows() - const asc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-ordering', - orderBy: 'score', - order: 'asc', - limit: 100 - }) - expect(asc).toHaveLength(3) - expect(asc.map((r: any) => r.id)).toEqual([ranked[0], ranked[1], missing]) - }) - - it('ties on the orderBy field break by id ascending, in BOTH directions', async () => { - const tiedIds: string[] = [] - for (let i = 0; i < 4; i++) { - tiedIds.push( - await brain.add({ - data: `tied ${i}`, - type: NounType.Person, - subtype: 'ns-law-ties', - metadata: { score: 5 } - }) - ) - } - const expectedOrder = [...tiedIds].sort() - - const asc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-ties', - orderBy: 'score', - order: 'asc', - limit: 100 - }) - expect(asc.map((r: any) => r.id)).toEqual(expectedOrder) - - const desc = await brain.find({ - type: NounType.Person, - subtype: 'ns-law-ties', - orderBy: 'score', - order: 'desc', - limit: 100 - }) - // Same tie-break ordering regardless of the primary direction — the - // contract states one universal rule ("id ascending"), not "reverse of - // the primary order". - expect(desc.map((r: any) => r.id)).toEqual(expectedOrder) - }) - - // ------------------------------------------------------------------- - // Rule 5 (options face) — unimplemented find() options refuse loudly - // instead of being accepted and silently ignored. - // ------------------------------------------------------------------- - - it('find({ cursor }) refuses with UnsupportedFindOptionError', async () => { - await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } }) - await expectRefusal(brain.find({ cursor: 'anything', limit: 10 }), UnsupportedFindOptionError) - }) - - it('find({ includeRelations }) refuses with UnsupportedFindOptionError', async () => { - await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } }) - await expectRefusal(brain.find({ includeRelations: true, limit: 10 }), UnsupportedFindOptionError) - }) - - it('find({ writeOnly }) refuses with UnsupportedFindOptionError', async () => { - await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } }) - await expectRefusal(brain.find({ writeOnly: true, limit: 10 }), UnsupportedFindOptionError) - }) -}) diff --git a/tests/integration/advanced-apis-regression.test.ts b/tests/integration/advanced-apis-regression.test.ts index 069d3e62..12c39112 100644 --- a/tests/integration/advanced-apis-regression.test.ts +++ b/tests/integration/advanced-apis-regression.test.ts @@ -164,19 +164,19 @@ describe('BR-ADV-FEATURES-BUN regression', () => { await b.close() }) - it('groupBy "system.type" resolves to the entity type, not null (the legacy "noun" alias is dead)', async () => { + it('groupBy "noun" resolves to the entity type, not null', async () => { const b: any = new Brainy({ requireSubtype: false, storage: { type: 'memory' } }) await b.init() await b.add({ data: 'p', type: NounType.Person }) b.defineAggregate({ name: 'byNoun', source: { type: NounType.Person }, - groupBy: ['system.type'], + groupBy: ['noun'], metrics: { count: { op: 'count' } } }) const rows: any[] = await b.find({ aggregate: 'byNoun' }) expect(rows.length).toBe(1) - expect(rows[0].groupKey['system.type']).toBe(NounType.Person) + expect(rows[0].groupKey.noun).toBe(NounType.Person) await b.close() }) }) diff --git a/tests/integration/aggregate-reserved-fields.test.ts b/tests/integration/aggregate-reserved-fields.test.ts index b8c11b4f..e81692d6 100644 --- a/tests/integration/aggregate-reserved-fields.test.ts +++ b/tests/integration/aggregate-reserved-fields.test.ts @@ -42,11 +42,8 @@ describe('aggregation + query field-resolution law', () => { it('reserved-field groupBy decrements on delete (the drift bug)', async () => { brain.defineAggregate({ name: 'by_subtype', - // system.subtype — subtype is an add() param (an engine scalar), never - // a user metadata field; bare 'subtype' now addresses the user's own - // metadata bag under the sealed field-addressing law. source: { type: NounType.Document }, - groupBy: ['system.subtype'], + groupBy: ['subtype'], metrics: { count: { op: 'count' } } }) @@ -63,7 +60,7 @@ describe('aggregation + query field-resolution law', () => { } let groups = await brain.queryAggregate('by_subtype') expect(groups).toHaveLength(1) - expect(groups[0].groupKey).toEqual({ 'system.subtype': 'note' }) + expect(groups[0].groupKey).toEqual({ subtype: 'note' }) expect(groups[0].metrics.count).toBe(5) await brain.remove(ids[0]) @@ -79,7 +76,7 @@ describe('aggregation + query field-resolution law', () => { brain.defineAggregate({ name: 'by_subtype', source: { type: NounType.Document }, - groupBy: ['system.subtype'], + groupBy: ['subtype'], metrics: { count: { op: 'count' } } }) const id = await brain.add({ @@ -91,7 +88,7 @@ describe('aggregation + query field-resolution law', () => { const groups = await brain.queryAggregate('by_subtype') const byKey = Object.fromEntries( - groups.map((g) => [String(g.groupKey['system.subtype']), g.metrics.count]) + groups.map((g) => [String(g.groupKey.subtype), g.metrics.count]) ) expect(byKey['published']).toBe(1) // The old group must be gone or zero — never still counting the entity. @@ -101,7 +98,7 @@ describe('aggregation + query field-resolution law', () => { it('source.where on a reserved field filters instead of matching nothing', async () => { brain.defineAggregate({ name: 'notes_only', - source: { type: NounType.Document, where: { 'system.subtype': 'note' } }, + source: { type: NounType.Document, where: { subtype: 'note' } }, groupBy: ['team'], metrics: { count: { op: 'count' } } }) diff --git a/tests/integration/all-apis-comprehensive.test.ts b/tests/integration/all-apis-comprehensive.test.ts index 7d82afea..d25f23c8 100644 --- a/tests/integration/all-apis-comprehensive.test.ts +++ b/tests/integration/all-apis-comprehensive.test.ts @@ -331,10 +331,8 @@ describe('Comprehensive All-APIs Test', () => { it('should handle metadata queries efficiently', async () => { const start = Date.now() - // system.type — the legacy where.type→noun alias is dead; bare 'type' - // in where now addresses the user's own metadata field. const results = await brain.find({ - where: { 'system.type': NounType.Document }, + where: { type: NounType.Document }, limit: 100 }) diff --git a/tests/integration/fact-log-dual-write.test.ts b/tests/integration/fact-log-dual-write.test.ts index 3c4eee4a..5ec66273 100644 --- a/tests/integration/fact-log-dual-write.test.ts +++ b/tests/integration/fact-log-dual-write.test.ts @@ -11,7 +11,7 @@ import { describe, it, expect, beforeEach, afterEach } from 'vitest' import * as fs from 'node:fs' import * as os from 'node:os' import * as path from 'node:path' -import { Brainy, ProtectedArtifactError, splitNounMetadataRecord, type CommitFact } from '../../src/index.js' +import { Brainy, ProtectedArtifactError, type CommitFact } from '../../src/index.js' async function allFacts(brain: any): Promise { const scan = brain.scanFacts() @@ -65,13 +65,7 @@ describe('fact log dual-write (memory adapter)', () => { const updateFact = facts[facts.length - 1] const op = updateFact.ops.find((o) => o.id === id)! expect(op.record).not.toBeNull() - // The fact log is byte-faithful: op.record.metadata is the RAW stored - // record (v2 nested-bag since the field-addressing law) — read the user - // field through the shape-aware split, like every other reader. - const { custom } = splitNounMetadataRecord( - op.record!.metadata as Record - ) - expect(custom.v).toBe('new') + expect((op.record!.metadata as any).v).toBe('new') }) it('a transact commits ONE fact carrying all its ops, with meta', async () => { diff --git a/tests/integration/history-repacking.test.ts b/tests/integration/history-repacking.test.ts deleted file mode 100644 index 2bcee038..00000000 --- a/tests/integration/history-repacking.test.ts +++ /dev/null @@ -1,186 +0,0 @@ -/** - * @module tests/integration/history-repacking - * @description The D1+D3 two-tier history lifecycle end-to-end on a real - * brain. Laws: (1) repacking is RE-REPRESENTATION — after folding, every - * asOf() read below the fold boundary answers exactly as before, across a - * cold reopen; (2) folded per-generation directories are physically gone - * (the file-count cure is real, not cosmetic); (3) repack + reclaim compose: - * bounded retention after repacking drops whole segments and asOf below the - * horizon throws GenerationCompactedError; (4) repackHistory is explicit - * API and time-bounded (spent budget = consistent no-op). - * - * Uses a tiny REPACK_LIVE_WINDOW override so a small history has a cold - * tier at all (the production window is 1024). - */ -import { describe, it, expect, afterEach } from 'vitest' -import * as fs from 'node:fs' -import * as path from 'node:path' -import * as os from 'node:os' -import { Brainy } from '../../src/brainy.js' -import { NounType } from '../../src/types/graphTypes.js' -import { GenerationStore } from '../../src/db/generationStore.js' -import { GenerationCompactedError } from '../../src/db/errors.js' -import { SEGMENTS_PREFIX } from '../../src/db/generationSegments.js' - -const stub = async (text: string): Promise => { - const h = text.split('').reduce((a, c) => a + c.charCodeAt(0), 0) - return new Array(384).fill(0).map((_, i) => Math.sin(h + i)) -} - -const openBrain = async (dir: string): Promise => { - const brain = new Brainy({ - requireSubtype: false, - storage: { type: 'filesystem', path: dir }, - embeddingFunction: stub - }) - await brain.init() - return brain -} - -describe('history repacking — the two-tier lifecycle', () => { - const dirs: string[] = [] - const tempDir = (): string => { - const d = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-repack-')) - dirs.push(d) - return d - } - const originalWindow = GenerationStore.REPACK_LIVE_WINDOW - - afterEach(() => { - ;(GenerationStore as any).REPACK_LIVE_WINDOW = originalWindow - for (const d of dirs.splice(0)) { - try { - fs.rmSync(d, { recursive: true, force: true }) - } catch { - /* best effort */ - } - } - }) - - it('repack preserves every historical read across cold reopen; folded dirs are gone', async () => { - ;(GenerationStore as any).REPACK_LIVE_WINDOW = 3 - const dir = tempDir() - const brain = await openBrain(dir) - - const id = await brain.add({ - data: 'versioned-entity', - type: NounType.Document, - metadata: { v: 0 } - }) - for (let v = 1; v <= 10; v++) await brain.update({ id, metadata: { v } }) - await brain.flush() - - // Ground truth BEFORE repacking: capture asOf views for early generations. - const before: Record = {} - for (const g of [2, 4, 6]) { - const db = await brain.asOf(g) - before[g] = (await db.get(id))?.metadata?.v as number - await db.release() - } - - const result = await brain.repackHistory() - expect(result.foldedGenerations).toBeGreaterThan(0) - expect(result.segmentsCreated).toBeGreaterThan(0) - - // The folded per-generation directories are PHYSICALLY gone… - const genDirs = fs - .readdirSync(path.join(dir, '_generations'), { withFileTypes: true }) - .filter((e) => e.isDirectory() && /^\d+$/.test(e.name)).length - expect(genDirs).toBeLessThanOrEqual(4) // live window (3) + at most the newest - // …and the segment tier exists (the filesystem adapter stores objects - // gzipped, so the manifest may live at either spelling). - const segDir = path.join(dir, SEGMENTS_PREFIX) - expect( - fs.existsSync(path.join(segDir, 'manifest.json')) || - fs.existsSync(path.join(segDir, 'manifest.json.gz')) - ).toBe(true) - expect(fs.readdirSync(segDir).some((f) => f.endsWith('.bgs'))).toBe(true) - - // Same asOf answers from the packed tier, same process… - for (const g of [2, 4, 6]) { - const db = await brain.asOf(g) - expect((await db.get(id))?.metadata?.v).toBe(before[g]) - await db.release() - } - await brain.close() - - // …and across a COLD REOPEN (manifest discovery, no live dirs to list). - const reopened = await openBrain(dir) - for (const g of [2, 4, 6]) { - const db = await reopened.asOf(g) - expect((await db.get(id))?.metadata?.v).toBe(before[g]) - await db.release() - } - expect((await reopened.get(id))?.metadata?.v).toBe(10) // live state untouched - await reopened.close() - }) - - it('repack + bounded reclaim compose: whole segments drop, horizon is loud', async () => { - ;(GenerationStore as any).REPACK_LIVE_WINDOW = 2 - const dir = tempDir() - const brain = await openBrain(dir) - const id = await brain.add({ data: 'reclaim-probe', type: NounType.Document, metadata: { v: 0 } }) - for (let v = 1; v <= 8; v++) await brain.update({ id, metadata: { v } }) - await brain.flush() - await brain.repackHistory() - - // Reclaim down to the 3 newest generations — packed segments below the - // horizon drop whole; asOf below throws loudly. - const res = await brain.compactHistory({ maxGenerations: 3 }) - expect(res.removedGenerations).toBeGreaterThan(0) - await expect(brain.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError) - expect((await brain.get(id))?.metadata?.v).toBe(8) - await brain.close() - }) - - it('generationDigest: reopen-stable, divergence-sensitive, loud below the horizon', async () => { - ;(GenerationStore as any).REPACK_LIVE_WINDOW = 2 - const dir = tempDir() - const brain = await openBrain(dir) - const id = await brain.add({ data: 'digest-probe', type: NounType.Document, metadata: { v: 0 } }) - for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } }) - await brain.flush() - await brain.repackHistory() - - const gen = brain.generation() - const atHead = await brain.generationDigest(gen) - const atMid = await brain.generationDigest(3) - expect(atHead).toMatch(/^[0-9a-f]{8}$/) - expect(atMid).not.toBe(atHead) // more history ⇒ different digest - await brain.close() - - // Reopen-stable: same history, same digests (packed prefix stability). - const reopened = await openBrain(dir) - expect(await reopened.generationDigest(gen)).toBe(atHead) - expect(await reopened.generationDigest(3)).toBe(atMid) - - // New history diverges the head digest. - await reopened.update({ id, metadata: { v: 7 } }) - await reopened.flush() - expect(await reopened.generationDigest(reopened.generation())).not.toBe(atHead) - - // Below the horizon: LOUD, never a silent pin of reclaimed history. - await reopened.compactHistory({ maxGenerations: 2 }) - await expect(reopened.generationDigest(1)).rejects.toBeInstanceOf(GenerationCompactedError) - await reopened.close() - }) - - it('a spent time budget is a consistent no-op; the next pass resumes', async () => { - ;(GenerationStore as any).REPACK_LIVE_WINDOW = 2 - const dir = tempDir() - const brain = await openBrain(dir) - const id = await brain.add({ data: 'budget-probe', type: NounType.Document, metadata: { v: 0 } }) - for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } }) - await brain.flush() - - const bounded = await brain.repackHistory({ timeBudgetMs: 0 }) - expect(bounded).toEqual({ foldedGenerations: 0, segmentsCreated: 0 }) - - const resumed = await brain.repackHistory() - expect(resumed.foldedGenerations).toBeGreaterThan(0) - const db = await brain.asOf(3) - expect((await db.get(id))?.metadata?.v).toBeDefined() - await db.release() - await brain.close() - }) -}) diff --git a/tests/integration/lens-consistency.test.ts b/tests/integration/lens-consistency.test.ts index 1b1cef81..64484a38 100644 --- a/tests/integration/lens-consistency.test.ts +++ b/tests/integration/lens-consistency.test.ts @@ -2,8 +2,8 @@ * @module tests/integration/lens-consistency * @description The three metadata "lenses" over one corpus must agree with * canonical ground truth id-for-id, warm AND after a cold reopen: - * - combined: find({ type: T, where: { 'system.subtype': S } }) - * - subtype-only: find({ where: { 'system.subtype': S } }) + * - combined: find({ type: T, where: { subtype: S } }) + * - subtype-only: find({ where: { subtype: S } }) * - type-only: find({ type: T }) * Ported from the fresh-brain probe that closed the type+subtype lens-drop * investigation (a restored pre-8.2.2 torn capture had entities visible to the @@ -63,11 +63,8 @@ async function assertAllLenses(brain: any): Promise { const subtypes = [...new Set(CORPUS.map((c) => c.subtype))] for (const { type, subtype } of CORPUS) { - // system.subtype — subtype is an add()/update() param (an engine scalar), - // never a user metadata field; bare 'subtype' now addresses the user's - // own metadata bag under the sealed field-addressing law. - const combined = idSet(await brain.find({ type, where: { 'system.subtype': subtype }, limit: 1000 })) - const subtypeOnly = idSet(await brain.find({ where: { 'system.subtype': subtype }, limit: 1000 })) + const combined = idSet(await brain.find({ type, where: { subtype }, limit: 1000 })) + const subtypeOnly = idSet(await brain.find({ where: { subtype }, limit: 1000 })) const truthPair = await groundTruth(brain, { type, subtype }) const truthSubtype = await groundTruth(brain, { subtype }) @@ -85,7 +82,7 @@ async function assertAllLenses(brain: any): Promise { // Count cross-check against the corpus definition itself. for (const subtype of subtypes) { const expected = CORPUS.filter((c) => c.subtype === subtype).reduce((s, c) => s + c.count, 0) - const got = (await brain.find({ where: { 'system.subtype': subtype }, limit: 1000 })).length + const got = (await brain.find({ where: { subtype }, limit: 1000 })).length expect(got).toBe(expected) } } @@ -126,16 +123,16 @@ describe('lens consistency — combined vs subtype-only vs canonical ground trut it('after an update() flips type AND subtype, every lens tracks the move exactly', async () => { // The historical cross-bucket-staleness path: change (concept, action) -> (task, review). - const victims = await brain.find({ type: 'concept', where: { 'system.subtype': 'action' }, limit: 1 }) + const victims = await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1 }) expect(victims.length).toBe(1) const id = victims[0].id await brain.update({ id, type: 'task', subtype: 'review' }) - const oldCombined = idSet(await brain.find({ type: 'concept', where: { 'system.subtype': 'action' }, limit: 1000 })) + const oldCombined = idSet(await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1000 })) expect(oldCombined.has(id)).toBe(false) // unposted from the old buckets - const newCombined = idSet(await brain.find({ type: 'task', where: { 'system.subtype': 'review' }, limit: 1000 })) + const newCombined = idSet(await brain.find({ type: 'task', where: { subtype: 'review' }, limit: 1000 })) expect(newCombined.has(id)).toBe(true) // posted to the new buckets - const subtypeOnly = idSet(await brain.find({ where: { 'system.subtype': 'review' }, limit: 1000 })) + const subtypeOnly = idSet(await brain.find({ where: { subtype: 'review' }, limit: 1000 })) expect(subtypeOnly.has(id)).toBe(true) }) }) diff --git a/tests/integration/level-field-shadow.test.ts b/tests/integration/level-field-shadow.test.ts deleted file mode 100644 index cfe34c13..00000000 --- a/tests/integration/level-field-shadow.test.ts +++ /dev/null @@ -1,194 +0,0 @@ -/** - * @module tests/integration/level-field-shadow - * @description The reserved-name shadow fix (VENUE-BRAINY-ORDERBY-NOOP, - * 2026-08-03): `level` is HNSW plumbing, not an entity field — it must never - * shadow user metadata of the same name. Pre-fix, STANDARD_ENTITY_FIELDS - * listed `level`, so every by-name read returned the engine's internal 0 - * (all-equal → stable sort → insertion order, silently), and the indexing - * views stamped level:0 into the same flattened column as user values - * (multi-valued [0, real] poison). Laws: - * (1) venue's exact repro sorts: three adds with metadata.level 3/9/6 → - * find({orderBy:'level'}) returns 9,6,3 desc and 3,6,9 asc; - * (2) where {level: N} matches through filter AND egress guard; - * (3) the index column carries the user value only (no 0 poison); - * (4) update() keeps `level` readable (the update indexing view is clean too); - * (5) the transact() update path never rewrites the noun record on a - * metadata-only patch (the planUpdate granularity completion). - */ -import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest' -import { Brainy } from '../../src/brainy.js' -import { NounType } from '../../src/types/graphTypes.js' -import { EXPECTED_INDEX_EPOCH } from '../../src/storage/brainFormat.js' - -const stubEmbedding = async (text: string): Promise => { - const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0) - return new Array(384).fill(0).map((_, i) => Math.sin(hash + i)) -} - -describe('level field shadow — user metadata named level is a real field', () => { - let brain: Brainy - - beforeEach(async () => { - brain = new Brainy({ - requireSubtype: false, - storage: { type: 'memory' as const }, - embeddingFunction: stubEmbedding - }) - await brain.init() - }) - - afterEach(async () => { - await brain.close() - }) - - async function addProbeRows(): Promise { - const ids: string[] = [] - for (const level of [3, 9, 6]) { - ids.push( - await brain.add({ - data: `probe character level ${level}`, - type: NounType.Person, - subtype: 'probe-char', - metadata: { name: `char-${level}`, level } - }) - ) - } - return ids - } - - it("venue's exact repro: orderBy 'level' sorts desc and asc", async () => { - await addProbeRows() - - const desc = await brain.find({ - type: NounType.Person, - subtype: 'probe-char', - orderBy: 'level', - order: 'desc', - limit: 100 - }) - expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3]) - - const asc = await brain.find({ - type: NounType.Person, - subtype: 'probe-char', - orderBy: 'level', - order: 'asc', - limit: 100 - }) - expect(asc.map((r: any) => r.metadata?.level)).toEqual([3, 6, 9]) - }) - - it('ordered reads are COMPLETE — no row dropped (the 2-of-3 face)', async () => { - const ids = await addProbeRows() - const desc = await brain.find({ - type: NounType.Person, - subtype: 'probe-char', - orderBy: 'level', - order: 'desc', - limit: 100 - }) - expect(desc).toHaveLength(3) - expect(new Set(desc.map((r: any) => r.id))).toEqual(new Set(ids)) - }) - - it('where {level: N} matches through the filter and the egress guard', async () => { - const ids = await addProbeRows() - const hit = await brain.find({ where: { level: 9 } }) - expect(hit).toHaveLength(1) - expect(hit[0].id).toBe(ids[1]) - expect(hit[0].metadata?.level).toBe(9) - }) - - it('the index column carries ONLY the user value (no 0 poison)', async () => { - const ids = await addProbeRows() - const metadataIndex = (brain as any).metadataIndex - const value = await metadataIndex.getFieldValueForEntity(ids[1], 'level') - expect(value).toBe(9) - - // Zero must not match anything — pre-fix every entity carried a phantom 0. - const phantom = await brain.find({ where: { level: 0 } }) - expect(phantom).toHaveLength(0) - }) - - it('update() keeps level readable (the update indexing view is clean)', async () => { - const ids = await addProbeRows() - await brain.update({ id: ids[0], metadata: { level: 12 } }) - const desc = await brain.find({ - type: NounType.Person, - subtype: 'probe-char', - orderBy: 'level', - order: 'desc', - limit: 100 - }) - expect(desc.map((r: any) => r.metadata?.level)).toEqual([12, 9, 6]) - }) - - it('transact() metadata-only update never rewrites the noun record', async () => { - const ids = await addProbeRows() - const storage = (brain as any).storage - const saveNounSpy = vi.spyOn(storage, 'saveNoun') - - await brain.transact([ - { op: 'update', id: ids[0], metadata: { level: 4 } }, - { op: 'update', id: ids[2], metadata: { level: 7 } } - ]) - - expect(saveNounSpy).not.toHaveBeenCalled() - saveNounSpy.mockRestore() - - const after = await brain.get(ids[0], { includeVectors: true }) - expect(after?.metadata?.level).toBe(4) - expect(Array.isArray(after?.vector) && after!.vector!.length).toBe(384) - }) - - it('this build runs index epoch 3 (the namespace-law key split rebuild)', () => { - expect(EXPECTED_INDEX_EPOCH).toBe(3) - }) -}) - -describe('noun-record writes never stamp over stored graph state', () => { - let brain: Brainy - - beforeEach(async () => { - brain = new Brainy({ - requireSubtype: false, - storage: { type: 'memory' as const }, - embeddingFunction: stubEmbedding - }) - await brain.init() - }) - - afterEach(async () => { - await brain.close() - }) - - it('a data-changing update preserves LEGACY inline connections in the record', async () => { - // Codec-era records carry an EMPTY connections field by design (the - // adjacency lives in a separate compressed blob) — the clobber window - // exists only for legacy pre-codec records whose adjacency is inline. - // Simulate one: write the record with inline connections directly. - const id = await brain.add({ - data: 'legacy-shaped node', - type: NounType.Concept, - metadata: { n: 1 } - }) - const storage = (brain as any).storage - const rec = await storage.getNoun(id) - const legacy = { - ...rec, - connections: new Map([[0, new Set(['00000000-0000-4000-8000-00000000aaaa'])]]), - level: 1 - } - await storage.saveNoun(legacy) - const before = await storage.getNoun(id) - expect(before.connections.size).toBeGreaterThan(0) - - // A data-changing update stages SaveNounOperation with placeholder - // adjacency — the legacy inline connections must survive the write. - await brain.update({ id, data: 'completely re-embedded text' }) - - const after = await storage.getNoun(id) - expect(after.connections.size).toBeGreaterThan(0) - expect(after.level).toBe(1) - }) -}) diff --git a/tests/integration/migration.test.ts b/tests/integration/migration.test.ts index f6c3741b..daa5fb2d 100644 --- a/tests/integration/migration.test.ts +++ b/tests/integration/migration.test.ts @@ -20,16 +20,6 @@ import { MigrationRunner, MIGRATIONS } from '../../src/migration/index.js' import type { Migration } from '../../src/migration/index.js' import { NounType, VerbType } from '../../src/types/graphTypes.js' -// THE VIEW CONTRACT (field-addressing law): transforms receive engine fields -// top-level and the USER's bag nested under `metadata` — user-field changes -// go inside the bag. These two helpers keep the one-liner migrations tidy. -const bagOf = (m: Record): Record => - m.metadata as Record -const withBag = ( - m: Record, - patch: Record -): Record => ({ ...m, metadata: { ...bagOf(m), ...patch } }) - // Helper to temporarily inject migrations into the MIGRATIONS array function withMigrations(migrations: Migration[], fn: () => Promise): Promise { const original = MIGRATIONS.splice(0, MIGRATIONS.length) @@ -88,11 +78,9 @@ describe('Migration System', () => { description: 'Add version field to entities with status', applies: 'nouns', transform: (m) => { - // Only transform entities that have our specific 'status' USER field - // (user fields live in the nested bag — the view contract). - const bag = m.metadata as Record - if ('status' in bag && !('version' in bag)) { - return { ...m, metadata: { ...bag, version: 1 } } + // Only transform entities that have our specific 'status' field + if ('status' in m && !('version' in m)) { + return { ...m, version: 1 } } return null } @@ -106,8 +94,7 @@ describe('Migration System', () => { // All 3 entities have 'status' metadata expect(p.affectedEntities).toBeGreaterThanOrEqual(3) expect(p.sampleChanges.length).toBeGreaterThan(0) - // Samples carry the VIEW shape: user fields inside `.metadata`. - expect(p.sampleChanges[0].after.metadata.version).toBe(1) + expect(p.sampleChanges[0].after.version).toBe(1) // Verify no data was modified (dry-run) const entity = await brain.get(id1) @@ -124,10 +111,9 @@ describe('Migration System', () => { description: 'Rename state to status', applies: 'nouns', transform: (m) => { - const bag = m.metadata as Record - if ('state' in bag) { - const { state, ...rest } = bag - return { ...m, metadata: { ...rest, status: state } } + if ('state' in m) { + const { state, ...rest } = m + return { ...rest, status: state } } return null } @@ -138,12 +124,11 @@ describe('Migration System', () => { const p = preview as any expect(p.sampleChanges.length).toBeGreaterThanOrEqual(1) - // Find the sample for our entity (it has the 'state' USER field — - // samples carry the VIEW shape, user fields inside `.metadata`) - const sample = p.sampleChanges.find((s: any) => s.before.metadata.state === 'draft') + // Find the sample for our entity (it has the 'state' field) + const sample = p.sampleChanges.find((s: any) => s.before.state === 'draft') expect(sample).toBeDefined() - expect(sample.after.metadata.status).toBe('draft') - expect(sample.after.metadata.state).toBeUndefined() + expect(sample.after.status).toBe('draft') + expect(sample.after.state).toBeUndefined() }) }) }) @@ -164,8 +149,8 @@ describe('Migration System', () => { description: 'Add migrated flag to entities with priority', applies: 'nouns', transform: (m) => { - if ('priority' in bagOf(m) && !('migrated' in bagOf(m))) { - return withBag(m, { migrated: true }) + if ('priority' in m && !('migrated' in m)) { + return { ...m, migrated: true } } return null } @@ -194,8 +179,8 @@ describe('Migration System', () => { description: 'Uppercase status field only when present', applies: 'nouns', transform: (m) => { - if (typeof bagOf(m).status === 'string') { - return withBag(m, { status: (bagOf(m).status as string).toUpperCase() }) + if (typeof m.status === 'string') { + return { ...m, status: (m.status as string).toUpperCase() } } return null } @@ -218,7 +203,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Double count', applies: 'nouns', - transform: (m) => typeof bagOf(m).count === 'number' ? withBag(m, { count: (bagOf(m).count as number) * 2 }) : null + transform: (m) => typeof m.count === 'number' ? { ...m, count: (m.count as number) * 2 } : null } const migration2: Migration = { @@ -226,7 +211,7 @@ describe('Migration System', () => { version: '1.1.0', description: 'Add 10 to count', applies: 'nouns', - transform: (m) => typeof bagOf(m).count === 'number' ? withBag(m, { count: (bagOf(m).count as number) + 10 }) : null + transform: (m) => typeof m.count === 'number' ? { ...m, count: (m.count as number) + 10 } : null } await withMigrations([migration1, migration2], async () => { @@ -244,7 +229,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Increment v', applies: 'nouns', - transform: (m) => typeof bagOf(m).v === 'number' ? withBag(m, { v: (bagOf(m).v as number) + 1 }) : null + transform: (m) => typeof m.v === 'number' ? { ...m, v: (m.v as number) + 1 } : null } await withMigrations([migration], async () => { @@ -281,7 +266,7 @@ describe('Migration System', () => { version: '2.0.0', description: 'Add y field to entities with x', applies: 'nouns', - transform: (m) => 'x' in bagOf(m) && !('y' in bagOf(m)) ? withBag(m, { y: 2 }) : null + transform: (m) => 'x' in m && !('y' in m) ? { ...m, y: 2 } : null } await withMigrations([migration], async () => { @@ -305,8 +290,8 @@ describe('Migration System', () => { description: 'Replace original with migrated', applies: 'nouns', transform: (m) => { - if (bagOf(m).original === true) { - return withBag(m, { original: false, migrated: true }) + if (m.original === true) { + return { ...m, original: false, migrated: true } } return null } @@ -338,7 +323,7 @@ describe('Migration System', () => { version: '4.0.0', description: 'Add field', applies: 'nouns', - transform: (m) => 'q' in bagOf(m) && !('r' in bagOf(m)) ? withBag(m, { r: 2 }) : null + transform: (m) => 'q' in m && !('r' in m) ? { ...m, r: 2 } : null } await withMigrations([migration], async () => { @@ -399,7 +384,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Auto migrate test', applies: 'nouns', - transform: (m) => 'legacy' in bagOf(m) ? withBag(m, { legacy: false, upgraded: true }) : null + transform: (m) => 'legacy' in m ? { ...m, legacy: false, upgraded: true } : null } await withMigrations([migration], async () => { @@ -425,7 +410,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Add y to entities with x', applies: 'nouns', - transform: (m) => 'x' in bagOf(m) ? withBag(m, { y: true }) : null + transform: (m) => 'x' in m ? { ...m, y: true } : null } const progressCalls: any[] = [] @@ -459,7 +444,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Increment v on entities that have it', applies: 'nouns', - transform: (m) => typeof bagOf(m).v === 'number' ? withBag(m, { v: (bagOf(m).v as number) + 1 }) : null + transform: (m) => typeof m.v === 'number' ? { ...m, v: (m.v as number) + 1 } : null } await withMigrations([migration], async () => { @@ -492,10 +477,9 @@ describe('Migration System', () => { description: 'Rename strength to intensity', applies: 'verbs', transform: (m) => { - const bag = bagOf(m) - if ('strength' in bag) { - const { strength, ...rest } = bag - return { ...m, metadata: { ...rest, intensity: strength } } + if ('strength' in m) { + const { strength, ...rest } = m + return { ...rest, intensity: strength } } return null } @@ -523,7 +507,7 @@ describe('Migration System', () => { version: '1.0.0', description: 'Update tag from old to new', applies: 'both', - transform: (m) => bagOf(m).tag === 'old' ? withBag(m, { tag: 'new' }) : null + transform: (m) => m.tag === 'old' ? { ...m, tag: 'new' } : null } await withMigrations([migration], async () => { @@ -593,11 +577,11 @@ describe('Migration System', () => { description: 'Transform that throws on non-number values', applies: 'nouns', transform: (m) => { - if ('value' in bagOf(m)) { - if (typeof bagOf(m).value !== 'number') { + if ('value' in m) { + if (typeof m.value !== 'number') { throw new Error('value must be a number') } - return withBag(m, { value: (bagOf(m).value as number) * 10 }) + return { ...m, value: (m.value as number) * 10 } } return null } @@ -631,7 +615,7 @@ describe('Migration System', () => { description: 'Always throws', applies: 'nouns', transform: (m) => { - if ('boom' in bagOf(m)) { + if ('boom' in m) { throw new Error('deliberate failure') } return null diff --git a/tests/integration/orderby-sort-bug.test.ts b/tests/integration/orderby-sort-bug.test.ts index aeb7ff66..9c28b1c9 100644 --- a/tests/integration/orderby-sort-bug.test.ts +++ b/tests/integration/orderby-sort-bug.test.ts @@ -56,7 +56,7 @@ describe('find({ orderBy }) sort bug regression', () => { const results = await brain.find({ type: NounType.Concept, - orderBy: 'system.createdAt', + orderBy: 'createdAt', order: 'desc', limit: 1 }) @@ -76,7 +76,7 @@ describe('find({ orderBy }) sort bug regression', () => { const results = await brain.find({ type: NounType.Concept, - orderBy: 'system.createdAt', + orderBy: 'createdAt', order: 'asc', limit: 1 }) @@ -94,7 +94,7 @@ describe('find({ orderBy }) sort bug regression', () => { const results = await brain.find({ type: NounType.Concept, - orderBy: 'system.createdAt', + orderBy: 'createdAt', order: 'desc' }) @@ -115,7 +115,7 @@ describe('find({ orderBy }) sort bug regression', () => { const results = await brain.find({ type: NounType.Concept, - orderBy: 'system.updatedAt', + orderBy: 'updatedAt', order: 'desc', limit: 1 }) @@ -136,7 +136,7 @@ describe('find({ orderBy }) sort bug regression', () => { const id3 = await brain.add({ data: 'third', type: NounType.Concept }) const results = await brain.find({ - orderBy: 'system.createdAt', + orderBy: 'createdAt', order: 'desc', limit: 2 }) @@ -215,6 +215,7 @@ describe('resolveEntityField helper', () => { 'id', 'vector', 'connections', + 'level', 'type', 'confidence', 'weight', @@ -227,9 +228,5 @@ describe('resolveEntityField helper', () => { for (const field of expected) { expect(STANDARD_ENTITY_FIELDS.has(field)).toBe(true) } - // `level` is deliberately NOT resolvable: it is HNSW plumbing, and listing - // it here shadowed user metadata named `level` in every by-name read - // (the reserved-name shadow bug). Plumbing stays out of the resolver. - expect(STANDARD_ENTITY_FIELDS.has('level')).toBe(false) }) }) diff --git a/tests/integration/update-write-granularity.test.ts b/tests/integration/update-write-granularity.test.ts deleted file mode 100644 index 234df334..00000000 --- a/tests/integration/update-write-granularity.test.ts +++ /dev/null @@ -1,126 +0,0 @@ -/** - * @module tests/integration/update-write-granularity - * @description Write-granularity law for update() (SELF-ENGINE-RESTART-GRIND, - * 2026-07-29): a metadata-only update must NEVER rewrite the noun record — - * the record carries the full vector, so an unconditional save turns every - * metadata touch into a whole-vector rewrite + fsync. Under a read-heavy - * consumer sweep bumping per-entity stats this amplified into disk saturation - * on a production deployment. Laws: - * (1) metadata-only update() → zero saveNoun calls (metadata leg only); - * (2) data/vector/type-changing update() → saveNoun runs (the vector leg and - * HNSW reindex still happen when the vector side actually changed); - * (3) the metadata-only path still lands: merged metadata readable, _rev - * bumped, find() by the new field sees the entity. - */ -import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest' -import { Brainy } from '../../src/brainy.js' -import { NounType } from '../../src/types/graphTypes.js' - -const stubEmbedding = async (text: string): Promise => { - const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0) - return new Array(384).fill(0).map((_, i) => Math.sin(hash + i)) -} - -describe('update() write granularity', () => { - let brain: Brainy - - beforeEach(async () => { - brain = new Brainy({ - requireSubtype: false, - storage: { type: 'memory' as const }, - embeddingFunction: stubEmbedding - }) - await brain.init() - }) - - afterEach(async () => { - await brain.close() - }) - - it('metadata-only update never rewrites the noun record (no vector rewrite)', async () => { - const id = await brain.add({ - data: 'granularity law subject', - type: NounType.Concept, - metadata: { touched: 0 } - }) - - const storage = (brain as any).storage - const saveNounSpy = vi.spyOn(storage, 'saveNoun') - - await brain.update({ id, metadata: { touched: 1 } }) - - expect(saveNounSpy).not.toHaveBeenCalled() - saveNounSpy.mockRestore() - - // The metadata leg still landed with full semantics. - const after = await brain.get(id, { includeVectors: true }) - expect(after?.metadata?.touched).toBe(1) - expect(after?._rev).toBe(2) - expect(Array.isArray(after?.vector) && after!.vector!.length).toBe(384) - - const found = await brain.find({ where: { touched: 1 } }) - expect(found.some((r: any) => r.id === id)).toBe(true) - }) - - it('confidence/weight/subtype-only updates also skip the noun record', async () => { - const id = await brain.add({ - data: 'reserved-field touch subject', - type: NounType.Concept, - metadata: {} - }) - - const storage = (brain as any).storage - const saveNounSpy = vi.spyOn(storage, 'saveNoun') - - await brain.update({ id, confidence: 0.5, weight: 2, subtype: 'note' }) - - expect(saveNounSpy).not.toHaveBeenCalled() - saveNounSpy.mockRestore() - - const after = await brain.get(id) - expect(after?.confidence).toBe(0.5) - expect(after?.subtype).toBe('note') - }) - - it('data-changing update still writes the noun record and reindexes', async () => { - const id = await brain.add({ - data: 'original embedded text', - type: NounType.Concept, - metadata: {} - }) - - const before = await brain.get(id, { includeVectors: true }) - - const storage = (brain as any).storage - const saveNounSpy = vi.spyOn(storage, 'saveNoun') - - await brain.update({ id, data: 'completely different embedded text' }) - - expect(saveNounSpy).toHaveBeenCalled() - saveNounSpy.mockRestore() - - const after = await brain.get(id, { includeVectors: true }) - expect(after?.data).toBe('completely different embedded text') - expect(after?.vector).not.toEqual(before?.vector) - }) - - it('explicit-vector update still writes the noun record', async () => { - const id = await brain.add({ - data: 'vector swap subject', - type: NounType.Concept, - metadata: {} - }) - - const storage = (brain as any).storage - const saveNounSpy = vi.spyOn(storage, 'saveNoun') - - const newVector = new Array(384).fill(0).map((_, i) => Math.cos(i)) - await brain.update({ id, vector: newVector }) - - expect(saveNounSpy).toHaveBeenCalled() - saveNounSpy.mockRestore() - - const after = await brain.get(id, { includeVectors: true }) - expect(after?.vector?.[0]).toBeCloseTo(1) // cos(0) - }) -}) diff --git a/tests/regression/metadata-index-cleanup.unit.test.ts b/tests/regression/metadata-index-cleanup.unit.test.ts index 3746e833..0984d727 100644 --- a/tests/regression/metadata-index-cleanup.unit.test.ts +++ b/tests/regression/metadata-index-cleanup.unit.test.ts @@ -244,10 +244,7 @@ describe('Metadata index cleanup after remove / removeMany', () => { const noConfidenceId = await addEntity({ type: 'thing' }) const withConfidenceId = await addEntity({ type: 'thing', confidence: 0.9 }) - // system.confidence — confidence is an engine scalar (an add() param), - // never a metadata field; bare 'confidence' now addresses the user's - // own metadata bag under the sealed field-addressing law. - const results = await brain.find({ where: { 'system.confidence': { exists: true } } }) + const results = await brain.find({ where: { confidence: { exists: true } } }) const ids = results.map(r => r.id) expect(ids).toContain(withConfidenceId) @@ -258,8 +255,7 @@ describe('Metadata index cleanup after remove / removeMany', () => { const noWeightId = await addEntity({ type: 'thing' }) const withWeightId = await addEntity({ type: 'thing', weight: 0.5 }) - // system.weight — same reasoning as system.confidence above. - const results = await brain.find({ where: { 'system.weight': { exists: true } } }) + const results = await brain.find({ where: { weight: { exists: true } } }) const ids = results.map(r => r.id) expect(ids).toContain(withWeightId) @@ -273,12 +269,11 @@ describe('Metadata index cleanup after remove / removeMany', () => { const id = await addEntity({ type: 'thing' }) await brain.remove(id) - // Entity must not appear in any confidence query. system.confidence — - // same addressing as the two tests above. - const existsTrue = await brain.find({ where: { 'system.confidence': { exists: true } } }) + // Entity must not appear in any confidence query + const existsTrue = await brain.find({ where: { confidence: { exists: true } } }) expect(existsTrue.map(r => r.id)).not.toContain(id) - const existsFalse = await brain.find({ where: { 'system.confidence': { exists: false } } }) + const existsFalse = await brain.find({ where: { confidence: { exists: false } } }) expect(existsFalse.map(r => r.id)).not.toContain(id) }) }) diff --git a/tests/unit/brainy/find-orderby-pagek.test.ts b/tests/unit/brainy/find-orderby-pagek.test.ts index 49fccb02..9a453f8d 100644 --- a/tests/unit/brainy/find-orderby-pagek.test.ts +++ b/tests/unit/brainy/find-orderby-pagek.test.ts @@ -42,8 +42,7 @@ describe('find({ where, orderBy }) bounds the sort to the page (CTX-BR-FIND-ORDE return real(f, ob, o, topK) } - // system.createdAt — entity age, not a user metadata field named 'createdAt'. - const results = await brain.find({ where: { bucket: 'x' }, orderBy: 'system.createdAt', order: 'desc', limit: 5 }) + const results = await brain.find({ where: { bucket: 'x' }, orderBy: 'createdAt', order: 'desc', limit: 5 }) expect(results).toHaveLength(5) // Page-bounded: ~ limit (5) + a small hidden-tier over-fetch — NOT all 50 matches. diff --git a/tests/unit/brainy/migration-deference.test.ts b/tests/unit/brainy/migration-deference.test.ts index 5968c620..6471b4ef 100644 --- a/tests/unit/brainy/migration-deference.test.ts +++ b/tests/unit/brainy/migration-deference.test.ts @@ -245,10 +245,7 @@ describe('rc.8 no-freeze migration deference (isMigrating / stampBrainFormat / b it('the brain-format marker module exports the compiled epoch + data-format constants', () => { // cor imports these from '@soulcraft/brainy/brain-format' (Hook 3) so both // sides share ONE source of truth — no duplicated constant to drift. - // Epoch 3: the namespace-law key split (bare user keys · literal - // 'system.' scalars, 2026-08-03) — every brain rebuilds onto the - // frozen keys at first open. (Epoch 2 same day: `level` indexability.) - expect(EXPECTED_INDEX_EPOCH).toBe(3) + expect(EXPECTED_INDEX_EPOCH).toBe(1) expect(CURRENT_DATA_FORMAT).toBe('8.0') }) }) diff --git a/tests/unit/brainy/reserved-field-policy.test.ts b/tests/unit/brainy/reserved-field-policy.test.ts new file mode 100644 index 00000000..c5f37af4 --- /dev/null +++ b/tests/unit/brainy/reserved-field-policy.test.ts @@ -0,0 +1,251 @@ +/** + * @module tests/unit/brainy/reserved-field-policy + * @description The 8.0 `reservedFieldPolicy` matrix — what happens when an + * untyped (JavaScript) caller smuggles a Brainy-reserved field INSIDE the + * `metadata` bag of a write call, past the compile-time guard. + * + * 8.0 is a clean break with no silent failures. The decided contract: + * - `'throw'` (DEFAULT): a reserved key in the bag throws a clear Error naming + * the offending key(s) and the correct write path. No remap, no data loss. + * - `'warn'`: legacy remap PLUS a one-shot (per method+field, per process) + * warning for EVERY reserved key found. + * - `'remap'`: the pre-8.0 silent remap, no warning. + * + * The deep correctness of the remap itself (top-level precedence, system-managed + * drops, transact()/with() mirrors, read-side splitting) lives in + * tests/unit/brainy/update-reserved-metadata-remap.test.ts (which now runs under + * `reservedFieldPolicy: 'remap'`). This file pins the POLICY SELECTION and the + * throw/warn behaviors. + * + * Compile-time callers can't write these shapes at all (see + * tests/unit/types/reserved-metadata-keys.test-d.ts); the `as object` widenings + * below simulate untyped callers. + */ + +import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest' +import { Brainy } from '../../../src/index.js' +import { NounType, VerbType } from '../../../src/types/graphTypes.js' +import { createTestConfig } from '../../helpers/test-factory.js' +import { prodLog } from '../../../src/utils/logger.js' + +describe('reservedFieldPolicy', () => { + describe("default policy is 'throw'", () => { + let brain: Brainy + + beforeEach(async () => { + // No reservedFieldPolicy override → resolves to 'throw'. + brain = new Brainy(createTestConfig()) + await brain.init() + }) + + afterEach(async () => { + await brain.close() + }) + + it('add() throws naming the offending key and the correct write path', async () => { + await expect( + brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'x', + metadata: { confidence: 0.8 } as object + }) + ).rejects.toThrow(/metadata\.confidence is a reserved field/) + + // The error names the right param and the reserved list for discoverability. + await expect( + brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'x', + metadata: { confidence: 0.8 } as object + }) + ).rejects.toThrow(/'confidence' param.*RESERVED_ENTITY_FIELDS/s) + }) + + it('add() lists EVERY offending key when several are present', async () => { + const err = await brain + .add({ + type: NounType.Person, + data: 'multi', + metadata: { confidence: 0.5, weight: 0.6, subtype: 'employee' } as object + }) + .catch((e) => e as Error) + expect(err).toBeInstanceOf(Error) + expect(err.message).toMatch(/confidence/) + expect(err.message).toMatch(/weight/) + expect(err.message).toMatch(/subtype/) + }) + + it('update() throws on a reserved key in the patch', async () => { + const id = await brain.add({ type: NounType.Concept, subtype: 'general', data: 'y' }) + await expect( + brain.update({ id, metadata: { confidence: 0.3 } as object }) + ).rejects.toThrow(/metadata\.confidence is a reserved field/) + }) + + it('relate() throws on a reserved key in the bag', async () => { + const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' }) + const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' }) + await expect( + brain.relate({ + from: a, + to: b, + type: VerbType.RelatedTo, + subtype: 'colleague', + metadata: { confidence: 0.4 } as object + }) + ).rejects.toThrow(/metadata\.confidence is a reserved field.*RESERVED_RELATION_FIELDS/s) + }) + + it('updateRelation() throws on a reserved key in the patch', async () => { + const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' }) + const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' }) + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.ReportsTo, + subtype: 'direct' + }) + await expect( + brain.updateRelation({ id: relId, metadata: { weight: 0.2 } as object }) + ).rejects.toThrow(/metadata\.weight is a reserved field/) + }) + + it('transact() add op throws on a reserved key in the bag', async () => { + await expect( + brain.transact([ + { + op: 'add', + type: NounType.Concept, + subtype: 'general', + data: 'tx', + metadata: { confidence: 0.7 } as object + } + ]) + ).rejects.toThrow(/metadata\.confidence is a reserved field/) + }) + + it('a custom (non-reserved) key in the bag does NOT throw', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'ok', + metadata: { status: 'draft', rating: 4 } + }) + const entity = await brain.get(id) + expect(entity?.metadata).toEqual({ status: 'draft', rating: 4 }) + }) + }) + + describe("'remap' policy remaps silently (no warning)", () => { + let brain: Brainy + let warnSpy: ReturnType + + beforeEach(async () => { + warnSpy = vi.spyOn(prodLog, 'warn').mockImplementation(() => {}) + brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' })) + await brain.init() + }) + + afterEach(async () => { + await brain.close() + warnSpy.mockRestore() + }) + + it('lifts user-mutable reserved fields to top-level without warning', async () => { + const id = await brain.add({ + type: NounType.Person, + data: 'remap lift', + metadata: { confidence: 0.8, weight: 0.6, subtype: 'employee', dept: 'eng' } as object + }) + const entity = await brain.get(id) + expect(entity?.confidence).toBe(0.8) + expect(entity?.weight).toBe(0.6) + expect(entity?.subtype).toBe('employee') + expect(entity?.metadata).toEqual({ dept: 'eng' }) + // 'remap' is silent about reserved fields (unrelated storage logs may fire, + // so assert specifically that no reserved-field warning was emitted). + const reservedWarned = warnSpy.mock.calls.some((c) => + String(c[0]).includes('reserved field') + ) + expect(reservedWarned).toBe(false) + }) + + it('preserves _originalId on natural-key ids through the remap path', async () => { + // A speculative view applies the same normalization and maps a natural-key + // id to a stable UUID, preserving the caller's original string. + const base = await brain.now() + const speculative = await base.with([ + { + op: 'add', + id: 'remap-spec-entity', + type: NounType.Concept, + subtype: 'general', + data: 'spec', + metadata: { confidence: 0.65, custom: 'spec' } as object + } + ]) + const entity = await speculative.get('remap-spec-entity') + expect(entity?.confidence).toBe(0.65) + expect(entity?.metadata).toEqual({ custom: 'spec', _originalId: 'remap-spec-entity' }) + await speculative.release() + await base.release() + }) + }) + + describe("'warn' policy remaps AND warns once per key", () => { + let brain: Brainy + let warnSpy: ReturnType + + beforeEach(async () => { + warnSpy = vi.spyOn(prodLog, 'warn').mockImplementation(() => {}) + brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'warn' })) + await brain.init() + }) + + afterEach(async () => { + await brain.close() + warnSpy.mockRestore() + }) + + it('remaps the value (same as remap) and emits a warning naming the field', async () => { + // Use a method+field combo unique to this test so the per-process one-shot + // registry has not already consumed it. + const id = await brain.add({ + type: NounType.Person, + data: 'warn lift', + // weight is user-mutable → remapped; this is the only 'warn'-policy + // add({ weight }) in the suite, so the one-shot warning fires here. + metadata: { weight: 0.42, dept: 'eng' } as object + }) + const entity = await brain.get(id) + // Value is honored (remap still happens under 'warn'). + expect(entity?.weight).toBe(0.42) + expect(entity?.metadata).toEqual({ dept: 'eng' }) + // And a warning was emitted naming the reserved field. + expect(warnSpy).toHaveBeenCalled() + const warned = warnSpy.mock.calls.some((c) => + String(c[0]).includes("'weight'") + ) + expect(warned).toBe(true) + }) + + it('warns for system-managed keys too (closes the historical gap)', async () => { + // Pre-8.0 only system-managed fields warned; 'warn' warns for every key. + // 'createdBy' (system-managed on update) is unique to this test. + const id = await brain.add({ type: NounType.Concept, subtype: 'general', data: 'sys' }) + warnSpy.mockClear() + await brain.update({ id, metadata: { createdBy: 'nope', keep: 'me' } as object }) + const entity = await brain.get(id) + // System-managed key dropped; custom field merged. + expect((entity?.metadata as Record)?.createdBy).toBeUndefined() + expect((entity?.metadata as Record)?.keep).toBe('me') + // A warning was emitted for the dropped system-managed key. + const warned = warnSpy.mock.calls.some((c) => + String(c[0]).includes("'createdBy'") + ) + expect(warned).toBe(true) + }) + }) +}) diff --git a/tests/unit/brainy/update-reserved-metadata-remap.test.ts b/tests/unit/brainy/update-reserved-metadata-remap.test.ts new file mode 100644 index 00000000..31713f99 --- /dev/null +++ b/tests/unit/brainy/update-reserved-metadata-remap.test.ts @@ -0,0 +1,403 @@ +/** + * @module tests/unit/brainy/update-reserved-metadata-remap + * @description Regression tests for the reserved-field metadata-bag trap, + * ported from the 7.x fix and extended to the full 8.0 contract. + * + * History: `add({metadata: {confidence}})` lifted reserved fields to their + * canonical top-level location, but `update({metadata: {confidence}})` + * silently dropped the same shape — the patch value survived the merge and + * was then clobbered by the preserve-existing spread. A production + * consumer's confidence-evolution writes no-oped for weeks before being + * caught by reading values back. + * + * These tests pin the LEGACY REMAP behavior, which in 8.0 is opt-in via + * `reservedFieldPolicy: 'remap'` (the default is `'throw'` — see the policy + * matrix in tests/unit/brainy/reserved-field-policy.test.ts). The brain in + * every test below is constructed with `reservedFieldPolicy: 'remap'` so these + * deep correctness assertions about the remap path stay exercised. + * + * Remap contract under test (every write path, entities AND relationships): + * - user-mutable reserved fields (`confidence`, `weight`, `subtype` — plus + * `service`/`createdBy` at add()/relate() time) remap from the metadata + * bag to their dedicated top-level param, with top-level winning when both + * are present; + * - system-managed reserved fields (`createdAt`, `_rev`, `noun`/`verb`, + * `data`, …) are dropped from the bag; + * - the same normalization applies to `transact()` operations and `with()` + * speculative views; + * - reads NEVER echo a reserved field inside `metadata`. + * + * TypeScript callers can't write these shapes at all (compile-time guard on + * the metadata param types — see tests/unit/types/reserved-metadata-keys.test-d.ts); + * these tests simulate untyped (JavaScript) callers, hence the `as object` + * widenings on the metadata literals. + */ + +import { describe, it, expect, beforeEach, afterEach } from 'vitest' +import { Brainy } from '../../../src/index.js' +import { NounType, VerbType } from '../../../src/types/graphTypes.js' +import { createTestConfig } from '../../helpers/test-factory.js' + +describe('reserved-field metadata remap (8.0 legacy remap path)', () => { + let brain: Brainy + + beforeEach(async () => { + // The remap path is opt-in in 8.0 (default policy is 'throw'). + brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' })) + await brain.init() + }) + + afterEach(async () => { + await brain.close() + }) + + describe('update() — the ported 7.x regression', () => { + it('remaps metadata.confidence to the top-level field (the production repro)', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'x', + metadata: { confidence: 0.8 } as object + }) + + // Top-level write works (always did) + await brain.update({ id, confidence: 0.42 }) + let entity = await brain.get(id) + expect(entity?.confidence).toBe(0.42) + + // Metadata-patch write — silently dropped pre-fix, remapped now + await brain.update({ id, metadata: { confidence: 0.33 } as object }) + entity = await brain.get(id) + expect(entity?.confidence).toBe(0.33) + // The reserved key must not linger inside the metadata bag + expect((entity?.metadata as Record)?.confidence).toBeUndefined() + }) + + it('remaps metadata.weight and metadata.subtype the same way', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'y', + metadata: {} + }) + + await brain.update({ id, metadata: { weight: 0.7, subtype: 'specialized' } as object }) + const entity = await brain.get(id) + expect(entity?.weight).toBe(0.7) + expect(entity?.subtype).toBe('specialized') + expect((entity?.metadata as Record)?.weight).toBeUndefined() + expect((entity?.metadata as Record)?.subtype).toBeUndefined() + }) + + it('top-level param wins when both top-level and metadata-patch carry the field', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'z', + metadata: { confidence: 0.5 } as object + }) + + await brain.update({ id, confidence: 0.9, metadata: { confidence: 0.1 } as object }) + const entity = await brain.get(id) + expect(entity?.confidence).toBe(0.9) + }) + + it('drops system-managed fields from patches without corrupting the entity', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'w', + metadata: { keep: 'me' } + }) + const before = await brain.get(id) + + await brain.update({ + id, + metadata: { createdAt: 1, _rev: 999, noun: 'organization', other: 'applied' } as object + }) + const after = await brain.get(id) + + expect(after?.createdAt).toBe(before?.createdAt) // immutable + expect(after?.type).toBe('concept') // noun patch ignored + expect(after?._rev).toBe((before?._rev ?? 1) + 1) // _rev patch ignored; normal bump applied + expect((after?.metadata as Record)?.other).toBe('applied') // custom fields still merge + expect((after?.metadata as Record)?.keep).toBe('me') + expect((after?.metadata as Record)?._rev).toBeUndefined() + expect((after?.metadata as Record)?.createdAt).toBeUndefined() + expect((after?.metadata as Record)?.noun).toBeUndefined() + }) + + it('custom (non-reserved) metadata patches are unaffected by the remap', async () => { + const id = await brain.add({ + type: NounType.Concept, + subtype: 'general', + data: 'v', + metadata: { status: 'draft' } + }) + + await brain.update({ id, metadata: { status: 'reviewed', rating: 4.5 } }) + const entity = await brain.get(id) + expect((entity?.metadata as Record)?.status).toBe('reviewed') + expect((entity?.metadata as Record)?.rating).toBe(4.5) + }) + }) + + describe('add() — explicit lift, identical contract', () => { + it('lifts confidence/weight/subtype out of the bag to top level', async () => { + const id = await brain.add({ + type: NounType.Person, + data: 'lift check', + metadata: { confidence: 0.8, weight: 0.6, subtype: 'employee', dept: 'eng' } as object + }) + + const entity = await brain.get(id) + expect(entity?.confidence).toBe(0.8) + expect(entity?.weight).toBe(0.6) + expect(entity?.subtype).toBe('employee') + expect(entity?.metadata).toEqual({ dept: 'eng' }) + }) + + it('lifts service (settable at add time) and lets the top-level param win', async () => { + const lifted = await brain.add({ + type: NounType.Person, + subtype: 'employee', + data: 'service lift', + metadata: { service: 'orders' } as object + }) + expect((await brain.get(lifted))?.service).toBe('orders') + + const topLevelWins = await brain.add({ + type: NounType.Person, + subtype: 'employee', + data: 'service precedence', + service: 'billing', + metadata: { service: 'orders' } as object + }) + const entity = await brain.get(topLevelWins) + expect(entity?.service).toBe('billing') + expect((entity?.metadata as Record)?.service).toBeUndefined() + }) + + it('a remapped subtype satisfies subtype enforcement like a top-level one', async () => { + brain.requireSubtype(NounType.Document) + + // Top-level missing, but the bag carries it — must not throw. + const id = await brain.add({ + type: NounType.Document, + data: 'enforcement via remap', + metadata: { subtype: 'invoice' } as object + }) + expect((await brain.get(id))?.subtype).toBe('invoice') + + // Neither place carries it — must throw. + await expect( + brain.add({ type: NounType.Document, data: 'no subtype anywhere' }) + ).rejects.toThrow(/subtype/) + }) + }) + + describe('transact() — same remap on add and update ops', () => { + it('normalizes reserved fields in transact add + update ops', async () => { + const db1 = await brain.transact([ + { + op: 'add', + type: NounType.Concept, + subtype: 'general', + data: 'tx', + metadata: { confidence: 0.7, custom: 'a' } as object + } + ]) + const id = db1.receipt!.ids[0] + + let entity = await brain.get(id) + expect(entity?.confidence).toBe(0.7) + expect(entity?.metadata).toEqual({ custom: 'a' }) + + await brain.transact([ + { op: 'update', id, metadata: { confidence: 0.25, custom: 'b' } as object } + ]) + entity = await brain.get(id) + expect(entity?.confidence).toBe(0.25) + expect(entity?.metadata).toEqual({ custom: 'b' }) + expect((entity?.metadata as Record)?.confidence).toBeUndefined() + }) + + it('historical asOf() reads surface reserved fields ONLY top-level', async () => { + const db1 = await brain.transact([ + { + op: 'add', + type: NounType.Concept, + subtype: 'general', + data: 'historical', + metadata: { confidence: 0.9, custom: 'past' } as object + } + ]) + const id = db1.receipt!.ids[0] + + // Move the world forward so generation db1 is historical. + await brain.transact([{ op: 'update', id, confidence: 0.1, metadata: { custom: 'now' } }]) + + const past = await brain.asOf(db1.generation) + const historical = await past.get(id) + expect(historical?.confidence).toBe(0.9) + expect(historical?.metadata).toEqual({ custom: 'past' }) + await past.release() + }) + + it('with() speculative views apply the same normalization', async () => { + const base = await brain.now() + const speculative = await base.with([ + { + op: 'add', + id: 'spec-entity', + type: NounType.Concept, + subtype: 'general', + data: 'spec', + metadata: { confidence: 0.65, custom: 'spec' } as object + } + ]) + + const entity = await speculative.get('spec-entity') + expect(entity?.confidence).toBe(0.65) + // 8.0 id normalization: a natural-key id is mapped to a stable UUID and + // the caller's original string is preserved under _originalId — surfaced + // here exactly as the durable transact()/add() paths do. + expect(entity?.metadata).toEqual({ custom: 'spec', _originalId: 'spec-entity' }) + await speculative.release() + await base.release() + }) + }) + + describe('read paths never echo reserved fields inside metadata', () => { + it('find() (storage pagination path) returns custom-only metadata with reserved fields top-level', async () => { + const id = await brain.add({ + type: NounType.Person, + subtype: 'employee', + data: 'pagination echo check', + confidence: 0.8, + weight: 0.6, + metadata: { dept: 'eng' } + }) + + // No query/filter → served by the direct storage pagination path + // (getNounsWithPagination), which historically echoed the full flat + // record (noun/subtype/createdAt/… inside metadata). + const results = await brain.find({ limit: 50 }) + const result = results.find((r) => r.id === id) + expect(result).toBeDefined() + expect(result?.entity.metadata).toEqual({ dept: 'eng' }) + expect(result?.entity.type).toBe(NounType.Person) + expect(result?.entity.subtype).toBe('employee') + expect(result?.entity.confidence).toBe(0.8) + expect(result?.entity.weight).toBe(0.6) + expect(typeof result?.entity.createdAt).toBe('number') + expect(result?.entity._rev).toBe(1) + }) + + it('related() by target surfaces reserved fields top-level, custom-only metadata', async () => { + const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'src' }) + const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'tgt' }) + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.ReportsTo, + subtype: 'direct', + confidence: 0.9, + weight: 0.5, + service: 'orders', + metadata: { note: 'target path' } + }) + + const relations = await brain.related({ to: b }) + const rel = relations.find((r) => r.id === relId) + expect(rel).toBeDefined() + expect(rel?.metadata).toEqual({ note: 'target path' }) + expect(rel?.subtype).toBe('direct') + expect(rel?.confidence).toBe(0.9) + expect(rel?.weight).toBe(0.5) + expect(rel?.service).toBe('orders') + expect(typeof rel?.createdAt).toBe('number') + }) + }) + + describe('relationships — relate() / updateRelation() mirror', () => { + let a: string + let b: string + + beforeEach(async () => { + a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' }) + b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' }) + }) + + it('relate() persists the top-level confidence and service params', async () => { + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.ReportsTo, + subtype: 'direct', + confidence: 0.77, + service: 'orders' + }) + + const relations = await brain.related({ from: a }) + const rel = relations.find((r) => r.id === relId) + expect(rel?.confidence).toBe(0.77) + expect(rel?.service).toBe('orders') + }) + + it('relate() remaps reserved fields out of the metadata bag', async () => { + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.RelatedTo, + subtype: 'colleague', + metadata: { confidence: 0.4, weight: 0.3, role: 'peer' } as object + }) + + const relations = await brain.related({ from: a }) + const rel = relations.find((r) => r.id === relId) + expect(rel?.confidence).toBe(0.4) + expect(rel?.weight).toBe(0.3) + expect(rel?.metadata).toEqual({ role: 'peer' }) + }) + + it('relation.metadata never echoes the verb type key', async () => { + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.RelatedTo, + subtype: 'colleague', + metadata: { note: 'no echo' } + }) + + const relations = await brain.related({ from: a }) + const rel = relations.find((r) => r.id === relId) + expect(rel?.type).toBe(VerbType.RelatedTo) + expect((rel?.metadata as Record)?.verb).toBeUndefined() + expect(rel?.metadata).toEqual({ note: 'no echo' }) + }) + + it('updateRelation() remaps the user-mutable trio and preserves service', async () => { + const relId = await brain.relate({ + from: a, + to: b, + type: VerbType.ReportsTo, + subtype: 'direct', + service: 'orders', + metadata: { keep: 'me' } + }) + + await brain.updateRelation({ + id: relId, + metadata: { confidence: 0.55, subtype: 'dotted-line', extra: 'applied' } as object + }) + + const relations = await brain.related({ from: a }) + const rel = relations.find((r) => r.id === relId) + expect(rel?.confidence).toBe(0.55) + expect(rel?.subtype).toBe('dotted-line') + expect(rel?.service).toBe('orders') // fixed at relate() time, never erased by updates + expect(rel?.metadata).toEqual({ keep: 'me', extra: 'applied' }) + }) + }) +}) diff --git a/tests/unit/brainy/visibility.test.ts b/tests/unit/brainy/visibility.test.ts index dd4540d7..a5a02422 100644 --- a/tests/unit/brainy/visibility.test.ts +++ b/tests/unit/brainy/visibility.test.ts @@ -198,47 +198,60 @@ describe('visibility (8.0 reserved field)', () => { expect(entity?.visibility).toBeUndefined() }) - it('metadata.visibility is the USER’s field (field-addressing law) — stored verbatim, never lifted to the engine tier', async () => { - const id = await brain.add({ - type: NounType.Concept, - data: 'y', - metadata: { visibility: 'internal', tag: 't' } as object - }) - const entity = await brain.get(id) - // The user's field lives in the bag, verbatim… - expect((entity?.metadata as Record)?.visibility).toBe('internal') - expect((entity?.metadata as Record)?.tag).toBe('t') - // …and the ENGINE tier is untouched: absent === public, so the entity - // stays visible on default reads (the engine tier is set only via the - // dedicated visibility param and reads at system.visibility). - expect(entity?.visibility).toBeUndefined() - const visible = await brain.find({ type: NounType.Concept, limit: 20 }) - expect(visible.map((r) => r.id)).toContain(id) + it('an untyped caller passing visibility inside metadata is normalized under reservedFieldPolicy:"remap" (lifted to top-level)', async () => { + // Simulate a JavaScript caller smuggling the reserved key past the compile-time guard. + // The legacy remap behavior is now opt-in (8.0 default is 'throw'). + const remapBrain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' })) + await remapBrain.init() + try { + const id = await remapBrain.add({ + type: NounType.Concept, + data: 'y', + metadata: { visibility: 'internal', tag: 't' } as object + }) + const entity = await remapBrain.get(id) + // Lifted to the top-level field… + expect(entity?.visibility).toBe('internal') + // …and stripped from the metadata bag. + expect((entity?.metadata as Record)?.visibility).toBeUndefined() + expect((entity?.metadata as Record)?.tag).toBe('t') + // It is excluded from the default count, exactly like a top-level internal write. + expect(await remapBrain.getNounCount()).toBe(0) + } finally { + await remapBrain.close() + } }) - it('a user field valued "system" cannot smuggle the Brainy-only tier — it is just user data', async () => { - const id = await brain.add({ - type: NounType.Concept, - data: 'z', - metadata: { visibility: 'system' } as object - }) - const entity = await brain.get(id) - // Engine tier unaffected → entity stays public (counted, visible); - // the string 'system' is ordinary user data in the bag. - expect(entity?.visibility).toBeUndefined() - expect((entity?.metadata as Record)?.visibility).toBe('system') - const found = await brain.find({ type: NounType.Concept, limit: 10 }) - expect(found.map((r) => r.id)).toContain(id) + it('a "system" value smuggled through metadata is dropped under reservedFieldPolicy:"remap", not honored', async () => { + // 'system' is Brainy-only; an untyped caller must not be able to set it. + const remapBrain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' })) + await remapBrain.init() + try { + const id = await remapBrain.add({ + type: NounType.Concept, + data: 'z', + metadata: { visibility: 'system' } as object + }) + const entity = await remapBrain.get(id) + // The smuggled 'system' was dropped → entity stays public (counted, visible). + expect(entity?.visibility).toBeUndefined() + expect(await remapBrain.getNounCount()).toBe(1) + const found = await remapBrain.find({ type: NounType.Concept, limit: 10 }) + expect(found.map((r) => r.id)).toContain(id) + } finally { + await remapBrain.close() + } }) - it('a forged system.visibility key in metadata refuses loudly at the write door', async () => { + it('an untyped caller passing visibility inside metadata throws under the default policy', async () => { + // 8.0 default: no silent remap — a reserved key in the bag is a loud error. await expect( brain.add({ type: NounType.Concept, data: 'throws', - metadata: { 'system.visibility': 'internal' } as object + metadata: { visibility: 'internal', tag: 't' } as object }) - ).rejects.toThrow(/system\./) + ).rejects.toThrow(/visibility.*reserved field/) }) }) }) diff --git a/tests/unit/db/fact-log.test.ts b/tests/unit/db/fact-log.test.ts index f1c226cc..abce2dc9 100644 --- a/tests/unit/db/fact-log.test.ts +++ b/tests/unit/db/fact-log.test.ts @@ -186,52 +186,4 @@ describe('fact log — round-trip, framing, reconcile, rotation, scan', () => { await log.sync() expect(log.segmentPaths()).toEqual([]) // only a tail exists — nothing sealed }) - - describe('scanFacts liveness contract (Stage-2 D1)', () => { - it('a wedged store fails LOUDLY within the first-batch bound — never a silent hang', async () => { - // Force a sealed segment (tiny rotateBytes) so the scan must READ from - // storage, then wedge that read: the exact production shape (a - // backlogged brain whose segment read never returned). - const mem: any = new MemoryStorage() - await mem.init() - const wedgeable = new FactLog(mem, { rotateBytes: 1 }) - await wedgeable.open(0) - await wedgeable.append(fact(1)) - await wedgeable.append(fact(2)) // second append rotates → seg 1 sealed - await wedgeable.sync() - - const realRead = mem.readRawBytes.bind(mem) - mem.readRawBytes = (p: string) => - p.includes('facts/seg-') ? new Promise(() => {}) : realRead(p) // hangs forever - - const scan = wedgeable.scanFacts({ firstBatchTimeoutMs: 200 }) - const started = Date.now() - await expect(scan.batches().next()).rejects.toThrow(/no first batch within 200ms/) - expect(Date.now() - started).toBeLessThan(5_000) // bound held, not a hang - }) - - it('a healthy scan is unaffected — first batch well inside the bound, all facts delivered', async () => { - for (let g = 1; g <= 5; g++) await log.append(fact(g)) - await log.sync() - const scan = log.scanFacts({ batchSize: 2 }) - const all: CommitFact[] = [] - for await (const b of scan.batches()) all.push(...b.facts) - expect(all.map((f) => f.generation)).toEqual([1, 2, 3, 4, 5]) - expect(scan.summary().factsYielded).toBe(5) - }) - - it('consumer think-time between pulls never counts against the producer', async () => { - for (let g = 1; g <= 4; g++) await log.append(fact(g)) - await log.sync() - // Bound tighter than the consumer's pause: only the FIRST pull is - // raced, so a slow consumer after batch 1 must not trip the deadline. - const gen = log.scanFacts({ batchSize: 2, firstBatchTimeoutMs: 150 }).batches() - const first = await gen.next() - expect(first.done).toBe(false) - await new Promise((r) => setTimeout(r, 400)) // dawdle past the bound - const second = await gen.next() - expect(second.done).toBe(false) - expect((await gen.next()).done).toBe(true) - }) - }) }) diff --git a/tests/unit/db/fieldAddressing.test.ts b/tests/unit/db/fieldAddressing.test.ts deleted file mode 100644 index 04110992..00000000 --- a/tests/unit/db/fieldAddressing.test.ts +++ /dev/null @@ -1,145 +0,0 @@ -/** - * @module tests/unit/db/fieldAddressing - * @description Unit pins for the one field-addressing law (ruled 2026-08-03). - * These pin the PURE half of the law — parsing, the ruled maps, plumbing - * invisibility, refusal text — including the RELATION map, which cannot be - * pinned through the public query API today (related() carries no - * field-addressing options): the verb mirror is contract-tested here at the - * module level so the two engines cannot drift on it. - */ -import { describe, it, expect } from 'vitest' -import { - SYSTEM_ENTITY_SCALARS, - SYSTEM_RELATION_SCALARS, - PLUMBING_FIELDS, - parseFieldAddress, - buildUnresolvableMessage, - InvalidFieldAddressError -} from '../../../src/db/fieldAddressing.js' - -describe('field-addressing law — pure module pins', () => { - it('the entity system map is EXACTLY the ruled ten scalars', () => { - expect([...SYSTEM_ENTITY_SCALARS].sort()).toEqual( - [ - 'confidence', - 'createdAt', - 'createdBy', - 'id', - 'service', - 'subtype', - 'type', - 'updatedAt', - 'visibility', - 'weight' - ].sort() - ) - }) - - it('the relation system map is the ruled verb mirror', () => { - expect([...SYSTEM_RELATION_SCALARS].sort()).toEqual( - [ - 'verb', - 'sourceId', - 'targetId', - 'confidence', - 'createdAt', - 'createdBy', - 'service', - 'subtype', - 'updatedAt', - 'visibility', - 'weight' - ].sort() - ) - }) - - it('plumbing is exactly the ruled five, and none of it leaks into a system map', () => { - expect([...PLUMBING_FIELDS].sort()).toEqual( - ['_rev', 'connections', 'data', 'level', 'vector'].sort() - ) - for (const field of PLUMBING_FIELDS) { - expect(SYSTEM_ENTITY_SCALARS.has(field)).toBe(false) - expect(SYSTEM_RELATION_SCALARS.has(field)).toBe(false) - } - }) - - it('bare names address user metadata — even when the name matches a system scalar', () => { - expect(parseFieldAddress('level', 'entity')).toEqual({ - scope: 'metadata', - field: 'level', - raw: 'level' - }) - expect(parseFieldAddress('confidence', 'entity').scope).toBe('metadata') - expect(parseFieldAddress('createdAt', 'entity').scope).toBe('metadata') - expect(parseFieldAddress('verb', 'relation').scope).toBe('metadata') - }) - - it('metadata.-prefix is the explicit spelling of the bare form', () => { - expect(parseFieldAddress('metadata.level', 'entity')).toEqual({ - scope: 'metadata', - field: 'level', - raw: 'metadata.level' - }) - }) - - it('system.-prefix reaches exactly the map — entity and relation', () => { - for (const field of SYSTEM_ENTITY_SCALARS) { - expect(parseFieldAddress(`system.${field}`, 'entity')).toEqual({ - scope: 'system', - field, - raw: `system.${field}` - }) - } - for (const field of SYSTEM_RELATION_SCALARS) { - expect(parseFieldAddress(`system.${field}`, 'relation').scope).toBe('system') - } - // The structural relation members are NOT entity scalars. - expect(() => parseFieldAddress('system.verb', 'entity')).toThrow(InvalidFieldAddressError) - expect(() => parseFieldAddress('system.sourceId', 'entity')).toThrow(InvalidFieldAddressError) - }) - - it('plumbing refuses in the system spelling, on both record kinds', () => { - for (const field of PLUMBING_FIELDS) { - expect(() => parseFieldAddress(`system.${field}`, 'entity')).toThrow( - InvalidFieldAddressError - ) - expect(() => parseFieldAddress(`system.${field}`, 'relation')).toThrow( - InvalidFieldAddressError - ) - } - }) - - it('refusal text carries the whole valid map — the fix lives in the message', () => { - try { - parseFieldAddress('system.level', 'entity') - expect.unreachable('should have thrown') - } catch (e) { - const msg = (e as Error).message - for (const field of SYSTEM_ENTITY_SCALARS) { - expect(msg).toContain(`system.${field}`) - } - expect(msg).toContain('plumbing') - } - }) - - it('malformed addresses refuse: empty name, bare metadata. prefix', () => { - expect(() => parseFieldAddress('', 'entity')).toThrow(InvalidFieldAddressError) - expect(() => parseFieldAddress('metadata.', 'entity')).toThrow(InvalidFieldAddressError) - }) - - it('the did-you-mean names BOTH candidates for a system-colliding bare name', () => { - const msg = buildUnresolvableMessage('createdAt', 'entity') - expect(msg).toContain('system.createdAt') - expect(msg).toContain('metadata.createdAt') - }) - - it('a non-colliding unknown bare name names both spellings — system. explicitly as NOT valid', () => { - // Cross-engine pin (cor's suite greps for both spellings in every - // refusal): the metadata candidate is the fix; the system spelling is - // named but HONESTLY marked invalid, never offered as a candidate. - const msg = buildUnresolvableMessage('scoore', 'entity') - expect(msg).toContain('metadata.scoore') - expect(msg).toContain('system.scoore') - expect(msg).toContain('NOT valid') - }) -}) diff --git a/tests/unit/db/generation-segments.test.ts b/tests/unit/db/generation-segments.test.ts deleted file mode 100644 index 27ab85cb..00000000 --- a/tests/unit/db/generation-segments.test.ts +++ /dev/null @@ -1,150 +0,0 @@ -/** - * @module tests/unit/db/generation-segments - * @description The generation-segment store (Stage-2 D1+D3 file format). - * Laws: (1) fold → read round-trips deltas and records byte-faithfully via - * sidecar point-reads; (2) the manifest is the ONLY discovery path — reopen - * reads one file, never a listing; (3) a lost/corrupt sidecar rebuilds from - * its segment loudly, a damaged SEGMENT fails loudly (never silent wrong - * data); (4) D3 reclaim drops whole segments only and bumps compactedBelow; - * (5) the packed digest is deterministic across reopen; (6) immutability — - * fold refuses overlap with sealed ranges. - */ -import { describe, it, expect, beforeEach } from 'vitest' -import { MemoryStorage } from '../../../src/storage/adapters/memoryStorage.js' -import { - GenerationSegmentStore, - SEGMENTS_PREFIX, - type FoldGeneration -} from '../../../src/db/generationSegments.js' - -const UUID = (n: number): string => `00000000-0000-4000-8000-${String(n).padStart(12, '0')}` - -const gen = (g: number, recordCount = 2): FoldGeneration => ({ - generation: g, - timestamp: 1_700_000_000_000 + g, - delta: { generation: g, nouns: [UUID(g)], verbs: [], bytes: 123 + g }, - records: Array.from({ length: recordCount }, (_, i) => ({ - kind: (i % 2 === 0 ? 'noun' : 'verb') as 'noun' | 'verb', - id: UUID(g * 100 + i), - record: { metadata: { noun: 'document', v: g }, vector: { v: [g, i] } } - })) -}) - -describe('db/GenerationSegmentStore — the D1+D3 packed tier', () => { - let storage: MemoryStorage - let store: GenerationSegmentStore - - beforeEach(async () => { - storage = new MemoryStorage() - await storage.init() - store = new GenerationSegmentStore(storage as any) - await store.open() - }) - - it('fold → read round-trips deltas and records via sidecar point-reads', async () => { - const meta = await store.fold([gen(1), gen(2), gen(3)]) - expect(meta).toMatchObject({ firstGeneration: 1, lastGeneration: 3, frames: 3 }) - expect(meta.checksum).toBeGreaterThan(0) - - expect(store.hasGeneration(2)).toBe(true) - expect(store.hasGeneration(4)).toBe(false) - - const d2 = await store.readDelta(2) - expect(d2?.delta).toEqual({ generation: 2, nouns: [UUID(2)], verbs: [], bytes: 125 }) - expect(d2?.timestamp).toBe(1_700_000_000_002) - - const records = await store.readRecords(3) - expect(records).toHaveLength(2) - expect(records![0]).toEqual({ - kind: 'noun', - id: UUID(300), - record: { metadata: { noun: 'document', v: 3 }, vector: { v: [3, 0] } } - }) - // Point read by id, both kinds. - expect(await store.readRecord(3, 'verb', UUID(301))).toEqual({ - metadata: { noun: 'document', v: 3 }, - vector: { v: [3, 1] } - }) - expect(await store.readRecord(3, 'noun', UUID(999))).toBeNull() - }) - - it('reopen discovers everything from the manifest alone — no listing', async () => { - await store.fold([gen(1), gen(2)]) - await store.fold([gen(3), gen(4)]) - - const reopened = new GenerationSegmentStore(storage as any) - await reopened.open() - expect(reopened.segments()).toHaveLength(2) - expect(reopened.hasGeneration(4)).toBe(true) - expect((await reopened.readDelta(1))?.timestamp).toBe(1_700_000_000_001) - }) - - it('a lost sidecar rebuilds from its segment; a damaged segment fails LOUDLY', async () => { - const meta = await store.fold([gen(1), gen(2)]) - const idxPath = `${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.idx` - await storage.deleteRawObject(idxPath) - - const reopened = new GenerationSegmentStore(storage as any) - await reopened.open() - // Rebuild path: still serves correct data. - expect((await reopened.readRecords(2))!).toHaveLength(2) - - // Now damage the SEGMENT itself: flip a payload byte → CRC mismatch, loud. - const segPath = `${SEGMENTS_PREFIX}/${meta.file}` - const bytes = (await storage.readRawBytes(segPath))! - bytes[bytes.length - 3] ^= 0xff - await storage.writeRawBytes(segPath, bytes) - const damaged = new GenerationSegmentStore(storage as any) - await damaged.open() - ;(damaged as any).sidecars.clear() - await storage.deleteRawObject(idxPath) // force the sequential rebuild over damaged bytes - await expect(damaged.readRecords(2)).rejects.toThrow(/CRC mismatch|damaged/) - }) - - it('D3 reclaim drops whole segments only and bumps compactedBelow', async () => { - await store.fold([gen(1), gen(2)]) - await store.fold([gen(3), gen(4)]) - await store.fold([gen(5), gen(6)]) - - // Horizon mid-segment-2 (below 4): only segment 1 is FULLY below → drops. - const r1 = await store.dropSegmentsBelow(4) - expect(r1).toEqual({ dropped: 1, compactedBelow: 3 }) - expect(store.hasGeneration(1)).toBe(false) - expect(store.hasGeneration(3)).toBe(true) // partial segment survives whole - - // Bytes actually gone. - expect(await storage.readRawBytes(`${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.bgs`)).toBeNull() - - // Horizon past everything: the rest drop; compactedBelow is durable. - const r2 = await store.dropSegmentsBelow(7) - expect(r2.dropped).toBe(2) - const reopened = new GenerationSegmentStore(storage as any) - await reopened.open() - expect(reopened.compactedBelow()).toBe(7) - expect(reopened.segments()).toHaveLength(0) - }) - - it('the packed digest is deterministic across reopen and changes with history', async () => { - await store.fold([gen(1), gen(2), gen(3)]) - const atSeal = await store.digestThroughPacked(3) - const midSegment = await store.digestThroughPacked(2) - expect(atSeal).not.toBeNull() - expect(midSegment).not.toBeNull() - expect(midSegment).not.toBe(atSeal) - - const reopened = new GenerationSegmentStore(storage as any) - await reopened.open() - expect(await reopened.digestThroughPacked(3)).toBe(atSeal) - expect(await reopened.digestThroughPacked(2)).toBe(midSegment) - - await reopened.fold([gen(4)]) - expect(await reopened.digestThroughPacked(4)).not.toBe(atSeal) - }) - - it('sealed segments are immutable — fold refuses overlap, requires ascending input', async () => { - await store.fold([gen(1), gen(2)]) - await expect(store.fold([gen(2), gen(3)])).rejects.toThrow(/overlaps the packed tier/) - await expect(store.fold([gen(4), gen(4)])).rejects.toThrow(/strictly ascending/) - await expect(store.fold([])).rejects.toThrow(/at least one generation/) - }) -}) diff --git a/tests/unit/db/whereMatcher.test.ts b/tests/unit/db/whereMatcher.test.ts index 6c0252d6..2223117c 100644 --- a/tests/unit/db/whereMatcher.test.ts +++ b/tests/unit/db/whereMatcher.test.ts @@ -32,7 +32,7 @@ function entity(overrides: Partial = {}): Entity { } describe('db/whereMatcher — resolveEntityField', () => { - it('system. resolves the entity scalar; bare/metadata. reads the metadata bag only (sealed 2026-08-03)', () => { + it('resolves standard top-level fields', () => { const e = entity({ subtype: 'invoice', service: 'billing', @@ -41,32 +41,17 @@ describe('db/whereMatcher — resolveEntityField', () => { _rev: 3, data: 'payload' }) - - // system. is the ONLY spelling that reaches an entity scalar. - expect(resolveEntityField(e, 'system.id')).toBe('e-1') - expect(resolveEntityField(e, 'system.type')).toBe(NounType.Document) - expect(resolveEntityField(e, 'system.subtype')).toBe('invoice') - expect(resolveEntityField(e, 'system.service')).toBe('billing') - expect(resolveEntityField(e, 'system.confidence')).toBe(0.9) - expect(resolveEntityField(e, 'system.weight')).toBe(0.5) - expect(resolveEntityField(e, 'system.createdAt')).toBe(1000) - expect(resolveEntityField(e, 'system.updatedAt')).toBe(2000) - - // Plumbing (_rev, data) is invisible even via system. — not in the - // ten-scalar map, so this internal resolver reads it as absent (the typed - // refusal for these lives one layer up, at the query-surface parser). - expect(resolveEntityField(e, 'system._rev')).toBeUndefined() - expect(resolveEntityField(e, 'system.data')).toBeUndefined() - - // Bare names are ALWAYS the user's metadata field — even when they share - // a spelling with an engine scalar, or with the now-dead 'noun' alias. - // This entity's metadata bag is empty, so every bare name below reads - // absent rather than silently falling back to the entity scalar. - expect(resolveEntityField(e, 'id')).toBeUndefined() - expect(resolveEntityField(e, 'type')).toBeUndefined() - expect(resolveEntityField(e, 'noun')).toBeUndefined() // legacy alias is dead - expect(resolveEntityField(e, 'subtype')).toBeUndefined() - expect(resolveEntityField(e, 'createdAt')).toBeUndefined() + expect(resolveEntityField(e, 'id')).toBe('e-1') + expect(resolveEntityField(e, 'type')).toBe(NounType.Document) + expect(resolveEntityField(e, 'noun')).toBe(NounType.Document) // alias + expect(resolveEntityField(e, 'subtype')).toBe('invoice') + expect(resolveEntityField(e, 'service')).toBe('billing') + expect(resolveEntityField(e, 'confidence')).toBe(0.9) + expect(resolveEntityField(e, 'weight')).toBe(0.5) + expect(resolveEntityField(e, '_rev')).toBe(3) + expect(resolveEntityField(e, 'createdAt')).toBe(1000) + expect(resolveEntityField(e, 'updatedAt')).toBe(2000) + expect(resolveEntityField(e, 'data')).toBe('payload') }) it('resolves custom fields from the metadata bag', () => { diff --git a/tests/unit/test-suite-coverage-guard.test.ts b/tests/unit/test-suite-coverage-guard.test.ts index 21f918f1..93db4421 100644 --- a/tests/unit/test-suite-coverage-guard.test.ts +++ b/tests/unit/test-suite-coverage-guard.test.ts @@ -30,9 +30,6 @@ function allTestFiles(dir: string, out: string[] = []): string[] { * conscious decision — a NEW orphan not listed here fails the guard below. */ const MANUAL_ONLY = new Set([ - // Conformance suites run as an explicit gate stage (both engines run them - // by direct invocation), never swept into the unit/integration configs. - 'tests/conformance/collider-fidelity.test.ts', 'tests/api/performance-benchmarks.test.ts', 'tests/critical-neural-validation.test.ts', 'tests/critical-performance-benchmark.test.ts', @@ -41,15 +38,7 @@ const MANUAL_ONLY = new Set([ 'tests/package-size-limit.test.ts', 'tests/performance/graph-scale-performance.test.ts', 'tests/performance/triple-intelligence-scale.test.ts', - 'tests/performance/typeAware.bench.test.ts', - // Cross-engine field-addressing conformance suite: pinned bit-for-bit against - // the native accelerator's implementation of the SAME contract, and invoked - // directly (`npx vitest run tests/conformance/namespace-law.test.ts`), never - // swept into the unit/integration gates — a run against a branch where the - // resolver hasn't landed yet must SKIP loudly (see the file's own SELF-SKIP - // doc), not silently pass/fail as a side effect of which gate happened to - // pick it up. - 'tests/conformance/namespace-law.test.ts' + 'tests/performance/typeAware.bench.test.ts' ]) function inGate(rel: string): boolean { diff --git a/tests/unit/types/nestedBagRecord.test.ts b/tests/unit/types/nestedBagRecord.test.ts deleted file mode 100644 index b8e9be46..00000000 --- a/tests/unit/types/nestedBagRecord.test.ts +++ /dev/null @@ -1,127 +0,0 @@ -/** - * @module tests/unit/types/nestedBagRecord - * @description Unit pins for the v2 (nested-bag) stored-record layer — the - * storage half of the field-addressing law. The write door accepts ANY user - * metadata name; what makes that lossless on disk is the record shape: - * engine fields top-level, the user bag NESTED verbatim, discriminated by - * the engine-written format stamp (never by names — names are the user's). - * These pins hold the builders, the discriminator, and the shape-aware - * split that every read path (live, batch, historical) routes through. - */ -import { describe, it, expect } from 'vitest' -import { - buildNounMetadataRecord, - buildVerbMetadataRecord, - splitNounMetadataRecord, - splitVerbMetadataRecord, - isNestedBagRecord, - METADATA_RECORD_FORMAT_KEY, - NESTED_BAG_FORMAT -} from '../../../src/types/reservedFields.js' - -const COLLIDER_BAG = { - confidence: 'user-confidence', - weight: 'user-weight', - subtype: 'user-subtype', - createdAt: 'user-createdAt', - service: 'user-service', - data: 'user-data', - noun: 'user-noun', - _rev: 'user-rev', - level: 7, - plain: 'control' -} - -describe('v2 nested-bag stored records — build / discriminate / split', () => { - it('build → split round-trips a fully colliding user bag VERBATIM', () => { - const record = buildNounMetadataRecord( - { noun: 'document', confidence: 0.25, createdAt: 111, updatedAt: 222, _rev: 1 }, - { ...COLLIDER_BAG } - ) - expect(isNestedBagRecord(record)).toBe(true) - expect(record[METADATA_RECORD_FORMAT_KEY]).toBe(NESTED_BAG_FORMAT) - - const { reserved, custom } = splitNounMetadataRecord(record) - // The engine half is exactly what the engine wrote… - expect(reserved.noun).toBe('document') - expect(reserved.confidence).toBe(0.25) - expect(reserved._rev).toBe(1) - // …and the user bag comes back byte-for-byte, colliders included. - expect(custom).toEqual(COLLIDER_BAG) - }) - - it('the verb mirror round-trips an edge collider bag verbatim', () => { - const record = buildVerbMetadataRecord( - { verb: 'relatedTo', weight: 1.0, confidence: 0.5, createdAt: 333 }, - { verb: 'user-verb', confidence: 'user-c', tag: 't' } - ) - expect(isNestedBagRecord(record)).toBe(true) - const { reserved, custom } = splitVerbMetadataRecord(record) - expect(reserved.verb).toBe('relatedTo') - expect(reserved.confidence).toBe(0.5) - expect(custom).toEqual({ verb: 'user-verb', confidence: 'user-c', tag: 't' }) - }) - - it('a LEGACY flat record (no stamp) splits BY NAME — sound because the pre-law door refused colliders', () => { - const legacy = { - noun: 'document', - confidence: 0.75, - createdAt: 111, - _rev: 2, - legacyField: 'legacy-value' - } - expect(isNestedBagRecord(legacy)).toBe(false) - const { reserved, custom } = splitNounMetadataRecord(legacy) - expect(reserved.confidence).toBe(0.75) - expect(reserved._rev).toBe(2) - expect(custom).toEqual({ legacyField: 'legacy-value' }) - }) - - it('the stamp is the discriminator, never the name: a legacy user OBJECT field named `metadata` does not fake a v2 record', () => { - // Pre-law, 'metadata' was never a reserved name — a flat record could - // legally carry a user object field spelled exactly 'metadata'. Without - // the engine-written stamp it must split as legacy, with that object - // preserved as an ordinary user field. - const legacyWithMetadataField = { - noun: 'document', - confidence: 0.5, - metadata: { nested: 'user-object' } - } - expect(isNestedBagRecord(legacyWithMetadataField)).toBe(false) - const { reserved, custom } = splitNounMetadataRecord(legacyWithMetadataField) - expect(reserved.confidence).toBe(0.5) - expect(custom).toEqual({ metadata: { nested: 'user-object' } }) - }) - - it('a malformed stamp (right key, wrong value / non-object bag) never discriminates as v2', () => { - expect( - isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: 999, metadata: {} }) - ).toBe(false) - expect( - isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, metadata: 'not-a-bag' }) - ).toBe(false) - expect( - isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, metadata: [1, 2] }) - ).toBe(false) - expect(isNestedBagRecord(null)).toBe(false) - expect(isNestedBagRecord(undefined)).toBe(false) - }) - - it('the v2 split never surfaces the stamp or the bag container as fields', () => { - const record = buildNounMetadataRecord({ noun: 'document', _rev: 1 }, { a: 1 }) - const { reserved, custom } = splitNounMetadataRecord(record) - expect(METADATA_RECORD_FORMAT_KEY in reserved).toBe(false) - expect(METADATA_RECORD_FORMAT_KEY in custom).toBe(false) - expect('metadata' in reserved).toBe(false) - expect(custom).toEqual({ a: 1 }) - }) - - it('builders copy the bag (no aliasing): later caller mutation cannot reach the record', () => { - const bag: Record = { a: 1 } - const record = buildNounMetadataRecord({ noun: 'document' }, bag) - bag.a = 999 - bag.b = 'sneaky' - expect((record.metadata as Record).a).toBe(1) - expect('b' in (record.metadata as Record)).toBe(false) - }) -}) diff --git a/tests/unit/types/reserved-metadata-keys.test-d.ts b/tests/unit/types/reserved-metadata-keys.test-d.ts new file mode 100644 index 00000000..37fceefa --- /dev/null +++ b/tests/unit/types/reserved-metadata-keys.test-d.ts @@ -0,0 +1,265 @@ +/** + * @module tests/unit/types/reserved-metadata-keys.test-d + * @description Compile-time tests for the reserved-field contract (layer 1 of + * three — see src/types/reservedFields.ts): a literal reserved key inside any + * `metadata` param is a TypeScript error, while the generic `T` ergonomics + * stay intact (typed bags, untyped brains, index-signature shapes, and the + * documented exemption for consumers who explicitly declare a reserved key in + * their own metadata type). + * + * Runs under vitest typecheck mode (`test.typecheck` in + * tests/configs/vitest.unit.config.ts) — these assertions are validated by + * `tsc`, never executed. The runtime half of the contract (the write-path + * remap for untyped callers) is pinned by + * tests/unit/brainy/update-reserved-metadata-remap.test.ts. + */ + +import { describe, it, assertType } from 'vitest' +import type { + AddParams, + UpdateParams, + RelateParams, + UpdateRelationParams, + TxOperation +} from '../../../src/index.js' +import { NounType, VerbType } from '../../../src/types/graphTypes.js' + +describe('reserved entity keys in metadata are compile errors', () => { + it('AddParams (untyped brain) rejects every reserved key but stays open for custom fields', () => { + // Custom fields of any shape remain legal — exactly the pre-8.0 latitude. + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + metadata: { dept: 'eng', level: 3, tags: ['a', 'b'], nested: { ok: true } } + }) + + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'noun' is reserved (the entity type travels via the top-level 'type' param) + metadata: { noun: 'organization' } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param) + metadata: { subtype: 'contractor' } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'createdAt' is reserved (system-managed) + metadata: { createdAt: Date.now() } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'updatedAt' is reserved (system-managed) + metadata: { updatedAt: Date.now() } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param) + metadata: { confidence: 0.8 } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'weight' is reserved (use the top-level 'weight' param) + metadata: { weight: 0.5 } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'service' is reserved (use the top-level 'service' param) + metadata: { service: 'orders' } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'data' is reserved (use the top-level 'data' param) + metadata: { data: 'content' } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'createdBy' is reserved (use the top-level 'createdBy' param) + metadata: { createdBy: { augmentation: 'importer', version: '1.0' } } + }) + assertType({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — '_rev' is reserved (system-managed revision counter) + metadata: { _rev: 7 } + }) + }) + + it('AddParams (typed brain) rejects reserved keys alongside the declared shape', () => { + interface EmployeeMeta { + dept: string + level: number + } + + assertType>({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + metadata: { dept: 'eng', level: 3 } + }) + + assertType>({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + // @ts-expect-error — 'confidence' is reserved even when T declares other fields + metadata: { dept: 'eng', level: 3, confidence: 0.8 } + }) + }) + + it('documented exemptions: T-declared reserved keys and index-signature shapes stay assignable', () => { + // A consumer who *explicitly* types a reserved key into their metadata + // shape keeps a working (if unwise) type — the guard exempts keyof T. + interface LegacyMeta { + confidence: number + note: string + } + assertType>({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + metadata: { confidence: 0.8, note: 'declared by the consumer type' } + }) + + // Index-signature metadata types (keyof T = string) remain fully open. + assertType>>({ + type: NounType.Person, + subtype: 'employee', + data: 'x', + metadata: { anything: 'goes', confidence: 0.8 } + }) + }) + + it('UpdateParams patch rejects reserved keys but accepts partial custom patches', () => { + interface EmployeeMeta { + dept: string + level: number + } + + // Partial patch of the declared shape is legal. + assertType>({ id: 'e1', metadata: { dept: 'sales' } }) + // Untyped patch with custom fields is legal. + assertType({ id: 'e1', metadata: { status: 'reviewed', rating: 4.5 } }) + + // @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param) + assertType({ id: 'e1', metadata: { confidence: 0.33 } }) + // @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param) + assertType({ id: 'e1', metadata: { subtype: 'specialized' } }) + // @ts-expect-error — '_rev' is reserved (pass 'ifRev' for optimistic concurrency) + assertType({ id: 'e1', metadata: { _rev: 3 } }) + // @ts-expect-error — 'confidence' is reserved even when T declares other fields + assertType>({ id: 'e1', metadata: { confidence: 0.1 } }) + }) +}) + +describe('reserved relationship keys in metadata are compile errors', () => { + it('RelateParams rejects reserved keys but stays open for custom edge fields', () => { + assertType({ + from: 'a', + to: 'b', + type: VerbType.ReportsTo, + subtype: 'direct', + metadata: { role: 'peer', since: 2024 } + }) + + assertType({ + from: 'a', + to: 'b', + type: VerbType.ReportsTo, + subtype: 'direct', + // @ts-expect-error — 'verb' is reserved (the relationship type travels via the top-level 'type' param) + metadata: { verb: 'relatedTo' } + }) + assertType({ + from: 'a', + to: 'b', + type: VerbType.ReportsTo, + subtype: 'direct', + // @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param) + metadata: { confidence: 0.9 } + }) + assertType({ + from: 'a', + to: 'b', + type: VerbType.ReportsTo, + subtype: 'direct', + // @ts-expect-error — 'weight' is reserved (use the top-level 'weight' param) + metadata: { weight: 0.4 } + }) + assertType({ + from: 'a', + to: 'b', + type: VerbType.ReportsTo, + subtype: 'direct', + // @ts-expect-error — 'service' is reserved (use the top-level 'service' param) + metadata: { service: 'orders' } + }) + }) + + it('UpdateRelationParams patch rejects reserved keys', () => { + assertType({ id: 'r1', metadata: { note: 'fine' } }) + + // @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param) + assertType({ id: 'r1', metadata: { confidence: 0.5 } }) + // @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param) + assertType({ id: 'r1', metadata: { subtype: 'dotted-line' } }) + // @ts-expect-error — 'createdAt' is reserved (system-managed) + assertType({ id: 'r1', metadata: { createdAt: 1 } }) + }) +}) + +describe('transact() operations inherit the same guard', () => { + it('TxOperation add/update/relate metadata rejects reserved keys', () => { + assertType({ + op: 'add', + type: NounType.Concept, + subtype: 'general', + data: 'tx', + metadata: { custom: 'a' } + }) + assertType({ + op: 'add', + type: NounType.Concept, + subtype: 'general', + data: 'tx', + // @ts-expect-error — 'confidence' is reserved on transact add ops too + metadata: { confidence: 0.7 } + }) + assertType({ + op: 'update', + id: 'e1', + // @ts-expect-error — 'weight' is reserved on transact update ops too + metadata: { weight: 0.2 } + }) + assertType({ + op: 'relate', + from: 'a', + to: 'b', + type: VerbType.RelatedTo, + subtype: 'colleague', + // @ts-expect-error — 'verb' is reserved on transact relate ops too + metadata: { verb: 'contains' } + }) + }) +}) diff --git a/tests/unit/utils/paramValidation.test.ts b/tests/unit/utils/paramValidation.test.ts index 7e5212b8..4dc83554 100644 --- a/tests/unit/utils/paramValidation.test.ts +++ b/tests/unit/utils/paramValidation.test.ts @@ -56,15 +56,11 @@ describe('Zero-Config Parameter Validation', () => { })).toThrow('cannot specify both query and vector') }) - it('should refuse cursor outright — even paired with offset — as an unimplemented option', () => { - // cursor is now a typed, unconditional refusal (UnsupportedFindOptionError): - // it used to be accepted-and-ignored, only conflicting when offset was also - // given. Accepted-and-ignored died as a class — cursor refuses on its own, - // so pairing it with offset refuses too, but with the SAME message. + it('should reject both cursor and offset', () => { expect(() => validateFindParams({ cursor: 'abc123', offset: 10 - })).toThrow("find() option 'cursor' is not implemented") + })).toThrow('cannot use both cursor and offset pagination') }) it('should validate vector dimensions', () => {