Compare commits
No commits in common. "v9.0.0" and "v8.11.0" have entirely different histories.
56 changed files with 2071 additions and 4812 deletions
47
CHANGELOG.md
47
CHANGELOG.md
|
|
@ -2,41 +2,6 @@
|
|||
|
||||
All notable changes to this project will be documented in this file. See [standard-version](https://github.com/conventional-changelog/standard-version) for commit guidelines.
|
||||
|
||||
### [9.0.0](https://source.soulcraft.com/soulcraft/brainy/compare/v8.11.0...v9.0.0) (2026-08-04)
|
||||
|
||||
- docs: 9.0 namespace-migration guide — the simple story + the mechanical sweep checklist, published for humans and tooling alike (61ab9db2)
|
||||
- fix(release): storefront leg republishes CI's exact forge artifact — byte-identity by construction, verified by cross-registry shasum before the ceremony reports success (d89df2ed)
|
||||
- docs: v9.0.0 release notes — the field-addressing law migration ledger; retitle the shipped 8.11.0 canonical-enumeration entry (header went stale at its cut) (55a7512c)
|
||||
- feat(namespace): merge the field-addressing law train — no special names, system.* scalars, nested-bag storage, epoch-3 index keys (19b477ae)
|
||||
- feat(namespace): NO SPECIAL NAMES + storage fidelity — the ruled completion of the field-addressing law (24bf6cdb)
|
||||
- feat(namespace): write-door forgery refusal (user metadata keys may never start 'system.') + refusal messages name both spellings in every branch (the non-colliding case marks system.<f> honestly as NOT valid) — cross-engine message pin alignment (48a6130a)
|
||||
- feat(namespace): conformance green 19/19 — data-aware did-you-mean on unindexed bare addresses, ordering contract on the column top-K path (never drop, nulls last, ties by id), shape-complete addressed reads (entity views AND raw storage shapes, shadow-proof both scopes), per-key source matching for dotted addresses; refusal classes unified under UnresolvableFieldError (8e962dab)
|
||||
- feat(namespace): aggregation reads under the law + epoch 3 (the key-split rebuild) + THE ARMING COMMIT — the capability constant, the law module, and the typed refusals export from the package root; both engines' conformance suites light on this signal (7492b6cb)
|
||||
- feat(namespace): egress guard + validation speak the law — whereMatcher's resolver reads system.* from the record and bare names from the metadata bag only (the bare-system switch is dead); validateFindParams refuses cursor/includeRelations/writeOnly typed (accepted-and-ignored dies as a class), validates order, and parses every orderBy address (c2fb28a2)
|
||||
- fix(namespace): noun-record updates preserve legacy inline HNSW adjacency — the placeholder-adjacency write stamped out pre-codec records' stored connections (crash-window unreachability); codec-era records were never at risk (empty field is the blob marker); pin covers the legacy shape (4679c894)
|
||||
- feat(namespace): find's own filter builders speak the frozen keys — params.type/subtype/service become system.* index keys at every construction site (three pipelines + the canonical buildMetadataFilter); the where.type→noun alias is dead (bare 'type' belongs to the user now) (7a28a946)
|
||||
- feat(namespace): the index speaks the frozen keys — record-frame scalars index under literal 'system.<field>' (legacy 'noun' spelling folds into system.type; plumbing never indexed from a record frame), user fields stay bare in every shape; filter + sorted paths route every address through parseFieldAddress; storage fallbacks read the addressed side of the record (11c724bc)
|
||||
- docs(namespace): the d.ts JSDoc wave — the sealed field-addressing law on the full find + aggregation surface, present-tense, with the refusal semantics and migration note inline (comment-only; verified zero code lines changed) (fcb24ab6)
|
||||
- test(namespace): unit pins for the pure law — the ruled maps verbatim (incl. the relation mirror, unpinnable via public API), plumbing refusals both kinds, did-you-mean text (5502abcd)
|
||||
- fix(namespace): the JS sorted fallback honors the ruled ordering contract — nulls last in BOTH directions (was nulls-first on desc) + deterministic id-ascending tie-break (56deb2e8)
|
||||
- test(namespace)+docs: the cross-engine conformance suite (self-arming — skips until the resolver exports land) + the public field-addressing docs page; sidebar order deconflicted to 7 (d8d0b55f)
|
||||
- feat(namespace): the one field-addressing law as a single source of truth — parseFieldAddress + the ruled ten-scalar system maps + plumbing invisibility + refusal builders (module only; query surfaces wire in next) (8f9a9989)
|
||||
- docs: port the 8.10.3 backport-release changelog entry to main (f6b14d21)
|
||||
- docs: port the 8.10.2 backport-release changelog entry to main — release branches carry the version bump, main carries the durable record (0b059ac5)
|
||||
- fix: user metadata named 'level' is a real field everywhere — the engine-internal node layer no longer shadows it in sort/filter/aggregation, and the indexing views stop stamping a phantom 0 into its column; index epoch 2 rebuilds existing brains at first open (1a09be06)
|
||||
- fix: metadata-only update() never rewrites the noun record — the unconditional whole-vector save turned per-entity stat touches into full rewrites+fsync, amplifying read-heavy sweeps into disk saturation on a production deployment (cb717be2)
|
||||
- fix(release): double the forge-publish poll budget — the runner executes jobs sequentially and the publish run queues behind the ci matrix (64049631)
|
||||
- Merge branch 'release/8.11.0' (1865f60a)
|
||||
- Merge branch 'release/8.10.1' (fc9f0d72)
|
||||
- chore: the forge is the address — retire the archived mirror from every live surface (415e824a)
|
||||
- Merge remote-tracking branch 'origin/main' (069a8894)
|
||||
- Merge branch 'release/8.10.0' (d918c060)
|
||||
- ci: run the pipeline on the forge (9a5a9ccc)
|
||||
- feat: two-tier history reads + the repacker + generationDigest — D1+D3 wired end-to-end (1201e255)
|
||||
- feat: generation-segment store — the D1+D3 packed-tier file format (d8acb377)
|
||||
- feat: scanFacts liveness contract — first batch or loud failure within a documented bound (f8e6da2b)
|
||||
|
||||
|
||||
### [8.11.0](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.1...v8.11.0) (2026-07-27)
|
||||
|
||||
- docs: the last two archived-host links point home (91ef1c8b)
|
||||
|
|
@ -46,18 +11,6 @@ All notable changes to this project will be documented in this file. See [standa
|
|||
- ci: run the pipeline on the forge (999d0ebb)
|
||||
|
||||
|
||||
### [8.10.3](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.2...v8.10.3) (2026-08-03)
|
||||
|
||||
- docs: dedupe the 8.10.2 release-notes entry the cherry doubled onto the branch (8c956608)
|
||||
- fix: user metadata named 'level' is a real field everywhere — the engine-internal node layer no longer shadows it in sort/filter/aggregation, and the indexing views stop stamping a phantom 0 into its column; index epoch 2 rebuilds existing brains at first open (958a0859)
|
||||
|
||||
|
||||
### [8.10.2](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.1...v8.10.2) (2026-07-29)
|
||||
|
||||
- docs: 8.10.2 consumer release notes — update() write granularity, PathResolver idle-log fix, graph-lsm key recognition (a0123b5b)
|
||||
- fix: metadata-only update() never rewrites the noun record — the unconditional whole-vector save turned per-entity stat touches into full rewrites+fsync, amplifying read-heavy sweeps into disk saturation on a production deployment (5b65eb82)
|
||||
|
||||
|
||||
### [8.10.1](https://source.soulcraft.com/soulcraft/brainy/compare/v8.10.0...v8.10.1) (2026-07-24)
|
||||
|
||||
- refactor: remove the orphaned transaction-result type left behind by the dead-path removal (edf123a5)
|
||||
|
|
|
|||
150
RELEASES.md
150
RELEASES.md
|
|
@ -31,7 +31,7 @@ is sometimes cited as a 7.x removal — those methods never existed on 7.x; the
|
|||
|
||||
---
|
||||
|
||||
## v8.11.0 — 2026-07-27 (canonical enumeration mode for export — storage-walked, canon-complete)
|
||||
## Unreleased (canonical enumeration mode for export — storage-walked, canon-complete)
|
||||
|
||||
From a fleet data-migration program's requirement for whole-brain exports that are
|
||||
provably canon-complete: `export()`'s default enumeration for a whole-brain/predicate
|
||||
|
|
@ -74,154 +74,6 @@ to the caller today.
|
|||
on CI**, triggered by the release tag, instead of PUTting the tarball from the laptop
|
||||
over WAN — no change to what gets published or how a consumer installs it.
|
||||
|
||||
## v9.0.0 — 2026-08-04 (the field-addressing law: your names and system.*, nothing in between)
|
||||
|
||||
**Major.** One law now governs every field name, on every surface:
|
||||
|
||||
> **Data is either in main space — where you can use ANY name — or it is in
|
||||
> `system.*`.**
|
||||
|
||||
Read `docs/concepts/field-addressing.md` (published on the docs site) for the
|
||||
full contract; this entry is the migration ledger.
|
||||
|
||||
### Breaking — query surfaces (`where` / `orderBy` / `groupBy` / aggregation)
|
||||
|
||||
- **A bare field name ALWAYS addresses your metadata.** `orderBy: 'createdAt'`
|
||||
no longer silently means the engine timestamp — it now refuses with a typed
|
||||
`UnresolvableFieldError` naming both candidates unless you actually have a
|
||||
user field of that name. Engine scalars are addressed explicitly:
|
||||
`system.id`, `system.type`, `system.subtype`, `system.createdAt`,
|
||||
`system.updatedAt`, `system.confidence`, `system.weight`,
|
||||
`system.visibility`, `system.service`, `system.createdBy` (relations mirror
|
||||
with `system.verb`/`system.sourceId`/`system.targetId`).
|
||||
**Sweep list:** `where: { subtype: … }` → `where: { 'system.subtype': … }` ·
|
||||
`orderBy: 'createdAt'` → `'system.createdAt'` · `groupBy: ['noun']` →
|
||||
`['system.type']` · any bare `visibility`/`service`/`confidence` filter that
|
||||
meant the engine value → its `system.*` spelling. Every missed site fails
|
||||
LOUDLY with the correction in the error message — nothing silently changes
|
||||
meaning without telling you.
|
||||
- **Unimplemented `find()` options refuse** (`cursor`, `includeRelations`,
|
||||
`writeOnly` → `UnsupportedFindOptionError`); `order` is validated;
|
||||
accepted-and-ignored is dead as a class.
|
||||
- **The ordering contract is pinned cross-engine:** missing/null `orderBy`
|
||||
values sort LAST in both directions, ties break by id ascending, and rows
|
||||
are never dropped from an ordered read.
|
||||
|
||||
### Breaking — write surfaces
|
||||
|
||||
- **There are no reserved metadata names anymore.** `metadata: { confidence,
|
||||
type, id, level, data, content, … }` are ordinary user fields — stored
|
||||
verbatim, indexed, filterable, sortable, aggregatable, faithful across
|
||||
restarts, index rebuilds, and `asOf()` time travel. The 8.x
|
||||
reserved-key-in-bag throw is GONE; code that relied on it (or on the
|
||||
`'warn'`/`'remap'` lift) must set engine scalars via their dedicated params
|
||||
(`confidence`, `weight`, `subtype`, `visibility`, …) — the bag never touches
|
||||
them now.
|
||||
- **`reservedFieldPolicy` is removed.** Passing it throws at construction with
|
||||
the migration note. `RESERVED_ENTITY_FIELDS`/`RESERVED_RELATION_FIELDS`
|
||||
remain exported but now describe the stored record's engine half, not a ban
|
||||
list; the `NoReservedEntityKeys`/`NoReservedRelationKeys` types are no-op
|
||||
(deprecated).
|
||||
- **The one refused spelling:** a metadata key literally starting `system.`
|
||||
(namespace forgery) — typed error on `add`/`update`/`relate`/`updateRelation`.
|
||||
- **Name-based index exclusions are gone.** Fields named `content`, `data`,
|
||||
`id`, `vector`, … in your bag now INDEX like everything else (they were
|
||||
silently un-indexed before — `where` on them returned `[]` with no error).
|
||||
Value-shape rules stay, uniform across all names: arrays >10 never become
|
||||
posting scalars; long values index hashed.
|
||||
- **Migration transforms receive one normalized view** (engine fields
|
||||
top-level, your bag nested under `metadata`) regardless of how old the
|
||||
stored record is, and must return the same shape — a stray non-engine
|
||||
top-level key refuses with the fix in the message.
|
||||
|
||||
### Storage format (automatic, no action)
|
||||
|
||||
- New/updated records persist as **nested-bag records** (engine fields
|
||||
top-level, your bag verbatim under `metadata`, sealed by a format stamp) —
|
||||
the shape that makes collider names lossless. Old flat records stay
|
||||
readable forever; nothing rewrites your data in place.
|
||||
- **Index epoch 3:** derived-index keys split the namespaces (bare user keys ·
|
||||
literal `system.<field>` keys; the legacy `noun` column is gone). Every
|
||||
brain rebuilds its derived indexes from canonical once, at first open —
|
||||
observable via `getIndexStatus()`, no manual step. Pair this release with
|
||||
the same-day native-accelerator release (its peer floor rises to `>=9`).
|
||||
- Raw-record consumers (fact-log scanners, export tooling): read bags through
|
||||
the exported shape-aware splitters (`splitNounMetadataRecord` /
|
||||
`splitVerbMetadataRecord`) — they handle both record eras.
|
||||
|
||||
### Fixed in the same train
|
||||
|
||||
- Default visibility exclusion was a silent no-op under the new addressing on
|
||||
pre-release builds (internal/system-tier rows could leak into default
|
||||
reads) — now pinned by conformance tests at every lifecycle boundary.
|
||||
- Per-type count surfaces (`getStats()`, count-by-type) read the new type
|
||||
column, with a legacy fallback for pre-rebuild reads.
|
||||
- Aggregation `source.where` evaluated dotted keys as nested paths — dotted
|
||||
addresses now match per-key, and the internal per-type counts aggregate
|
||||
rebuilds itself onto the new keys automatically.
|
||||
|
||||
### Conformance
|
||||
|
||||
Both engines ship a shared self-arming conformance suite (the law cases, the
|
||||
ordering contract, and the reopen-collider fidelity case: every collider name
|
||||
written as user data, verified verbatim through live reads, reopen, a forced
|
||||
epoch rebuild, and time travel). Capability signal:
|
||||
`FIELD_ADDRESSING_CAPABILITY = 'field-addressing/v1'` plus the typed error
|
||||
classes, exported from the package root.
|
||||
|
||||
## v8.10.3 — 2026-08-03, 8.10-line backport (natural field names stop colliding with engine internals)
|
||||
|
||||
From a production report: sorting by a user metadata field named `level` silently
|
||||
returned insertion order — the engine's internal HNSW node layer (also called
|
||||
`level`) shadowed the user's field in every by-name read, and the indexing path
|
||||
stamped a hardcoded `0` into the same index column (multi-valued poison). `level`
|
||||
is a perfectly natural field name (game characters, priorities, floors); the
|
||||
engine was wrong, not the caller.
|
||||
|
||||
- **`level` is user data now, everywhere.** Engine plumbing no longer resolves by
|
||||
name, never shadows metadata, and never enters the indexed views. `orderBy:
|
||||
'level'`, `where: { level: 9 }`, `groupBy: ['level']` all read YOUR field.
|
||||
Regression pins: `tests/integration/level-field-shadow.test.ts` (the reporting
|
||||
consumer's exact repro rows).
|
||||
- **Index epoch 2.** The derived posting set changed, so every existing brain
|
||||
rebuilds its metadata index from canonical at first open — poisoned columns
|
||||
heal automatically; no manual step. First open after upgrade pays one rebuild
|
||||
(observable via `getIndexStatus()`); pair this release with the same-day
|
||||
native-accelerator release, which makes `level` indexable on the native path.
|
||||
- **`transact()` metadata-only updates stop rewriting the vector record** — the
|
||||
v8.10.2 write-granularity law now covers the batch/plan path too (it was
|
||||
fixed for `update()` but the transact plan builder still staged the
|
||||
unconditional save). If you batch stat touches through `transact()`, this is
|
||||
your write-amplification fix.
|
||||
- (The "coming next" note this entry carried shipped as v9.0.0 — the
|
||||
field-addressing law above.)
|
||||
|
||||
---
|
||||
|
||||
## v8.10.2 — 2026-07-29 (metadata-only updates stop rewriting the vector record)
|
||||
|
||||
From a production incident on a large deployment: a read-heavy sweep that bumped
|
||||
per-entity stats (metadata-only `update()` calls) saturated the disk — 5.8GB written
|
||||
in 40 minutes — because every `update()` unconditionally re-persisted the WHOLE noun
|
||||
record, unchanged vector included, fsynced.
|
||||
|
||||
- **`update()` write granularity fixed at the core.** A metadata-only update (no new
|
||||
`data`, `vector`, or `type`) now writes the metadata leg and index deltas ONLY —
|
||||
the vector-bearing noun record is never rewritten. Vector-side writes and HNSW
|
||||
reindexing still happen exactly when the vector side actually changed. Regression
|
||||
pins: `tests/integration/update-write-granularity.test.ts`.
|
||||
- **Consumer guidance:** per-entity stat touches are now cheap, but batch them anyway
|
||||
(one `transact()` instead of N `update()` calls) — granularity fixes the cost per
|
||||
touch; batching fixes the count.
|
||||
- Idle VFS `PathResolver` no longer logs `NaN% hit rate` once a minute (stats log
|
||||
only on new traffic, at debug level).
|
||||
- Native graph providers' `graph-lsm-*` storage keys are recognized as system
|
||||
resources — the per-boot `Unknown key format` warning for them is gone.
|
||||
|
||||
Pairs with the native accelerator's same-day patch release; adopt as one bump.
|
||||
|
||||
---
|
||||
|
||||
## v8.10.1 — 2026-07-24 (the no-hot-retry contract + warm()'s metadata surface under native providers)
|
||||
|
||||
From a production incident: a native-provider op ground 38-40s inside a transaction,
|
||||
|
|
|
|||
|
|
@ -1,233 +0,0 @@
|
|||
---
|
||||
title: Field addressing: your fields and system fields
|
||||
slug: concepts/field-addressing
|
||||
public: true
|
||||
category: concepts
|
||||
template: concept
|
||||
order: 7
|
||||
description: The one rule for every query-surface field name — a bare name always means your metadata, system.<field> reaches the ten engine scalars explicitly, and anything else refuses by name.
|
||||
next:
|
||||
- guides/namespace-migration
|
||||
- concepts/consistency-model
|
||||
---
|
||||
|
||||
# Field addressing: your fields and system fields
|
||||
|
||||
Every query surface in Brainy — `find()`'s `where`, `orderBy`, aggregation
|
||||
`groupBy`, and aggregation `source.where` — resolves field names by one rule,
|
||||
with no exceptions:
|
||||
|
||||
> **A bare field name always means your metadata. `system.<field>` reaches an
|
||||
> engine scalar, and only when you spell it explicitly.**
|
||||
|
||||
```typescript
|
||||
await brain.find({ orderBy: 'level' }) // reads entity.metadata.level — YOUR field
|
||||
await brain.find({ orderBy: 'system.createdAt' }) // reads the engine's createdAt scalar
|
||||
await brain.find({ orderBy: 'metadata.level' }) // identical to bare 'level' — explicit scope
|
||||
```
|
||||
|
||||
There is no priority list, no "try the system field, fall back to metadata"
|
||||
behavior, and no name that resolves differently depending on what else
|
||||
happens to exist on your entities. A field called `level`, `score`,
|
||||
`createdAt`, or `type` in your own `metadata` is read as *your* field, every
|
||||
time, by its bare name.
|
||||
|
||||
## Why this rule exists
|
||||
|
||||
An internal report from a production deployment found that a user metadata
|
||||
field literally named `level` was being silently shadowed by the engine's
|
||||
own internal index layer field of the same name — every sort by `level`
|
||||
returned insertion order, with no error raised. This rule makes that class of
|
||||
bug structurally impossible: bare names belong to you, unconditionally, and
|
||||
anything that isn't yours has to be spelled out.
|
||||
|
||||
## The system scalars
|
||||
|
||||
`system.<field>` addresses exactly ten scalars on an entity — no more, no
|
||||
fewer:
|
||||
|
||||
| System field | What it is |
|
||||
|---|---|
|
||||
| `system.id` | The entity's id |
|
||||
| `system.type` | The entity's `NounType` |
|
||||
| `system.subtype` | The per-app sub-classification passed to `add()` |
|
||||
| `system.createdAt` | When the entity was created |
|
||||
| `system.updatedAt` | When the entity was last written |
|
||||
| `system.confidence` | The `confidence` param (0–1) |
|
||||
| `system.weight` | The `weight` param |
|
||||
| `system.visibility` | `'public'` / `'internal'` (see the visibility tiers in [Consistency Model](./consistency-model.md)) |
|
||||
| `system.service` | The multi-tenancy `service` tag |
|
||||
| `system.createdBy` | Who/what created the entity |
|
||||
|
||||
Relationships mirror the same eight shared scalars (`subtype`, `createdAt`,
|
||||
`updatedAt`, `confidence`, `weight`, `visibility`, `service`, `createdBy`)
|
||||
plus three of their own:
|
||||
|
||||
| System field (relationship) | What it is |
|
||||
|---|---|
|
||||
| `system.verb` | The relationship's `VerbType` |
|
||||
| `system.sourceId` | The id of the entity the relationship starts from |
|
||||
| `system.targetId` | The id of the entity the relationship points to |
|
||||
|
||||
Anything not on these two lists is not a system scalar — `system.<name>` for
|
||||
any other name refuses (see "Refusal semantics" below), even if that name
|
||||
sounds like it should be engine-owned.
|
||||
|
||||
## Invisible plumbing — never addressable, in either spelling
|
||||
|
||||
Five names are pure engine internals. They are not reachable as a bare name,
|
||||
and not reachable as `system.<name>` either — they simply have no place on
|
||||
the query surface:
|
||||
|
||||
- **`vector`** — the stored embedding. It participates in similarity search
|
||||
(`query`, `near`, vector `find()`), never in `where`/`orderBy`/`groupBy`.
|
||||
- **`connections`** — graph adjacency. Reached through `connected` and
|
||||
`brain.related()`, not through field addressing.
|
||||
- **`level`** — the internal index layer number used by the nearest-neighbor
|
||||
graph. It is pure index plumbing with no query-surface meaning at all —
|
||||
which is exactly why a user field of the same name must never be shadowed
|
||||
by it. `level` as a bare name is always yours; there is no engine-owned
|
||||
spelling of it to compete with.
|
||||
- **`data`** — your entity's content payload, not a scalar. It can be a
|
||||
string, a number, or an arbitrary object, so sorting or filtering it as a
|
||||
single comparable value would lie about its actual shape. Content is
|
||||
reached through the content/text-search APIs (`query`, `searchMode:
|
||||
'text'`), not through `where`/`orderBy`.
|
||||
- **`_rev`** — the per-entity revision counter used for optimistic
|
||||
concurrency (`ifRev`). It is a CAS token, not a queryable dimension.
|
||||
|
||||
`system.level`, `system.vector`, and `system.data` all refuse for the same
|
||||
reason: they are not in the ten-scalar system map, full stop.
|
||||
|
||||
## `metadata.<field>` — the explicit spelling of "mine"
|
||||
|
||||
Prefix any field with `metadata.` to say the same thing a bare name already
|
||||
says, spelled out. The two are interchangeable everywhere a field name is
|
||||
accepted, including `orderBy`:
|
||||
|
||||
```typescript
|
||||
await brain.find({ where: { 'customer.tier': 'gold' } })
|
||||
await brain.find({ where: { 'metadata.customer.tier': 'gold' } }) // identical
|
||||
await brain.find({ orderBy: 'metadata.score', order: 'desc' }) // identical to orderBy: 'score'
|
||||
```
|
||||
|
||||
Reach for the explicit spelling when it reads more clearly next to a
|
||||
`system.` field in the same query — for example, sorting by your own `score`
|
||||
while filtering on `system.confidence`.
|
||||
|
||||
## No special names — the write side
|
||||
|
||||
The same law governs writes:
|
||||
|
||||
> **Data is either in main space, where developers can use anything, or it
|
||||
> is in `system.*`.**
|
||||
|
||||
There are **no reserved metadata names**. A field called `confidence`,
|
||||
`type`, `id`, `data`, `content`, or anything else inside your `metadata` bag
|
||||
is an ordinary user field: it is stored verbatim, indexed, filterable,
|
||||
sortable, aggregatable, and it survives restarts, index rebuilds, and
|
||||
time-travel (`asOf`) reads exactly as written — even when an engine scalar
|
||||
shares its spelling. The engine's values are written only through their
|
||||
dedicated params (`confidence`, `weight`, `subtype`, `visibility`, …) and
|
||||
read at `system.<field>`; your bag can never touch them and they can never
|
||||
shadow your bag.
|
||||
|
||||
```typescript
|
||||
const id = await brain.add({
|
||||
data: 'Ada Lovelace',
|
||||
type: NounType.Person,
|
||||
confidence: 0.9, // the ENGINE scalar
|
||||
metadata: { confidence: 'self-rated' } // YOUR field, same spelling — both live
|
||||
})
|
||||
|
||||
await brain.find({ where: { confidence: 'self-rated' } }) // finds it (yours)
|
||||
await brain.find({ where: { 'system.confidence': 0.9 } }) // finds it (engine's)
|
||||
```
|
||||
|
||||
The one spelling a write refuses is a metadata key that literally starts
|
||||
with `system.` — the explicit address namespace cannot be forged as a user
|
||||
field name. That refusal is typed and names the fix.
|
||||
|
||||
Value **shape** rules still apply uniformly to every name (they are not name
|
||||
carve-outs): arrays longer than 10 elements are not turned into posting-list
|
||||
scalars, and very long values are indexed by hash.
|
||||
|
||||
## Refusal semantics
|
||||
|
||||
A name that resolves to neither your metadata nor a system scalar is a typed
|
||||
refusal, not a silent empty result and not a guess. Refusals name **both**
|
||||
candidates, so the fix is always in the error text:
|
||||
|
||||
```typescript
|
||||
await brain.find({ orderBy: 'createdAt' })
|
||||
// UnresolvableFieldError: no metadata field 'createdAt' — did you mean
|
||||
// system.createdAt or metadata.createdAt?
|
||||
```
|
||||
|
||||
`UnresolvableFieldError` is exported from the package root:
|
||||
|
||||
```typescript
|
||||
import { UnresolvableFieldError } from '@soulcraft/brainy'
|
||||
|
||||
try {
|
||||
await brain.find({ orderBy: 'createdAt' })
|
||||
} catch (err) {
|
||||
if (err instanceof UnresolvableFieldError) {
|
||||
// err.message names both candidates — usually enough to fix the call site.
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
A handful of `find()` options are not implemented yet: `cursor`,
|
||||
`includeRelations`, and `writeOnly`. Rather than accepting them and quietly
|
||||
ignoring the option, `find()` refuses with `UnsupportedFindOptionError` —
|
||||
also exported from the package root — so a call site can never believe an
|
||||
unimplemented option took effect when it didn't.
|
||||
|
||||
## The ordering contract
|
||||
|
||||
`orderBy` behaves identically regardless of which engine (the pure-TypeScript
|
||||
path or a native accelerator) is serving the query:
|
||||
|
||||
- An entity missing the `orderBy` field, or holding `null` on it, sorts
|
||||
**LAST — in both `asc` and `desc`**. It is never treated as "smaller than
|
||||
everything" in one direction and "larger than everything" in the other; it
|
||||
is simply last, either way.
|
||||
- Rows are **never dropped** from an ordered read because they lack the
|
||||
field — a missing value changes position, never presence.
|
||||
- Ties on the `orderBy` field break by **id ascending**, regardless of the
|
||||
primary sort direction.
|
||||
|
||||
```typescript
|
||||
// employees: [{ score: 9 }, { score: 5 }, { /* no score field */ }]
|
||||
await brain.find({ orderBy: 'score', order: 'desc' }) // [9, 5, missing] — missing is last
|
||||
await brain.find({ orderBy: 'score', order: 'asc' }) // [5, 9, missing] — missing is STILL last
|
||||
```
|
||||
|
||||
## Migrating existing call sites
|
||||
|
||||
If you have call sites written before this rule shipped that rely on a bare
|
||||
system name — `orderBy: 'createdAt'`, `where: { confidence: { greaterThan:
|
||||
0.8 } }`, and similar — they now refuse instead of silently resolving to the
|
||||
engine field. The fix is always in the error: swap the bare name for
|
||||
`system.<field>` (or `metadata.<field>` if you actually meant your own field
|
||||
of that name, and it happens to share a name with a system scalar):
|
||||
|
||||
```typescript
|
||||
// Before: bare 'createdAt' silently meant the engine's timestamp.
|
||||
await brain.find({ orderBy: 'createdAt' })
|
||||
|
||||
// After: say which one you meant.
|
||||
await brain.find({ orderBy: 'system.createdAt' }) // the engine timestamp
|
||||
await brain.find({ orderBy: 'metadata.createdAt' }) // your own field named createdAt, if you have one
|
||||
```
|
||||
|
||||
There is no silent migration path by design — every ambiguous call site
|
||||
surfaces as a refusal naming its own fix, once, the first time it runs
|
||||
against the new rule.
|
||||
|
||||
## Where to go next
|
||||
|
||||
- [Consistency Model](./consistency-model.md) — visibility tiers, revision
|
||||
counters, and the rest of the read/write contract this page's
|
||||
read-time addressing rule.
|
||||
|
|
@ -1,99 +0,0 @@
|
|||
---
|
||||
title: Migrating to 9.0 — your fields and system fields
|
||||
slug: guides/namespace-migration
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 1
|
||||
description: The simple story of the 9.0 field-addressing change and the mechanical checklist for updating your call sites — every miss fails loudly with the fix in the error.
|
||||
next:
|
||||
- concepts/field-addressing
|
||||
---
|
||||
|
||||
# Migrating to 9.0 — your fields and system fields
|
||||
|
||||
The one-sentence version: **your data's field names are now completely
|
||||
yours, the engine's own fields all live behind one `system.` prefix, and
|
||||
nothing in between can silently go wrong anymore.**
|
||||
|
||||
## What changed, simply
|
||||
|
||||
**1. Any field name just works.** Before 9.0 the engine quietly owned
|
||||
certain names. A field called `level` could be shadowed by the engine's
|
||||
internal index layer of the same name (sorts silently returned insertion
|
||||
order); names like `confidence` or `subtype` were rejected inside
|
||||
`metadata`; names like `content` or `id` were silently never indexed, so
|
||||
filtering on them returned nothing. All of that is gone. Any name —
|
||||
`level`, `confidence`, `type`, `id`, `content`, anything — is stored
|
||||
exactly as written and works with every feature: filtering, sorting,
|
||||
grouping, aggregation, search, and time-travel reads.
|
||||
|
||||
**2. The engine's fields moved behind `system.`.** The engine still keeps
|
||||
its own per-record bookkeeping — creation time, type, confidence, and so
|
||||
on. Those are reached one way only now: spelled out, e.g.
|
||||
`system.createdAt`, `system.type`. They are just as queryable and sortable
|
||||
as before. `orderBy: 'createdAt'` means *your* field named `createdAt`;
|
||||
`orderBy: 'system.createdAt'` means the engine's timestamp. No guessing,
|
||||
no priority rules.
|
||||
|
||||
**3. Storage keeps the two physically separate.** New records store your
|
||||
metadata in its own nested compartment, so a user field named
|
||||
`confidence` and the engine's confidence live side by side, both intact,
|
||||
through restarts, index rebuilds, and `asOf()` history. Old records stay
|
||||
readable forever; nothing rewrites your data.
|
||||
|
||||
**4. Mistakes are loud.** An ambiguous or unknown field name is a typed
|
||||
error naming the fix. Unimplemented options refuse instead of being
|
||||
ignored. The only forbidden name in your metadata is one literally
|
||||
starting with `system.`.
|
||||
|
||||
## The mechanical checklist
|
||||
|
||||
Every missed site fails **loudly** with the correction in the error
|
||||
message — nothing silently changes meaning. Sweep these patterns:
|
||||
|
||||
| Before (8.x) | After (9.0) |
|
||||
|---|---|
|
||||
| `orderBy: 'createdAt'` (meaning the engine timestamp) | `orderBy: 'system.createdAt'` |
|
||||
| `where: { subtype: 'invoice' }` (the engine subtype) | `where: { 'system.subtype': 'invoice' }` |
|
||||
| `where: { confidence: { greaterThan: 0.8 } }` (the engine scalar) | `where: { 'system.confidence': { greaterThan: 0.8 } }` |
|
||||
| `groupBy: ['noun']` or `groupBy: ['type']` | `groupBy: ['system.type']` |
|
||||
| `where: { visibility: 'internal' }` / `{ service: … }` (engine values) | `'system.visibility'` / `'system.service'` |
|
||||
| `metadata: { confidence: 0.9 }` expecting a throw or a lift to the engine scalar | it is YOUR field now — set the engine scalar via the `confidence` param |
|
||||
| `new Brainy({ reservedFieldPolicy: … })` | remove the option (it throws with this note) |
|
||||
| `find({ cursor })` / `includeRelations` / `writeOnly` | refuse with `UnsupportedFindOptionError` — they were silently ignored before |
|
||||
|
||||
If a bare name in a query was genuinely *your* field all along (`orderBy:
|
||||
'score'`, `where: { status: 'active' }`), **change nothing** — bare names
|
||||
mean your fields, always.
|
||||
|
||||
## What happens at first open
|
||||
|
||||
Each existing database rebuilds its derived indexes once, automatically,
|
||||
at the first open on 9.0 (index epoch 3 — the index keys split the two
|
||||
namespaces). One-time cost, observable via `getIndexStatus()`; no manual
|
||||
step, and your stored data is not modified.
|
||||
|
||||
## For tooling and raw-record readers
|
||||
|
||||
If you read raw stored records (fact-log scanners, export tooling), use
|
||||
the exported shape-aware splitters — they handle both record eras:
|
||||
|
||||
```typescript
|
||||
import { splitNounMetadataRecord } from '@soulcraft/brainy'
|
||||
const { reserved, custom } = splitNounMetadataRecord(rawRecord)
|
||||
// reserved = engine fields · custom = the user's bag, ANY names
|
||||
```
|
||||
|
||||
Feature detection (never version-sniff):
|
||||
|
||||
```typescript
|
||||
import * as brainy from '@soulcraft/brainy'
|
||||
const lawActive = 'FIELD_ADDRESSING_CAPABILITY' in brainy // 'field-addressing/v1'
|
||||
```
|
||||
|
||||
## Where to go next
|
||||
|
||||
- [Field addressing](../concepts/field-addressing.md) — the full contract:
|
||||
the ten system scalars, the relation mirror, refusal semantics, and the
|
||||
cross-engine ordering guarantees.
|
||||
4
package-lock.json
generated
4
package-lock.json
generated
|
|
@ -1,12 +1,12 @@
|
|||
{
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "9.0.0",
|
||||
"version": "8.11.0",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "9.0.0",
|
||||
"version": "8.11.0",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@msgpack/msgpack": "^3.1.2",
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
{
|
||||
"name": "@soulcraft/brainy",
|
||||
"version": "9.0.0",
|
||||
"version": "8.11.0",
|
||||
"description": "Universal Knowledge Protocol™ - World's first Triple Intelligence database unifying vector, graph, and document search in one API. Stage 3 CANONICAL: 42 nouns × 127 verbs covering 96-97% of all human knowledge.",
|
||||
"main": "dist/index.js",
|
||||
"module": "dist/index.js",
|
||||
|
|
|
|||
|
|
@ -189,7 +189,7 @@ echo -e "${GREEN}✅ Pushed to origin${NC}\n"
|
|||
# the forge/npmjs pair enough to publish the storefront leg.
|
||||
FORGE_NPM_REG="https://source.soulcraft.com/api/packages/soulcraft/npm/"
|
||||
FORGE_POLL_INTERVAL_S=15
|
||||
FORGE_POLL_MAX_ATTEMPTS=80 # 80 × 15s = 20 minutes — the runner is sequential; the publish run queues behind ci.yml jobs
|
||||
FORGE_POLL_MAX_ATTEMPTS=40 # 40 × 15s = 10 minutes
|
||||
echo -e "${BLUE}9️⃣ Waiting for CI to publish v${NEW_VERSION} to the forge registry (home)...${NC}"
|
||||
FORGE_LANDED=false
|
||||
for ((attempt = 1; attempt <= FORGE_POLL_MAX_ATTEMPTS; attempt++)); do
|
||||
|
|
@ -212,28 +212,10 @@ else
|
|||
fi
|
||||
|
||||
echo -e "${BLUE}9️⃣½ Publishing to npmjs (storefront, dist-tag: ${NPM_TAG})...${NC}"
|
||||
# BYTE-IDENTITY LAW: the storefront republishes CI's EXACT artifact — download
|
||||
# the tarball the forge serves and publish that file, never a fresh local pack
|
||||
# (a local rebuild can differ byte-wise, and the fleet verifies the pair by
|
||||
# shasum across registries).
|
||||
STOREFRONT_TMP="$(mktemp -d)"
|
||||
(cd "$STOREFRONT_TMP" && npm pack "@soulcraft/brainy@${NEW_VERSION}" "--@soulcraft:registry=${FORGE_NPM_REG}" >/dev/null)
|
||||
FORGE_TARBALL="$(ls "$STOREFRONT_TMP"/soulcraft-brainy-*.tgz)"
|
||||
echo -e "${BLUE} forge artifact: $(sha256sum "$FORGE_TARBALL" | cut -d' ' -f1)${NC}"
|
||||
npm publish "$FORGE_TARBALL" --tag "$NPM_TAG" "--@soulcraft:registry=https://registry.npmjs.org/"
|
||||
rm -rf "$STOREFRONT_TMP"
|
||||
npm publish --tag "$NPM_TAG" "--@soulcraft:registry=https://registry.npmjs.org/"
|
||||
# Brainy is the only PUBLIC @soulcraft package — verify visibility after every publish.
|
||||
npm access get status @soulcraft/brainy "--@soulcraft:registry=https://registry.npmjs.org/" || true
|
||||
# Verify the pair is byte-identical by registry-reported shasum — divergence here
|
||||
# means the storefront leg must be treated as failed, loudly.
|
||||
FORGE_SHA=$(npm view "@soulcraft/brainy@${NEW_VERSION}" dist.shasum "--@soulcraft:registry=${FORGE_NPM_REG}" 2>/dev/null || echo "forge-unavailable")
|
||||
NPMJS_SHA=$(npm view "@soulcraft/brainy@${NEW_VERSION}" dist.shasum "--@soulcraft:registry=https://registry.npmjs.org/" 2>/dev/null || echo "npmjs-unavailable")
|
||||
if [ "$FORGE_SHA" = "$NPMJS_SHA" ]; then
|
||||
echo -e "${GREEN}✅ Published to npmjs — byte-identical pair (shasum ${NPMJS_SHA})${NC}\n"
|
||||
else
|
||||
echo -e "${RED}❌ REGISTRY DIVERGENCE: forge shasum ${FORGE_SHA} != npmjs shasum ${NPMJS_SHA} — investigate before announcing${NC}\n"
|
||||
exit 1
|
||||
fi
|
||||
echo -e "${GREEN}✅ Published to npmjs${NC}\n"
|
||||
|
||||
# Step 11: Release object on the forge (presentational — the tag, CHANGELOG,
|
||||
# and RELEASES.md are the record; this just gives the forge UI a release page).
|
||||
|
|
|
|||
|
|
@ -14,22 +14,7 @@
|
|||
*/
|
||||
|
||||
import type { StorageAdapter, HNSWNounWithMetadata } from '../coreTypes.js'
|
||||
import { parseFieldAddress, readEntityFieldAddress } from '../db/fieldAddressing.js'
|
||||
import type { HNSWNounWithMetadata as AddressedEntity } from '../coreTypes.js'
|
||||
|
||||
/**
|
||||
* Read a user-supplied field name under the one addressing law (sealed
|
||||
* 2026-08-03): bare / `metadata.` = the user's metadata field, `system.<x>` =
|
||||
* the ruled engine scalar, malformed = typed refusal. The aggregation engine
|
||||
* NEVER resolves names any other way — the pre-law resolver made bare
|
||||
* `subtype`/`confidence` read engine scalars, silently shadowing user fields.
|
||||
*/
|
||||
function readAddressed(e: unknown, name: string): unknown {
|
||||
return readEntityFieldAddress(
|
||||
e as AddressedEntity,
|
||||
parseFieldAddress(name, 'entity')
|
||||
)
|
||||
}
|
||||
import { resolveEntityField } from '../coreTypes.js'
|
||||
import type {
|
||||
AggregateDefinition,
|
||||
AggregateGroupState,
|
||||
|
|
@ -110,15 +95,11 @@ function matchesSource(entity: Record<string, unknown>, source: AggregateDefinit
|
|||
// live in the custom bag, so those filters could never match anything.
|
||||
if (source.where && Object.keys(source.where).length > 0) {
|
||||
const e = entity as unknown as HNSWNounWithMetadata
|
||||
for (const [key, condition] of Object.entries(source.where)) {
|
||||
// Evaluate ONE field at a time under a neutral key: the address may be
|
||||
// dotted ('system.subtype'), and the filter evaluator would otherwise
|
||||
// walk dots as a nested path instead of treating the key as an address.
|
||||
const value = readAddressed(e, key)
|
||||
if (!matchesMetadataFilter({ v: value }, { v: condition } as Record<string, unknown>)) {
|
||||
return false
|
||||
}
|
||||
const resolved: Record<string, unknown> = {}
|
||||
for (const key of Object.keys(source.where)) {
|
||||
resolved[key] = resolveEntityField(e, key)
|
||||
}
|
||||
if (!matchesMetadataFilter(resolved, source.where)) return false
|
||||
}
|
||||
|
||||
return true
|
||||
|
|
@ -148,11 +129,11 @@ function computeGroupKeys(
|
|||
|
||||
for (const dim of groupBy) {
|
||||
if (typeof dim === 'string') {
|
||||
const val = readAddressed(e, dim)
|
||||
const val = resolveEntityField(e, dim)
|
||||
const v = val !== undefined && val !== null ? String(val) : '__null__'
|
||||
for (const k of keys) k[dim] = v
|
||||
} else if ('unnest' in dim) {
|
||||
const val = readAddressed(e, dim.field)
|
||||
const val = resolveEntityField(e, dim.field)
|
||||
const raw = Array.isArray(val) ? val : val !== undefined && val !== null ? [val] : []
|
||||
// Distinct elements: an entity with duplicate tags counts once per distinct tag.
|
||||
const elems = Array.from(new Set(raw.map(x => String(x))))
|
||||
|
|
@ -164,7 +145,7 @@ function computeGroupKeys(
|
|||
keys = next
|
||||
} else {
|
||||
// Time-windowed field
|
||||
const val = readAddressed(e, dim.field)
|
||||
const val = resolveEntityField(e, dim.field)
|
||||
const v = typeof val === 'number' ? bucketTimestamp(val, dim.window) : '__null__'
|
||||
for (const k of keys) k[dim.field] = v
|
||||
}
|
||||
|
|
@ -193,7 +174,7 @@ function computeGroupKey(
|
|||
* in metadata are both handled in one place.
|
||||
*/
|
||||
function getNumericField(entity: Record<string, unknown>, field: string): number | undefined {
|
||||
const val = readAddressed(entity as unknown as HNSWNounWithMetadata, field)
|
||||
const val = resolveEntityField(entity as unknown as HNSWNounWithMetadata, field)
|
||||
if (typeof val === 'number' && !isNaN(val)) return val
|
||||
if (typeof val === 'string') {
|
||||
const num = parseFloat(val)
|
||||
|
|
@ -1009,7 +990,7 @@ export class AggregationIndex {
|
|||
// distinctCount tracks distinct values of ANY type (strings, numbers, booleans),
|
||||
// keyed by their string form — NOT numeric-coerced, since its primary use is
|
||||
// categorical (distinct categories / users / tags), not numeric columns.
|
||||
const raw = readAddressed(entity as unknown as HNSWNounWithMetadata, metricDef.field!)
|
||||
const raw = resolveEntityField(entity as unknown as HNSWNounWithMetadata, metricDef.field!)
|
||||
if (raw !== undefined && raw !== null) {
|
||||
if (!state.valueCounts) state.valueCounts = {}
|
||||
const key = String(raw)
|
||||
|
|
@ -1053,7 +1034,7 @@ export class AggregationIndex {
|
|||
state.count = Math.max(0, state.count - 1)
|
||||
state.sum = Math.max(0, state.sum - 1)
|
||||
} else if (metricDef.op === 'distinctCount') {
|
||||
const raw = readAddressed(entity as unknown as HNSWNounWithMetadata, metricDef.field!)
|
||||
const raw = resolveEntityField(entity as unknown as HNSWNounWithMetadata, metricDef.field!)
|
||||
if (raw !== undefined && raw !== null && state.valueCounts) {
|
||||
const key = String(raw)
|
||||
const c = state.valueCounts[key]
|
||||
|
|
|
|||
891
src/brainy.ts
891
src/brainy.ts
File diff suppressed because it is too large
Load diff
|
|
@ -284,12 +284,7 @@ export const STANDARD_ENTITY_FIELDS: ReadonlySet<string> = new Set([
|
|||
'id',
|
||||
'vector',
|
||||
'connections',
|
||||
// 'level' is deliberately ABSENT: it is HNSW plumbing, not an entity field.
|
||||
// Listing it here made every by-name read of a user metadata field called
|
||||
// `level` resolve to the engine's internal node layer instead — a silent
|
||||
// shadow that broke sort/filter/aggregation on a perfectly natural field
|
||||
// name (VENUE-BRAINY-ORDERBY-NOOP). Engine plumbing is invisible to the
|
||||
// query surface; a bare `level` reads `entity.metadata.level`.
|
||||
'level',
|
||||
'type',
|
||||
'subtype',
|
||||
'visibility',
|
||||
|
|
|
|||
67
src/db/db.ts
67
src/db/db.ts
|
|
@ -59,6 +59,10 @@ import type {
|
|||
import type { StorageAdapter } from '../coreTypes.js'
|
||||
import { exportGraph } from './portableGraph.js'
|
||||
import type { ExportSelector, ExportOptions, PortableGraph } from './portableGraph.js'
|
||||
import {
|
||||
splitNounMetadataRecord,
|
||||
splitVerbMetadataRecord
|
||||
} from '../types/reservedFields.js'
|
||||
import { v4 as uuidv4 } from '../universal/uuid.js'
|
||||
import { coerceNewEntityId, resolveEntityId, ORIGINAL_ID_KEY } from '../utils/idNormalization.js'
|
||||
import { EntityNotFoundError } from '../errors/notFound.js'
|
||||
|
|
@ -701,15 +705,23 @@ export class Db<T = any> {
|
|||
for (const op of ops) {
|
||||
switch (op.op) {
|
||||
case 'add': {
|
||||
// Field-addressing law: the metadata bag is the user's, VERBATIM —
|
||||
// no reserved-name lift, no drops. Engine scalars come ONLY from
|
||||
// their dedicated op fields; a bag field named `confidence` is an
|
||||
// ordinary user field, exactly as on the committed write path.
|
||||
const custom = { ...(op.metadata as Record<string, unknown> | undefined) }
|
||||
const confidence = op.confidence
|
||||
const weight = op.weight
|
||||
const subtype = op.subtype
|
||||
const service = op.service
|
||||
// Reserved-field normalization — mirror of the brain.transact()
|
||||
// write path: user-settable fields lift to their dedicated field
|
||||
// (top-level wins), system-managed fields drop, and the entity's
|
||||
// metadata bag carries ONLY custom fields. Speculative views skip
|
||||
// the one-shot warnings — committing the same ops through
|
||||
// `brain.transact()` warns on the real write path.
|
||||
const { reserved, custom } = splitNounMetadataRecord(
|
||||
op.metadata as Record<string, unknown> | undefined
|
||||
)
|
||||
const confidence =
|
||||
op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined)
|
||||
const weight =
|
||||
op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined)
|
||||
const subtype =
|
||||
op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined)
|
||||
const service =
|
||||
op.service ?? (typeof reserved.service === 'string' ? reserved.service : undefined)
|
||||
|
||||
// Id normalization (8.0) — mirror of the committed transact() add
|
||||
// path: a natural key coerces to a STABLE UUID (v5), preserving the
|
||||
|
|
@ -747,12 +759,16 @@ export class Db<T = any> {
|
|||
`with(): entity ${updateId} not found at generation ${this.gen}`
|
||||
)
|
||||
}
|
||||
// Field-addressing law — mirror of the add case: the patch bag is
|
||||
// the user's verbatim; engine scalars only from dedicated op fields.
|
||||
const custom = { ...(op.metadata as Record<string, unknown> | undefined) }
|
||||
const confidence = op.confidence
|
||||
const weight = op.weight
|
||||
const subtype = op.subtype
|
||||
// Same reserved-field normalization as the committed update path.
|
||||
const { reserved, custom } = splitNounMetadataRecord(
|
||||
op.metadata as Record<string, unknown> | undefined
|
||||
)
|
||||
const confidence =
|
||||
op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined)
|
||||
const weight =
|
||||
op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined)
|
||||
const subtype =
|
||||
op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined)
|
||||
const mergedMetadata =
|
||||
op.merge !== false
|
||||
? ({ ...(base.metadata as object), ...custom } as T)
|
||||
|
|
@ -814,14 +830,19 @@ export class Db<T = any> {
|
|||
}
|
||||
if (duplicate) break
|
||||
|
||||
// Field-addressing law — relationship mirror of the add case: the
|
||||
// edge bag is the user's verbatim; engine scalars only from
|
||||
// dedicated op fields.
|
||||
const custom = { ...(op.metadata as Record<string, unknown> | undefined) }
|
||||
const confidence = op.confidence
|
||||
const weight = op.weight
|
||||
const subtype = op.subtype
|
||||
const service = op.service
|
||||
// Reserved-field normalization — relationship mirror of the add
|
||||
// op above (and of the committed relate() path).
|
||||
const { reserved, custom } = splitVerbMetadataRecord(
|
||||
op.metadata as Record<string, unknown> | undefined
|
||||
)
|
||||
const confidence =
|
||||
op.confidence ?? (typeof reserved.confidence === 'number' ? reserved.confidence : undefined)
|
||||
const weight =
|
||||
op.weight ?? (typeof reserved.weight === 'number' ? reserved.weight : undefined)
|
||||
const subtype =
|
||||
op.subtype ?? (typeof reserved.subtype === 'string' ? reserved.subtype : undefined)
|
||||
const service =
|
||||
op.service ?? (typeof reserved.service === 'string' ? reserved.service : undefined)
|
||||
|
||||
const id = uuidv4()
|
||||
overlay.verbs.set(id, {
|
||||
|
|
|
|||
|
|
@ -102,26 +102,12 @@ export interface FactScanBatch {
|
|||
segmentId: string
|
||||
}
|
||||
|
||||
/**
|
||||
* Liveness bound on a scan's FIRST batch (Stage-2 co-freeze, D1 contract):
|
||||
* `batches()` must yield its first batch — or fail loudly — within this many
|
||||
* ms of the first pull. A backlogged or damaged store may be SLOW, but it may
|
||||
* never be SILENT: a consumer awaiting the first batch is otherwise
|
||||
* indistinguishable from a wedge (the exact failure shape a production heal
|
||||
* hit against a generations-backlogged brain).
|
||||
*/
|
||||
export const SCANFACTS_FIRST_BATCH_MS = 10_000
|
||||
|
||||
/** The telemetry a scan OPEN returns (frozen shape). */
|
||||
export interface FactScanHandle {
|
||||
headGeneration: number
|
||||
segmentCount: number
|
||||
approxFactCount: number
|
||||
/**
|
||||
* Ordered batches; a detected gap aborts LOUDLY, never a silent skip.
|
||||
* Liveness contract: the FIRST batch resolves or rejects within
|
||||
* {@link SCANFACTS_FIRST_BATCH_MS} of the first pull — never a silent hang.
|
||||
*/
|
||||
/** Ordered batches; a detected gap aborts LOUDLY, never a silent skip. */
|
||||
batches: () => AsyncGenerator<FactScanBatch>
|
||||
/** Close telemetry — the invariant cross-check, valid after iteration ends. */
|
||||
summary: () => { factsYielded: number; segmentsRead: number }
|
||||
|
|
@ -454,8 +440,6 @@ export class FactLog {
|
|||
toGeneration?: number
|
||||
kinds?: Array<'noun' | 'verb'>
|
||||
batchSize?: number
|
||||
/** Test override for the first-batch liveness bound (default {@link SCANFACTS_FIRST_BATCH_MS}). */
|
||||
firstBatchTimeoutMs?: number
|
||||
}): FactScanHandle {
|
||||
const from = options?.fromGeneration ?? 1
|
||||
const to = options?.toGeneration ?? this.head
|
||||
|
|
@ -530,43 +514,11 @@ export class FactLog {
|
|||
}
|
||||
}
|
||||
|
||||
// Liveness wrapper: the FIRST pull races the contract deadline. Only the
|
||||
// first — the bound is time-to-first-batch (proof the producer is alive),
|
||||
// not per-batch pacing; and it runs only while a pull is actually pending,
|
||||
// so consumer think-time between pulls never counts against the producer.
|
||||
const firstBatchTimeoutMs = options?.firstBatchTimeoutMs ?? SCANFACTS_FIRST_BATCH_MS
|
||||
async function* batchesWithLiveness(this: void): AsyncGenerator<FactScanBatch> {
|
||||
const inner = batches()
|
||||
let timer: NodeJS.Timeout | undefined
|
||||
try {
|
||||
const deadline = new Promise<never>((_, reject) => {
|
||||
timer = setTimeout(
|
||||
() =>
|
||||
reject(
|
||||
new Error(
|
||||
`fact log: scanFacts produced no first batch within ${firstBatchTimeoutMs}ms ` +
|
||||
`(liveness contract) — the store is wedged or unreadably slow; aborting scan LOUDLY ` +
|
||||
`instead of hanging the consumer.`
|
||||
)
|
||||
),
|
||||
firstBatchTimeoutMs
|
||||
)
|
||||
timer.unref?.()
|
||||
})
|
||||
const first = await Promise.race([inner.next(), deadline])
|
||||
if (first.done) return
|
||||
yield first.value
|
||||
} finally {
|
||||
clearTimeout(timer)
|
||||
}
|
||||
yield* inner
|
||||
}
|
||||
|
||||
return {
|
||||
headGeneration: this.head,
|
||||
segmentCount: segments.length + (tailSnapshot.length > 0 ? 1 : 0),
|
||||
approxFactCount,
|
||||
batches: batchesWithLiveness,
|
||||
batches,
|
||||
summary: () => ({ factsYielded, segmentsRead })
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,323 +0,0 @@
|
|||
/**
|
||||
* @module db/fieldAddressing
|
||||
* @description The one field-addressing law for every query surface (find()'s
|
||||
* `where` / `orderBy` / `groupBy`, aggregation `source.where`), ruled
|
||||
* 2026-08-03 after a production incident in which a user metadata field
|
||||
* named `level` was silently shadowed by the engine's internal HNSW node
|
||||
* layer (VENUE-BRAINY-ORDERBY-NOOP — thread id kept verbatim as the audit
|
||||
* key; it names no product):
|
||||
*
|
||||
* 1. A BARE field name addresses the user's metadata field. Always.
|
||||
* No priority resolution, no fallback chain — `orderBy: 'level'`
|
||||
* reads `entity.metadata.level`, full stop.
|
||||
* 2. `system.<field>` addresses an engine scalar, reachable ONLY with the
|
||||
* explicit prefix. The entity map is exactly ten scalars; the relation
|
||||
* map mirrors it with `verb`/`sourceId`/`targetId` as the structural
|
||||
* members.
|
||||
* 3. Engine plumbing (`vector`, `connections`, `level`, `data`, `_rev`) is
|
||||
* INVISIBLE to the query surface in either spelling — `system.level`
|
||||
* refuses; bare `level` is the user's field.
|
||||
* 4. `metadata.<field>` is the explicit spelling of the bare form —
|
||||
* identical semantics on every path.
|
||||
* 5. Anything unresolvable refuses with a TYPED error naming both
|
||||
* candidate spellings — an accepted name either works or refuses;
|
||||
* there is no third state.
|
||||
*
|
||||
* This module is the SINGLE source of truth for the law: parsing, the maps,
|
||||
* and the refusal builders live here so the JS engine, the provider seams,
|
||||
* and the cross-engine conformance suite can never drift on the contract.
|
||||
*/
|
||||
|
||||
import type { HNSWNounWithMetadata, HNSWVerbWithMetadata } from '../coreTypes.js'
|
||||
|
||||
/**
|
||||
* @description The entity-side `system.*` map — EXACTLY the ten engine
|
||||
* scalars David ruled queryable (2026-08-03). Adding a name here is a
|
||||
* cross-engine contract change: the native accelerator's conformance suite
|
||||
* pins this list verbatim, so any edit must ship as a paired release.
|
||||
*/
|
||||
export const SYSTEM_ENTITY_SCALARS: ReadonlySet<string> = new Set([
|
||||
'id',
|
||||
'type',
|
||||
'subtype',
|
||||
'createdAt',
|
||||
'updatedAt',
|
||||
'confidence',
|
||||
'weight',
|
||||
'visibility',
|
||||
'service',
|
||||
'createdBy'
|
||||
])
|
||||
|
||||
/**
|
||||
* @description The relation-side `system.*` map — the verb mirror of
|
||||
* {@link SYSTEM_ENTITY_SCALARS}: `verb`, `sourceId`, `targetId` are the
|
||||
* structural members beside the eight shared scalars. Same one law, same
|
||||
* pairing rule for edits.
|
||||
*/
|
||||
export const SYSTEM_RELATION_SCALARS: ReadonlySet<string> = new Set([
|
||||
'verb',
|
||||
'sourceId',
|
||||
'targetId',
|
||||
'subtype',
|
||||
'createdAt',
|
||||
'updatedAt',
|
||||
'confidence',
|
||||
'weight',
|
||||
'visibility',
|
||||
'service',
|
||||
'createdBy'
|
||||
])
|
||||
|
||||
/**
|
||||
* @description Engine plumbing — never addressable from the query surface in
|
||||
* ANY spelling. `level` is the HNSW node layer (the incident field: listing
|
||||
* it as resolvable shadowed real user data); `data` is the payload container,
|
||||
* not a scalar — content is reached through the content/text-search APIs,
|
||||
* and addressing it as a sortable field would lie about its shape.
|
||||
*/
|
||||
export const PLUMBING_FIELDS: ReadonlySet<string> = new Set([
|
||||
'vector',
|
||||
'connections',
|
||||
'level',
|
||||
'data',
|
||||
'_rev'
|
||||
])
|
||||
|
||||
/** @description Which record kind a field address is being resolved against. */
|
||||
export type FieldAddressKind = 'entity' | 'relation'
|
||||
|
||||
/**
|
||||
* @description A parsed, law-valid field address. `scope` says which side of
|
||||
* the record the name lives on; `field` is the unprefixed name to read.
|
||||
*/
|
||||
export interface FieldAddress {
|
||||
/** 'metadata' = the user's field (bare or `metadata.`-prefixed); 'system' = an engine scalar. */
|
||||
scope: 'metadata' | 'system'
|
||||
/** The field name with any scope prefix removed. */
|
||||
field: string
|
||||
/** The exact spelling the caller used — preserved for error text and telemetry. */
|
||||
raw: string
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a query-surface field name under the one law. Pure and data-blind:
|
||||
* this validates the ADDRESS (spelling + map membership), not whether any
|
||||
* row actually carries the field — data-aware refusals (the did-you-mean
|
||||
* for a bare system-scalar name no row carries) belong to the query layer,
|
||||
* which calls {@link buildUnresolvableMessage} with index knowledge.
|
||||
*
|
||||
* @param raw - The field name as the caller wrote it (`level`,
|
||||
* `metadata.level`, `system.createdAt`, …)
|
||||
* @param kind - Entity or relation resolution (selects the system map)
|
||||
* @returns The parsed {@link FieldAddress}
|
||||
* @throws {InvalidFieldAddressError} for a `system.*` name outside the ruled
|
||||
* map (including every plumbing field) or a malformed spelling — the error
|
||||
* text enumerates the valid system scalars so the fix is in the message.
|
||||
*
|
||||
* @example
|
||||
* parseFieldAddress('level', 'entity') // { scope: 'metadata', field: 'level' }
|
||||
* parseFieldAddress('metadata.level', 'entity') // { scope: 'metadata', field: 'level' }
|
||||
* parseFieldAddress('system.createdAt', 'entity') // { scope: 'system', field: 'createdAt' }
|
||||
* parseFieldAddress('system.level', 'entity') // throws — plumbing is invisible
|
||||
*/
|
||||
export function parseFieldAddress(
|
||||
raw: string,
|
||||
kind: FieldAddressKind
|
||||
): FieldAddress {
|
||||
const systemMap =
|
||||
kind === 'entity' ? SYSTEM_ENTITY_SCALARS : SYSTEM_RELATION_SCALARS
|
||||
|
||||
if (raw.startsWith('system.')) {
|
||||
const field = raw.slice('system.'.length)
|
||||
if (!systemMap.has(field)) {
|
||||
throw new InvalidFieldAddressError(raw, kind, systemMap)
|
||||
}
|
||||
return { scope: 'system', field, raw }
|
||||
}
|
||||
|
||||
if (raw.startsWith('metadata.')) {
|
||||
const field = raw.slice('metadata.'.length)
|
||||
if (field.length === 0) {
|
||||
throw new InvalidFieldAddressError(raw, kind, systemMap)
|
||||
}
|
||||
return { scope: 'metadata', field, raw }
|
||||
}
|
||||
|
||||
if (raw.length === 0) {
|
||||
throw new InvalidFieldAddressError(raw, kind, systemMap)
|
||||
}
|
||||
|
||||
// Bare name = the user's metadata field. Always. Even when the same name
|
||||
// exists in the system map — `confidence` as a bare name is the user's
|
||||
// metadata field named confidence; the engine scalar is system.confidence.
|
||||
return { scope: 'metadata', field: raw, raw }
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the addressed value off an entity. The ONLY sanctioned way a query
|
||||
* surface turns a {@link FieldAddress} into a value — direct property reads
|
||||
* against records re-create the shadow class this module exists to kill.
|
||||
*
|
||||
* @returns The value, or `undefined` when the record does not carry it
|
||||
* (missing values sort LAST in both directions per the ordering contract —
|
||||
* they are never grounds for dropping a row).
|
||||
*/
|
||||
export function readEntityFieldAddress(
|
||||
entity: HNSWNounWithMetadata,
|
||||
address: FieldAddress
|
||||
): unknown {
|
||||
const rec = entity as unknown as Record<string, unknown>
|
||||
const bag =
|
||||
rec.metadata && typeof rec.metadata === 'object'
|
||||
? (rec.metadata as Record<string, unknown>)
|
||||
: null
|
||||
|
||||
if (address.scope === 'system') {
|
||||
// System scalars live at the record's top level, NEVER in the user's
|
||||
// bag — a user field named `confidence` must be unreachable from
|
||||
// system.confidence (and vice versa). Entity views carry the scalars
|
||||
// top-level directly; record-derived views spell the type `noun`.
|
||||
const top = rec[address.field]
|
||||
if (top !== undefined) return top
|
||||
if (address.field === 'type') return rec.noun
|
||||
return undefined
|
||||
}
|
||||
|
||||
// User scope: the bag IS the user's namespace, authoritative — EVERY name
|
||||
// reads from it, engine spellings included (`bag.confidence` is the user's
|
||||
// confidence field under the field-addressing law).
|
||||
if (bag) return bag[address.field]
|
||||
|
||||
// No bag at all: a LEGACY flat record (pre-nested-bag storage). Its keys
|
||||
// matching system/plumbing names are the ENGINE's — the pre-law write door
|
||||
// refused user colliders — so a bare system name reads as ABSENT rather
|
||||
// than resurrecting the shadow this module exists to kill. Same for the
|
||||
// legacy 'noun' spelling.
|
||||
if (
|
||||
SYSTEM_ENTITY_SCALARS.has(address.field) ||
|
||||
PLUMBING_FIELDS.has(address.field) ||
|
||||
address.field === 'noun'
|
||||
) {
|
||||
return undefined
|
||||
}
|
||||
return rec[address.field]
|
||||
}
|
||||
|
||||
/**
|
||||
* Relation twin of {@link readEntityFieldAddress}. The stored flat record
|
||||
* keys the relation type under `verb`; public Relation shapes may carry it
|
||||
* as `type` — both spellings of the record are read, the ADDRESS is always
|
||||
* `system.verb`.
|
||||
*/
|
||||
export function readRelationFieldAddress(
|
||||
verb: HNSWVerbWithMetadata,
|
||||
address: FieldAddress
|
||||
): unknown {
|
||||
if (address.scope === 'system') {
|
||||
const rec = verb as unknown as Record<string, unknown>
|
||||
if (address.field === 'verb') return rec.verb ?? rec.type
|
||||
return rec[address.field]
|
||||
}
|
||||
return verb.metadata?.[address.field]
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the ruled did-you-mean refusal text for a bare name that resolved to
|
||||
* metadata but is UNKNOWN to the index — the data-aware half of the law,
|
||||
* called by the query layer once it has consulted the known-field set:
|
||||
*
|
||||
* "no metadata field 'createdAt' — did you mean system.createdAt or
|
||||
* metadata.createdAt?"
|
||||
*
|
||||
* When the bare name is NOT a system scalar the system candidate is omitted
|
||||
* (there is only one thing the caller could have meant; the refusal exists
|
||||
* because refusing beats silently sorting nothing).
|
||||
*/
|
||||
export function buildUnresolvableMessage(
|
||||
raw: string,
|
||||
kind: FieldAddressKind
|
||||
): string {
|
||||
const systemMap =
|
||||
kind === 'entity' ? SYSTEM_ENTITY_SCALARS : SYSTEM_RELATION_SCALARS
|
||||
if (systemMap.has(raw)) {
|
||||
return (
|
||||
`no metadata field '${raw}' — did you mean system.${raw} or metadata.${raw}? ` +
|
||||
`(bare names always address your metadata; engine fields need the system. prefix)`
|
||||
)
|
||||
}
|
||||
return (
|
||||
`no metadata field '${raw}' on this store — nothing carries it, so an ordered or ` +
|
||||
`filtered read against it cannot mean anything. Spell it metadata.${raw} once the ` +
|
||||
`field exists, or check the field name (system.${raw} is NOT valid — '${raw}' is ` +
|
||||
`not one of the engine's system scalars).`
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Refusal for a syntactically valid address that resolves to
|
||||
* NOTHING — a bare name no user field carries. Carries the did-you-mean
|
||||
* (both candidate spellings when the name collides with a system scalar) so
|
||||
* the fix ships inside the error. Thrown by the query layer with index
|
||||
* knowledge, never by the pure parser.
|
||||
*/
|
||||
export class UnresolvableFieldError extends Error {
|
||||
public readonly raw: string
|
||||
public readonly kind: FieldAddressKind
|
||||
|
||||
constructor(raw: string, kind: FieldAddressKind, messageOverride?: string) {
|
||||
super(messageOverride ?? buildUnresolvableMessage(raw, kind))
|
||||
this.name = 'UnresolvableFieldError'
|
||||
this.raw = raw
|
||||
this.kind = kind
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Refusal for a malformed or out-of-map field ADDRESS —
|
||||
* `system.<anything-not-in-the-map>` (including all plumbing), an empty
|
||||
* name, or a bare `metadata.` prefix. The message carries the full valid
|
||||
* system map so the fix never needs a docs lookup.
|
||||
*/
|
||||
export class InvalidFieldAddressError extends UnresolvableFieldError {
|
||||
constructor(raw: string, kind: FieldAddressKind, systemMap: ReadonlySet<string>) {
|
||||
const valid = [...systemMap].map((f) => `system.${f}`).join(', ')
|
||||
super(
|
||||
raw,
|
||||
kind,
|
||||
`'${raw}' is not an addressable ${kind} field. Bare names address your own ` +
|
||||
`metadata fields; engine fields are exactly: ${valid}. Engine plumbing ` +
|
||||
`(vector, connections, level, data, _rev) is not part of the query surface.`
|
||||
)
|
||||
this.name = 'InvalidFieldAddressError'
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
/**
|
||||
* @description Refusal for a find() option that is accepted by the type
|
||||
* surface but NOT implemented — an accepted option must work or refuse;
|
||||
* accepted-and-ignored died as a class (sealed 2026-08-03). Names the
|
||||
* option and the honest state so nobody discovers a no-op by measurement.
|
||||
*/
|
||||
export class UnsupportedFindOptionError extends Error {
|
||||
public readonly option: string
|
||||
|
||||
constructor(option: string) {
|
||||
super(
|
||||
`find() option '${option}' is not implemented — it used to be silently ` +
|
||||
`ignored, which read as working. Remove it from the call (or track the ` +
|
||||
`feature request); it will be honored or refused, never swallowed.`
|
||||
)
|
||||
this.name = 'UnsupportedFindOptionError'
|
||||
this.option = option
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description The capability signal both engines' conformance suites arm on
|
||||
* (never a version guess): its presence at the package root means the one
|
||||
* field-addressing law is LIVE on every query surface — bare = user metadata,
|
||||
* `system.*` = the ruled scalars, plumbing invisible, refusals typed.
|
||||
*/
|
||||
export const FIELD_ADDRESSING_CAPABILITY = 'field-addressing/v1'
|
||||
|
|
@ -1,459 +0,0 @@
|
|||
/**
|
||||
* @module db/generationSegments
|
||||
* @description The generation-segment store — Stage-2 D1+D3+repacking's file
|
||||
* format (co-frozen 2026-07-19; design: the d1-d3-repacking spec).
|
||||
*
|
||||
* Packs CONSECUTIVE cold generations' record-sets (before-images + delta)
|
||||
* into append-once segment files with derived sidecar indexes, so history
|
||||
* scales in SEGMENTS (tens) instead of FILES-PER-GENERATION (hundreds of
|
||||
* thousands), and cold-open reads ONE manifest instead of listing the
|
||||
* backlog. Layout under `_generations/segments/`:
|
||||
*
|
||||
* - `seg-<firstGen, zero-padded 20>.bgs` — magic "BGS1", then one frame per
|
||||
* generation: `u32 payloadLen | u32 crc32c | msgpack payload`. Payload is
|
||||
* POSITIONAL: `[generation, timestamp, delta, records[], flags]` with
|
||||
* records `[kindByte, id, record]`. `flags` reserves encoding evolution
|
||||
* (bit 0 = compressed payload — v1 always 0; a future writer upgrade,
|
||||
* never a format break). Sealed segments are IMMUTABLE — the fact log's
|
||||
* own law, generalized.
|
||||
* - `seg-<firstGen>.idx` — DERIVED sidecar (msgpack): per-generation frame
|
||||
* offsets (point reads = one ranged read, never a listing) + per-id
|
||||
* generation postings (per-id chain rebuilds read only what they need).
|
||||
* Corrupt/missing → rebuilt from its segment in one sequential read,
|
||||
* loudly.
|
||||
* - `manifest.json` — the segment catalogue + `compactedBelow` (D3's
|
||||
* horizon marker). Cold-open reads THIS; the packed backlog is never
|
||||
* listed.
|
||||
*
|
||||
* D3 semantics carried here: bounded-retention reclaim drops WHOLE segments
|
||||
* at boundaries (O(1) per segment, no rewrite); under the archival profile
|
||||
* (`retention: 'all'`) nothing here is ever dropped — folding is the only
|
||||
* transform (re-representation, never deletion).
|
||||
*/
|
||||
|
||||
import { encode as msgpackEncode, decode as msgpackDecode } from '@msgpack/msgpack'
|
||||
import { crc32c } from '../utils/crc32c.js'
|
||||
import type { FactLogStorage } from './factLog.js'
|
||||
import { prodLog } from '../utils/logger.js'
|
||||
|
||||
/** Directory for segment files + manifest, under the generations prefix. */
|
||||
export const SEGMENTS_PREFIX = '_generations/segments'
|
||||
|
||||
/** Target sealed-segment size (co-freeze proposal; tunable on evidence). */
|
||||
export const SEGMENT_TARGET_BYTES = 64 * 1024 * 1024
|
||||
|
||||
const MAGIC = new TextEncoder().encode('BGS1')
|
||||
const FRAME_PREFIX_BYTES = 8 // u32 payloadLen + u32 crc32c
|
||||
const MANIFEST_PATH = `${SEGMENTS_PREFIX}/manifest.json`
|
||||
|
||||
/** One generation's fold input — exactly what the live tier holds for it. */
|
||||
export interface FoldGeneration {
|
||||
generation: number
|
||||
timestamp: number
|
||||
/** The tx.json delta object, carried verbatim. */
|
||||
delta: unknown
|
||||
/** The before-image record-set (empty for record-less generations). */
|
||||
records: Array<{ kind: 'noun' | 'verb'; id: string; record: unknown }>
|
||||
}
|
||||
|
||||
/** Manifest entry for one sealed segment. */
|
||||
export interface SegmentMeta {
|
||||
file: string
|
||||
firstGeneration: number
|
||||
lastGeneration: number
|
||||
frames: number
|
||||
bytes: number
|
||||
/** crc32c of the full segment byte stream — the digest chain's link. */
|
||||
checksum: number
|
||||
}
|
||||
|
||||
interface SegmentManifest {
|
||||
version: 1
|
||||
compactedBelow: number
|
||||
segments: SegmentMeta[]
|
||||
}
|
||||
|
||||
interface SidecarIndex {
|
||||
version: 1
|
||||
/** [generation, frameOffset, frameLen] ascending by generation. */
|
||||
generations: Array<[number, number, number]>
|
||||
/** `${kindByte}:${id}` → ascending generations holding a record for it. */
|
||||
ids: Record<string, number[]>
|
||||
}
|
||||
|
||||
const segmentFileName = (firstGeneration: number): string =>
|
||||
`seg-${String(firstGeneration).padStart(20, '0')}.bgs`
|
||||
const sidecarFileName = (firstGeneration: number): string =>
|
||||
`seg-${String(firstGeneration).padStart(20, '0')}.idx`
|
||||
|
||||
/**
|
||||
* The generation-segment store. Owns the packed tier ONLY — the live
|
||||
* per-generation tier and the routing between tiers belong to
|
||||
* `GenerationStore`. All mutating entry points here are called under the
|
||||
* generation store's commit mutex.
|
||||
*/
|
||||
export class GenerationSegmentStore {
|
||||
private readonly storage: FactLogStorage
|
||||
private manifest: SegmentManifest = { version: 1, compactedBelow: 0, segments: [] }
|
||||
/** Sidecar cache — segments are immutable, so entries never invalidate. */
|
||||
private readonly sidecars = new Map<string, SidecarIndex>()
|
||||
|
||||
constructor(storage: FactLogStorage) {
|
||||
this.storage = storage
|
||||
}
|
||||
|
||||
/** Load the manifest (ONE read — never a directory listing). */
|
||||
async open(): Promise<void> {
|
||||
const raw = (await this.storage.readRawObject(MANIFEST_PATH)) as SegmentManifest | null
|
||||
if (raw) {
|
||||
if (raw.version !== 1) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] manifest version ${String(raw.version)} is newer than this ` +
|
||||
`engine understands — refusing to serve partial history. Upgrade the engine.`
|
||||
)
|
||||
}
|
||||
this.manifest = raw
|
||||
}
|
||||
}
|
||||
|
||||
/** The packed tier's catalogue (ascending, immutable snapshot). */
|
||||
segments(): readonly SegmentMeta[] {
|
||||
return this.manifest.segments
|
||||
}
|
||||
|
||||
/** D3's horizon marker: generations below this were reclaimed (bounded profiles only). */
|
||||
compactedBelow(): number {
|
||||
return this.manifest.compactedBelow
|
||||
}
|
||||
|
||||
/** The covering sealed segment for `gen`, or null if it lives outside the packed tier. */
|
||||
private coveringSegment(gen: number): SegmentMeta | null {
|
||||
// Manifest is ascending and ranges never overlap — binary search.
|
||||
const segs = this.manifest.segments
|
||||
let lo = 0
|
||||
let hi = segs.length - 1
|
||||
while (lo <= hi) {
|
||||
const mid = (lo + hi) >> 1
|
||||
const s = segs[mid]
|
||||
if (gen < s.firstGeneration) hi = mid - 1
|
||||
else if (gen > s.lastGeneration) lo = mid + 1
|
||||
else return s
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
/** True when `gen` is packed (readable from this tier). */
|
||||
hasGeneration(gen: number): boolean {
|
||||
return this.coveringSegment(gen) !== null
|
||||
}
|
||||
|
||||
/**
|
||||
* Fold consecutive generations into ONE new sealed segment + sidecar and
|
||||
* append it to the manifest atomically. Caller guarantees: `gens` is
|
||||
* ascending, contiguous with the packed tier (first = last packed + 1 when
|
||||
* segments exist), and already durable in the live tier. Crash between the
|
||||
* segment write and the caller's live-tier delete leaves a DUPLICATE
|
||||
* representation — resolved live-tier-wins by the reader; never a gap.
|
||||
*/
|
||||
async fold(gens: FoldGeneration[]): Promise<SegmentMeta> {
|
||||
if (gens.length === 0) {
|
||||
throw new Error('[GenerationSegments] fold() requires at least one generation')
|
||||
}
|
||||
for (let i = 1; i < gens.length; i++) {
|
||||
if (gens[i].generation <= gens[i - 1].generation) {
|
||||
throw new Error('[GenerationSegments] fold() input must be strictly ascending')
|
||||
}
|
||||
}
|
||||
const last = this.manifest.segments[this.manifest.segments.length - 1]
|
||||
if (last && gens[0].generation <= last.lastGeneration) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] fold() overlaps the packed tier: ${gens[0].generation} ≤ ` +
|
||||
`sealed ${last.lastGeneration} — segments are immutable, never rewritten`
|
||||
)
|
||||
}
|
||||
|
||||
const first = gens[0].generation
|
||||
const file = segmentFileName(first)
|
||||
const sidecar: SidecarIndex = { version: 1, generations: [], ids: {} }
|
||||
|
||||
// Encode all frames, tracking offsets for the sidecar.
|
||||
const parts: Uint8Array[] = [MAGIC]
|
||||
let offset = MAGIC.length
|
||||
for (const g of gens) {
|
||||
const payload = msgpackEncode([
|
||||
g.generation,
|
||||
g.timestamp,
|
||||
g.delta,
|
||||
g.records.map((r) => [r.kind === 'noun' ? 0 : 1, r.id, r.record]),
|
||||
0 // flags: v1 = uncompressed
|
||||
])
|
||||
const frame = new Uint8Array(FRAME_PREFIX_BYTES + payload.length)
|
||||
const view = new DataView(frame.buffer)
|
||||
view.setUint32(0, payload.length, true)
|
||||
view.setUint32(4, crc32c(payload), true)
|
||||
frame.set(payload, FRAME_PREFIX_BYTES)
|
||||
sidecar.generations.push([g.generation, offset, frame.length])
|
||||
for (const r of g.records) {
|
||||
const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}`
|
||||
;(sidecar.ids[key] ??= []).push(g.generation)
|
||||
}
|
||||
parts.push(frame)
|
||||
offset += frame.length
|
||||
}
|
||||
const total = parts.reduce((n, p) => n + p.length, 0)
|
||||
const bytes = new Uint8Array(total)
|
||||
let at = 0
|
||||
for (const p of parts) {
|
||||
bytes.set(p, at)
|
||||
at += p.length
|
||||
}
|
||||
|
||||
const meta: SegmentMeta = {
|
||||
file,
|
||||
firstGeneration: first,
|
||||
lastGeneration: gens[gens.length - 1].generation,
|
||||
frames: gens.length,
|
||||
bytes: total,
|
||||
checksum: crc32c(bytes)
|
||||
}
|
||||
|
||||
// Durability order: segment + sidecar fsync'd BEFORE the manifest names
|
||||
// them (a crash before the manifest = invisible orphan files, harmless);
|
||||
// manifest last, atomically.
|
||||
const segPath = `${SEGMENTS_PREFIX}/${file}`
|
||||
const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(first)}`
|
||||
await this.storage.writeRawBytes(segPath, bytes)
|
||||
await this.storage.writeRawBytes(idxPath, msgpackEncode(sidecar))
|
||||
await this.storage.syncRawObjects([segPath, idxPath])
|
||||
const next: SegmentManifest = {
|
||||
...this.manifest,
|
||||
segments: [...this.manifest.segments, meta]
|
||||
}
|
||||
await this.storage.writeRawObject(MANIFEST_PATH, next)
|
||||
await this.storage.syncRawObjects([MANIFEST_PATH])
|
||||
this.manifest = next
|
||||
this.sidecars.set(file, sidecar)
|
||||
return meta
|
||||
}
|
||||
|
||||
/** Load (or rebuild, loudly) a segment's sidecar. */
|
||||
private async sidecarFor(meta: SegmentMeta): Promise<SidecarIndex> {
|
||||
const cached = this.sidecars.get(meta.file)
|
||||
if (cached) return cached
|
||||
const idxPath = `${SEGMENTS_PREFIX}/${sidecarFileName(meta.firstGeneration)}`
|
||||
const raw = await this.storage.readRawBytes(idxPath)
|
||||
if (raw) {
|
||||
try {
|
||||
const idx = msgpackDecode(raw) as SidecarIndex
|
||||
if (idx.version === 1) {
|
||||
this.sidecars.set(meta.file, idx)
|
||||
return idx
|
||||
}
|
||||
} catch {
|
||||
// fall through to rebuild
|
||||
}
|
||||
}
|
||||
// Sidecars are DERIVED: rebuild from the segment, loudly — never serve
|
||||
// wrong offsets silently.
|
||||
prodLog.warn(
|
||||
`[GenerationSegments] sidecar for ${meta.file} missing or unreadable — rebuilding from the segment`
|
||||
)
|
||||
const rebuilt = await this.rebuildSidecar(meta)
|
||||
await this.storage.writeRawBytes(idxPath, msgpackEncode(rebuilt))
|
||||
this.sidecars.set(meta.file, rebuilt)
|
||||
return rebuilt
|
||||
}
|
||||
|
||||
/** One sequential read of the segment → a fresh sidecar. Verifies every frame CRC. */
|
||||
private async rebuildSidecar(meta: SegmentMeta): Promise<SidecarIndex> {
|
||||
const frames = await this.readAllFrames(meta)
|
||||
const idx: SidecarIndex = { version: 1, generations: [], ids: {} }
|
||||
for (const f of frames) {
|
||||
idx.generations.push([f.generation, f.offset, f.frameLen])
|
||||
for (const r of f.records) {
|
||||
const key = `${r.kind === 'noun' ? 0 : 1}:${r.id}`
|
||||
;(idx.ids[key] ??= []).push(f.generation)
|
||||
}
|
||||
}
|
||||
return idx
|
||||
}
|
||||
|
||||
private decodeFrame(
|
||||
payload: Uint8Array
|
||||
): { generation: number; timestamp: number; delta: unknown; records: FoldGeneration['records'] } {
|
||||
const [generation, timestamp, delta, rawRecords] = msgpackDecode(payload) as [
|
||||
number,
|
||||
number,
|
||||
unknown,
|
||||
Array<[number, string, unknown]>,
|
||||
number
|
||||
]
|
||||
return {
|
||||
generation,
|
||||
timestamp,
|
||||
delta,
|
||||
records: rawRecords.map(([kindByte, id, record]) => ({
|
||||
kind: kindByte === 0 ? ('noun' as const) : ('verb' as const),
|
||||
id,
|
||||
record
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
private async readAllFrames(meta: SegmentMeta): Promise<
|
||||
Array<ReturnType<GenerationSegmentStore['decodeFrame']> & { offset: number; frameLen: number }>
|
||||
> {
|
||||
const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`)
|
||||
if (!bytes) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] sealed segment ${meta.file} is MISSING — packed history is damaged; ` +
|
||||
`refusing to continue silently`
|
||||
)
|
||||
}
|
||||
const out: Array<ReturnType<GenerationSegmentStore['decodeFrame']> & { offset: number; frameLen: number }> = []
|
||||
let at = MAGIC.length
|
||||
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength)
|
||||
while (at + FRAME_PREFIX_BYTES <= bytes.length) {
|
||||
const payloadLen = view.getUint32(at, true)
|
||||
const crc = view.getUint32(at + 4, true)
|
||||
const payload = bytes.subarray(at + FRAME_PREFIX_BYTES, at + FRAME_PREFIX_BYTES + payloadLen)
|
||||
if (payload.length !== payloadLen || crc32c(payload) !== crc) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] frame CRC mismatch in ${meta.file} at offset ${at} — ` +
|
||||
`packed history is damaged; refusing to serve it`
|
||||
)
|
||||
}
|
||||
out.push({ ...this.decodeFrame(payload), offset: at, frameLen: FRAME_PREFIX_BYTES + payloadLen })
|
||||
at += FRAME_PREFIX_BYTES + payloadLen
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
/** Read one packed generation's frame via its sidecar offset (one ranged read). */
|
||||
private async readFrame(
|
||||
gen: number
|
||||
): Promise<ReturnType<GenerationSegmentStore['decodeFrame']> | null> {
|
||||
const meta = this.coveringSegment(gen)
|
||||
if (!meta) return null
|
||||
const idx = await this.sidecarFor(meta)
|
||||
// generations ascending → binary search.
|
||||
const gens = idx.generations
|
||||
let lo = 0
|
||||
let hi = gens.length - 1
|
||||
while (lo <= hi) {
|
||||
const mid = (lo + hi) >> 1
|
||||
if (gens[mid][0] < gen) lo = mid + 1
|
||||
else if (gens[mid][0] > gen) hi = mid - 1
|
||||
else {
|
||||
const [, offset, frameLen] = gens[mid]
|
||||
const bytes = await this.storage.readRawBytes(`${SEGMENTS_PREFIX}/${meta.file}`)
|
||||
if (!bytes) {
|
||||
throw new Error(`[GenerationSegments] sealed segment ${meta.file} is MISSING`)
|
||||
}
|
||||
const frame = bytes.subarray(offset, offset + frameLen)
|
||||
const view = new DataView(frame.buffer, frame.byteOffset, frame.byteLength)
|
||||
const payloadLen = view.getUint32(0, true)
|
||||
const crc = view.getUint32(4, true)
|
||||
const payload = frame.subarray(FRAME_PREFIX_BYTES, FRAME_PREFIX_BYTES + payloadLen)
|
||||
if (payload.length !== payloadLen || crc32c(payload) !== crc) {
|
||||
throw new Error(
|
||||
`[GenerationSegments] frame CRC mismatch for generation ${gen} in ${meta.file} — ` +
|
||||
`packed history is damaged; refusing to serve it`
|
||||
)
|
||||
}
|
||||
return this.decodeFrame(payload)
|
||||
}
|
||||
}
|
||||
// In the covering range but not present: the packed tier is dense by
|
||||
// construction (fold packs every generation it is handed, including
|
||||
// record-less ones) — absence inside a sealed range is damage.
|
||||
throw new Error(
|
||||
`[GenerationSegments] generation ${gen} is inside sealed segment ${meta.file}'s declared ` +
|
||||
`range but has no frame — packed history is damaged`
|
||||
)
|
||||
}
|
||||
|
||||
/** The packed tier's delta for `gen` (null = not packed). */
|
||||
async readDelta(gen: number): Promise<{ delta: unknown; timestamp: number } | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
return frame ? { delta: frame.delta, timestamp: frame.timestamp } : null
|
||||
}
|
||||
|
||||
/** The packed tier's full record-set for `gen` (null = not packed). */
|
||||
async readRecords(gen: number): Promise<FoldGeneration['records'] | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
return frame ? frame.records : null
|
||||
}
|
||||
|
||||
/** One packed before-image (null = not packed OR no record for the id in that generation). */
|
||||
async readRecord(gen: number, kind: 'noun' | 'verb', id: string): Promise<unknown | null> {
|
||||
const frame = await this.readFrame(gen)
|
||||
if (!frame) return null
|
||||
const hit = frame.records.find((r) => r.kind === kind && r.id === id)
|
||||
return hit ? hit.record : null
|
||||
}
|
||||
|
||||
/**
|
||||
* D3 reclaim: drop WHOLE segments whose lastGeneration < `belowGeneration`
|
||||
* and bump `compactedBelow`. Partial segments are never dropped — the
|
||||
* boundary waits. NEVER called under the archival profile (the caller
|
||||
* enforces retention semantics; this method only executes boundary drops).
|
||||
*/
|
||||
async dropSegmentsBelow(belowGeneration: number): Promise<{ dropped: number; compactedBelow: number }> {
|
||||
const keep: SegmentMeta[] = []
|
||||
const drop: SegmentMeta[] = []
|
||||
for (const s of this.manifest.segments) {
|
||||
;(s.lastGeneration < belowGeneration ? drop : keep).push(s)
|
||||
}
|
||||
if (drop.length === 0) {
|
||||
return { dropped: 0, compactedBelow: this.manifest.compactedBelow }
|
||||
}
|
||||
const compactedBelow = Math.max(
|
||||
this.manifest.compactedBelow,
|
||||
drop[drop.length - 1].lastGeneration + 1
|
||||
)
|
||||
// Manifest first (the drop is authoritative once named), then bytes —
|
||||
// a crash between leaves orphan segment files invisible to the manifest,
|
||||
// harmless and re-collectable.
|
||||
const next: SegmentManifest = { ...this.manifest, compactedBelow, segments: keep }
|
||||
await this.storage.writeRawObject(MANIFEST_PATH, next)
|
||||
await this.storage.syncRawObjects([MANIFEST_PATH])
|
||||
this.manifest = next
|
||||
for (const s of drop) {
|
||||
await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${s.file}`)
|
||||
await this.storage.deleteRawObject(`${SEGMENTS_PREFIX}/${sidecarFileName(s.firstGeneration)}`)
|
||||
this.sidecars.delete(s.file)
|
||||
}
|
||||
return { dropped: drop.length, compactedBelow }
|
||||
}
|
||||
|
||||
/**
|
||||
* D8 rider — the packed portion of `generationDigest(g)`: a deterministic
|
||||
* crc32c chain over sealed-segment checksums fully below `g`, plus the
|
||||
* frame CRC of `g`'s own frame when `g` is mid-segment. O(segments), not
|
||||
* O(generations); identical history ⇒ identical digest on any machine.
|
||||
* The live-tier portion is composed by the caller.
|
||||
*/
|
||||
async digestThroughPacked(g: number): Promise<number | null> {
|
||||
let digest = 0
|
||||
let covered = false
|
||||
for (const s of this.manifest.segments) {
|
||||
if (s.lastGeneration <= g) {
|
||||
digest = crc32c(new TextEncoder().encode(`${digest}:${s.checksum}`))
|
||||
if (s.lastGeneration === g) covered = true
|
||||
} else if (s.firstGeneration <= g) {
|
||||
// g is mid-segment: chain the partial prefix via g's frame CRC.
|
||||
const frame = await this.readFrame(g)
|
||||
if (frame === null) return null
|
||||
const idx = await this.sidecarFor(s)
|
||||
const upTo = idx.generations.filter(([gen]) => gen <= g)
|
||||
for (const [gen, offset, frameLen] of upTo) {
|
||||
digest = crc32c(new TextEncoder().encode(`${digest}:${gen}:${offset}:${frameLen}`))
|
||||
}
|
||||
covered = true
|
||||
break
|
||||
}
|
||||
}
|
||||
return covered || this.manifest.segments.length > 0 ? digest : null
|
||||
}
|
||||
}
|
||||
|
|
@ -46,8 +46,6 @@ import type {
|
|||
TxLogEntry
|
||||
} from './types.js'
|
||||
import { FactLog, storageSupportsFactLog, type CommitFact, type FactOp } from './factLog.js'
|
||||
import { GenerationSegmentStore, type FoldGeneration } from './generationSegments.js'
|
||||
import { crc32c } from '../utils/crc32c.js'
|
||||
|
||||
/**
|
||||
* The byte-identical before-images of every id a commit touches, read UNDER
|
||||
|
|
@ -268,21 +266,6 @@ export class GenerationStore {
|
|||
*/
|
||||
private historyBytesTotal: number | null = null
|
||||
|
||||
/**
|
||||
* The packed tier (D1+D3): sealed segments holding folded cold
|
||||
* generations. Null until {@link open} wires it (and on storage adapters
|
||||
* without raw-byte primitives — the live tier then carries everything,
|
||||
* exactly as before the packed tier existed).
|
||||
*/
|
||||
private segments: GenerationSegmentStore | null = null
|
||||
|
||||
/**
|
||||
* Live-tier window: generations newer than `committed - REPACK_LIVE_WINDOW`
|
||||
* are never folded — the hot tail stays in the per-generation layout the
|
||||
* write path owns. Matches the resident chain window's scale.
|
||||
*/
|
||||
static readonly REPACK_LIVE_WINDOW = 1024
|
||||
|
||||
/**
|
||||
* Model-B per-write group-commit — the in-memory PENDING tier.
|
||||
*
|
||||
|
|
@ -450,33 +433,6 @@ export class GenerationStore {
|
|||
this.factLog = null
|
||||
}
|
||||
|
||||
// PACKED TIER (D1+D3): same capability gate as the fact log. Opening
|
||||
// reads ONE manifest — never a listing of the packed backlog — and seeds
|
||||
// committedRanges with the sealed ranges so packed generations resolve
|
||||
// exactly like live ones.
|
||||
if (storageSupportsFactLog(this.storage)) {
|
||||
this.segments = new GenerationSegmentStore(this.storage)
|
||||
await this.segments.open()
|
||||
const packedRanges = this.segments
|
||||
.segments()
|
||||
.map((s): [number, number] => [s.firstGeneration, Math.min(s.lastGeneration, this.committed)])
|
||||
.filter(([lo, hi]) => lo <= hi)
|
||||
if (packedRanges.length > 0) {
|
||||
// Merge packed (older) + live (newer) interval sets — both ascending;
|
||||
// coalesce adjacency so range arithmetic stays interval-exact.
|
||||
const merged: Array<[number, number]> = []
|
||||
for (const r of [...packedRanges, ...this.committedRanges].sort((a, b) => a[0] - b[0])) {
|
||||
const last = merged[merged.length - 1]
|
||||
if (last && r[0] <= last[1] + 1) last[1] = Math.max(last[1], r[1])
|
||||
else merged.push([r[0], r[1]])
|
||||
}
|
||||
this.committedRanges = merged
|
||||
}
|
||||
this.horizonGen = Math.max(this.horizonGen, this.segments.compactedBelow() - 1)
|
||||
} else {
|
||||
this.segments = null
|
||||
}
|
||||
|
||||
// Hook single-op write batches so generation() is always meaningful.
|
||||
// Suppressed while a transact batch executes (the batch is ONE generation).
|
||||
if (!options?.readOnly) {
|
||||
|
|
@ -544,51 +500,6 @@ export class GenerationStore {
|
|||
* deltas (cache-bounded reads).
|
||||
* @returns Counts, bytes, generation range, and the compaction horizon.
|
||||
*/
|
||||
/**
|
||||
* @description D8 (gate-to-generation provenance): a deterministic content
|
||||
* digest of the generation log THROUGH `g` — identical history ⇒ identical
|
||||
* digest on any machine; any divergence (different records, different
|
||||
* order, reclaimed range) ⇒ different digest. Composed from the packed
|
||||
* tier's sealed-segment checksum chain (O(segments)) plus the live tier's
|
||||
* per-generation delta digests (O(live window at most)). Release gates pin
|
||||
* {generation, digest} and verify both at execution time.
|
||||
* @param g - The generation to digest through (≤ committed).
|
||||
* @returns A hex digest string, stable across reopen and repacking states
|
||||
* ONLY for fully-packed prefixes — repacking changes representation, so
|
||||
* the composed digest is defined over CONTENT: live-tier gens hash their
|
||||
* delta + record ids, packed gens hash via frame CRCs. A gate should pin
|
||||
* after a repack pass for long-term stability, or re-pin on repack.
|
||||
*/
|
||||
async generationDigest(g: number): Promise<string> {
|
||||
if (!Number.isInteger(g) || g < 1 || g > this.committed) {
|
||||
throw new RangeError(
|
||||
`generationDigest(): generation ${g} is out of range [1, ${this.committed}]`
|
||||
)
|
||||
}
|
||||
if (g <= this.horizonGen) {
|
||||
throw new GenerationCompactedError(g, this.horizonGen)
|
||||
}
|
||||
let digest = 0
|
||||
const enc = new TextEncoder()
|
||||
if (this.segments) {
|
||||
const packed = await this.segments.digestThroughPacked(g)
|
||||
if (packed !== null) digest = packed
|
||||
}
|
||||
// Live-tier composition: every committed gen ≤ g not covered by a sealed
|
||||
// segment hashes its delta content in ascending order.
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
if (gen > g) break
|
||||
if (this.segments?.hasGeneration(gen)) continue
|
||||
const delta = await this.getDelta(gen)
|
||||
digest = crc32c(
|
||||
enc.encode(
|
||||
`${digest}:${gen}:${delta.timestamp}:${[...delta.nouns].sort().join(',')}:${[...delta.verbs].sort().join(',')}`
|
||||
)
|
||||
)
|
||||
}
|
||||
return digest.toString(16).padStart(8, '0')
|
||||
}
|
||||
|
||||
async historyStats(): Promise<{
|
||||
generations: number
|
||||
bytes: number
|
||||
|
|
@ -627,17 +538,14 @@ export class GenerationStore {
|
|||
try {
|
||||
paths = await this.storage.listRawObjects(`${GENERATIONS_PREFIX}/${gen}/prev`)
|
||||
} catch {
|
||||
paths = []
|
||||
return []
|
||||
}
|
||||
const records: GenerationRecord[] = []
|
||||
for (const p of paths) {
|
||||
const record = (await this.storage.readRawObject(p)) as GenerationRecord | null
|
||||
if (record) records.push(record)
|
||||
}
|
||||
if (records.length > 0) return records
|
||||
// Two-tier: folded generations serve their record-set from the segment.
|
||||
const packed = await this.segments?.readRecords(gen)
|
||||
return packed ? (packed.map((r) => r.record) as GenerationRecord[]) : []
|
||||
return records
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -1875,15 +1783,9 @@ export class GenerationStore {
|
|||
if (pending) {
|
||||
return (kind === 'noun' ? pending.nouns : pending.verbs).get(id) ?? null
|
||||
}
|
||||
const live = (await this.storage.readRawObject(
|
||||
return (await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/prev/${id}.json`
|
||||
)) as GenerationRecord | null
|
||||
if (live) return live
|
||||
// Two-tier: the packed tier serves folded generations (live-tier-wins).
|
||||
if (this.segments?.hasGeneration(gen)) {
|
||||
return (await this.segments.readRecord(gen, kind, id)) as GenerationRecord | null
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -2230,21 +2132,6 @@ export class GenerationStore {
|
|||
`${GENERATIONS_PREFIX}/${gen}/tx.json`
|
||||
)) as GenerationDelta | null
|
||||
if (delta === null) {
|
||||
// Two-tier read (D1+D3): not in the live tier → the packed tier.
|
||||
// Live-tier-wins ordering (a crash mid-fold leaves a duplicate, never
|
||||
// a gap), so the segment lookup runs only after the live miss.
|
||||
const packed = await this.segments?.readDelta(gen)
|
||||
if (packed) {
|
||||
const d = packed.delta as GenerationDelta
|
||||
const entry = {
|
||||
nouns: new Set(d.nouns),
|
||||
verbs: new Set(d.verbs),
|
||||
timestamp: packed.timestamp,
|
||||
bytes: d.bytes ?? 0
|
||||
}
|
||||
this.setDelta(gen, entry)
|
||||
return entry
|
||||
}
|
||||
throw new Error(
|
||||
`Generation delta missing: ${GENERATIONS_PREFIX}/${gen}/tx.json ` +
|
||||
`(store corrupted or records removed outside compactHistory())`
|
||||
|
|
@ -2326,94 +2213,6 @@ export class GenerationStore {
|
|||
* @param options - Retention caps (see {@link CompactHistoryOptions}).
|
||||
* @returns Count of removed record-sets and the new horizon.
|
||||
*/
|
||||
/**
|
||||
* @description The REPACKER (D1+D3+repacking): fold cold live-tier
|
||||
* generations into sealed segments — re-representation, never deletion.
|
||||
* Every record and delta stays readable (asOf/chains unchanged); the
|
||||
* per-generation directories are deleted only AFTER their segment is
|
||||
* durable (crash between = duplicate representation, resolved
|
||||
* live-tier-wins by every reader; never a gap). This is the transform that
|
||||
* takes a 70k-file history to tens of segment files, and the ONLY history
|
||||
* transform permitted under the archival profile.
|
||||
*
|
||||
* Folds oldest-first, contiguous from the packed boundary, in batches, and
|
||||
* stops at the live window ({@link GenerationStore.REPACK_LIVE_WINDOW})
|
||||
* or when `timeBudgetMs` is spent — an early stop is a consistent prefix;
|
||||
* the next pass resumes.
|
||||
*/
|
||||
async repackHistory(options?: { timeBudgetMs?: number; batchGenerations?: number }): Promise<{
|
||||
foldedGenerations: number
|
||||
segmentsCreated: number
|
||||
}> {
|
||||
if (!this.segments) return { foldedGenerations: 0, segmentsCreated: 0 }
|
||||
const segments = this.segments
|
||||
return this.withMutex(async () => {
|
||||
const deadline =
|
||||
options?.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined
|
||||
const batchSize = options?.batchGenerations ?? 512
|
||||
const coldCeiling = this.committed - GenerationStore.REPACK_LIVE_WINDOW
|
||||
const packedThrough =
|
||||
segments.segments().length > 0
|
||||
? segments.segments()[segments.segments().length - 1].lastGeneration
|
||||
: 0
|
||||
|
||||
// Cold, unpacked, committed generations — ascending, contiguous scan.
|
||||
const eligible: number[] = []
|
||||
for (const gen of this.committedGensAsc()) {
|
||||
if (gen > coldCeiling) break
|
||||
if (gen <= packedThrough) continue // already packed (dup fold barred)
|
||||
if (this.pendingBuffer.has(gen)) continue // un-flushed = live by definition
|
||||
eligible.push(gen)
|
||||
}
|
||||
|
||||
let folded = 0
|
||||
let segmentsCreated = 0
|
||||
for (let i = 0; i < eligible.length; i += batchSize) {
|
||||
if (deadline !== undefined && Date.now() >= deadline) break
|
||||
const batch = eligible.slice(i, i + batchSize)
|
||||
const foldInput: FoldGeneration[] = []
|
||||
for (const gen of batch) {
|
||||
const delta = (await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/tx.json`
|
||||
)) as GenerationDelta | null
|
||||
if (delta === null) {
|
||||
// Already folded by a prior crashed pass whose dirs were removed,
|
||||
// or damage — getDelta's two-tier read decides which, loudly,
|
||||
// when someone asks. Skip; never fold a generation we cannot read.
|
||||
continue
|
||||
}
|
||||
const records: FoldGeneration['records'] = []
|
||||
for (const [kind, ids] of [
|
||||
['noun', delta.nouns] as const,
|
||||
['verb', delta.verbs] as const
|
||||
]) {
|
||||
for (const id of ids) {
|
||||
const record = await this.storage.readRawObject(
|
||||
`${GENERATIONS_PREFIX}/${gen}/prev/${id}.json`
|
||||
)
|
||||
if (record) records.push({ kind, id, record })
|
||||
}
|
||||
}
|
||||
foldInput.push({ generation: gen, timestamp: delta.timestamp, delta, records })
|
||||
}
|
||||
if (foldInput.length === 0) continue
|
||||
await segments.fold(foldInput)
|
||||
segmentsCreated++
|
||||
// Segment + manifest durable → the live copies retire.
|
||||
for (const g of foldInput) {
|
||||
await this.storage.removeRawPrefix(`${GENERATIONS_PREFIX}/${g.generation}`)
|
||||
}
|
||||
folded += foldInput.length
|
||||
}
|
||||
if (folded > 0) {
|
||||
prodLog.info(
|
||||
`[GenerationStore] repacked ${folded} cold generation(s) into ${segmentsCreated} segment(s) — history preserved, file count reduced`
|
||||
)
|
||||
}
|
||||
return { foldedGenerations: folded, segmentsCreated }
|
||||
})
|
||||
}
|
||||
|
||||
async compact(options?: CompactHistoryOptions): Promise<CompactHistoryResult> {
|
||||
return this.withMutex(async () => {
|
||||
const minPinned = this.minPinnedGeneration()
|
||||
|
|
@ -2505,16 +2304,6 @@ export class GenerationStore {
|
|||
// Reclaimed generations leave the per-id chains stale → rebuild on next read.
|
||||
this.invalidateChains()
|
||||
this.horizonGen = Math.max(this.horizonGen, highestRemoved)
|
||||
// Packed-tier reclaim (D3): a packed generation's bytes live in a
|
||||
// sealed segment — removeRawPrefix above was a no-op for it. Drop
|
||||
// WHOLE segments now fully below the horizon; a partially-reclaimed
|
||||
// segment keeps its bytes until the boundary passes it (the frozen
|
||||
// partial-segments-wait rule; logical reclamation above still holds —
|
||||
// the generations left committedRanges and asOf below the horizon
|
||||
// throws regardless).
|
||||
if (this.segments) {
|
||||
await this.segments.dropSegmentsBelow(this.horizonGen + 1)
|
||||
}
|
||||
const manifest: GenerationManifest = {
|
||||
version: 1,
|
||||
generation: this.committed,
|
||||
|
|
|
|||
|
|
@ -61,44 +61,41 @@ export class UnsupportedWhereOperatorError extends Error {
|
|||
* @returns The field's value, or `undefined` when absent.
|
||||
*/
|
||||
export function resolveEntityField(entity: Entity, field: string): unknown {
|
||||
// THE ONE ADDRESSING LAW (sealed 2026-08-03): `system.<field>` reads the
|
||||
// entity scalar; bare and `metadata.`-prefixed names read the user's
|
||||
// metadata bag (dotted paths traverse INSIDE the bag). The old bare-name
|
||||
// switch over system fields is dead — bare `createdAt` is the user's own
|
||||
// field now; the engine scalar is `system.createdAt`. Plumbing (vector,
|
||||
// connections, level, data, _rev) is invisible: no spelling reaches it.
|
||||
if (field.startsWith('system.')) {
|
||||
switch (field.slice('system.'.length)) {
|
||||
case 'type':
|
||||
return entity.type
|
||||
case 'subtype':
|
||||
return entity.subtype
|
||||
case 'id':
|
||||
return entity.id
|
||||
case 'createdAt':
|
||||
return entity.createdAt
|
||||
case 'updatedAt':
|
||||
return entity.updatedAt
|
||||
case 'service':
|
||||
return entity.service
|
||||
case 'createdBy':
|
||||
return entity.createdBy
|
||||
case 'confidence':
|
||||
return entity.confidence
|
||||
case 'weight':
|
||||
return entity.weight
|
||||
case 'visibility':
|
||||
return (entity as unknown as Record<string, unknown>).visibility
|
||||
}
|
||||
// Out-of-map system spelling: parse refuses these upstream with a typed
|
||||
// error; reaching here (internal callers only) reads as absent.
|
||||
return undefined
|
||||
switch (field) {
|
||||
case 'noun':
|
||||
case 'type':
|
||||
return entity.type
|
||||
case 'subtype':
|
||||
return entity.subtype
|
||||
case 'id':
|
||||
return entity.id
|
||||
case 'createdAt':
|
||||
return entity.createdAt
|
||||
case 'updatedAt':
|
||||
return entity.updatedAt
|
||||
case 'service':
|
||||
return entity.service
|
||||
case 'createdBy':
|
||||
return entity.createdBy
|
||||
case 'confidence':
|
||||
return entity.confidence
|
||||
case 'weight':
|
||||
return entity.weight
|
||||
case '_rev':
|
||||
return entity._rev
|
||||
case 'data':
|
||||
return entity.data
|
||||
}
|
||||
|
||||
const path = field.startsWith('metadata.') ? field.slice('metadata.'.length) : field
|
||||
const bag = (entity.metadata ?? {}) as Record<string, unknown>
|
||||
if (!path.includes('.')) return bag[path]
|
||||
return resolvePath(bag, path)
|
||||
if (field.includes('.')) {
|
||||
// Dotted path: resolve against the whole entity first (`metadata.x`),
|
||||
// then against the metadata bag (`address.city` on nested metadata).
|
||||
const fromEntity = resolvePath(entity as unknown as Record<string, unknown>, field)
|
||||
if (fromEntity !== undefined) return fromEntity
|
||||
return resolvePath((entity.metadata ?? {}) as Record<string, unknown>, field)
|
||||
}
|
||||
|
||||
return ((entity.metadata ?? {}) as Record<string, unknown>)[field]
|
||||
}
|
||||
|
||||
/** Walk a dotted path through nested plain objects. */
|
||||
|
|
|
|||
|
|
@ -22,6 +22,7 @@ import { SmartYAMLImporter } from '../importers/SmartYAMLImporter.js'
|
|||
import { SmartDOCXImporter } from '../importers/SmartDOCXImporter.js'
|
||||
import { VFSStructureGenerator } from '../importers/VFSStructureGenerator.js'
|
||||
import { NounType, VerbType } from '../types/graphTypes.js'
|
||||
import { splitNounMetadataRecord, splitVerbMetadataRecord } from '../types/reservedFields.js'
|
||||
import { v4 as uuidv4 } from '../universal/uuid.js'
|
||||
import * as fs from 'fs'
|
||||
import * as path from 'path'
|
||||
|
|
@ -870,18 +871,35 @@ export class ImportCoordinator {
|
|||
}
|
||||
|
||||
/**
|
||||
* Normalize an extractor/consumer metadata bag for spreading — the
|
||||
* field-addressing law: the bag is the user's, VERBATIM. No name is
|
||||
* reserved anymore ('confidence', 'subtype', 'type', … in a source bag
|
||||
* import as ordinary user fields); the old reserved-key strip was data
|
||||
* loss under the law and is gone. A forged 'system.'-prefixed key still
|
||||
* refuses loudly at the write door (`rejectForgedSystemKeys`).
|
||||
* Strip Brainy-reserved entity keys out of an extractor-supplied metadata bag.
|
||||
*
|
||||
* Extractors (and consumer `customMetadata`) can carry reserved keys
|
||||
* (`confidence`, `subtype`, `weight`, …) inside `metadata`. Brainy 8.0's
|
||||
* default `reservedFieldPolicy` is `'throw'`, so spreading such a bag into
|
||||
* `add({ metadata })` would reject the whole import. The import pipeline owns
|
||||
* the correct write path: user-mutable reserved values are passed as dedicated
|
||||
* `AddParams` params (see the call sites), so here we simply drop the reserved
|
||||
* half of the bag and keep only the custom fields that belong in `metadata`.
|
||||
*
|
||||
* @param bag - The extractor/consumer metadata bag (may be undefined).
|
||||
* @returns The bag itself, or `{}` for non-object inputs.
|
||||
* @returns The custom-only metadata (reserved keys removed).
|
||||
*/
|
||||
private bagVerbatim(bag: Record<string, any> | undefined | null): Record<string, any> {
|
||||
private stripReservedFromBag(bag: Record<string, any> | undefined | null): Record<string, any> {
|
||||
if (!bag || typeof bag !== 'object') return {}
|
||||
return bag
|
||||
return splitNounMetadataRecord(bag).custom
|
||||
}
|
||||
|
||||
/**
|
||||
* Relationship mirror of {@link stripReservedFromBag} — strips reserved verb
|
||||
* keys (`verb`, `confidence`, `weight`, `subtype`, …) out of an edge metadata
|
||||
* bag so it carries only custom fields. Reserved values that have a dedicated
|
||||
* `RelateParams` param are passed there by the call site instead.
|
||||
* @param bag - The extractor/consumer edge metadata bag (may be undefined).
|
||||
* @returns The custom-only edge metadata (reserved keys removed).
|
||||
*/
|
||||
private stripReservedFromRelationBag(bag: Record<string, any> | undefined | null): Record<string, any> {
|
||||
if (!bag || typeof bag !== 'object') return {}
|
||||
return splitVerbMetadataRecord(bag).custom
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -999,7 +1017,7 @@ export class ImportCoordinator {
|
|||
importedAt: trackingContext.importedAt,
|
||||
importFormat: trackingContext.importFormat,
|
||||
importSource: trackingContext.importSource,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
})
|
||||
|
|
@ -1027,11 +1045,13 @@ export class ImportCoordinator {
|
|||
data: entity.description || entity.name,
|
||||
type: entity.type,
|
||||
subtype: entity.subtype ?? options.defaultSubtype ?? 'imported',
|
||||
// Engine confidence rides its dedicated param; the bag below is
|
||||
// the user's verbatim (no name is reserved — field-addressing law).
|
||||
// `confidence` is a reserved field — pass it as the dedicated param,
|
||||
// never inside the metadata bag (8.0 reservedFieldPolicy defaults to 'throw').
|
||||
confidence: entity.confidence,
|
||||
metadata: {
|
||||
...this.bagVerbatim(entity.metadata),
|
||||
// Extractor/consumer bags may smuggle reserved keys — strip them so
|
||||
// the bag carries only custom fields.
|
||||
...this.stripReservedFromBag(entity.metadata),
|
||||
name: entity.name,
|
||||
vfsPath: vfsFile?.path,
|
||||
importedFrom: 'import-coordinator',
|
||||
|
|
@ -1044,7 +1064,7 @@ export class ImportCoordinator {
|
|||
importSource: trackingContext.importSource,
|
||||
sourceRow: row.rowNumber,
|
||||
sourceSheet: row.sheet,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
@ -1125,7 +1145,7 @@ export class ImportCoordinator {
|
|||
importIds: [trackingContext.importId],
|
||||
projectId: trackingContext.projectId,
|
||||
importFormat: trackingContext.importFormat,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromRelationBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
@ -1160,7 +1180,7 @@ export class ImportCoordinator {
|
|||
confidence: entity.confidence,
|
||||
metadata: {
|
||||
// Strip any reserved keys an extractor smuggled into the bag.
|
||||
...this.bagVerbatim(entity.metadata),
|
||||
...this.stripReservedFromBag(entity.metadata),
|
||||
name: entity.name,
|
||||
vfsPath: vfsFile?.path,
|
||||
importedFrom: 'import-coordinator',
|
||||
|
|
@ -1174,7 +1194,7 @@ export class ImportCoordinator {
|
|||
importSource: trackingContext.importSource,
|
||||
sourceRow: row.rowNumber,
|
||||
sourceSheet: row.sheet,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
})
|
||||
|
|
@ -1214,7 +1234,7 @@ export class ImportCoordinator {
|
|||
importIds: [trackingContext.importId],
|
||||
projectId: trackingContext.projectId,
|
||||
importFormat: trackingContext.importFormat,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromRelationBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
})
|
||||
|
|
@ -1269,7 +1289,7 @@ export class ImportCoordinator {
|
|||
projectId: trackingContext.projectId,
|
||||
importedAt: trackingContext.importedAt,
|
||||
importFormat: trackingContext.importFormat,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
})
|
||||
|
|
@ -1299,7 +1319,7 @@ export class ImportCoordinator {
|
|||
projectId: trackingContext.projectId,
|
||||
importedAt: trackingContext.importedAt,
|
||||
importFormat: trackingContext.importFormat,
|
||||
...this.bagVerbatim(trackingContext.customMetadata)
|
||||
...this.stripReservedFromRelationBag(trackingContext.customMetadata)
|
||||
})
|
||||
}
|
||||
})
|
||||
|
|
@ -1402,7 +1422,7 @@ export class ImportCoordinator {
|
|||
...(typeof (rel as any).confidence === 'number' && { confidence: (rel as any).confidence }),
|
||||
...(typeof (rel as any).weight === 'number' && { weight: (rel as any).weight }),
|
||||
metadata: {
|
||||
...this.bagVerbatim(rel.metadata),
|
||||
...this.stripReservedFromRelationBag(rel.metadata),
|
||||
relationshipType: 'semantic', // Distinguish from VFS/provenance
|
||||
inferredType: verbType !== rel.type, // Track if type was enhanced
|
||||
originalType: rel.type
|
||||
|
|
|
|||
27
src/index.ts
27
src/index.ts
|
|
@ -89,12 +89,7 @@ export {
|
|||
RESERVED_ENTITY_FIELDS,
|
||||
RESERVED_RELATION_FIELDS,
|
||||
splitNounMetadataRecord,
|
||||
splitVerbMetadataRecord,
|
||||
buildNounMetadataRecord,
|
||||
buildVerbMetadataRecord,
|
||||
isNestedBagRecord,
|
||||
METADATA_RECORD_FORMAT_KEY,
|
||||
NESTED_BAG_FORMAT
|
||||
splitVerbMetadataRecord
|
||||
} from './types/reservedFields.js'
|
||||
export type {
|
||||
ReservedEntityField,
|
||||
|
|
@ -111,25 +106,6 @@ export type {
|
|||
// Export Aggregation Engine
|
||||
export { AggregationIndex, AggregateMaterializer, bucketTimestamp, parseBucketRange } from './aggregation/index.js'
|
||||
|
||||
// THE ONE FIELD-ADDRESSING LAW (sealed 2026-08-03) — the arming surface both
|
||||
// engines' conformance suites detect: bare names = user metadata, system.* =
|
||||
// the ten ruled scalars, plumbing invisible, refusals typed with the fix in
|
||||
// the message. See docs/concepts/field-addressing.md.
|
||||
export {
|
||||
FIELD_ADDRESSING_CAPABILITY,
|
||||
SYSTEM_ENTITY_SCALARS,
|
||||
SYSTEM_RELATION_SCALARS,
|
||||
PLUMBING_FIELDS,
|
||||
parseFieldAddress,
|
||||
readEntityFieldAddress,
|
||||
readRelationFieldAddress,
|
||||
buildUnresolvableMessage,
|
||||
InvalidFieldAddressError,
|
||||
UnresolvableFieldError,
|
||||
UnsupportedFindOptionError
|
||||
} from './db/fieldAddressing.js'
|
||||
export type { FieldAddress, FieldAddressKind } from './db/fieldAddressing.js'
|
||||
|
||||
// Export Neural Import (AI data understanding)
|
||||
export { NeuralImport } from './neural/neuralImport.js'
|
||||
export type {
|
||||
|
|
@ -247,7 +223,6 @@ export type {
|
|||
CommitFact,
|
||||
FactOp,
|
||||
FactScanBatch,
|
||||
SCANFACTS_FIRST_BATCH_MS,
|
||||
FactScanHandle
|
||||
} from './db/factLog.js'
|
||||
// The generalized family stamp — which source generation a projection
|
||||
|
|
|
|||
|
|
@ -9,67 +9,6 @@ import type { BaseStorage } from '../storage/baseStorage.js'
|
|||
import type { NounMetadata, VerbMetadata } from '../coreTypes.js'
|
||||
import type { Migration, MigrationState, MigrationPreview, MigrationResult, MigrateOptions, MigrationError } from './types.js'
|
||||
import { MIGRATIONS } from './migrations.js'
|
||||
import {
|
||||
splitNounMetadataRecord,
|
||||
splitVerbMetadataRecord,
|
||||
buildNounMetadataRecord,
|
||||
buildVerbMetadataRecord,
|
||||
RESERVED_ENTITY_FIELDS,
|
||||
RESERVED_RELATION_FIELDS
|
||||
} from '../types/reservedFields.js'
|
||||
|
||||
const RESERVED_NOUN_SET: ReadonlySet<string> = new Set(RESERVED_ENTITY_FIELDS)
|
||||
const RESERVED_VERB_SET: ReadonlySet<string> = new Set(RESERVED_RELATION_FIELDS)
|
||||
|
||||
/**
|
||||
* Normalize a stored record (either era: legacy flat OR v2 nested-bag) into
|
||||
* THE transform view — the one shape every migration transform receives:
|
||||
* engine fields top-level, the user's metadata bag nested under `metadata`.
|
||||
* Transforms never see the storage era; a migration written today works on
|
||||
* a brain of any age.
|
||||
*/
|
||||
function toTransformView(
|
||||
record: Record<string, unknown>,
|
||||
kind: 'noun' | 'verb'
|
||||
): Record<string, unknown> {
|
||||
const { reserved, custom } =
|
||||
kind === 'noun' ? splitNounMetadataRecord(record) : splitVerbMetadataRecord(record)
|
||||
return { ...reserved, metadata: { ...custom } }
|
||||
}
|
||||
|
||||
/**
|
||||
* Convert a transform's returned view back into a stamped v2 stored record.
|
||||
* LOUD CONTRACT: user fields belong inside `.metadata` — a stray top-level
|
||||
* key that is not an engine field is a migration bug under the
|
||||
* field-addressing law (pre-law transforms wrote user fields flat), and it
|
||||
* refuses with the fix in the message rather than silently dropping or
|
||||
* silently storing it as an engine key.
|
||||
*/
|
||||
function fromTransformView(
|
||||
view: Record<string, unknown>,
|
||||
kind: 'noun' | 'verb'
|
||||
): Record<string, unknown> {
|
||||
const reservedSet = kind === 'noun' ? RESERVED_NOUN_SET : RESERVED_VERB_SET
|
||||
const engine: Record<string, unknown> = {}
|
||||
for (const [key, value] of Object.entries(view)) {
|
||||
if (key === 'metadata') continue
|
||||
if (!reservedSet.has(key)) {
|
||||
throw new Error(
|
||||
`migration transform returned a top-level key '${key}' that is not an ` +
|
||||
`engine field — under the field-addressing law user fields live inside ` +
|
||||
`.metadata (return { ...view, metadata: { ...view.metadata, ${key}: … } }).`
|
||||
)
|
||||
}
|
||||
engine[key] = value
|
||||
}
|
||||
const bag =
|
||||
view.metadata && typeof view.metadata === 'object' && !Array.isArray(view.metadata)
|
||||
? (view.metadata as Record<string, unknown>)
|
||||
: {}
|
||||
return kind === 'noun'
|
||||
? buildNounMetadataRecord(engine, bag)
|
||||
: buildVerbMetadataRecord(engine, bag)
|
||||
}
|
||||
|
||||
const MIGRATION_STATE_KEY = '__migration_state__'
|
||||
const PREVIEW_SAMPLE_SIZE = 5
|
||||
|
|
@ -186,16 +125,14 @@ export class MigrationRunner {
|
|||
const entityMeta = metadataBatch.get(entity.id)
|
||||
if (!entityMeta) continue
|
||||
|
||||
// Transforms see THE view (engine fields + nested user bag),
|
||||
// never the raw storage era.
|
||||
const view = toTransformView(entityMeta as Record<string, unknown>, 'noun')
|
||||
const result = this.applyTransforms(view, nounMigrations)
|
||||
const metadata = entityMeta as Record<string, unknown>
|
||||
const result = this.applyTransforms(metadata, nounMigrations)
|
||||
if (result !== null) {
|
||||
affectedEntities++
|
||||
if (sampleChanges.length < PREVIEW_SAMPLE_SIZE) {
|
||||
sampleChanges.push({
|
||||
id: entity.id,
|
||||
before: view,
|
||||
before: { ...metadata },
|
||||
after: result
|
||||
})
|
||||
}
|
||||
|
|
@ -220,14 +157,14 @@ export class MigrationRunner {
|
|||
const verbMeta = await this.storage.getVerbMetadata(verb.id)
|
||||
if (!verbMeta) continue
|
||||
|
||||
const view = toTransformView(verbMeta as Record<string, unknown>, 'verb')
|
||||
const result = this.applyTransforms(view, verbMigrations)
|
||||
const metadata = verbMeta as Record<string, unknown>
|
||||
const result = this.applyTransforms(metadata, verbMigrations)
|
||||
if (result !== null) {
|
||||
affectedEntities++
|
||||
if (sampleChanges.length < PREVIEW_SAMPLE_SIZE) {
|
||||
sampleChanges.push({
|
||||
id: verb.id,
|
||||
before: view,
|
||||
before: { ...metadata },
|
||||
after: result
|
||||
})
|
||||
}
|
||||
|
|
@ -352,16 +289,9 @@ export class MigrationRunner {
|
|||
if (!entityMeta) continue
|
||||
|
||||
try {
|
||||
const transformed = migration.transform(
|
||||
toTransformView(entityMeta as Record<string, unknown>, 'noun')
|
||||
)
|
||||
const transformed = migration.transform(entityMeta as Record<string, unknown>)
|
||||
if (transformed !== null) {
|
||||
// Re-stamp as a v2 record (also upgrades legacy records touched
|
||||
// by a migration onto the nested-bag shape).
|
||||
await this.storage.saveNounMetadata(
|
||||
entity.id,
|
||||
fromTransformView(transformed, 'noun') as NounMetadata
|
||||
)
|
||||
await this.storage.saveNounMetadata(entity.id, transformed as NounMetadata)
|
||||
modified++
|
||||
}
|
||||
} catch (err) {
|
||||
|
|
@ -427,14 +357,9 @@ export class MigrationRunner {
|
|||
if (!metadata) continue
|
||||
|
||||
try {
|
||||
const transformed = migration.transform(
|
||||
toTransformView(metadata as Record<string, unknown>, 'verb')
|
||||
)
|
||||
const transformed = migration.transform(metadata as Record<string, unknown>)
|
||||
if (transformed !== null) {
|
||||
await this.storage.saveVerbMetadata(
|
||||
verb.id,
|
||||
fromTransformView(transformed, 'verb') as VerbMetadata
|
||||
)
|
||||
await this.storage.saveVerbMetadata(verb.id, transformed as VerbMetadata)
|
||||
modified++
|
||||
}
|
||||
} catch (err) {
|
||||
|
|
|
|||
|
|
@ -14,19 +14,7 @@ export interface Migration {
|
|||
description: string
|
||||
/** Which entity types this migration applies to */
|
||||
applies: 'nouns' | 'verbs' | 'both'
|
||||
/**
|
||||
* Return the transformed record view, or null if no change needed.
|
||||
*
|
||||
* THE VIEW CONTRACT (field-addressing law): the transform receives ONE
|
||||
* normalized shape regardless of how old the stored record is — engine
|
||||
* fields top-level (`noun`/`verb`, `subtype`, `confidence`, `weight`,
|
||||
* timestamps, `_rev`, …) and the USER's metadata bag nested under
|
||||
* `metadata` (where every name is the user's, engine spellings included).
|
||||
* Return the same shape: user-field changes go inside `.metadata`; a
|
||||
* stray non-engine top-level key in the returned object refuses loudly
|
||||
* (it is the pre-law flat habit, and silently guessing its namespace
|
||||
* would corrupt data).
|
||||
*/
|
||||
/** Return transformed metadata, or null if no change needed */
|
||||
transform: (metadata: Record<string, unknown>) => Record<string, unknown> | null
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@
|
|||
|
||||
import { Brainy } from '../brainy.js'
|
||||
import { NounType, VerbType } from '../types/graphTypes.js'
|
||||
import { splitNounMetadataRecord, splitVerbMetadataRecord } from '../types/reservedFields.js'
|
||||
import * as fs from '../universal/fs.js'
|
||||
import * as path from '../universal/path.js'
|
||||
// @ts-ignore
|
||||
|
|
@ -802,14 +803,12 @@ export class NeuralImport {
|
|||
data: this.extractMainText(entity.originalData),
|
||||
type: entity.nounType as NounType,
|
||||
subtype: entity.subtype ?? options.defaultSubtype ?? 'extracted',
|
||||
// Engine confidence rides its dedicated param; the source object
|
||||
// imports as the user's bag VERBATIM — no name is reserved
|
||||
// (field-addressing law).
|
||||
// `confidence` is a reserved field — dedicated param, not metadata
|
||||
// (8.0 reservedFieldPolicy defaults to 'throw').
|
||||
confidence: entity.confidence,
|
||||
metadata: {
|
||||
...(typeof entity.originalData === 'object' && entity.originalData !== null
|
||||
? entity.originalData
|
||||
: {}),
|
||||
// Strip any reserved keys the source data smuggled into the bag.
|
||||
...splitNounMetadataRecord(entity.originalData).custom,
|
||||
id: entity.suggestedId
|
||||
}
|
||||
})
|
||||
|
|
@ -823,13 +822,11 @@ export class NeuralImport {
|
|||
type: relationship.verbType as VerbType,
|
||||
subtype: relationship.subtype ?? options.defaultSubtype ?? 'extracted',
|
||||
weight: relationship.weight,
|
||||
confidence: relationship.confidence, // engine confidence — dedicated param
|
||||
confidence: relationship.confidence, // reserved field — dedicated param, not metadata
|
||||
metadata: {
|
||||
context: relationship.context,
|
||||
// The edge bag imports verbatim — no name is reserved.
|
||||
...(typeof relationship.metadata === 'object' && relationship.metadata !== null
|
||||
? relationship.metadata
|
||||
: {})
|
||||
// Strip any reserved keys smuggled into the edge metadata bag.
|
||||
...splitVerbMetadataRecord(relationship.metadata).custom
|
||||
}
|
||||
})
|
||||
}
|
||||
|
|
|
|||
|
|
@ -133,11 +133,6 @@ export class MemoryStorage extends BaseStorage {
|
|||
*/
|
||||
protected async deleteObjectFromPath(path: string): Promise<void> {
|
||||
this.objectStore.delete(path)
|
||||
// Filesystem parity: on disk, objects and raw BYTE files are both just
|
||||
// files — unlink removes whichever exists. Without this, deleteRawObject
|
||||
// on a raw-bytes path (fact-log/generation segments) silently no-ops on
|
||||
// memory storage: the delete "succeeds" and the bytes remain.
|
||||
this.rawBytesStore.delete(path)
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
|
|||
|
|
@ -36,8 +36,7 @@ import { BrainyError, ProtectedArtifactError, DerivedArtifactMissingError } from
|
|||
import { MetadataWriteBuffer } from '../utils/metadataWriteBuffer.js'
|
||||
import {
|
||||
splitNounMetadataRecord,
|
||||
splitVerbMetadataRecord,
|
||||
isNestedBagRecord
|
||||
splitVerbMetadataRecord
|
||||
} from '../types/reservedFields.js'
|
||||
|
||||
/**
|
||||
|
|
@ -383,10 +382,6 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
// identical to the unknown-key fallback these keys hit
|
||||
// before being listed here — this only kills the
|
||||
// per-boot "Unknown key format" warning)
|
||||
id.startsWith('graph-lsm-') || // Graph-LSM store manifests written through storage by
|
||||
// an active native graph provider — same
|
||||
// warn-then-route fallback as above; listing the family
|
||||
// silences the per-boot warning on provider-backed brains
|
||||
isSingletonSystemKey(id) // Known singletons (e.g. brainy:entityIdMapper) hit the
|
||||
// same warn-then-route fallback without this — the
|
||||
// routing below already handles them identically
|
||||
|
|
@ -1014,14 +1009,8 @@ export abstract class BaseStorage extends BaseStorageAdapter {
|
|||
const hashes: string[] = []
|
||||
for (const record of records) {
|
||||
if (record.kind !== 'noun') continue
|
||||
// The VFS blob pointer (`storage: {type:'blob', hash}`) is a USER-bag
|
||||
// field: in a v2 nested-bag record it lives inside `metadata`, in a
|
||||
// legacy flat record it sits at the top level — read shape-aware.
|
||||
const raw = record.metadata as Record<string, unknown> | null
|
||||
const bag = isNestedBagRecord(raw)
|
||||
? (raw!.metadata as Record<string, unknown>)
|
||||
: raw
|
||||
const storage = (bag as { storage?: { type?: string; hash?: unknown } } | null)?.storage
|
||||
const storage = (record.metadata as { storage?: { type?: string; hash?: unknown } } | null)
|
||||
?.storage
|
||||
if (storage?.type === 'blob' && typeof storage.hash === 'string') {
|
||||
hashes.push(storage.hash)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -69,15 +69,7 @@ export const BRAIN_FORMAT_PATH = '_system/brain-format.json'
|
|||
* (the 8.0 GA baseline). An on-disk `indexEpoch` that differs from this — or an
|
||||
* absent marker — triggers a full derived-index rebuild on open.
|
||||
*/
|
||||
// Epoch 3 (2026-08-03, the namespace-law pair): the index key format split
|
||||
// the two namespaces — user fields keep bare flattened keys, the ten system
|
||||
// scalars moved to literal 'system.<field>' keys (the legacy 'noun' column
|
||||
// spelling died with them). Every brain rebuilds its derived indexes from
|
||||
// canonical at first open onto the frozen keys.
|
||||
// Epoch 2 (2026-08-03, same day, the interim pair): user metadata fields
|
||||
// named `level` became indexable on both engines; poisoned multi-valued
|
||||
// `level` columns healed through the rebuild.
|
||||
export const EXPECTED_INDEX_EPOCH = 3
|
||||
export const EXPECTED_INDEX_EPOCH = 1
|
||||
|
||||
/**
|
||||
* @description The data-layer format string this build writes and runs as.
|
||||
|
|
|
|||
|
|
@ -77,27 +77,8 @@ export class SaveNounOperation implements Operation {
|
|||
? null
|
||||
: await this.storage.getNoun(this.noun.id)
|
||||
|
||||
// PRESERVE stored graph state on updates. Callers stage this op with
|
||||
// placeholder adjacency ({connections: empty, level: 0}) because the
|
||||
// vector index owns those values and persists them at flush. Codec-era
|
||||
// records (2.4.0+) carry an empty connections field by design (adjacency
|
||||
// lives in a separate compressed blob — the placeholder is harmless), but
|
||||
// LEGACY pre-codec records store adjacency INLINE: writing the
|
||||
// placeholder over one stamped out its stored connections, leaving a
|
||||
// crash window (until the next flush) where a reload found the node
|
||||
// unreachable. Stale adjacency in that window is tolerable — HNSW
|
||||
// self-corrects at the reindex flush; EMPTY adjacency is silent recall
|
||||
// loss. The read above is already paid for rollback; preservation is free.
|
||||
const toSave: HNSWNoun =
|
||||
previousNoun && this.noun.connections.size === 0
|
||||
? {
|
||||
...this.noun,
|
||||
connections: previousNoun.connections || this.noun.connections,
|
||||
level: previousNoun.level ?? this.noun.level
|
||||
}
|
||||
: this.noun
|
||||
|
||||
await this.storage.saveNoun(toSave)
|
||||
// Save new noun
|
||||
await this.storage.saveNoun(this.noun)
|
||||
|
||||
// Return rollback action
|
||||
return async () => {
|
||||
|
|
|
|||
|
|
@ -320,18 +320,15 @@ export interface AddParams<T = any> {
|
|||
*/
|
||||
visibility?: 'public' | 'internal'
|
||||
/**
|
||||
* Structured queryable fields — indexed by MetadataIndex, used in `where`
|
||||
* filters, `orderBy`, and aggregation.
|
||||
* Structured queryable fields — indexed by MetadataIndex, used in `where` filters.
|
||||
*
|
||||
* THE FIELD-ADDRESSING LAW: every name here is YOURS. There are no
|
||||
* reserved metadata names — `confidence`, `type`, `id`, `level`, `data`,
|
||||
* `content`, … are ordinary user fields that index, filter, sort, and
|
||||
* aggregate like any other, and survive faithfully across restarts and
|
||||
* rebuilds. Engine scalars are set only via their dedicated params
|
||||
* (`confidence`, `weight`, `subtype`, …) and are queried explicitly as
|
||||
* `system.<field>` (`where: { 'system.confidence': … }`). The ONE illegal
|
||||
* spelling is a key starting `'system.'` — the engine's explicit address
|
||||
* namespace cannot be forged; such a write refuses with a typed error.
|
||||
* Reserved entity fields (`RESERVED_ENTITY_FIELDS` — `noun`, `subtype`, `visibility`,
|
||||
* `createdAt`, `updatedAt`, `confidence`, `weight`, `service`, `data`, `createdBy`,
|
||||
* `_rev`) may NOT appear here — they have dedicated top-level params and the type makes
|
||||
* a literal reserved key a compile error. Untyped (JavaScript) callers that pass one
|
||||
* anyway are normalized at write time: user-settable fields remap to their top-level
|
||||
* param (top-level wins when both are supplied), system-managed fields are dropped with
|
||||
* a one-shot warning.
|
||||
*/
|
||||
metadata?: EntityMetadataInput<T>
|
||||
/** Custom entity ID. When omitted, a time-ordered UUID v7 is generated; a supplied natural-key string is normalized to a stable UUID v5. */
|
||||
|
|
@ -389,11 +386,12 @@ export interface UpdateParams<T = any> {
|
|||
*/
|
||||
visibility?: EntityVisibility
|
||||
/**
|
||||
* Metadata fields to merge (or replace when `merge: false`). Every name is
|
||||
* the user's (the field-addressing law) — a patch field named `confidence`
|
||||
* updates YOUR field of that name, never the engine scalar (use the
|
||||
* dedicated `confidence` param for that). Keys spelled `'system.…'` refuse
|
||||
* with a typed error (namespace forgery).
|
||||
* Metadata fields to merge (or replace when `merge: false`). Reserved entity
|
||||
* fields (`RESERVED_ENTITY_FIELDS`) may NOT appear here — `confidence` /
|
||||
* `weight` / `subtype` / `visibility` have dedicated params on this call, and the rest
|
||||
* are system-managed. A literal reserved key is a compile error; untyped callers
|
||||
* are normalized at write time (remap user-settable, drop system-managed
|
||||
* with a one-shot warning).
|
||||
*/
|
||||
metadata?: EntityMetadataPatch<T>
|
||||
merge?: boolean // Merge or replace metadata (default: true)
|
||||
|
|
@ -446,11 +444,11 @@ export interface RelateParams<T = any> {
|
|||
/** Content for the relationship (optional — overrides auto-computed vector) */
|
||||
data?: any
|
||||
/**
|
||||
* Structured queryable fields on the edge. Every name is the user's (the
|
||||
* field-addressing law) — `verb`, `confidence`, `weight`, … in this bag are
|
||||
* ordinary user fields; engine scalars ride their dedicated params and are
|
||||
* addressed as `system.<field>`. Keys spelled `'system.…'` refuse with a
|
||||
* typed error (namespace forgery).
|
||||
* Structured queryable fields on the edge. Reserved relationship fields
|
||||
* (`RESERVED_RELATION_FIELDS` — `verb`, `subtype`, `visibility`, `createdAt`,
|
||||
* `updatedAt`, `confidence`, `weight`, `service`, `data`, `createdBy`, `_rev`) may NOT
|
||||
* appear here — they have dedicated params. A literal reserved key is a
|
||||
* compile error; untyped callers are normalized at write time.
|
||||
*/
|
||||
metadata?: RelationMetadataInput<T>
|
||||
/** Create reverse edge too (default: false) */
|
||||
|
|
@ -480,9 +478,10 @@ export interface UpdateRelationParams<T = any> {
|
|||
confidence?: number // New confidence (0-1)
|
||||
data?: any // New content
|
||||
/**
|
||||
* Metadata fields to merge (or replace when `merge: false`). Every name is
|
||||
* the user's (the field-addressing law); engine scalars ride their
|
||||
* dedicated params. Keys spelled `'system.…'` refuse with a typed error.
|
||||
* Metadata fields to merge (or replace when `merge: false`). Reserved
|
||||
* relationship fields (`RESERVED_RELATION_FIELDS`) may NOT appear here —
|
||||
* a literal reserved key is a compile error; untyped callers are
|
||||
* normalized at write time.
|
||||
*/
|
||||
metadata?: RelationMetadataPatch<T>
|
||||
merge?: boolean // Merge or replace metadata
|
||||
|
|
@ -499,43 +498,6 @@ export interface UpdateRelationParams<T = any> {
|
|||
* - **Graph:** `connected` for relationship traversal (via GraphAdjacencyIndex)
|
||||
*
|
||||
* See also: [Query Operators](../../docs/QUERY_OPERATORS.md) for all `where` operators.
|
||||
*
|
||||
* @remarks
|
||||
* **Field-addressing law.** Governs every query-surface field name — `where`
|
||||
* and `orderBy` on this interface, plus `AggregateSource.where` and
|
||||
* `AggregateDefinition.groupBy` in the aggregation engine:
|
||||
*
|
||||
* 1. A bare name (e.g. `'level'`, `'rank'`, `'score'`) always means the
|
||||
* caller's own metadata field — it reads `entity.metadata.<name>`. There
|
||||
* is no fallback to an engine-internal field of the same name and no
|
||||
* priority resolution between the two; metadata wins unconditionally.
|
||||
* 2. `system.<field>` reaches an engine scalar, explicitly, and only for
|
||||
* these ten: `id`, `type`, `subtype`, `createdAt`, `updatedAt`,
|
||||
* `confidence`, `weight`, `visibility`, `service`, `createdBy`.
|
||||
* 3. `vector`, `connections`, `level` (the engine-internal node field — a
|
||||
* different thing from a user metadata field also named `level`),
|
||||
* `data`, and `_rev` are invisible plumbing: neither spelling can
|
||||
* address them from a query surface.
|
||||
* 4. `metadata.<field>` is the explicit spelling of the bare form and means
|
||||
* exactly the same thing as rule 1.
|
||||
* 5. A name that matches none of the above — most often a bare name that
|
||||
* collides with one of the ten system-scalar names in rule 2 — REFUSES
|
||||
* with a typed {@link UnresolvableFieldError} naming both candidates,
|
||||
* e.g. `no metadata field 'createdAt' — did you mean system.createdAt or
|
||||
* metadata.createdAt?`. The same loud-refusal principle covers whole
|
||||
* options: the previously accepted-and-silently-ignored `cursor`,
|
||||
* `includeRelations`, and `writeOnly` now throw
|
||||
* {@link UnsupportedFindOptionError} instead of doing nothing.
|
||||
* 6. **Ordering contract** (identical on the pure-JS engine and the native
|
||||
* accelerator): rows missing or `null` on the `orderBy` field sort LAST
|
||||
* in BOTH `asc` and `desc` order and are never dropped from the result;
|
||||
* ties break by `id` ascending.
|
||||
*
|
||||
* Migration note: a call site written against the old rule — e.g.
|
||||
* `orderBy: 'createdAt'` or `where: { visibility: 'internal' }` meaning the
|
||||
* engine scalar — now refuses instead of silently reading the wrong field.
|
||||
* The thrown error names the exact fix (`system.createdAt`). A loud
|
||||
* refusal with the fix in hand beats a silent behavior flip.
|
||||
*/
|
||||
export interface FindParams<T = any> {
|
||||
// Vector Intelligence
|
||||
|
|
@ -554,18 +516,7 @@ export interface FindParams<T = any> {
|
|||
* `{ exists: true }`, `{ missing: true }`) use `where: { subtype: { …operators… } }`.
|
||||
*/
|
||||
subtype?: string | string[]
|
||||
/**
|
||||
* Metadata filters using BFO operators (e.g., `{ year: { greaterThan: 2020 } }`).
|
||||
* Field names follow the field-addressing law — see the `@remarks` on
|
||||
* {@link FindParams}: a bare key is always the caller's metadata field;
|
||||
* an engine scalar needs the explicit `system.<field>` form.
|
||||
*
|
||||
* @example
|
||||
* ```typescript
|
||||
* await brain.find({ where: { level: { greaterThan: 5 } } }) // metadata.level
|
||||
* await brain.find({ where: { 'system.visibility': 'internal' } }) // engine scalar
|
||||
* ```
|
||||
*/
|
||||
/** Metadata filters using BFO operators (e.g., `{ year: { greaterThan: 2020 } }`) */
|
||||
where?: Partial<T>
|
||||
|
||||
// Visibility
|
||||
|
|
@ -597,49 +548,13 @@ export interface FindParams<T = any> {
|
|||
// Control options
|
||||
limit?: number // Max results (default: 10)
|
||||
offset?: number // Skip N results
|
||||
/**
|
||||
* @deprecated Not implemented. Passing `cursor` throws
|
||||
* {@link UnsupportedFindOptionError} — it used to be accepted and
|
||||
* silently ignored, which masked that no cursor pagination ever ran. Use
|
||||
* `offset` / `limit` until cursor pagination ships.
|
||||
*/
|
||||
cursor?: string // Cursor-based pagination
|
||||
|
||||
// Sorting
|
||||
/**
|
||||
* Field to sort by. Follows the field-addressing law (see the `@remarks`
|
||||
* on {@link FindParams}): a bare name (`'level'`, `'rank'`, `'score'`, …)
|
||||
* always sorts by that metadata field; the ten engine scalars sort only
|
||||
* via the explicit `system.<field>` form (e.g. `'system.createdAt'`); a
|
||||
* name that resolves to neither throws {@link UnresolvableFieldError}
|
||||
* naming the fix.
|
||||
*
|
||||
* Ordering contract (identical on the pure-JS engine and the native
|
||||
* accelerator): rows missing or `null` on this field sort LAST in BOTH
|
||||
* `asc` and `desc` order and are never dropped from the result; ties
|
||||
* break by `id` ascending.
|
||||
*
|
||||
* @example
|
||||
* ```typescript
|
||||
* await brain.find({ orderBy: 'level', order: 'desc' }) // metadata.level
|
||||
* await brain.find({ orderBy: 'system.createdAt', order: 'desc' }) // engine scalar
|
||||
* ```
|
||||
*/
|
||||
orderBy?: string
|
||||
/**
|
||||
* Sort direction: `'asc'` (default) or `'desc'`. Per the ordering
|
||||
* contract on `orderBy`, rows missing/`null` on the sorted field sort
|
||||
* LAST in both directions — `order` never moves them to the front.
|
||||
*/
|
||||
orderBy?: string // Field to sort by (e.g., 'createdAt', 'title', 'metadata.priority')
|
||||
order?: 'asc' | 'desc' // Sort direction: 'asc' (default) or 'desc'
|
||||
|
||||
// Advanced options
|
||||
/**
|
||||
* @deprecated Not implemented. Passing `includeRelations` throws
|
||||
* {@link UnsupportedFindOptionError} — it used to be accepted and
|
||||
* silently ignored, so no relationships were ever attached. Fetch
|
||||
* relationships separately via `brain.related()`.
|
||||
*/
|
||||
includeRelations?: boolean // Include entity relationships
|
||||
excludeVFS?: boolean // Exclude VFS entities from results (default: false - VFS included)
|
||||
service?: string // Multi-tenancy filter
|
||||
|
|
@ -672,11 +587,6 @@ export interface FindParams<T = any> {
|
|||
}
|
||||
|
||||
// Performance options
|
||||
/**
|
||||
* @deprecated Not implemented. Passing `writeOnly` throws
|
||||
* {@link UnsupportedFindOptionError} — it used to be accepted and
|
||||
* silently ignored, so validation was never actually skipped.
|
||||
*/
|
||||
writeOnly?: boolean // Skip validation for high-speed ingestion
|
||||
|
||||
// Aggregation
|
||||
|
|
@ -1426,10 +1336,7 @@ export type GroupByDimension =
|
|||
export interface AggregateSource {
|
||||
/** Filter by entity type(s) */
|
||||
type?: NounType | NounType[]
|
||||
/**
|
||||
* Metadata filter — same syntax and field-addressing law as find()'s
|
||||
* `where` (see the `@remarks` on {@link FindParams}).
|
||||
*/
|
||||
/** Metadata filter (same syntax as find({ where })) */
|
||||
where?: Record<string, unknown>
|
||||
/** Multi-tenancy service filter */
|
||||
service?: string
|
||||
|
|
@ -1443,11 +1350,7 @@ export interface AggregateDefinition {
|
|||
name: string
|
||||
/** Which entities contribute to this aggregate */
|
||||
source: AggregateSource
|
||||
/**
|
||||
* Dimensions to group by — field names follow the same field-addressing
|
||||
* law as find()'s `where` / `orderBy` (see the `@remarks` on
|
||||
* {@link FindParams}).
|
||||
*/
|
||||
/** Dimensions to group by */
|
||||
groupBy: GroupByDimension[]
|
||||
/** Named metrics to compute */
|
||||
metrics: Record<string, AggregateMetricDef>
|
||||
|
|
@ -1506,25 +1409,16 @@ export interface AggregateGroupState {
|
|||
export interface AggregateQueryParams {
|
||||
/** Name of the aggregate to query */
|
||||
name: string
|
||||
/**
|
||||
* Filter aggregate groups by their key values — same field-addressing
|
||||
* law as find() (see the `@remarks` on {@link FindParams}).
|
||||
*/
|
||||
/** Filter aggregate groups by their key values */
|
||||
where?: Record<string, unknown>
|
||||
/**
|
||||
* Filter groups by their computed METRIC values (SQL HAVING). Same BFO operators as
|
||||
* `where`, but applied to the derived metric results plus `count`, e.g.
|
||||
* `{ revenue: { greaterThan: 1000 } }`. Evaluated per group (O(groups), independent of
|
||||
* entity count), before sort/pagination. Metric names and `count` are looked up
|
||||
* directly, not field-addressed; a group-KEY field used here follows the same
|
||||
* field-addressing law as find() (see the `@remarks` on {@link FindParams}).
|
||||
* entity count), before sort/pagination.
|
||||
*/
|
||||
having?: Record<string, unknown>
|
||||
/**
|
||||
* Sort by metric name (a key from `metrics`, looked up directly) or by a
|
||||
* group key field — a group key field follows the same field-addressing
|
||||
* law as find()'s `orderBy` (see the `@remarks` on {@link FindParams}).
|
||||
*/
|
||||
/** Sort by metric name or group key field */
|
||||
orderBy?: string
|
||||
/** Sort direction */
|
||||
order?: 'asc' | 'desc'
|
||||
|
|
@ -2028,6 +1922,32 @@ export interface BrainyConfig {
|
|||
*/
|
||||
force?: boolean
|
||||
|
||||
/**
|
||||
* How write paths react when an untyped (JavaScript) caller smuggles a
|
||||
* Brainy-reserved field (`RESERVED_ENTITY_FIELDS` / `RESERVED_RELATION_FIELDS`
|
||||
* — `confidence`, `weight`, `subtype`, `visibility`, `service`, `createdBy`,
|
||||
* `noun`/`verb`, `data`, `createdAt`, `updatedAt`, `_rev`) **inside the
|
||||
* `metadata` bag** of `add()` / `update()` / `relate()` / `updateRelation()`
|
||||
* (and their `transact()` / `with()` mirrors). TypeScript callers can't write
|
||||
* these shapes at all — the compile-time guard on the metadata param types
|
||||
* (`NoReservedEntityKeys` / `NoReservedRelationKeys`) rejects a literal
|
||||
* reserved key — so this policy only governs untyped callers that slip one
|
||||
* past the compiler.
|
||||
*
|
||||
* - `'throw'` (**default, 8.0**): a reserved key in the bag throws a clear
|
||||
* `Error` naming the offending key(s) and the correct write path. No silent
|
||||
* remap, no data loss, no surprise. This is the 8.0 "no silent failures"
|
||||
* contract.
|
||||
* - `'warn'`: legacy remapping with a loud, one-shot (per key, per process)
|
||||
* warning for EVERY reserved key found — user-mutable fields are remapped to
|
||||
* their dedicated top-level param (top-level wins when both are supplied),
|
||||
* system-managed fields are dropped. Use while migrating untyped call sites.
|
||||
* - `'remap'`: the pre-8.0 silent remapping, no warning. Last-resort
|
||||
* compatibility hatch for code that intentionally relies on the bag path.
|
||||
*
|
||||
* @default 'throw'
|
||||
*/
|
||||
reservedFieldPolicy?: 'throw' | 'warn' | 'remap'
|
||||
}
|
||||
|
||||
// ============= Neural API Types =============
|
||||
|
|
|
|||
|
|
@ -1,54 +1,35 @@
|
|||
/**
|
||||
* @module types/reservedFields
|
||||
* @description The stored-record layout contract — ONE place that defines
|
||||
* which keys of a persisted metadata record belong to the ENGINE (top-level
|
||||
* entity/relationship fields) and how the USER's metadata bag is kept apart
|
||||
* from them, faithfully, across flush / reopen / rebuild / time travel.
|
||||
* @description The canonical reserved-field contract — ONE place that defines
|
||||
* which keys belong to Brainy (top-level entity/relationship fields) and may
|
||||
* therefore never live inside a `metadata` bag.
|
||||
*
|
||||
* THE FIELD-ADDRESSING LAW (ruled 2026-08-03, VENUE-BRAINY-ORDERBY-NOOP):
|
||||
* data is either in main space — where developers can use ANY name, and it
|
||||
* all works with every database function — or it is in `system.*`. There are
|
||||
* NO reserved user-facing metadata names anymore: `confidence`, `type`,
|
||||
* `level`, `data`, `id`, `content` … inside a metadata bag are ordinary user
|
||||
* fields. The only refused write is a user metadata key literally starting
|
||||
* with `'system.'` (namespace forgery — see `rejectForgedSystemKeys`).
|
||||
* Three layers enforce the contract, all driven by the constants below:
|
||||
*
|
||||
* That law makes name-based storage discrimination unsound for NEW records
|
||||
* (a user field named `confidence` may now legally sit beside the engine's
|
||||
* confidence scalar), so persisted metadata records carry the user bag
|
||||
* NESTED, shape-discriminated by a format stamp:
|
||||
* 1. **Compile time** — `AddParams.metadata`, `UpdateParams.metadata`,
|
||||
* `RelateParams.metadata` and `UpdateRelationParams.metadata` are typed so
|
||||
* a literal reserved key is a TypeScript error (see
|
||||
* {@link EntityMetadataInput} / {@link RelationMetadataInput}).
|
||||
* 2. **Write time** — for untyped (JavaScript) callers that smuggle a
|
||||
* reserved key past the compiler anyway, every write path normalizes the
|
||||
* bag: user-mutable fields are remapped to their dedicated top-level
|
||||
* param (top-level wins when both are supplied) and system-managed fields
|
||||
* are dropped with a one-shot warning naming the correct write path.
|
||||
* 3. **Read time** — every read path splits the stored flat record through
|
||||
* {@link splitNounMetadataRecord} / {@link splitVerbMetadataRecord}, so a
|
||||
* reserved field is surfaced ONLY at top level and `entity.metadata` /
|
||||
* `relation.metadata` contain ONLY custom fields, always — live reads,
|
||||
* batch reads, and historical (`asOf`) reads alike.
|
||||
*
|
||||
* - **v2 (nested-bag)** — `{ …engine fields…, [METADATA_RECORD_FORMAT_KEY]:
|
||||
* NESTED_BAG_FORMAT, metadata: { …user bag, verbatim… } }`. Built ONLY by
|
||||
* {@link buildNounMetadataRecord} / {@link buildVerbMetadataRecord}; the
|
||||
* engine half and the user bag can never collide because they never share
|
||||
* a level.
|
||||
* - **legacy (flat)** — engine fields and user fields mixed at one level,
|
||||
* discriminated BY NAME through the RESERVED_* lists. Sound for legacy
|
||||
* records precisely because the pre-law write door REFUSED user metadata
|
||||
* carrying those names — a flat key matching a reserved name IS the
|
||||
* engine's value in any record the old door admitted.
|
||||
*
|
||||
* {@link splitNounMetadataRecord} / {@link splitVerbMetadataRecord} read
|
||||
* BOTH shapes (stamp first, name split as the legacy fallback) and are the
|
||||
* single read-side choke point for live, batch, AND historical (`asOf`)
|
||||
* reads — the generation store snapshots whole records, so time travel
|
||||
* rides the same split.
|
||||
*
|
||||
* The RESERVED_* lists therefore no longer describe a user-facing ban — they
|
||||
* describe the ENGINE HALF of the stored record layout (and drive the legacy
|
||||
* split). The write-door remap machinery and the compile-time metadata key
|
||||
* bans that used to enforce the old contract are gone.
|
||||
* Documented for consumers in `docs/concepts/consistency-model.md`
|
||||
* ("Reserved fields").
|
||||
*/
|
||||
|
||||
/**
|
||||
* @description Entity (noun) field names owned by the ENGINE in a stored
|
||||
* metadata record. In v2 (nested-bag) records these are the legal TOP-LEVEL
|
||||
* keys beside the nested `metadata` bag; in legacy flat records they drive
|
||||
* the by-name split. They are NOT a user-facing ban list: since the
|
||||
* field-addressing law, a user metadata field may carry any of these names
|
||||
* and remains the user's — it lives inside the nested bag, never at the
|
||||
* record's top level.
|
||||
* @description Entity (noun) field names reserved by Brainy. These keys are
|
||||
* stored in the flat per-entity metadata record alongside custom fields, but
|
||||
* they belong to Brainy: every read path extracts them to top-level
|
||||
* `Entity` fields, and no write path accepts them inside `metadata`.
|
||||
*
|
||||
* | Key | Canonical write path |
|
||||
* |-----|----------------------|
|
||||
|
|
@ -138,54 +119,68 @@ export type ReservedRelationField = (typeof RESERVED_RELATION_FIELDS)[number]
|
|||
type IsAny<T> = 0 extends 1 & T ? true : false
|
||||
|
||||
/**
|
||||
* @deprecated The compile-time reserved-key ban died with the
|
||||
* field-addressing law: every name is legal user metadata now. Kept as an
|
||||
* empty (no-op) guard so external type references keep compiling; it bans
|
||||
* nothing.
|
||||
* @description Compile-time tripwire: marks every reserved entity key as
|
||||
* `never` so an object literal carrying one fails to type-check. Keys that
|
||||
* `T` itself declares (including via an index signature, where
|
||||
* `keyof T = string`) are exempted — a consumer who *explicitly* types a
|
||||
* reserved key into their metadata shape keeps a working (if unwise) type,
|
||||
* and index-signature metadata types remain assignable.
|
||||
*/
|
||||
export type NoReservedEntityKeys<T> = unknown
|
||||
export type NoReservedEntityKeys<T> = {
|
||||
readonly [K in ReservedEntityField as K extends keyof T ? never : K]?: never
|
||||
}
|
||||
|
||||
/**
|
||||
* @deprecated Relationship mirror of {@link NoReservedEntityKeys} — no-op
|
||||
* for the same reason.
|
||||
* @description Relationship mirror of {@link NoReservedEntityKeys}.
|
||||
*/
|
||||
export type NoReservedRelationKeys<T> = unknown
|
||||
export type NoReservedRelationKeys<T> = {
|
||||
readonly [K in ReservedRelationField as K extends keyof T ? never : K]?: never
|
||||
}
|
||||
|
||||
/**
|
||||
* @description The metadata bag shape for untyped brains (`T = any`): an
|
||||
* open index signature (any custom key, any value — exactly the pre-8.0
|
||||
* latitude) intersected with the reserved-key guard, whose declared
|
||||
* `?: never` properties take precedence over the index signature so a
|
||||
* literal reserved key is still a compile error.
|
||||
*/
|
||||
type OpenBag<Guard> = { [key: string]: any } & Guard
|
||||
|
||||
/**
|
||||
* @description The type of `AddParams.metadata`: the consumer's metadata
|
||||
* shape `T`, open. Under the field-addressing law EVERY key is a legal user
|
||||
* field (engine scalars are written only via their dedicated params and read
|
||||
* at `system.*`), so no name is banned at compile time. The one illegal
|
||||
* spelling — a key starting `'system.'` — cannot be expressed as a mapped
|
||||
* type ban and is refused at runtime (`rejectForgedSystemKeys`).
|
||||
* shape `T` with reserved entity keys forbidden at compile time. For untyped
|
||||
* brains (`T = any`) the bag stays open ({@link OpenBag}), so arbitrary
|
||||
* custom fields remain legal while literal reserved keys still error.
|
||||
*/
|
||||
export type EntityMetadataInput<T> = IsAny<T> extends true
|
||||
? { [key: string]: any }
|
||||
: T
|
||||
? OpenBag<NoReservedEntityKeys<object>>
|
||||
: T & NoReservedEntityKeys<T>
|
||||
|
||||
/**
|
||||
* @description The type of `UpdateParams.metadata`: a partial patch of the
|
||||
* consumer's metadata shape. Same openness as {@link EntityMetadataInput}.
|
||||
* consumer's metadata shape with reserved entity keys forbidden at compile
|
||||
* time. Same `T = any` handling as {@link EntityMetadataInput}.
|
||||
*/
|
||||
export type EntityMetadataPatch<T> = IsAny<T> extends true
|
||||
? { [key: string]: any }
|
||||
: Partial<T>
|
||||
? OpenBag<NoReservedEntityKeys<object>>
|
||||
: Partial<T> & NoReservedEntityKeys<T>
|
||||
|
||||
/**
|
||||
* @description The type of `RelateParams.metadata`: the consumer's edge
|
||||
* metadata shape, open — the relation mirror of {@link EntityMetadataInput}.
|
||||
* metadata shape with reserved relationship keys forbidden at compile time.
|
||||
*/
|
||||
export type RelationMetadataInput<T> = IsAny<T> extends true
|
||||
? { [key: string]: any }
|
||||
: T
|
||||
? OpenBag<NoReservedRelationKeys<object>>
|
||||
: T & NoReservedRelationKeys<T>
|
||||
|
||||
/**
|
||||
* @description The type of `UpdateRelationParams.metadata`: a partial patch
|
||||
* of the consumer's edge metadata shape, open.
|
||||
* of the consumer's edge metadata shape with reserved relationship keys
|
||||
* forbidden at compile time.
|
||||
*/
|
||||
export type RelationMetadataPatch<T> = IsAny<T> extends true
|
||||
? { [key: string]: any }
|
||||
: Partial<T>
|
||||
? OpenBag<NoReservedRelationKeys<object>>
|
||||
: Partial<T> & NoReservedRelationKeys<T>
|
||||
|
||||
/**
|
||||
* @description Result of splitting a stored flat metadata record into its
|
||||
|
|
@ -201,103 +196,6 @@ export interface SplitMetadataRecord<F extends string> {
|
|||
const RESERVED_ENTITY_SET: ReadonlySet<string> = new Set(RESERVED_ENTITY_FIELDS)
|
||||
const RESERVED_RELATION_SET: ReadonlySet<string> = new Set(RESERVED_RELATION_FIELDS)
|
||||
|
||||
/**
|
||||
* @description The format-stamp key of a persisted metadata record. Its
|
||||
* presence with the exact value {@link NESTED_BAG_FORMAT} marks a v2
|
||||
* (nested-bag) record; its absence marks a legacy flat record. The stamp is
|
||||
* what makes the shape check collision-proof against legacy user data: a
|
||||
* pre-law record COULD carry a user field named `metadata` (the name was
|
||||
* never reserved), but it cannot also carry this engine-written stamp.
|
||||
*/
|
||||
export const METADATA_RECORD_FORMAT_KEY = '_fmt'
|
||||
|
||||
/**
|
||||
* @description The nested-bag record format stamp (v2, the field-addressing
|
||||
* law's storage shape, 2026-08-03): engine fields at top level, the user's
|
||||
* metadata bag NESTED verbatim under `metadata`. Cross-engine: the native
|
||||
* provider discriminates record shapes by the same stamp.
|
||||
*/
|
||||
export const NESTED_BAG_FORMAT = 2
|
||||
|
||||
/**
|
||||
* @description `true` when a persisted record carries the v2 nested-bag
|
||||
* stamp (and a structurally valid nested bag).
|
||||
*/
|
||||
export function isNestedBagRecord(
|
||||
record: Record<string, unknown> | null | undefined
|
||||
): boolean {
|
||||
return (
|
||||
record !== null &&
|
||||
record !== undefined &&
|
||||
typeof record === 'object' &&
|
||||
record[METADATA_RECORD_FORMAT_KEY] === NESTED_BAG_FORMAT &&
|
||||
typeof record.metadata === 'object' &&
|
||||
record.metadata !== null &&
|
||||
!Array.isArray(record.metadata)
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Build a v2 (nested-bag) entity metadata record — THE only
|
||||
* sanctioned way to construct a persisted noun metadata record. The engine
|
||||
* half goes top-level; the user bag nests verbatim under `metadata`; the
|
||||
* format stamp seals the shape. Because the two halves never share a level,
|
||||
* a user field named `confidence` (or any other engine spelling) survives
|
||||
* flush / reopen / rebuild / time travel exactly as written.
|
||||
* @param engineFields - The engine-owned half (keys from
|
||||
* {@link RESERVED_ENTITY_FIELDS} — `noun`, timestamps, `_rev`, …).
|
||||
* @param userBag - The consumer's metadata bag, stored verbatim.
|
||||
* @returns The stamped v2 record.
|
||||
*/
|
||||
export function buildNounMetadataRecord(
|
||||
engineFields: Partial<Record<ReservedEntityField, unknown>>,
|
||||
userBag: Record<string, unknown> | undefined
|
||||
): Record<string, unknown> {
|
||||
return {
|
||||
...engineFields,
|
||||
[METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT,
|
||||
metadata: { ...(userBag ?? {}) }
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Build a v2 (nested-bag) relationship metadata record — the
|
||||
* verb mirror of {@link buildNounMetadataRecord}.
|
||||
* @param engineFields - The engine-owned half (keys from
|
||||
* {@link RESERVED_RELATION_FIELDS} — `verb`, `weight`, timestamps, …).
|
||||
* @param userBag - The consumer's edge metadata bag, stored verbatim.
|
||||
* @returns The stamped v2 record.
|
||||
*/
|
||||
export function buildVerbMetadataRecord(
|
||||
engineFields: Partial<Record<ReservedRelationField, unknown>>,
|
||||
userBag: Record<string, unknown> | undefined
|
||||
): Record<string, unknown> {
|
||||
return {
|
||||
...engineFields,
|
||||
[METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT,
|
||||
metadata: { ...(userBag ?? {}) }
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Shape-first split of a v2 record: the engine half is the top
|
||||
* level filtered through the reserved list (belt — the builders only ever
|
||||
* write reserved names there), the user bag is `record.metadata` verbatim.
|
||||
*/
|
||||
function splitNestedRecord<F extends string>(
|
||||
record: Record<string, unknown>,
|
||||
reservedSet: ReadonlySet<string>
|
||||
): SplitMetadataRecord<F> {
|
||||
const reserved: Record<string, unknown> = {}
|
||||
for (const [key, value] of Object.entries(record)) {
|
||||
if (reservedSet.has(key)) reserved[key] = value
|
||||
}
|
||||
return {
|
||||
reserved: reserved as Partial<Record<F, unknown>>,
|
||||
custom: { ...(record.metadata as Record<string, unknown>) }
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Shared splitter — partitions a record's keys against a
|
||||
* reserved-name set. `null`/`undefined` records split to two empty objects.
|
||||
|
|
@ -324,45 +222,33 @@ function splitRecord<F extends string>(
|
|||
}
|
||||
|
||||
/**
|
||||
* @description Split a stored entity (noun) metadata record into engine
|
||||
* fields and the user's metadata bag — THE canonical read-side split, shape
|
||||
* aware. v2 (nested-bag) records split by SHAPE: engine half top-level, bag
|
||||
* = `record.metadata` verbatim (user collider names survive faithfully).
|
||||
* Legacy flat records split BY NAME through the reserved list — sound for
|
||||
* them because the pre-law write door refused user metadata carrying those
|
||||
* names. Every entity read path (live `get()`, batch reads, paginated
|
||||
* listings, and historical `asOf()` materialization — the generation store
|
||||
* snapshots whole records) goes through this function, so the two shapes
|
||||
* can never drift between read paths.
|
||||
* @param record - The stored metadata record (either shape).
|
||||
* @returns `reserved` (engine-owned fields) and `custom` (the consumer's metadata bag).
|
||||
* @description Split a stored entity (noun) flat metadata record into
|
||||
* reserved fields and custom metadata — THE canonical read-side split. Every
|
||||
* entity read path (live `get()`, batch reads, paginated listings, and
|
||||
* historical `asOf()` materialization) goes through this function, so the
|
||||
* reserved list can never drift between read paths.
|
||||
* @param record - The stored flat metadata record.
|
||||
* @returns `reserved` (Brainy-owned fields) and `custom` (the consumer's metadata bag).
|
||||
* @example
|
||||
* const { reserved, custom } = splitNounMetadataRecord(stored)
|
||||
* // reserved.noun → entity.type, reserved.confidence → entity.confidence, …
|
||||
* // custom → entity.metadata (the user's fields only, always — ANY names)
|
||||
* // custom → entity.metadata (custom fields only, always)
|
||||
*/
|
||||
export function splitNounMetadataRecord(
|
||||
record: Record<string, unknown> | null | undefined
|
||||
): SplitMetadataRecord<ReservedEntityField> {
|
||||
if (isNestedBagRecord(record)) {
|
||||
return splitNestedRecord(record as Record<string, unknown>, RESERVED_ENTITY_SET)
|
||||
}
|
||||
return splitRecord(record, RESERVED_ENTITY_SET)
|
||||
}
|
||||
|
||||
/**
|
||||
* @description Split a stored relationship (verb) metadata record into
|
||||
* engine fields and the user's edge metadata bag — the verb mirror of
|
||||
* {@link splitNounMetadataRecord}, shape aware, used by every relationship
|
||||
* read path.
|
||||
* @param record - The stored metadata record (either shape).
|
||||
* @returns `reserved` (engine-owned fields) and `custom` (the consumer's metadata bag).
|
||||
* @description Split a stored relationship (verb) flat metadata record into
|
||||
* reserved fields and custom metadata — the verb mirror of
|
||||
* {@link splitNounMetadataRecord}, used by every relationship read path.
|
||||
* @param record - The stored flat metadata record.
|
||||
* @returns `reserved` (Brainy-owned fields) and `custom` (the consumer's metadata bag).
|
||||
*/
|
||||
export function splitVerbMetadataRecord(
|
||||
record: Record<string, unknown> | null | undefined
|
||||
): SplitMetadataRecord<ReservedRelationField> {
|
||||
if (isNestedBagRecord(record)) {
|
||||
return splitNestedRecord(record as Record<string, unknown>, RESERVED_RELATION_SET)
|
||||
}
|
||||
return splitRecord(record, RESERVED_RELATION_SET)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -5,7 +5,6 @@
|
|||
*/
|
||||
|
||||
import { StorageAdapter, resolveEntityField, NounMetadata, VerbMetadata } from '../coreTypes.js'
|
||||
import { SYSTEM_ENTITY_SCALARS, parseFieldAddress, UnresolvableFieldError } from '../db/fieldAddressing.js'
|
||||
import { ColumnStore } from '../indexes/columnStore/ColumnStore.js'
|
||||
import type { MetadataIndexProvider } from '../plugin.js'
|
||||
import { MetadataIndexCache, MetadataIndexCacheConfig } from './metadataIndexCache.js'
|
||||
|
|
@ -44,8 +43,8 @@ import { BrainyError } from '../errors/brainyError.js'
|
|||
* bucketed field is added (e.g. a compressed float), add it here too.
|
||||
*/
|
||||
const BUCKETED_INDEX_FIELDS: ReadonlySet<string> = new Set([
|
||||
'system.createdAt',
|
||||
'system.updatedAt'
|
||||
'createdAt',
|
||||
'updatedAt'
|
||||
])
|
||||
|
||||
export interface MetadataIndexEntry {
|
||||
|
|
@ -73,11 +72,8 @@ export interface MetadataIndexConfig {
|
|||
maxIndexSize?: number // Max number of entries per field value (default: 10000)
|
||||
rebuildThreshold?: number // Rebuild if index is this % stale (default: 0.1)
|
||||
autoOptimize?: boolean // Auto-cleanup unused entries (default: true)
|
||||
// NOTE: the name-based indexedFields/excludeFields knobs died with the
|
||||
// field-addressing law ("no special names"): EVERY user field indexes,
|
||||
// whatever its name. Bulk-payload protection is value-SHAPE based and
|
||||
// uniform across all names (large arrays never become posting scalars;
|
||||
// long values index hashed) — shape is not a name carve-out.
|
||||
indexedFields?: string[] // Only index these fields (default: all)
|
||||
excludeFields?: string[] // Never index these fields
|
||||
}
|
||||
|
||||
export interface MetadataIndexOptions {
|
||||
|
|
@ -188,12 +184,31 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
this.config = {
|
||||
maxIndexSize: config.maxIndexSize ?? 10000,
|
||||
rebuildThreshold: config.rebuildThreshold ?? 0.1,
|
||||
autoOptimize: config.autoOptimize ?? true
|
||||
// No name-based exclude/allow lists — the field-addressing law: every
|
||||
// user field indexes, whatever its name ('content', 'data', 'id',
|
||||
// 'vector', … included). Bulk payloads are kept out by uniform value-
|
||||
// SHAPE rules in extractIndexableFields (arrays >10 never become
|
||||
// posting scalars; >100-char values index hashed), never by name.
|
||||
autoOptimize: config.autoOptimize ?? true,
|
||||
indexedFields: config.indexedFields ?? [],
|
||||
excludeFields: config.excludeFields ?? [
|
||||
// ONLY exclude truly un-indexable fields (binary data, large content)
|
||||
// Timestamps are NOW indexed with automatic bucketing (prevents pollution)
|
||||
|
||||
// Vectors and embeddings (binary data, already have HNSW indexes)
|
||||
'embedding',
|
||||
'vector',
|
||||
'embeddings',
|
||||
'vectors',
|
||||
|
||||
// Large content fields (too large for metadata indexing)
|
||||
'content',
|
||||
'data',
|
||||
'originalData',
|
||||
'_data',
|
||||
|
||||
// Primary keys (use direct lookups instead)
|
||||
'id'
|
||||
|
||||
// NOTE: 'accessed', 'modified', 'createdAt', etc. are NO LONGER excluded!
|
||||
// They are now indexed with automatic 1-minute bucketing to prevent file pollution
|
||||
// This enables range queries like: modified > yesterday
|
||||
]
|
||||
}
|
||||
|
||||
// Initialize metadata cache with similar config to search cache
|
||||
|
|
@ -285,7 +300,7 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
}
|
||||
|
||||
// Warm the cache with common fields (lazy loading optimization)
|
||||
// This loads the type column ('system.type') needed for type counts
|
||||
// This loads the 'noun' sparse index which is needed for type counts
|
||||
await this.warmCache()
|
||||
|
||||
// Load type counts AFTER warmCache (sparse index is now cached)
|
||||
|
|
@ -334,9 +349,8 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
* Target: >80% cache hit rate for typical workloads
|
||||
*/
|
||||
async warmCache(): Promise<void> {
|
||||
// Common columns used in most queries — the frozen system keys, plus
|
||||
// legacy spellings for a pre-epoch-3 brain read before its rebuild runs.
|
||||
const commonFields = ['system.type', 'system.service', 'system.createdAt', 'noun']
|
||||
// Common fields used in most queries
|
||||
const commonFields = ['noun', 'type', 'service', 'createdAt']
|
||||
|
||||
prodLog.debug(`🔥 Warming metadata cache with common fields: ${commonFields.join(', ')}`)
|
||||
|
||||
|
|
@ -522,11 +536,9 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
}
|
||||
|
||||
/**
|
||||
* Lazy load entity counts from the type column (O(n) where n = number of
|
||||
* types). The frozen key is 'system.type' (epoch 3); the legacy 'noun'
|
||||
* column is read as a fallback for a pre-epoch-3 brain observed before its
|
||||
* rebuild has run (e.g. a reader-mode open against an old writer).
|
||||
* Lazy load entity counts from the 'noun' field sparse index (O(n) where n = number of types)
|
||||
* FIX: Previously read from stats.nounCount which was SERVICE-keyed, not TYPE-keyed
|
||||
* Now computes counts from the sparse index which has the correct type information
|
||||
*/
|
||||
private async lazyLoadCounts(): Promise<void> {
|
||||
try {
|
||||
|
|
@ -536,31 +548,23 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
this.entityCountsByTypeFixed.fill(0)
|
||||
this.verbCountsByTypeFixed.fill(0)
|
||||
|
||||
// PRIMARY (8.0+): rehydrate per-type counts from the column store's
|
||||
// type column — the authoritative on-disk source after a cold reopen.
|
||||
// Frozen key first ('system.type', epoch 3), legacy 'noun' as the
|
||||
// pre-rebuild fallback.
|
||||
// PRIMARY (8.0+): rehydrate per-type counts from the column store's 'noun'
|
||||
// field — the authoritative on-disk source after a cold reopen.
|
||||
//
|
||||
// The chunked sparse-index WRITE path was removed in 7.20.0 (commit
|
||||
// 11be039): new workspaces persist the type column ONLY to the column
|
||||
// store, never to a sparse-index blob. So the legacy sparse path below
|
||||
// finds nothing and leaves every count at 0 — which is exactly why
|
||||
// counts.byType/byTypeEnum/topTypes/allNounTypeCounts all read empty
|
||||
// 11be039): new workspaces persist the 'noun' field ONLY to the column
|
||||
// store, never to a `__sparse_index__noun` blob. So the legacy sparse
|
||||
// path below finds nothing and leaves every count at 0 — which is exactly
|
||||
// why counts.byType/byTypeEnum/topTypes/allNounTypeCounts all read empty
|
||||
// after close()+reopen while find()/getNounCount() (different sources)
|
||||
// stay correct. The column store's per-value cardinality matches the warm
|
||||
// `updateTypeFieldAffinity` counts EXACTLY because both are driven from the
|
||||
// same `addToIndex` field set, in lockstep, with no visibility gate on
|
||||
// either — so this rehydration reproduces the warm values precisely.
|
||||
const indexedCols = this.columnStore ? this.columnStore.getIndexedFields() : []
|
||||
const typeCol = indexedCols.includes('system.type')
|
||||
? 'system.type'
|
||||
: indexedCols.includes('noun')
|
||||
? 'noun'
|
||||
: null
|
||||
if (this.columnStore && typeCol) {
|
||||
const nounValues = await this.columnStore.getFilterValues(typeCol)
|
||||
if (this.columnStore && this.columnStore.getIndexedFields().includes('noun')) {
|
||||
const nounValues = await this.columnStore.getFilterValues('noun')
|
||||
for (const value of nounValues) {
|
||||
const bitmap = await this.columnStore.filter(typeCol, value)
|
||||
const bitmap = await this.columnStore.filter('noun', value)
|
||||
if (bitmap.size > 0) {
|
||||
// Use the stored value directly as the key (the legacy sparse path
|
||||
// did the same): it is already the normalized type string that
|
||||
|
|
@ -575,17 +579,16 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
}
|
||||
|
||||
// LEGACY FALLBACK (pre-7.20.0 workspaces still on the chunked sparse index).
|
||||
const sparseCol = (await this.loadSparseIndex('system.type')) ? 'system.type' : 'noun'
|
||||
const nounSparseIndex = await this.loadSparseIndex(sparseCol)
|
||||
const nounSparseIndex = await this.loadSparseIndex('noun')
|
||||
if (!nounSparseIndex) {
|
||||
// No column-store type column and no sparse index yet — counts will be
|
||||
// No column-store 'noun' field and no sparse index yet — counts will be
|
||||
// populated as entities are added.
|
||||
return
|
||||
}
|
||||
|
||||
// Iterate through all chunks and sum up bitmap sizes by type
|
||||
for (const chunkId of nounSparseIndex.getAllChunkIds()) {
|
||||
const chunk = await this.chunkManager.loadChunk(sparseCol, chunkId)
|
||||
const chunk = await this.chunkManager.loadChunk('noun', chunkId)
|
||||
if (chunk) {
|
||||
for (const [type, bitmap] of chunk.entries) {
|
||||
const currentCount = this.totalEntitiesByType.get(type) || 0
|
||||
|
|
@ -1175,86 +1178,76 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
return `__HASH_${Math.abs(hash).toString(36)}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if field should be indexed
|
||||
*/
|
||||
private shouldIndexField(field: string): boolean {
|
||||
if (this.config.excludeFields.includes(field)) return false
|
||||
if (this.config.indexedFields.length > 0) {
|
||||
return this.config.indexedFields.includes(field)
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract indexable field-value pairs from entity or metadata
|
||||
*
|
||||
* Handles BOTH entity structure (with top-level fields) AND record shapes
|
||||
* - Record-frame system scalars index under literal 'system.<field>' keys
|
||||
* - The user's metadata bag indexes under bare keys — EVERY name (the
|
||||
* field-addressing law: no special names; 'level', 'data', 'id',
|
||||
* 'content', 'vector' in a bag are ordinary user fields)
|
||||
* - Record-frame plumbing (vector, connections, level, data, _rev, id)
|
||||
* never indexes — that is namespace routing, not a name carve-out
|
||||
* - Value-SHAPE rules apply uniformly to all names: arrays >10 never
|
||||
* become posting scalars; purely numeric key names (array indices)
|
||||
* skip; >100-char values index hashed (normalizeValue)
|
||||
* Now handles BOTH entity structure (with top-level fields) AND plain metadata
|
||||
* - Extracts from top-level fields (confidence, weight, timestamps, type, service, etc.)
|
||||
* - Also extracts from nested metadata field (custom user fields)
|
||||
* - Skips HNSW-specific fields (vector, connections, level, id)
|
||||
* - Maps 'type' → 'noun' for backward compatibility with existing indexes
|
||||
*
|
||||
* BUG FIX: Exclude vector embeddings and large arrays from indexing
|
||||
* BUG FIX: Also exclude purely numeric field names (array indices)
|
||||
* - Vector fields (384+ dimensions) were creating 825K chunk files for 1,144 entities
|
||||
* - Arrays converted to objects with numeric keys were still being indexed
|
||||
*/
|
||||
private extractIndexableFields(data: any): Array<{ field: string, value: any }> {
|
||||
const fields: Array<{ field: string, value: any }> = []
|
||||
|
||||
// RECORD-FRAME-ONLY plumbing guard: on an entity/stored-record frame
|
||||
// these keys are the engine's structural payloads (the 384-dim vector,
|
||||
// embeddings, the adjacency list, the identity field) and never index.
|
||||
// This set is NEVER applied inside the user's metadata bag — under the
|
||||
// field-addressing law every user name indexes; a real vector-sized
|
||||
// value in a bag is kept out by the uniform array-size shape guard, not
|
||||
// by its name.
|
||||
const RECORD_PLUMBING = new Set(['vector', 'embedding', 'embeddings', 'connections', 'id'])
|
||||
// Fields that should NEVER be indexed: bulk structural payloads that would
|
||||
// blow up the index (the 384-dim vector, embeddings, the adjacency list).
|
||||
// These are also caught by the array-size guard below, but naming them is
|
||||
// belt-and-suspenders. NOTE: `level` was previously here (an HNSW node's
|
||||
// layer) but it never actually reaches this path — every caller passes a
|
||||
// metadata bag or Entity record, neither of which carries the node's
|
||||
// `level` — so its only effect was to silently drop a legitimate USER
|
||||
// metadata field named `level` (log level, skill level, access level…),
|
||||
// making `where: { level: … }` return nothing. Removed. (`id` stays: it is
|
||||
// the reserved entity-identity field, resolved specially by find().)
|
||||
const NEVER_INDEX = new Set(['vector', 'embedding', 'embeddings', 'connections', 'id'])
|
||||
|
||||
// THE FROZEN INDEX KEY FORMAT (cross-engine, sealed 2026-08-03; the native
|
||||
// accelerator keys identically — epoch 3 rebuilds every brain onto it):
|
||||
// user fields index under BARE keys exactly as the caller wrote them;
|
||||
// the ten system scalars index under literal 'system.<field>' keys — the
|
||||
// key IS the query address, so the two namespaces can never collide
|
||||
// inside the index again.
|
||||
// Frame kinds: 'entity-record' = entityForIndexing shape / v2 nested-bag
|
||||
// stored record (user fields nested under `metadata`; stray top-level
|
||||
// keys are DROPPED, not guessed); 'flat-record' = the LEGACY stored
|
||||
// metadata-record shape (user fields flat beside the engine's — sound to
|
||||
// split by name because the pre-law write door refused user metadata
|
||||
// carrying engine names, so a flat key matching a system name IS the
|
||||
// system value); 'user' = inside the metadata bag, where EVERY key is
|
||||
// the user's and indexes bare — collider names included.
|
||||
type Frame = 'entity-record' | 'flat-record' | 'user'
|
||||
const extract = (obj: any, prefix = '', frame: Frame = 'entity-record'): void => {
|
||||
const extract = (obj: any, prefix = ''): void => {
|
||||
for (const [key, value] of Object.entries(obj)) {
|
||||
let fullKey = prefix ? `${prefix}.${key}` : key
|
||||
const fullKey = prefix ? `${prefix}.${key}` : key
|
||||
|
||||
if (!prefix && frame !== 'user') {
|
||||
if (key === 'metadata' && typeof value === 'object' && value !== null && !Array.isArray(value)) {
|
||||
extract(value, '', 'user') // the user's namespace: bare keys
|
||||
continue
|
||||
}
|
||||
if (key === 'type' || key === 'noun') {
|
||||
fullKey = 'system.type' // legacy 'noun' spelling folds into the frozen key
|
||||
} else if (SYSTEM_ENTITY_SCALARS.has(key) && key !== 'id') {
|
||||
fullKey = `system.${key}`
|
||||
} else if (
|
||||
key === 'data' || key === '_rev' || key === 'level' || key === '_fmt' ||
|
||||
RECORD_PLUMBING.has(key)
|
||||
) {
|
||||
continue // plumbing / identity / format stamp — never indexed from a record frame
|
||||
} else if (frame === 'entity-record') {
|
||||
continue // stray entity-frame key: dropped, not guessed
|
||||
}
|
||||
// flat-record fallthrough: a non-system, non-plumbing key IS a user
|
||||
// field (flat beside the engine's, legacy shape) — indexes bare.
|
||||
}
|
||||
// User frame: NO name-based skips — every user field indexes, whatever
|
||||
// its name (the field-addressing law). Only the uniform value-shape
|
||||
// guards below apply.
|
||||
// Skip fields in never-index list (CRITICAL: prevents vector indexing bug + HNSW fields)
|
||||
if (!prefix && NEVER_INDEX.has(key)) continue
|
||||
|
||||
// Skip purely numeric field names (array indices converted to object keys)
|
||||
// Legitimate field names should never be purely numeric
|
||||
// This catches vectors stored as objects: {0: 0.1, 1: 0.2, ...}
|
||||
if (/^\d+$/.test(key)) continue
|
||||
|
||||
// Skip fields based on user configuration
|
||||
if (!this.shouldIndexField(fullKey)) continue
|
||||
|
||||
// Special handling for metadata field at top level
|
||||
// Flatten metadata fields to top-level (no prefix) for cleaner queries
|
||||
// Standard fields are already at top-level, custom fields go in metadata
|
||||
// By flattening here, queries can use { category: 'B' } instead of { 'metadata.category': 'B' }
|
||||
if (key === 'metadata' && !prefix && typeof value === 'object' && !Array.isArray(value)) {
|
||||
extract(value, '') // Flatten to top-level, no prefix
|
||||
continue
|
||||
}
|
||||
|
||||
// Skip large arrays (> 10 elements) - likely vectors or bulk data
|
||||
if (Array.isArray(value) && value.length > 10) continue
|
||||
|
||||
if (value && typeof value === 'object' && !Array.isArray(value)) {
|
||||
// Recurse into nested objects (but not arrays), keeping the frame
|
||||
extract(value, fullKey, frame)
|
||||
// Recurse into nested objects (but not arrays)
|
||||
extract(value, fullKey)
|
||||
} else if (Array.isArray(value) && value.length <= 10) {
|
||||
// Small arrays: index as multi-value field (all with same field name)
|
||||
// Example: tags: ["javascript", "node"] → field="tags", value="javascript" + field="tags", value="node"
|
||||
|
|
@ -1265,21 +1258,16 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
}
|
||||
}
|
||||
} else {
|
||||
// Primitive value: index it under the frozen key computed above.
|
||||
// (The legacy 'type'→'noun' remap is gone — 'noun' columns die at
|
||||
// the epoch-3 rebuild; system.type is the one spelling.)
|
||||
fields.push({ field: fullKey, value })
|
||||
// Primitive value: index it
|
||||
// Map 'type' → 'noun' for backward compatibility
|
||||
const indexField = (!prefix && key === 'type') ? 'noun' : fullKey
|
||||
fields.push({ field: indexField, value })
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (data && typeof data === 'object') {
|
||||
// Shape detection for the top frame: an object carrying a nested
|
||||
// `metadata` bag is the entityForIndexing shape; anything else is the
|
||||
// flat stored-record shape (user fields flat beside reserved ones).
|
||||
const entityShaped =
|
||||
'metadata' in data && typeof data.metadata === 'object' && data.metadata !== null
|
||||
extract(data, '', entityShaped ? 'entity-record' : 'flat-record')
|
||||
extract(data)
|
||||
}
|
||||
|
||||
// Extract words for hybrid text search
|
||||
|
|
@ -1481,11 +1469,10 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
prodLog.debug(`Entity ${id} has ${wordFields.length} indexed words (large document)`)
|
||||
}
|
||||
|
||||
// Sort fields to process the type column first for type-field affinity
|
||||
// tracking ('system.type' is the frozen key; 'noun' died at epoch 3).
|
||||
// Sort fields to process 'noun' field first for type-field affinity tracking
|
||||
fields.sort((a, b) => {
|
||||
if (a.field === 'system.type') return -1
|
||||
if (b.field === 'system.type') return 1
|
||||
if (a.field === 'noun') return -1
|
||||
if (b.field === 'noun') return 1
|
||||
return 0
|
||||
})
|
||||
|
||||
|
|
@ -1924,15 +1911,22 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
// Skip logical operators
|
||||
if (rawField === 'allOf' || rawField === 'anyOf' || rawField === 'not') continue
|
||||
|
||||
// THE ONE ADDRESSING LAW (sealed 2026-08-03): every filter key routes
|
||||
// through parseFieldAddress — bare and 'metadata.'-prefixed spellings
|
||||
// address the user's fields (indexed under BARE keys), 'system.<field>'
|
||||
// addresses the ten engine scalars (indexed under their literal
|
||||
// 'system.<field>' keys). A malformed address (system.<not-in-map>,
|
||||
// plumbing in the system spelling) throws typed BEFORE any index read —
|
||||
// an accepted name either works or refuses.
|
||||
const address = parseFieldAddress(rawField, 'entity')
|
||||
const field = address.scope === 'system' ? `system.${address.field}` : address.field
|
||||
// Metadata is FLATTENED at index time (metadata.entry.title indexes as
|
||||
// entry.title), so a `metadata.`-prefixed where key is almost always
|
||||
// the caller spelling the STORAGE shape rather than the index shape.
|
||||
// Accept both spellings: when the key as spelled is unindexed but its
|
||||
// stripped spelling is, query the stripped one. A literal nested
|
||||
// custom key named `metadata` still wins when indexed as spelled
|
||||
// (checked first), so that rare shape keeps working.
|
||||
let field = rawField
|
||||
if (
|
||||
rawField.startsWith('metadata.') &&
|
||||
this.columnStore &&
|
||||
!this.columnStore.hasField(rawField) &&
|
||||
this.columnStore.hasField(rawField.slice('metadata.'.length))
|
||||
) {
|
||||
field = rawField.slice('metadata.'.length)
|
||||
}
|
||||
|
||||
let fieldResults: string[] = []
|
||||
|
||||
|
|
@ -2213,30 +2207,9 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
order: 'asc' | 'desc' = 'asc',
|
||||
topK?: number
|
||||
): Promise<string[]> {
|
||||
// THE ONE ADDRESSING LAW — the orderBy address routes through the same
|
||||
// parse the filter path uses (the historical asymmetry where the filter
|
||||
// path understood 'metadata.' but the sorted path never did is dead).
|
||||
// Bare / 'metadata.' → the user's bare index key; 'system.<field>' → the
|
||||
// literal frozen key; malformed addresses throw typed before any read.
|
||||
const orderAddress = parseFieldAddress(orderBy, 'entity')
|
||||
const orderKey =
|
||||
orderAddress.scope === 'system' ? `system.${orderAddress.field}` : orderAddress.field
|
||||
|
||||
// DATA-AWARE REFUSAL (the did-you-mean): a bare address no user field
|
||||
// carries cannot mean anything as a sort key — and when the name collides
|
||||
// with a system scalar the caller almost certainly meant system.<field>.
|
||||
// Refusing loudly with both candidates beats silently sorting nothing.
|
||||
if (
|
||||
orderAddress.scope === 'metadata' &&
|
||||
!(this.columnStore && this.columnStore.hasField(orderKey)) &&
|
||||
!(await this.loadSparseIndex(orderKey))
|
||||
) {
|
||||
throw new UnresolvableFieldError(orderAddress.raw, 'entity')
|
||||
}
|
||||
|
||||
// Column store path: O(K log S) sort via k-way merge across segments.
|
||||
// No per-entity storage reads, no precision loss from bucketing.
|
||||
if (this.columnStore && this.columnStore.hasField(orderKey)) {
|
||||
if (this.columnStore && this.columnStore.hasField(orderBy)) {
|
||||
// Get filtered IDs from existing roaring bitmap path
|
||||
const hasFilter = filter && Object.keys(filter).length > 0
|
||||
const filteredIds = hasFilter ? await this.getIdsForFilter(filter) : []
|
||||
|
|
@ -2256,41 +2229,20 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
// log K) heap, not a full sort materialization.
|
||||
const k = topK !== undefined ? Math.min(topK, filteredIds.length) : filteredIds.length
|
||||
sortedIntIds = await this.columnStore.filteredSortTopK(
|
||||
filterBitmap, orderKey, order, k
|
||||
filterBitmap, orderBy, order, k
|
||||
)
|
||||
} else {
|
||||
// Unfiltered sort — column store handles the full entity set efficiently
|
||||
sortedIntIds = await this.columnStore.sortTopK(
|
||||
orderKey, order, topK !== undefined ? Math.min(topK, this.idMapper.size) : this.idMapper.size
|
||||
orderBy, order, topK !== undefined ? Math.min(topK, this.idMapper.size) : this.idMapper.size
|
||||
)
|
||||
}
|
||||
|
||||
// Convert int IDs back to UUIDs. Number() narrowing is lossless — the
|
||||
// shipped EntityIdSpaceExceeded guard caps the JS mapper at u32.
|
||||
const sortedUuids = sortedIntIds
|
||||
return sortedIntIds
|
||||
.map(intId => this.idMapper.getUuid(Number(intId)))
|
||||
.filter((uuid): uuid is string => uuid !== undefined)
|
||||
|
||||
// ORDERING CONTRACT (cross-engine, sealed): rows missing the field are
|
||||
// NEVER dropped — they sort LAST in both directions — and ties break by
|
||||
// id ascending. The column only contains rows that HAVE the field, so
|
||||
// (1) re-sort the page deterministically (value, then id) with K cheap
|
||||
// value reads, and (2) append the filtered rows the column omitted,
|
||||
// id-ascending, filling any remaining page budget.
|
||||
const page = await Promise.all(
|
||||
sortedUuids.map(async id => ({ id, value: await this.getFieldValueForEntity(id, orderKey) }))
|
||||
)
|
||||
page.sort((a, b) => this.compareAddressedValues(a.value, b.value, a.id, b.id, order))
|
||||
let result = page.map(p => p.id)
|
||||
|
||||
if (hasFilter) {
|
||||
const present = new Set(sortedUuids)
|
||||
if (topK === undefined || result.length < topK) {
|
||||
const missing = filteredIds.filter(id => !present.has(id)).sort()
|
||||
result = result.concat(missing)
|
||||
}
|
||||
}
|
||||
return topK !== undefined ? result.slice(0, topK) : result
|
||||
}
|
||||
|
||||
// Fallback: sparse index path (for fields not yet in column store).
|
||||
|
|
@ -2303,11 +2255,27 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
|
||||
const idValuePairs: Array<{ id: string, value: any }> = []
|
||||
for (const id of filteredIds) {
|
||||
const value = await this.getFieldValueForEntity(id, orderKey)
|
||||
const value = await this.getFieldValueForEntity(id, orderBy)
|
||||
idValuePairs.push({ id, value })
|
||||
}
|
||||
|
||||
idValuePairs.sort((a, b) => this.compareAddressedValues(a.value, b.value, a.id, b.id, order))
|
||||
idValuePairs.sort((a, b) => {
|
||||
if (a.value == null && b.value == null) return 0
|
||||
if (a.value == null) return order === 'asc' ? 1 : -1
|
||||
if (b.value == null) return order === 'asc' ? -1 : 1
|
||||
if (a.value === b.value) return 0
|
||||
// Numbers compare numerically; everything else by code-point (UTF-8 byte) order.
|
||||
// This makes the JS fallback sort match cor's native column store exactly
|
||||
// (numeric i64/f64 vs code-point strings) and stay deterministic across
|
||||
// environments, unlike the `<` operator's UTF-16 ordering for strings.
|
||||
let comparison: number
|
||||
if (typeof a.value === 'number' && typeof b.value === 'number') {
|
||||
comparison = a.value < b.value ? -1 : 1
|
||||
} else {
|
||||
comparison = compareCodePoints(String(a.value), String(b.value))
|
||||
}
|
||||
return order === 'asc' ? comparison : -comparison
|
||||
})
|
||||
|
||||
const sorted = idValuePairs.map(p => p.id)
|
||||
return topK !== undefined ? sorted.slice(0, topK) : sorted
|
||||
|
|
@ -2341,51 +2309,11 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
*
|
||||
* @public (called from brainy.ts for sorted queries)
|
||||
*/
|
||||
/**
|
||||
* The cross-engine ordering contract in one comparator (sealed 2026-08-03):
|
||||
* missing/null values sort LAST in BOTH directions — the direction flip
|
||||
* never moves them to the front — and ties break by id ascending, so an
|
||||
* ordered read is deterministic and identical on both engines. Numbers
|
||||
* compare numerically; everything else by code-point (UTF-8 byte) order,
|
||||
* matching the native column store exactly.
|
||||
*/
|
||||
private compareAddressedValues(
|
||||
aVal: any,
|
||||
bVal: any,
|
||||
aId: string,
|
||||
bId: string,
|
||||
order: 'asc' | 'desc'
|
||||
): number {
|
||||
const aNull = aVal == null
|
||||
const bNull = bVal == null
|
||||
if (aNull || bNull) {
|
||||
if (aNull && bNull) return aId < bId ? -1 : aId > bId ? 1 : 0
|
||||
return aNull ? 1 : -1
|
||||
}
|
||||
let comparison = 0
|
||||
if (aVal !== bVal) {
|
||||
if (typeof aVal === 'number' && typeof bVal === 'number') {
|
||||
comparison = aVal < bVal ? -1 : 1
|
||||
} else {
|
||||
comparison = compareCodePoints(String(aVal), String(bVal))
|
||||
}
|
||||
}
|
||||
if (comparison === 0) return aId < bId ? -1 : aId > bId ? 1 : 0
|
||||
return order === 'asc' ? comparison : -comparison
|
||||
}
|
||||
|
||||
async getFieldValueForEntity(entityId: string, field: string): Promise<any> {
|
||||
// `field` arrives as a FROZEN INDEX KEY (bare = user metadata;
|
||||
// 'system.<field>' = engine scalar). Storage fallbacks read the matching
|
||||
// side of the record — a system key reads the record scalar, a bare key
|
||||
// reads the user's metadata bag; the two can never shadow each other.
|
||||
const systemInner = field.startsWith('system.') ? field.slice('system.'.length) : null
|
||||
|
||||
// Path 1: Bucketed fields need the actual (un-bucketed) value from storage.
|
||||
// Path 1: Bucketed fields need the actual value from storage.
|
||||
if (BUCKETED_INDEX_FIELDS.has(field)) {
|
||||
const noun = await this.storage.getNoun(entityId)
|
||||
if (!noun) return undefined
|
||||
return (noun as unknown as Record<string, unknown>)[systemInner as string]
|
||||
return noun ? resolveEntityField(noun, field) : undefined
|
||||
}
|
||||
|
||||
// Path 3 precondition: entity must be in the id mapper for bitmap lookup.
|
||||
|
|
@ -2402,11 +2330,7 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
// yet indexed. resolveEntityField handles the shape contract.
|
||||
if (!sparseIndex) {
|
||||
const noun = await this.storage.getNoun(entityId)
|
||||
if (!noun) return undefined
|
||||
if (systemInner !== null) {
|
||||
return (noun as unknown as Record<string, unknown>)[systemInner]
|
||||
}
|
||||
return (noun as { metadata?: Record<string, unknown> }).metadata?.[field]
|
||||
return noun ? resolveEntityField(noun, field) : undefined
|
||||
}
|
||||
|
||||
// Path 3: Search sparse index chunks for this entity's value.
|
||||
|
|
@ -2833,17 +2757,6 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
// VFS Statistics Methods (uses existing Roaring bitmap infrastructure)
|
||||
// ============================================================================
|
||||
|
||||
/**
|
||||
* Read the type column's bitmap for one type value — frozen key first
|
||||
* ('system.type', epoch 3), legacy 'noun' as the pre-rebuild fallback.
|
||||
*/
|
||||
private async getTypeBitmap(type: string): Promise<RoaringBitmap32 | null> {
|
||||
return (
|
||||
(await this.getBitmapFromChunks('system.type', type)) ??
|
||||
(await this.getBitmapFromChunks('noun', type))
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* Get VFS entity count for a specific type using Roaring bitmap intersection
|
||||
* Uses hardware-accelerated SIMD operations (AVX2/SSE4.2)
|
||||
|
|
@ -2852,7 +2765,7 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
*/
|
||||
async getVFSEntityCountByType(type: string): Promise<number> {
|
||||
const vfsBitmap = await this.getBitmapFromChunks('isVFSEntity', true)
|
||||
const typeBitmap = await this.getTypeBitmap(type)
|
||||
const typeBitmap = await this.getBitmapFromChunks('noun', type)
|
||||
|
||||
if (!vfsBitmap || !typeBitmap) return 0
|
||||
|
||||
|
|
@ -2875,7 +2788,7 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
|
||||
// Iterate through all known types and compute VFS count via intersection
|
||||
for (const type of this.totalEntitiesByType.keys()) {
|
||||
const typeBitmap = await this.getTypeBitmap(type)
|
||||
const typeBitmap = await this.getBitmapFromChunks('noun', type)
|
||||
if (typeBitmap) {
|
||||
const intersection = RoaringBitmap32.and(vfsBitmap, typeBitmap)
|
||||
if (intersection.size > 0) {
|
||||
|
|
@ -3469,21 +3382,18 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
* Tracks which fields commonly appear with which entity types
|
||||
*/
|
||||
private updateTypeFieldAffinity(entityId: string, field: string, value: any, operation: 'add' | 'remove', metadata?: any): void {
|
||||
// Only track affinity for user fields (plus the type column itself,
|
||||
// which drives detection). Engine columns carry the literal 'system.'
|
||||
// prefix under the frozen key format.
|
||||
if (field.startsWith('system.') && field !== 'system.type') return
|
||||
// Only track affinity for non-system fields (but allow 'noun' for type detection)
|
||||
if (this.config.excludeFields.includes(field) && field !== 'noun') return
|
||||
|
||||
// For the type column ('system.type'), the value IS the entity type
|
||||
// For the 'noun' field, the value IS the entity type
|
||||
let entityType: string | null = null
|
||||
|
||||
if (field === 'system.type') {
|
||||
if (field === 'noun') {
|
||||
// This is the type definition itself
|
||||
entityType = this.normalizeValue(value, field) // Pass field for bucketing!
|
||||
} else if (metadata && (metadata.noun ?? metadata.type)) {
|
||||
// Extract entity type from the source shape: stored records carry it
|
||||
// under 'noun', entity-for-indexing views under 'type'.
|
||||
entityType = this.normalizeValue(metadata.noun ?? metadata.type, 'system.type')
|
||||
} else if (metadata && metadata.noun) {
|
||||
// Extract entity type from metadata
|
||||
entityType = this.normalizeValue(metadata.noun, 'noun')
|
||||
} else {
|
||||
// No type information available, skip affinity tracking
|
||||
return
|
||||
|
|
@ -3506,9 +3416,8 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
const currentCount = typeFields.get(field) || 0
|
||||
typeFields.set(field, currentCount + 1)
|
||||
|
||||
// Update total entities of this type (only count once per entity —
|
||||
// the type column appears exactly once per entity)
|
||||
if (field === 'system.type') {
|
||||
// Update total entities of this type (only count once per entity)
|
||||
if (field === 'noun') {
|
||||
const newCount = this.totalEntitiesByType.get(entityType)! + 1
|
||||
this.totalEntitiesByType.set(entityType, newCount)
|
||||
|
||||
|
|
@ -3531,7 +3440,7 @@ export class MetadataIndexManager implements MetadataIndexProvider {
|
|||
}
|
||||
|
||||
// Update total entities of this type
|
||||
if (field === 'system.type') {
|
||||
if (field === 'noun') {
|
||||
const total = this.totalEntitiesByType.get(entityType)!
|
||||
if (total > 1) {
|
||||
const newCount = total - 1
|
||||
|
|
|
|||
|
|
@ -17,7 +17,6 @@ import { findCallerLocation } from './callerLocation.js'
|
|||
// fallback branches that no supported runtime can reach.
|
||||
import * as os from 'node:os'
|
||||
import * as fs from 'node:fs'
|
||||
import { parseFieldAddress, UnsupportedFindOptionError } from '../db/fieldAddressing.js'
|
||||
|
||||
const getSystemMemory = (): number => {
|
||||
if (os) {
|
||||
|
|
@ -467,31 +466,9 @@ export function validateFindParams(params: FindParams): void {
|
|||
throw new Error('cannot specify both query and vector - they are mutually exclusive')
|
||||
}
|
||||
|
||||
// ACCEPTED-AND-IGNORED DIED AS A CLASS (sealed 2026-08-03): options the
|
||||
// engine does not implement REFUSE with a typed error instead of silently
|
||||
// doing nothing — a production consumer discovered a no-op by measurement
|
||||
// once; never again.
|
||||
if (params.cursor !== undefined) {
|
||||
throw new UnsupportedFindOptionError('cursor')
|
||||
}
|
||||
if ((params as Record<string, unknown>).includeRelations !== undefined) {
|
||||
throw new UnsupportedFindOptionError('includeRelations')
|
||||
}
|
||||
if ((params as Record<string, unknown>).writeOnly !== undefined) {
|
||||
throw new UnsupportedFindOptionError('writeOnly')
|
||||
}
|
||||
|
||||
// THE ONE ADDRESSING LAW: the orderBy address must PARSE (bare/metadata. =
|
||||
// user field, system.<field> = the ruled map, anything else refuses typed
|
||||
// with the valid map in the message) and order must be a real direction.
|
||||
if (params.orderBy !== undefined) {
|
||||
if (typeof params.orderBy !== 'string') {
|
||||
throw new Error('orderBy must be a string field address')
|
||||
}
|
||||
parseFieldAddress(params.orderBy, 'entity') // throws InvalidFieldAddressError on a bad address
|
||||
}
|
||||
if (params.order !== undefined && params.order !== 'asc' && params.order !== 'desc') {
|
||||
throw new Error(`order must be 'asc' or 'desc', got '${String(params.order)}'`)
|
||||
// Universal truth: can't use both cursor and offset pagination
|
||||
if (params.cursor !== undefined && params.offset !== undefined) {
|
||||
throw new Error('cannot use both cursor and offset pagination simultaneously')
|
||||
}
|
||||
|
||||
// Auto-limit query length based on memory
|
||||
|
|
@ -518,28 +495,7 @@ export function validateFindParams(params: FindParams): void {
|
|||
/**
|
||||
* Validate add parameters
|
||||
*/
|
||||
|
||||
/**
|
||||
* The namespace cannot be forged: a USER metadata key literally spelled
|
||||
* 'system.<anything>' would collide with the engine's explicit address
|
||||
* namespace at read time — refuse it at the write door, loudly, with the
|
||||
* fix in the message (sealed 2026-08-03).
|
||||
*/
|
||||
function rejectForgedSystemKeys(metadata: Record<string, unknown> | undefined, site: string): void {
|
||||
if (!metadata) return
|
||||
for (const key of Object.keys(metadata)) {
|
||||
if (key.startsWith('system.')) {
|
||||
throw new Error(
|
||||
`${site}: metadata key '${key}' is not allowed — the 'system.' prefix is the ` +
|
||||
`engine's explicit address namespace and cannot be used as a user field name. ` +
|
||||
`Rename the field (e.g. '${key.slice('system.'.length)}').`
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
export function validateAddParams(params: AddParams): void {
|
||||
rejectForgedSystemKeys(params.metadata as Record<string, unknown> | undefined, 'add()')
|
||||
// Universal truth: must have data or vector
|
||||
if (!params.data && !params.vector) {
|
||||
throw new Error(
|
||||
|
|
@ -580,7 +536,6 @@ export function validateAddParams(params: AddParams): void {
|
|||
* Validate update parameters
|
||||
*/
|
||||
export function validateUpdateParams(params: UpdateParams): void {
|
||||
rejectForgedSystemKeys(params.metadata as Record<string, unknown> | undefined, 'update()')
|
||||
// Universal truth: must have an ID
|
||||
if (!params.id) {
|
||||
throw new Error('id is required for update')
|
||||
|
|
@ -618,7 +573,6 @@ export function validateUpdateParams(params: UpdateParams): void {
|
|||
* Validate relate parameters
|
||||
*/
|
||||
export function validateRelateParams(params: RelateParams): void {
|
||||
rejectForgedSystemKeys(params.metadata as Record<string, unknown> | undefined, 'relate()')
|
||||
// 8.0 verb-id contract (L.7): verb ids are UUIDs, generated by brainy.
|
||||
// RelateParams has no `id` field — an untyped caller passing one would
|
||||
// previously have it silently ignored (a generated UUID was used instead).
|
||||
|
|
@ -667,7 +621,6 @@ export function validateRelateParams(params: RelateParams): void {
|
|||
* accepts type/subtype/weight/confidence/data/metadata changes.
|
||||
*/
|
||||
export function validateUpdateRelationParams(params: UpdateRelationParams): void {
|
||||
rejectForgedSystemKeys(params.metadata as Record<string, unknown> | undefined, 'updateRelation()')
|
||||
if (!params.id) {
|
||||
throw new Error('id is required for updateRelation')
|
||||
}
|
||||
|
|
|
|||
|
|
@ -57,7 +57,6 @@ export class PathResolver {
|
|||
// Statistics
|
||||
private cacheHits = 0
|
||||
private cacheMisses = 0
|
||||
private lastLoggedLookups = 0 // last total the maintenance tick logged stats at
|
||||
private metadataIndexHits = 0
|
||||
private metadataIndexMisses = 0
|
||||
private graphTraversalFallbacks = 0
|
||||
|
|
@ -520,14 +519,10 @@ export class PathResolver {
|
|||
}
|
||||
}
|
||||
|
||||
// Log cache statistics only when there is new traffic to report — an
|
||||
// idle resolver stays silent. 0/0 lookups previously rendered
|
||||
// "NaN% hit rate" (and the %1000 gate passes at zero), which spammed
|
||||
// production journals once a minute on every idle VFS.
|
||||
const totalLookups = this.cacheHits + this.cacheMisses
|
||||
if (totalLookups > 0 && totalLookups !== this.lastLoggedLookups && totalLookups % 1000 === 0) {
|
||||
this.lastLoggedLookups = totalLookups
|
||||
prodLog.debug(`[PathResolver] Cache stats: ${Math.round((this.cacheHits / totalLookups) * 100)}% hit rate, ${this.pathCache.size} entries, ${this.hotPaths.size} hot paths`)
|
||||
// Log cache statistics (in production, send to monitoring)
|
||||
const hitRate = this.cacheHits / (this.cacheHits + this.cacheMisses)
|
||||
if ((this.cacheHits + this.cacheMisses) % 1000 === 0) {
|
||||
console.log(`[PathResolver] Cache stats: ${Math.round(hitRate * 100)}% hit rate, ${this.pathCache.size} entries, ${this.hotPaths.size} hot paths`)
|
||||
}
|
||||
}, 60000) // Every minute
|
||||
// Cache maintenance must never keep the host process alive.
|
||||
|
|
|
|||
|
|
@ -1,307 +0,0 @@
|
|||
/**
|
||||
* @module tests/conformance/collider-fidelity
|
||||
* @description THE REOPEN-COLLIDER CONFORMANCE CASE (required cross-engine
|
||||
* before any RC counts as gates-green — ruled 2026-08-03). The
|
||||
* field-addressing law's fidelity half: user metadata may carry ANY name —
|
||||
* including every engine spelling (`confidence`, `type`, `id`, `createdAt`,
|
||||
* …) and every plumbing name (`level`, `data`, `vector`, `_rev`) — and the
|
||||
* value survives, verbatim and reachable, across the FULL lifecycle: live
|
||||
* reads, where/orderBy, flush, close+reopen, a forced epoch rebuild, and
|
||||
* time travel. The engine scalars stay separately reachable at `system.*`
|
||||
* the whole way. No halfway states.
|
||||
*
|
||||
* Self-arming like the namespace-law suite: skips loudly until the arming
|
||||
* exports are present, so the suite can sit on a branch ahead of the build.
|
||||
*/
|
||||
import { describe, it, expect, beforeAll, afterAll } from 'vitest'
|
||||
import { mkdtempSync, rmSync } from 'node:fs'
|
||||
import { tmpdir } from 'node:os'
|
||||
import { join } from 'node:path'
|
||||
import * as brainyExports from '../../src/index.js'
|
||||
import { Brainy, NounType, VerbType } from '../../src/index.js'
|
||||
import {
|
||||
BRAIN_FORMAT_PATH,
|
||||
EXPECTED_INDEX_EPOCH
|
||||
} from '../../src/storage/brainFormat.js'
|
||||
|
||||
const ARMED = 'UnresolvableFieldError' in brainyExports
|
||||
const suite = ARMED ? describe : describe.skip
|
||||
if (!ARMED) {
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn(
|
||||
'[collider-fidelity] SKIPPING: package root does not export the ' +
|
||||
'field-addressing law surface yet (UnresolvableFieldError absent).'
|
||||
)
|
||||
}
|
||||
|
||||
/** Every entity system scalar name written as a USER metadata field, with
|
||||
* unmistakable user values, plus the plumbing names and naturals. */
|
||||
const COLLIDER_BAG = {
|
||||
// the ten entity system scalars, as user fields
|
||||
id: 'user-id',
|
||||
type: 'user-type',
|
||||
subtype: 'user-subtype',
|
||||
createdAt: 'user-createdAt',
|
||||
updatedAt: 'user-updatedAt',
|
||||
confidence: 'user-confidence',
|
||||
weight: 'user-weight',
|
||||
visibility: 'user-visibility',
|
||||
service: 'user-service',
|
||||
createdBy: 'user-createdBy',
|
||||
// plumbing names, as user fields
|
||||
level: 7,
|
||||
data: 'user-data',
|
||||
vector: 'user-vector',
|
||||
_rev: 'user-rev',
|
||||
// naturals previously silently un-indexed by name
|
||||
content: 'user-content',
|
||||
// a plain control field
|
||||
plain: 'control'
|
||||
} as const
|
||||
|
||||
|
||||
suite('collider fidelity — the reopen-collider case (both suites, ruled)', () => {
|
||||
let dir: string
|
||||
let brain: Brainy
|
||||
let colliderId: string
|
||||
|
||||
const open = async (): Promise<Brainy> => {
|
||||
const b = new Brainy({
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
requireSubtype: false
|
||||
})
|
||||
await b.init()
|
||||
return b
|
||||
}
|
||||
|
||||
/** The full read battery — run at every lifecycle boundary. */
|
||||
const verifyColliderTruth = async (label: string): Promise<void> => {
|
||||
// 1. get(): the bag comes back verbatim; engine scalars stay engine.
|
||||
const entity = await brain.get(colliderId)
|
||||
expect(entity, `${label}: entity readable`).toBeTruthy()
|
||||
for (const [k, v] of Object.entries(COLLIDER_BAG)) {
|
||||
expect(
|
||||
(entity!.metadata as Record<string, unknown>)[k],
|
||||
`${label}: bag.${k} verbatim`
|
||||
).toEqual(v)
|
||||
}
|
||||
expect(entity!.type, `${label}: engine type intact`).toBe(NounType.Document)
|
||||
expect(entity!.confidence, `${label}: engine confidence intact`).toBe(0.25)
|
||||
|
||||
// 2. where on collider names (bare = the user's field, always).
|
||||
for (const [k, v] of [
|
||||
['confidence', 'user-confidence'],
|
||||
['type', 'user-type'],
|
||||
['id', 'user-id'],
|
||||
['content', 'user-content'],
|
||||
['data', 'user-data'],
|
||||
['level', 7]
|
||||
] as const) {
|
||||
const rows = await brain.find({ where: { [k]: v }, limit: 10 })
|
||||
expect(
|
||||
rows.map((r) => r.id),
|
||||
`${label}: where {${k}} finds the collider row`
|
||||
).toContain(colliderId)
|
||||
}
|
||||
|
||||
// 3. system.* keeps reading the ENGINE values.
|
||||
const byEngine = await brain.find({
|
||||
where: { 'system.confidence': 0.25 },
|
||||
limit: 10
|
||||
})
|
||||
expect(
|
||||
byEngine.map((r) => r.id),
|
||||
`${label}: system.confidence reads the engine scalar`
|
||||
).toContain(colliderId)
|
||||
const byUserSpelledSystem = await brain.find({
|
||||
where: { 'system.confidence': 'user-confidence' },
|
||||
limit: 10
|
||||
})
|
||||
expect(
|
||||
byUserSpelledSystem.map((r) => r.id),
|
||||
`${label}: the user's value is NOT reachable via system.*`
|
||||
).not.toContain(colliderId)
|
||||
|
||||
// 4. orderBy a collider name orders by the USER values.
|
||||
const ordered = await brain.find({
|
||||
type: NounType.Document,
|
||||
orderBy: 'level',
|
||||
order: 'desc',
|
||||
limit: 10
|
||||
})
|
||||
expect(ordered.length, `${label}: ordered read complete`).toBe(3)
|
||||
expect(
|
||||
(ordered[0].metadata as Record<string, unknown>).plain,
|
||||
`${label}: user level orders desc (7 first)`
|
||||
).toBe('control')
|
||||
}
|
||||
|
||||
beforeAll(async () => {
|
||||
dir = mkdtempSync(join(tmpdir(), 'brainy-collider-'))
|
||||
brain = await open()
|
||||
|
||||
colliderId = await brain.add({
|
||||
data: 'the collider probe document',
|
||||
type: NounType.Document,
|
||||
confidence: 0.25,
|
||||
metadata: { ...COLLIDER_BAG }
|
||||
})
|
||||
// two ordering companions with smaller user `level`s
|
||||
await brain.add({
|
||||
data: 'ordering companion low',
|
||||
type: NounType.Document,
|
||||
metadata: { level: 3, plain: 'low' }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'ordering companion mid',
|
||||
type: NounType.Document,
|
||||
metadata: { level: 5, plain: 'mid' }
|
||||
})
|
||||
}, 120000)
|
||||
|
||||
afterAll(async () => {
|
||||
await brain.close().catch(() => {})
|
||||
rmSync(dir, { recursive: true, force: true })
|
||||
})
|
||||
|
||||
it('LIVE: colliders are the user’s, verbatim and fully queryable', async () => {
|
||||
await verifyColliderTruth('live')
|
||||
})
|
||||
|
||||
it('REOPEN: the restart boundary loses nothing', async () => {
|
||||
await brain.flush()
|
||||
await brain.close()
|
||||
brain = await open()
|
||||
await verifyColliderTruth('reopen')
|
||||
})
|
||||
|
||||
it('REBUILD: a forced epoch rebuild re-indexes the colliders from canonical', async () => {
|
||||
await brain.close()
|
||||
// Simulate epoch drift: a missing marker forces the full derived-index
|
||||
// rebuild at open — the exact path every pre-law brain takes once.
|
||||
rmSync(join(dir, BRAIN_FORMAT_PATH), { force: true })
|
||||
brain = await open()
|
||||
await verifyColliderTruth('rebuild')
|
||||
// And the rebuild re-stamps the current epoch.
|
||||
const marker = await (
|
||||
brain as unknown as {
|
||||
storage: { readRawObject(p: string): Promise<{ indexEpoch?: number } | null> }
|
||||
}
|
||||
).storage.readRawObject(BRAIN_FORMAT_PATH)
|
||||
expect(marker?.indexEpoch).toBe(EXPECTED_INDEX_EPOCH)
|
||||
})
|
||||
|
||||
it('TIME TRAVEL: asOf reads historical collider values faithfully', async () => {
|
||||
const gen = brain.generation()
|
||||
await brain.update({ id: colliderId, metadata: { confidence: 'user-confidence-v2' } })
|
||||
const now = await brain.get(colliderId)
|
||||
expect((now!.metadata as Record<string, unknown>).confidence).toBe('user-confidence-v2')
|
||||
|
||||
const past = await brain.asOf(gen)
|
||||
try {
|
||||
const then = await past.get(colliderId)
|
||||
expect(
|
||||
(then!.metadata as Record<string, unknown>).confidence,
|
||||
'asOf reads the pre-update USER value'
|
||||
).toBe('user-confidence')
|
||||
} finally {
|
||||
await past.release()
|
||||
}
|
||||
// engine scalar untouched throughout
|
||||
expect(now!.confidence).toBe(0.25)
|
||||
})
|
||||
|
||||
it('RELATION MIRROR: edge collider bags survive write → read → reopen', async () => {
|
||||
const a = await brain.add({ data: 'edge endpoint a', type: NounType.Person, metadata: { plain: 'a' } })
|
||||
const b = await brain.add({ data: 'edge endpoint b', type: NounType.Person, metadata: { plain: 'b' } })
|
||||
const edgeBag = {
|
||||
verb: 'user-verb',
|
||||
confidence: 'user-edge-confidence',
|
||||
weight: 'user-edge-weight',
|
||||
subtype: 'user-edge-subtype',
|
||||
createdAt: 'user-edge-createdAt',
|
||||
service: 'user-edge-service'
|
||||
}
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.RelatedTo,
|
||||
confidence: 0.5,
|
||||
metadata: { ...edgeBag }
|
||||
})
|
||||
|
||||
const check = async (label: string): Promise<void> => {
|
||||
const rels = await brain.related({ from: a, type: VerbType.RelatedTo })
|
||||
const rel = rels.find((r) => r.id === relId)
|
||||
expect(rel, `${label}: relation readable`).toBeTruthy()
|
||||
for (const [k, v] of Object.entries(edgeBag)) {
|
||||
expect(
|
||||
(rel!.metadata as Record<string, unknown>)[k],
|
||||
`${label}: edge bag.${k} verbatim`
|
||||
).toEqual(v)
|
||||
}
|
||||
expect(rel!.confidence, `${label}: engine edge confidence intact`).toBe(0.5)
|
||||
expect(rel!.type, `${label}: engine verb intact`).toBe(VerbType.RelatedTo)
|
||||
}
|
||||
|
||||
await check('live')
|
||||
await brain.flush()
|
||||
await brain.close()
|
||||
brain = await open()
|
||||
await check('reopen')
|
||||
})
|
||||
|
||||
it('FORGERY: user metadata keys spelled system.* refuse at every write door', async () => {
|
||||
await expect(
|
||||
brain.add({ data: 'forged', type: NounType.Document, metadata: { 'system.confidence': 1 } })
|
||||
).rejects.toThrow(/system\./)
|
||||
await expect(
|
||||
brain.update({ id: colliderId, metadata: { 'system.type': 'x' } })
|
||||
).rejects.toThrow(/system\./)
|
||||
const a = await brain.add({ data: 'forgery endpoint a', type: NounType.Person, metadata: {} })
|
||||
const b = await brain.add({ data: 'forgery endpoint b', type: NounType.Person, metadata: {} })
|
||||
await expect(
|
||||
brain.relate({ from: a, to: b, type: VerbType.RelatedTo, metadata: { 'system.verb': 'x' } })
|
||||
).rejects.toThrow(/system\./)
|
||||
})
|
||||
|
||||
it('CONFIG: the dead reservedFieldPolicy option refuses loudly, never ignored', () => {
|
||||
expect(
|
||||
() => new Brainy({ storage: { type: 'memory' }, reservedFieldPolicy: 'throw' } as never)
|
||||
).toThrow(/field-addressing law/)
|
||||
})
|
||||
|
||||
it('LEGACY: a pre-law flat record still reads with engine fields top-level', async () => {
|
||||
const storage = (
|
||||
brain as unknown as {
|
||||
storage: {
|
||||
saveNoun(n: unknown): Promise<void>
|
||||
saveNounMetadata(id: string, m: Record<string, unknown>): Promise<void>
|
||||
}
|
||||
}
|
||||
).storage
|
||||
const legacyId = '00000000-0000-4000-8000-00000000f1a7'
|
||||
await storage.saveNoun({ id: legacyId, vector: new Array(384).fill(0.01), connections: new Map(), level: 0 })
|
||||
// Legacy FLAT shape: engine + user keys mixed at one level, NO _fmt stamp.
|
||||
// Sound to split by name — the pre-law door refused user colliders.
|
||||
await storage.saveNounMetadata(legacyId, {
|
||||
noun: NounType.Document,
|
||||
confidence: 0.75,
|
||||
createdAt: 1700000000000,
|
||||
updatedAt: 1700000000000,
|
||||
_rev: 1,
|
||||
legacyField: 'legacy-value'
|
||||
})
|
||||
const entity = await brain.get(legacyId)
|
||||
expect(entity).toBeTruthy()
|
||||
expect(entity!.confidence, 'legacy flat confidence = engine').toBe(0.75)
|
||||
expect(
|
||||
(entity!.metadata as Record<string, unknown>).legacyField,
|
||||
'legacy custom field = user bag'
|
||||
).toBe('legacy-value')
|
||||
expect(
|
||||
(entity!.metadata as Record<string, unknown>).confidence,
|
||||
'legacy flat engine key never leaks into the bag'
|
||||
).toBeUndefined()
|
||||
})
|
||||
})
|
||||
|
|
@ -1,486 +0,0 @@
|
|||
/**
|
||||
* @module tests/conformance/namespace-law
|
||||
* @description Conformance suite for the ruled field-addressing contract
|
||||
* announced in RELEASES.md ("Coming next... one field-addressing law — bare
|
||||
* names = user metadata, `system.<field>` for engine fields, typed refusals
|
||||
* for unresolvable names"). This suite is the drift-proof shared by this
|
||||
* engine and its native accelerator: both must satisfy every test here
|
||||
* bit-for-bit, because they implement the SAME contract independently.
|
||||
*
|
||||
* The rule, in full:
|
||||
* 1. A bare field name in `where` / `orderBy` / `groupBy` / aggregation
|
||||
* `source.where` ALWAYS means the caller's own `metadata` field. No
|
||||
* priority resolution, no engine fallback — ever.
|
||||
* 2. `system.<field>` reaches an engine scalar, and ONLY an engine scalar,
|
||||
* and ONLY when spelled explicitly. The addressable entity map is exactly
|
||||
* ten names: id, type, subtype, createdAt, updatedAt, confidence, weight,
|
||||
* visibility, service, createdBy. The relationship map is system.verb,
|
||||
* system.sourceId, system.targetId, plus the eight scalars shared with
|
||||
* entities.
|
||||
* 3. Some names are invisible plumbing and are never addressable in either
|
||||
* spelling: vector, connections, level, data, _rev. `system.level`,
|
||||
* `system.vector`, and `system.data` all refuse — they are not in the
|
||||
* system map. Bare `level` is a perfectly ordinary user field.
|
||||
* 4. `metadata.<field>` is the explicit-user-scope spelling: identical
|
||||
* semantics to the bare spelling, valid everywhere the bare spelling is.
|
||||
* 5. Anything that resolves to neither a user field nor a system scalar is a
|
||||
* typed refusal naming both candidates (`UnresolvableFieldError`).
|
||||
* Unimplemented `find()` options (`cursor`, `includeRelations`,
|
||||
* `writeOnly`) refuse with `UnsupportedFindOptionError` instead of being
|
||||
* silently accepted and ignored.
|
||||
* 6. Ordering is identical on both engines: rows missing/null on the
|
||||
* `orderBy` field sort LAST in BOTH directions and are never dropped;
|
||||
* ties break by id ascending.
|
||||
*
|
||||
* The motivating incident (told generically — see CLAUDE.md naming rule): an
|
||||
* internal report from a production deployment showed a user metadata field
|
||||
* literally named `level` silently shadowed by the engine's internal HNSW
|
||||
* node layer, breaking sort order with zero errors raised. This contract
|
||||
* makes that class of bug impossible, and testable forever.
|
||||
*
|
||||
* SELF-SKIP: the resolver this suite pins is being built in a parallel
|
||||
* session and has not landed on every branch yet. Rather than going red on
|
||||
* a branch that simply hasn't caught up, the suite detects whether the
|
||||
* contract is live by the one thing any conformant implementation must
|
||||
* export — `UnresolvableFieldError` from the package root — and skips
|
||||
* loudly (never silently) until it does. This is the house pattern: a
|
||||
* sibling engine's gate once went red because a test armed before its
|
||||
* feature existed.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
import * as brainyExports from '../../src/index.js'
|
||||
|
||||
const stubEmbedding = async (text: string): Promise<number[]> => {
|
||||
const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(hash + i))
|
||||
}
|
||||
|
||||
// Detected purely by the exported error-class NAME — never by reaching into
|
||||
// implementation internals. Both engines building this contract must export
|
||||
// it from the package root, so this is a legitimate, implementation-agnostic
|
||||
// readiness probe.
|
||||
const lawActive = 'UnresolvableFieldError' in brainyExports
|
||||
const UnresolvableFieldError = (brainyExports as Record<string, unknown>).UnresolvableFieldError as new (
|
||||
...args: any[]
|
||||
) => Error
|
||||
const UnsupportedFindOptionError = (brainyExports as Record<string, unknown>)
|
||||
.UnsupportedFindOptionError as new (...args: any[]) => Error
|
||||
|
||||
// Always runs, regardless of lawActive — the loud signal that the rest of
|
||||
// this file was skipped, and why.
|
||||
it('namespace law armed?', () => {
|
||||
if (!lawActive) {
|
||||
console.warn(
|
||||
'[conformance] namespace-law suite SKIPPED — UnresolvableFieldError not exported yet; arms when the resolver lands'
|
||||
)
|
||||
}
|
||||
expect(true).toBe(true)
|
||||
})
|
||||
|
||||
/**
|
||||
* Awaits `promise`, asserting it rejects with an instance of `ErrorClass`
|
||||
* whose `.message` contains every string in `mustContain`. Fails loudly if
|
||||
* the promise resolves instead of rejecting.
|
||||
*/
|
||||
async function expectRefusal(
|
||||
promise: Promise<unknown>,
|
||||
ErrorClass: new (...args: any[]) => Error,
|
||||
...mustContain: string[]
|
||||
): Promise<void> {
|
||||
let threw = false
|
||||
try {
|
||||
await promise
|
||||
} catch (err) {
|
||||
threw = true
|
||||
expect(err).toBeInstanceOf(ErrorClass)
|
||||
for (const fragment of mustContain) {
|
||||
expect((err as Error).message).toContain(fragment)
|
||||
}
|
||||
}
|
||||
expect(threw).toBe(true)
|
||||
}
|
||||
|
||||
describe.skipIf(!lawActive)('namespace law — bare/system/metadata field addressing', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
/** The star case from the motivating incident: metadata.level 3/9/6. */
|
||||
async function addLevelRows(): Promise<string[]> {
|
||||
const ids: string[] = []
|
||||
for (const level of [3, 9, 6]) {
|
||||
ids.push(
|
||||
await brain.add({
|
||||
data: `probe level ${level}`,
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-level',
|
||||
metadata: { name: `p-${level}`, level }
|
||||
})
|
||||
)
|
||||
}
|
||||
return ids
|
||||
}
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 1 — bare field name = the user's metadata field, always.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it("bare orderBy 'level' reads user metadata, desc and asc (the star case)", async () => {
|
||||
await addLevelRows()
|
||||
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-level',
|
||||
orderBy: 'level',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3])
|
||||
|
||||
const asc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-level',
|
||||
orderBy: 'level',
|
||||
order: 'asc',
|
||||
limit: 100
|
||||
})
|
||||
expect(asc.map((r: any) => r.metadata?.level)).toEqual([3, 6, 9])
|
||||
})
|
||||
|
||||
it("bare where { level: N } matches the user's field", async () => {
|
||||
const ids = await addLevelRows()
|
||||
const hit = await brain.find({ type: NounType.Person, subtype: 'ns-law-level', where: { level: 9 } })
|
||||
expect(hit).toHaveLength(1)
|
||||
expect(hit[0].id).toBe(ids[1])
|
||||
expect(hit[0].metadata?.level).toBe(9)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 4 — metadata.<field> is the explicit-user-scope spelling,
|
||||
// identical semantics to bare, valid on every path including orderBy.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it("'metadata.level' resolves identically to bare 'level'", async () => {
|
||||
await addLevelRows()
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-level',
|
||||
orderBy: 'metadata.level',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3])
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 2 — system.<field> reaches an engine scalar explicitly.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it('system.createdAt sorts by entity age', async () => {
|
||||
const ids: string[] = []
|
||||
for (const name of ['first', 'second', 'third']) {
|
||||
ids.push(
|
||||
await brain.add({
|
||||
data: `aged ${name}`,
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-aged',
|
||||
metadata: { name }
|
||||
})
|
||||
)
|
||||
// Guarantee distinct createdAt timestamps between adds.
|
||||
await new Promise((resolve) => setTimeout(resolve, 5))
|
||||
}
|
||||
|
||||
const asc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-aged',
|
||||
orderBy: 'system.createdAt',
|
||||
order: 'asc',
|
||||
limit: 100
|
||||
})
|
||||
expect(asc.map((r: any) => r.id)).toEqual(ids)
|
||||
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-aged',
|
||||
orderBy: 'system.createdAt',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc.map((r: any) => r.id)).toEqual([...ids].reverse())
|
||||
})
|
||||
|
||||
it('where on system.confidence filters by the engine scalar', async () => {
|
||||
const highId = await brain.add({
|
||||
data: 'high confidence row',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-confidence',
|
||||
confidence: 0.95,
|
||||
metadata: { name: 'hi' }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'low confidence row',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-confidence',
|
||||
confidence: 0.4,
|
||||
metadata: { name: 'lo' }
|
||||
})
|
||||
|
||||
const hit = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-confidence',
|
||||
where: { 'system.confidence': 0.95 }
|
||||
})
|
||||
expect(hit).toHaveLength(1)
|
||||
expect(hit[0].id).toBe(highId)
|
||||
})
|
||||
|
||||
it('groupBy on system.subtype groups by the engine scalar, not user metadata', async () => {
|
||||
await brain.add({ data: 'i1', type: NounType.Document, subtype: 'invoice' })
|
||||
await brain.add({ data: 'i2', type: NounType.Document, subtype: 'invoice' })
|
||||
await brain.add({ data: 'r1', type: NounType.Document, subtype: 'receipt' })
|
||||
|
||||
brain.defineAggregate({
|
||||
name: 'ns_law_by_subtype_system',
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['system.subtype'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
|
||||
const groups = await brain.queryAggregate('ns_law_by_subtype_system')
|
||||
const invoiceGroup = groups.find((g) => Object.values(g.groupKey).includes('invoice'))
|
||||
const receiptGroup = groups.find((g) => Object.values(g.groupKey).includes('receipt'))
|
||||
expect(invoiceGroup?.metrics.count).toBe(2)
|
||||
expect(receiptGroup?.metrics.count).toBe(1)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 1 (groupBy face) — bare groupBy dimensions read user metadata,
|
||||
// never the engine's own notion of the same-sounding name.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it('groupBy on a bare user metadata field groups by that field', async () => {
|
||||
await brain.add({
|
||||
data: 'd1',
|
||||
type: NounType.Document,
|
||||
subtype: 'ns-law-group-bare',
|
||||
metadata: { team: 'alpha' }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'd2',
|
||||
type: NounType.Document,
|
||||
subtype: 'ns-law-group-bare',
|
||||
metadata: { team: 'alpha' }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'd3',
|
||||
type: NounType.Document,
|
||||
subtype: 'ns-law-group-bare',
|
||||
metadata: { team: 'beta' }
|
||||
})
|
||||
|
||||
brain.defineAggregate({
|
||||
name: 'ns_law_by_team_bare',
|
||||
// system.subtype — bare 'subtype' would address user metadata under the
|
||||
// law (the exact migration every fleet consumer's aggregates make).
|
||||
source: { type: NounType.Document, where: { 'system.subtype': 'ns-law-group-bare' } },
|
||||
groupBy: ['team'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
|
||||
const groups = await brain.queryAggregate('ns_law_by_team_bare')
|
||||
const alphaGroup = groups.find((g) => Object.values(g.groupKey).includes('alpha'))
|
||||
const betaGroup = groups.find((g) => Object.values(g.groupKey).includes('beta'))
|
||||
expect(alphaGroup?.metrics.count).toBe(2)
|
||||
expect(betaGroup?.metrics.count).toBe(1)
|
||||
})
|
||||
|
||||
it('where on a bare user metadata field filters normally (score, not a system name)', async () => {
|
||||
await brain.add({
|
||||
data: 'high score',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-score',
|
||||
metadata: { score: 42 }
|
||||
})
|
||||
await brain.add({
|
||||
data: 'low score',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-score',
|
||||
metadata: { score: 7 }
|
||||
})
|
||||
|
||||
const hit = await brain.find({ type: NounType.Person, subtype: 'ns-law-score', where: { score: 42 } })
|
||||
expect(hit).toHaveLength(1)
|
||||
expect(hit[0].metadata?.score).toBe(42)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 5 — typed refusals, naming both candidates.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it("bare orderBy 'createdAt' refuses when no such metadata field exists — names both candidates", async () => {
|
||||
await brain.add({
|
||||
data: 'no metadata.createdAt here',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-refuse-createdAt',
|
||||
metadata: { name: 'x' }
|
||||
})
|
||||
|
||||
await expectRefusal(
|
||||
brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-refuse-createdAt',
|
||||
orderBy: 'createdAt',
|
||||
limit: 10
|
||||
}),
|
||||
UnresolvableFieldError,
|
||||
'system.createdAt',
|
||||
'metadata.createdAt'
|
||||
)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 3 — invisible plumbing refuses in either spelling; system.<name>
|
||||
// for a name that isn't in the ten-scalar map is unresolvable.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it('system.level refuses — level is invisible plumbing, never a system scalar', async () => {
|
||||
await brain.add({
|
||||
data: 'has a level metadata field',
|
||||
type: NounType.Person,
|
||||
metadata: { level: 5 }
|
||||
})
|
||||
await expectRefusal(brain.find({ orderBy: 'system.level', limit: 10 }), UnresolvableFieldError)
|
||||
})
|
||||
|
||||
it('system.vector refuses — vector is invisible plumbing, never a system scalar', async () => {
|
||||
await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } })
|
||||
await expectRefusal(brain.find({ orderBy: 'system.vector', limit: 10 }), UnresolvableFieldError)
|
||||
})
|
||||
|
||||
it('system.data refuses — data is a payload container, never a system scalar', async () => {
|
||||
await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } })
|
||||
await expectRefusal(brain.find({ orderBy: 'system.data', limit: 10 }), UnresolvableFieldError)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 6 — the ordering contract.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
async function addOrderingProbeRows(): Promise<{ ranked: string[]; missing: string }> {
|
||||
const low = await brain.add({
|
||||
data: 'low score',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ordering',
|
||||
metadata: { score: 5 }
|
||||
})
|
||||
const high = await brain.add({
|
||||
data: 'high score',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ordering',
|
||||
metadata: { score: 9 }
|
||||
})
|
||||
const missing = await brain.add({
|
||||
data: 'no score field at all',
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ordering',
|
||||
metadata: { name: 'no-score' }
|
||||
})
|
||||
return { ranked: [low, high], missing }
|
||||
}
|
||||
|
||||
it('a row missing the orderBy field sorts LAST in desc — and is never dropped', async () => {
|
||||
const { ranked, missing } = await addOrderingProbeRows()
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ordering',
|
||||
orderBy: 'score',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc).toHaveLength(3)
|
||||
expect(desc.map((r: any) => r.id)).toEqual([ranked[1], ranked[0], missing])
|
||||
})
|
||||
|
||||
it('a row missing the orderBy field sorts LAST in asc too — and is never dropped', async () => {
|
||||
const { ranked, missing } = await addOrderingProbeRows()
|
||||
const asc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ordering',
|
||||
orderBy: 'score',
|
||||
order: 'asc',
|
||||
limit: 100
|
||||
})
|
||||
expect(asc).toHaveLength(3)
|
||||
expect(asc.map((r: any) => r.id)).toEqual([ranked[0], ranked[1], missing])
|
||||
})
|
||||
|
||||
it('ties on the orderBy field break by id ascending, in BOTH directions', async () => {
|
||||
const tiedIds: string[] = []
|
||||
for (let i = 0; i < 4; i++) {
|
||||
tiedIds.push(
|
||||
await brain.add({
|
||||
data: `tied ${i}`,
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ties',
|
||||
metadata: { score: 5 }
|
||||
})
|
||||
)
|
||||
}
|
||||
const expectedOrder = [...tiedIds].sort()
|
||||
|
||||
const asc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ties',
|
||||
orderBy: 'score',
|
||||
order: 'asc',
|
||||
limit: 100
|
||||
})
|
||||
expect(asc.map((r: any) => r.id)).toEqual(expectedOrder)
|
||||
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'ns-law-ties',
|
||||
orderBy: 'score',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
// Same tie-break ordering regardless of the primary direction — the
|
||||
// contract states one universal rule ("id ascending"), not "reverse of
|
||||
// the primary order".
|
||||
expect(desc.map((r: any) => r.id)).toEqual(expectedOrder)
|
||||
})
|
||||
|
||||
// -------------------------------------------------------------------
|
||||
// Rule 5 (options face) — unimplemented find() options refuse loudly
|
||||
// instead of being accepted and silently ignored.
|
||||
// -------------------------------------------------------------------
|
||||
|
||||
it('find({ cursor }) refuses with UnsupportedFindOptionError', async () => {
|
||||
await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } })
|
||||
await expectRefusal(brain.find({ cursor: 'anything', limit: 10 }), UnsupportedFindOptionError)
|
||||
})
|
||||
|
||||
it('find({ includeRelations }) refuses with UnsupportedFindOptionError', async () => {
|
||||
await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } })
|
||||
await expectRefusal(brain.find({ includeRelations: true, limit: 10 }), UnsupportedFindOptionError)
|
||||
})
|
||||
|
||||
it('find({ writeOnly }) refuses with UnsupportedFindOptionError', async () => {
|
||||
await brain.add({ data: 'row', type: NounType.Person, metadata: { name: 'x' } })
|
||||
await expectRefusal(brain.find({ writeOnly: true, limit: 10 }), UnsupportedFindOptionError)
|
||||
})
|
||||
})
|
||||
|
|
@ -164,19 +164,19 @@ describe('BR-ADV-FEATURES-BUN regression', () => {
|
|||
await b.close()
|
||||
})
|
||||
|
||||
it('groupBy "system.type" resolves to the entity type, not null (the legacy "noun" alias is dead)', async () => {
|
||||
it('groupBy "noun" resolves to the entity type, not null', async () => {
|
||||
const b: any = new Brainy({ requireSubtype: false, storage: { type: 'memory' } })
|
||||
await b.init()
|
||||
await b.add({ data: 'p', type: NounType.Person })
|
||||
b.defineAggregate({
|
||||
name: 'byNoun',
|
||||
source: { type: NounType.Person },
|
||||
groupBy: ['system.type'],
|
||||
groupBy: ['noun'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
const rows: any[] = await b.find({ aggregate: 'byNoun' })
|
||||
expect(rows.length).toBe(1)
|
||||
expect(rows[0].groupKey['system.type']).toBe(NounType.Person)
|
||||
expect(rows[0].groupKey.noun).toBe(NounType.Person)
|
||||
await b.close()
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -42,11 +42,8 @@ describe('aggregation + query field-resolution law', () => {
|
|||
it('reserved-field groupBy decrements on delete (the drift bug)', async () => {
|
||||
brain.defineAggregate({
|
||||
name: 'by_subtype',
|
||||
// system.subtype — subtype is an add() param (an engine scalar), never
|
||||
// a user metadata field; bare 'subtype' now addresses the user's own
|
||||
// metadata bag under the sealed field-addressing law.
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['system.subtype'],
|
||||
groupBy: ['subtype'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
|
||||
|
|
@ -63,7 +60,7 @@ describe('aggregation + query field-resolution law', () => {
|
|||
}
|
||||
let groups = await brain.queryAggregate('by_subtype')
|
||||
expect(groups).toHaveLength(1)
|
||||
expect(groups[0].groupKey).toEqual({ 'system.subtype': 'note' })
|
||||
expect(groups[0].groupKey).toEqual({ subtype: 'note' })
|
||||
expect(groups[0].metrics.count).toBe(5)
|
||||
|
||||
await brain.remove(ids[0])
|
||||
|
|
@ -79,7 +76,7 @@ describe('aggregation + query field-resolution law', () => {
|
|||
brain.defineAggregate({
|
||||
name: 'by_subtype',
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['system.subtype'],
|
||||
groupBy: ['subtype'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
const id = await brain.add({
|
||||
|
|
@ -91,7 +88,7 @@ describe('aggregation + query field-resolution law', () => {
|
|||
|
||||
const groups = await brain.queryAggregate('by_subtype')
|
||||
const byKey = Object.fromEntries(
|
||||
groups.map((g) => [String(g.groupKey['system.subtype']), g.metrics.count])
|
||||
groups.map((g) => [String(g.groupKey.subtype), g.metrics.count])
|
||||
)
|
||||
expect(byKey['published']).toBe(1)
|
||||
// The old group must be gone or zero — never still counting the entity.
|
||||
|
|
@ -101,7 +98,7 @@ describe('aggregation + query field-resolution law', () => {
|
|||
it('source.where on a reserved field filters instead of matching nothing', async () => {
|
||||
brain.defineAggregate({
|
||||
name: 'notes_only',
|
||||
source: { type: NounType.Document, where: { 'system.subtype': 'note' } },
|
||||
source: { type: NounType.Document, where: { subtype: 'note' } },
|
||||
groupBy: ['team'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
|
|
|
|||
|
|
@ -331,10 +331,8 @@ describe('Comprehensive All-APIs Test', () => {
|
|||
it('should handle metadata queries efficiently', async () => {
|
||||
const start = Date.now()
|
||||
|
||||
// system.type — the legacy where.type→noun alias is dead; bare 'type'
|
||||
// in where now addresses the user's own metadata field.
|
||||
const results = await brain.find({
|
||||
where: { 'system.type': NounType.Document },
|
||||
where: { type: NounType.Document },
|
||||
limit: 100
|
||||
})
|
||||
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
|||
import * as fs from 'node:fs'
|
||||
import * as os from 'node:os'
|
||||
import * as path from 'node:path'
|
||||
import { Brainy, ProtectedArtifactError, splitNounMetadataRecord, type CommitFact } from '../../src/index.js'
|
||||
import { Brainy, ProtectedArtifactError, type CommitFact } from '../../src/index.js'
|
||||
|
||||
async function allFacts(brain: any): Promise<CommitFact[]> {
|
||||
const scan = brain.scanFacts()
|
||||
|
|
@ -65,13 +65,7 @@ describe('fact log dual-write (memory adapter)', () => {
|
|||
const updateFact = facts[facts.length - 1]
|
||||
const op = updateFact.ops.find((o) => o.id === id)!
|
||||
expect(op.record).not.toBeNull()
|
||||
// The fact log is byte-faithful: op.record.metadata is the RAW stored
|
||||
// record (v2 nested-bag since the field-addressing law) — read the user
|
||||
// field through the shape-aware split, like every other reader.
|
||||
const { custom } = splitNounMetadataRecord(
|
||||
op.record!.metadata as Record<string, unknown>
|
||||
)
|
||||
expect(custom.v).toBe('new')
|
||||
expect((op.record!.metadata as any).v).toBe('new')
|
||||
})
|
||||
|
||||
it('a transact commits ONE fact carrying all its ops, with meta', async () => {
|
||||
|
|
|
|||
|
|
@ -1,186 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/history-repacking
|
||||
* @description The D1+D3 two-tier history lifecycle end-to-end on a real
|
||||
* brain. Laws: (1) repacking is RE-REPRESENTATION — after folding, every
|
||||
* asOf() read below the fold boundary answers exactly as before, across a
|
||||
* cold reopen; (2) folded per-generation directories are physically gone
|
||||
* (the file-count cure is real, not cosmetic); (3) repack + reclaim compose:
|
||||
* bounded retention after repacking drops whole segments and asOf below the
|
||||
* horizon throws GenerationCompactedError; (4) repackHistory is explicit
|
||||
* API and time-bounded (spent budget = consistent no-op).
|
||||
*
|
||||
* Uses a tiny REPACK_LIVE_WINDOW override so a small history has a cold
|
||||
* tier at all (the production window is 1024).
|
||||
*/
|
||||
import { describe, it, expect, afterEach } from 'vitest'
|
||||
import * as fs from 'node:fs'
|
||||
import * as path from 'node:path'
|
||||
import * as os from 'node:os'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
import { GenerationStore } from '../../src/db/generationStore.js'
|
||||
import { GenerationCompactedError } from '../../src/db/errors.js'
|
||||
import { SEGMENTS_PREFIX } from '../../src/db/generationSegments.js'
|
||||
|
||||
const stub = async (text: string): Promise<number[]> => {
|
||||
const h = text.split('').reduce((a, c) => a + c.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(h + i))
|
||||
}
|
||||
|
||||
const openBrain = async (dir: string): Promise<Brainy> => {
|
||||
const brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'filesystem', path: dir },
|
||||
embeddingFunction: stub
|
||||
})
|
||||
await brain.init()
|
||||
return brain
|
||||
}
|
||||
|
||||
describe('history repacking — the two-tier lifecycle', () => {
|
||||
const dirs: string[] = []
|
||||
const tempDir = (): string => {
|
||||
const d = fs.mkdtempSync(path.join(os.tmpdir(), 'brainy-repack-'))
|
||||
dirs.push(d)
|
||||
return d
|
||||
}
|
||||
const originalWindow = GenerationStore.REPACK_LIVE_WINDOW
|
||||
|
||||
afterEach(() => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = originalWindow
|
||||
for (const d of dirs.splice(0)) {
|
||||
try {
|
||||
fs.rmSync(d, { recursive: true, force: true })
|
||||
} catch {
|
||||
/* best effort */
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
it('repack preserves every historical read across cold reopen; folded dirs are gone', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 3
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
|
||||
const id = await brain.add({
|
||||
data: 'versioned-entity',
|
||||
type: NounType.Document,
|
||||
metadata: { v: 0 }
|
||||
})
|
||||
for (let v = 1; v <= 10; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
|
||||
// Ground truth BEFORE repacking: capture asOf views for early generations.
|
||||
const before: Record<number, number> = {}
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await brain.asOf(g)
|
||||
before[g] = (await db.get(id))?.metadata?.v as number
|
||||
await db.release()
|
||||
}
|
||||
|
||||
const result = await brain.repackHistory()
|
||||
expect(result.foldedGenerations).toBeGreaterThan(0)
|
||||
expect(result.segmentsCreated).toBeGreaterThan(0)
|
||||
|
||||
// The folded per-generation directories are PHYSICALLY gone…
|
||||
const genDirs = fs
|
||||
.readdirSync(path.join(dir, '_generations'), { withFileTypes: true })
|
||||
.filter((e) => e.isDirectory() && /^\d+$/.test(e.name)).length
|
||||
expect(genDirs).toBeLessThanOrEqual(4) // live window (3) + at most the newest
|
||||
// …and the segment tier exists (the filesystem adapter stores objects
|
||||
// gzipped, so the manifest may live at either spelling).
|
||||
const segDir = path.join(dir, SEGMENTS_PREFIX)
|
||||
expect(
|
||||
fs.existsSync(path.join(segDir, 'manifest.json')) ||
|
||||
fs.existsSync(path.join(segDir, 'manifest.json.gz'))
|
||||
).toBe(true)
|
||||
expect(fs.readdirSync(segDir).some((f) => f.endsWith('.bgs'))).toBe(true)
|
||||
|
||||
// Same asOf answers from the packed tier, same process…
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await brain.asOf(g)
|
||||
expect((await db.get(id))?.metadata?.v).toBe(before[g])
|
||||
await db.release()
|
||||
}
|
||||
await brain.close()
|
||||
|
||||
// …and across a COLD REOPEN (manifest discovery, no live dirs to list).
|
||||
const reopened = await openBrain(dir)
|
||||
for (const g of [2, 4, 6]) {
|
||||
const db = await reopened.asOf(g)
|
||||
expect((await db.get(id))?.metadata?.v).toBe(before[g])
|
||||
await db.release()
|
||||
}
|
||||
expect((await reopened.get(id))?.metadata?.v).toBe(10) // live state untouched
|
||||
await reopened.close()
|
||||
})
|
||||
|
||||
it('repack + bounded reclaim compose: whole segments drop, horizon is loud', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'reclaim-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 8; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
await brain.repackHistory()
|
||||
|
||||
// Reclaim down to the 3 newest generations — packed segments below the
|
||||
// horizon drop whole; asOf below throws loudly.
|
||||
const res = await brain.compactHistory({ maxGenerations: 3 })
|
||||
expect(res.removedGenerations).toBeGreaterThan(0)
|
||||
await expect(brain.asOf(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
expect((await brain.get(id))?.metadata?.v).toBe(8)
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('generationDigest: reopen-stable, divergence-sensitive, loud below the horizon', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'digest-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
await brain.repackHistory()
|
||||
|
||||
const gen = brain.generation()
|
||||
const atHead = await brain.generationDigest(gen)
|
||||
const atMid = await brain.generationDigest(3)
|
||||
expect(atHead).toMatch(/^[0-9a-f]{8}$/)
|
||||
expect(atMid).not.toBe(atHead) // more history ⇒ different digest
|
||||
await brain.close()
|
||||
|
||||
// Reopen-stable: same history, same digests (packed prefix stability).
|
||||
const reopened = await openBrain(dir)
|
||||
expect(await reopened.generationDigest(gen)).toBe(atHead)
|
||||
expect(await reopened.generationDigest(3)).toBe(atMid)
|
||||
|
||||
// New history diverges the head digest.
|
||||
await reopened.update({ id, metadata: { v: 7 } })
|
||||
await reopened.flush()
|
||||
expect(await reopened.generationDigest(reopened.generation())).not.toBe(atHead)
|
||||
|
||||
// Below the horizon: LOUD, never a silent pin of reclaimed history.
|
||||
await reopened.compactHistory({ maxGenerations: 2 })
|
||||
await expect(reopened.generationDigest(1)).rejects.toBeInstanceOf(GenerationCompactedError)
|
||||
await reopened.close()
|
||||
})
|
||||
|
||||
it('a spent time budget is a consistent no-op; the next pass resumes', async () => {
|
||||
;(GenerationStore as any).REPACK_LIVE_WINDOW = 2
|
||||
const dir = tempDir()
|
||||
const brain = await openBrain(dir)
|
||||
const id = await brain.add({ data: 'budget-probe', type: NounType.Document, metadata: { v: 0 } })
|
||||
for (let v = 1; v <= 6; v++) await brain.update({ id, metadata: { v } })
|
||||
await brain.flush()
|
||||
|
||||
const bounded = await brain.repackHistory({ timeBudgetMs: 0 })
|
||||
expect(bounded).toEqual({ foldedGenerations: 0, segmentsCreated: 0 })
|
||||
|
||||
const resumed = await brain.repackHistory()
|
||||
expect(resumed.foldedGenerations).toBeGreaterThan(0)
|
||||
const db = await brain.asOf(3)
|
||||
expect((await db.get(id))?.metadata?.v).toBeDefined()
|
||||
await db.release()
|
||||
await brain.close()
|
||||
})
|
||||
})
|
||||
|
|
@ -2,8 +2,8 @@
|
|||
* @module tests/integration/lens-consistency
|
||||
* @description The three metadata "lenses" over one corpus must agree with
|
||||
* canonical ground truth id-for-id, warm AND after a cold reopen:
|
||||
* - combined: find({ type: T, where: { 'system.subtype': S } })
|
||||
* - subtype-only: find({ where: { 'system.subtype': S } })
|
||||
* - combined: find({ type: T, where: { subtype: S } })
|
||||
* - subtype-only: find({ where: { subtype: S } })
|
||||
* - type-only: find({ type: T })
|
||||
* Ported from the fresh-brain probe that closed the type+subtype lens-drop
|
||||
* investigation (a restored pre-8.2.2 torn capture had entities visible to the
|
||||
|
|
@ -63,11 +63,8 @@ async function assertAllLenses(brain: any): Promise<void> {
|
|||
const subtypes = [...new Set(CORPUS.map((c) => c.subtype))]
|
||||
|
||||
for (const { type, subtype } of CORPUS) {
|
||||
// system.subtype — subtype is an add()/update() param (an engine scalar),
|
||||
// never a user metadata field; bare 'subtype' now addresses the user's
|
||||
// own metadata bag under the sealed field-addressing law.
|
||||
const combined = idSet(await brain.find({ type, where: { 'system.subtype': subtype }, limit: 1000 }))
|
||||
const subtypeOnly = idSet(await brain.find({ where: { 'system.subtype': subtype }, limit: 1000 }))
|
||||
const combined = idSet(await brain.find({ type, where: { subtype }, limit: 1000 }))
|
||||
const subtypeOnly = idSet(await brain.find({ where: { subtype }, limit: 1000 }))
|
||||
const truthPair = await groundTruth(brain, { type, subtype })
|
||||
const truthSubtype = await groundTruth(brain, { subtype })
|
||||
|
||||
|
|
@ -85,7 +82,7 @@ async function assertAllLenses(brain: any): Promise<void> {
|
|||
// Count cross-check against the corpus definition itself.
|
||||
for (const subtype of subtypes) {
|
||||
const expected = CORPUS.filter((c) => c.subtype === subtype).reduce((s, c) => s + c.count, 0)
|
||||
const got = (await brain.find({ where: { 'system.subtype': subtype }, limit: 1000 })).length
|
||||
const got = (await brain.find({ where: { subtype }, limit: 1000 })).length
|
||||
expect(got).toBe(expected)
|
||||
}
|
||||
}
|
||||
|
|
@ -126,16 +123,16 @@ describe('lens consistency — combined vs subtype-only vs canonical ground trut
|
|||
|
||||
it('after an update() flips type AND subtype, every lens tracks the move exactly', async () => {
|
||||
// The historical cross-bucket-staleness path: change (concept, action) -> (task, review).
|
||||
const victims = await brain.find({ type: 'concept', where: { 'system.subtype': 'action' }, limit: 1 })
|
||||
const victims = await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1 })
|
||||
expect(victims.length).toBe(1)
|
||||
const id = victims[0].id
|
||||
await brain.update({ id, type: 'task', subtype: 'review' })
|
||||
|
||||
const oldCombined = idSet(await brain.find({ type: 'concept', where: { 'system.subtype': 'action' }, limit: 1000 }))
|
||||
const oldCombined = idSet(await brain.find({ type: 'concept', where: { subtype: 'action' }, limit: 1000 }))
|
||||
expect(oldCombined.has(id)).toBe(false) // unposted from the old buckets
|
||||
const newCombined = idSet(await brain.find({ type: 'task', where: { 'system.subtype': 'review' }, limit: 1000 }))
|
||||
const newCombined = idSet(await brain.find({ type: 'task', where: { subtype: 'review' }, limit: 1000 }))
|
||||
expect(newCombined.has(id)).toBe(true) // posted to the new buckets
|
||||
const subtypeOnly = idSet(await brain.find({ where: { 'system.subtype': 'review' }, limit: 1000 }))
|
||||
const subtypeOnly = idSet(await brain.find({ where: { subtype: 'review' }, limit: 1000 }))
|
||||
expect(subtypeOnly.has(id)).toBe(true)
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -1,194 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/level-field-shadow
|
||||
* @description The reserved-name shadow fix (VENUE-BRAINY-ORDERBY-NOOP,
|
||||
* 2026-08-03): `level` is HNSW plumbing, not an entity field — it must never
|
||||
* shadow user metadata of the same name. Pre-fix, STANDARD_ENTITY_FIELDS
|
||||
* listed `level`, so every by-name read returned the engine's internal 0
|
||||
* (all-equal → stable sort → insertion order, silently), and the indexing
|
||||
* views stamped level:0 into the same flattened column as user values
|
||||
* (multi-valued [0, real] poison). Laws:
|
||||
* (1) venue's exact repro sorts: three adds with metadata.level 3/9/6 →
|
||||
* find({orderBy:'level'}) returns 9,6,3 desc and 3,6,9 asc;
|
||||
* (2) where {level: N} matches through filter AND egress guard;
|
||||
* (3) the index column carries the user value only (no 0 poison);
|
||||
* (4) update() keeps `level` readable (the update indexing view is clean too);
|
||||
* (5) the transact() update path never rewrites the noun record on a
|
||||
* metadata-only patch (the planUpdate granularity completion).
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
import { EXPECTED_INDEX_EPOCH } from '../../src/storage/brainFormat.js'
|
||||
|
||||
const stubEmbedding = async (text: string): Promise<number[]> => {
|
||||
const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(hash + i))
|
||||
}
|
||||
|
||||
describe('level field shadow — user metadata named level is a real field', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
async function addProbeRows(): Promise<string[]> {
|
||||
const ids: string[] = []
|
||||
for (const level of [3, 9, 6]) {
|
||||
ids.push(
|
||||
await brain.add({
|
||||
data: `probe character level ${level}`,
|
||||
type: NounType.Person,
|
||||
subtype: 'probe-char',
|
||||
metadata: { name: `char-${level}`, level }
|
||||
})
|
||||
)
|
||||
}
|
||||
return ids
|
||||
}
|
||||
|
||||
it("venue's exact repro: orderBy 'level' sorts desc and asc", async () => {
|
||||
await addProbeRows()
|
||||
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'probe-char',
|
||||
orderBy: 'level',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc.map((r: any) => r.metadata?.level)).toEqual([9, 6, 3])
|
||||
|
||||
const asc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'probe-char',
|
||||
orderBy: 'level',
|
||||
order: 'asc',
|
||||
limit: 100
|
||||
})
|
||||
expect(asc.map((r: any) => r.metadata?.level)).toEqual([3, 6, 9])
|
||||
})
|
||||
|
||||
it('ordered reads are COMPLETE — no row dropped (the 2-of-3 face)', async () => {
|
||||
const ids = await addProbeRows()
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'probe-char',
|
||||
orderBy: 'level',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc).toHaveLength(3)
|
||||
expect(new Set(desc.map((r: any) => r.id))).toEqual(new Set(ids))
|
||||
})
|
||||
|
||||
it('where {level: N} matches through the filter and the egress guard', async () => {
|
||||
const ids = await addProbeRows()
|
||||
const hit = await brain.find({ where: { level: 9 } })
|
||||
expect(hit).toHaveLength(1)
|
||||
expect(hit[0].id).toBe(ids[1])
|
||||
expect(hit[0].metadata?.level).toBe(9)
|
||||
})
|
||||
|
||||
it('the index column carries ONLY the user value (no 0 poison)', async () => {
|
||||
const ids = await addProbeRows()
|
||||
const metadataIndex = (brain as any).metadataIndex
|
||||
const value = await metadataIndex.getFieldValueForEntity(ids[1], 'level')
|
||||
expect(value).toBe(9)
|
||||
|
||||
// Zero must not match anything — pre-fix every entity carried a phantom 0.
|
||||
const phantom = await brain.find({ where: { level: 0 } })
|
||||
expect(phantom).toHaveLength(0)
|
||||
})
|
||||
|
||||
it('update() keeps level readable (the update indexing view is clean)', async () => {
|
||||
const ids = await addProbeRows()
|
||||
await brain.update({ id: ids[0], metadata: { level: 12 } })
|
||||
const desc = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: 'probe-char',
|
||||
orderBy: 'level',
|
||||
order: 'desc',
|
||||
limit: 100
|
||||
})
|
||||
expect(desc.map((r: any) => r.metadata?.level)).toEqual([12, 9, 6])
|
||||
})
|
||||
|
||||
it('transact() metadata-only update never rewrites the noun record', async () => {
|
||||
const ids = await addProbeRows()
|
||||
const storage = (brain as any).storage
|
||||
const saveNounSpy = vi.spyOn(storage, 'saveNoun')
|
||||
|
||||
await brain.transact([
|
||||
{ op: 'update', id: ids[0], metadata: { level: 4 } },
|
||||
{ op: 'update', id: ids[2], metadata: { level: 7 } }
|
||||
])
|
||||
|
||||
expect(saveNounSpy).not.toHaveBeenCalled()
|
||||
saveNounSpy.mockRestore()
|
||||
|
||||
const after = await brain.get(ids[0], { includeVectors: true })
|
||||
expect(after?.metadata?.level).toBe(4)
|
||||
expect(Array.isArray(after?.vector) && after!.vector!.length).toBe(384)
|
||||
})
|
||||
|
||||
it('this build runs index epoch 3 (the namespace-law key split rebuild)', () => {
|
||||
expect(EXPECTED_INDEX_EPOCH).toBe(3)
|
||||
})
|
||||
})
|
||||
|
||||
describe('noun-record writes never stamp over stored graph state', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('a data-changing update preserves LEGACY inline connections in the record', async () => {
|
||||
// Codec-era records carry an EMPTY connections field by design (the
|
||||
// adjacency lives in a separate compressed blob) — the clobber window
|
||||
// exists only for legacy pre-codec records whose adjacency is inline.
|
||||
// Simulate one: write the record with inline connections directly.
|
||||
const id = await brain.add({
|
||||
data: 'legacy-shaped node',
|
||||
type: NounType.Concept,
|
||||
metadata: { n: 1 }
|
||||
})
|
||||
const storage = (brain as any).storage
|
||||
const rec = await storage.getNoun(id)
|
||||
const legacy = {
|
||||
...rec,
|
||||
connections: new Map([[0, new Set(['00000000-0000-4000-8000-00000000aaaa'])]]),
|
||||
level: 1
|
||||
}
|
||||
await storage.saveNoun(legacy)
|
||||
const before = await storage.getNoun(id)
|
||||
expect(before.connections.size).toBeGreaterThan(0)
|
||||
|
||||
// A data-changing update stages SaveNounOperation with placeholder
|
||||
// adjacency — the legacy inline connections must survive the write.
|
||||
await brain.update({ id, data: 'completely re-embedded text' })
|
||||
|
||||
const after = await storage.getNoun(id)
|
||||
expect(after.connections.size).toBeGreaterThan(0)
|
||||
expect(after.level).toBe(1)
|
||||
})
|
||||
})
|
||||
|
|
@ -20,16 +20,6 @@ import { MigrationRunner, MIGRATIONS } from '../../src/migration/index.js'
|
|||
import type { Migration } from '../../src/migration/index.js'
|
||||
import { NounType, VerbType } from '../../src/types/graphTypes.js'
|
||||
|
||||
// THE VIEW CONTRACT (field-addressing law): transforms receive engine fields
|
||||
// top-level and the USER's bag nested under `metadata` — user-field changes
|
||||
// go inside the bag. These two helpers keep the one-liner migrations tidy.
|
||||
const bagOf = (m: Record<string, unknown>): Record<string, unknown> =>
|
||||
m.metadata as Record<string, unknown>
|
||||
const withBag = (
|
||||
m: Record<string, unknown>,
|
||||
patch: Record<string, unknown>
|
||||
): Record<string, unknown> => ({ ...m, metadata: { ...bagOf(m), ...patch } })
|
||||
|
||||
// Helper to temporarily inject migrations into the MIGRATIONS array
|
||||
function withMigrations(migrations: Migration[], fn: () => Promise<void>): Promise<void> {
|
||||
const original = MIGRATIONS.splice(0, MIGRATIONS.length)
|
||||
|
|
@ -88,11 +78,9 @@ describe('Migration System', () => {
|
|||
description: 'Add version field to entities with status',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
// Only transform entities that have our specific 'status' USER field
|
||||
// (user fields live in the nested bag — the view contract).
|
||||
const bag = m.metadata as Record<string, unknown>
|
||||
if ('status' in bag && !('version' in bag)) {
|
||||
return { ...m, metadata: { ...bag, version: 1 } }
|
||||
// Only transform entities that have our specific 'status' field
|
||||
if ('status' in m && !('version' in m)) {
|
||||
return { ...m, version: 1 }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -106,8 +94,7 @@ describe('Migration System', () => {
|
|||
// All 3 entities have 'status' metadata
|
||||
expect(p.affectedEntities).toBeGreaterThanOrEqual(3)
|
||||
expect(p.sampleChanges.length).toBeGreaterThan(0)
|
||||
// Samples carry the VIEW shape: user fields inside `.metadata`.
|
||||
expect(p.sampleChanges[0].after.metadata.version).toBe(1)
|
||||
expect(p.sampleChanges[0].after.version).toBe(1)
|
||||
|
||||
// Verify no data was modified (dry-run)
|
||||
const entity = await brain.get(id1)
|
||||
|
|
@ -124,10 +111,9 @@ describe('Migration System', () => {
|
|||
description: 'Rename state to status',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
const bag = m.metadata as Record<string, unknown>
|
||||
if ('state' in bag) {
|
||||
const { state, ...rest } = bag
|
||||
return { ...m, metadata: { ...rest, status: state } }
|
||||
if ('state' in m) {
|
||||
const { state, ...rest } = m
|
||||
return { ...rest, status: state }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -138,12 +124,11 @@ describe('Migration System', () => {
|
|||
const p = preview as any
|
||||
expect(p.sampleChanges.length).toBeGreaterThanOrEqual(1)
|
||||
|
||||
// Find the sample for our entity (it has the 'state' USER field —
|
||||
// samples carry the VIEW shape, user fields inside `.metadata`)
|
||||
const sample = p.sampleChanges.find((s: any) => s.before.metadata.state === 'draft')
|
||||
// Find the sample for our entity (it has the 'state' field)
|
||||
const sample = p.sampleChanges.find((s: any) => s.before.state === 'draft')
|
||||
expect(sample).toBeDefined()
|
||||
expect(sample.after.metadata.status).toBe('draft')
|
||||
expect(sample.after.metadata.state).toBeUndefined()
|
||||
expect(sample.after.status).toBe('draft')
|
||||
expect(sample.after.state).toBeUndefined()
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
@ -164,8 +149,8 @@ describe('Migration System', () => {
|
|||
description: 'Add migrated flag to entities with priority',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
if ('priority' in bagOf(m) && !('migrated' in bagOf(m))) {
|
||||
return withBag(m, { migrated: true })
|
||||
if ('priority' in m && !('migrated' in m)) {
|
||||
return { ...m, migrated: true }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -194,8 +179,8 @@ describe('Migration System', () => {
|
|||
description: 'Uppercase status field only when present',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
if (typeof bagOf(m).status === 'string') {
|
||||
return withBag(m, { status: (bagOf(m).status as string).toUpperCase() })
|
||||
if (typeof m.status === 'string') {
|
||||
return { ...m, status: (m.status as string).toUpperCase() }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -218,7 +203,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Double count',
|
||||
applies: 'nouns',
|
||||
transform: (m) => typeof bagOf(m).count === 'number' ? withBag(m, { count: (bagOf(m).count as number) * 2 }) : null
|
||||
transform: (m) => typeof m.count === 'number' ? { ...m, count: (m.count as number) * 2 } : null
|
||||
}
|
||||
|
||||
const migration2: Migration = {
|
||||
|
|
@ -226,7 +211,7 @@ describe('Migration System', () => {
|
|||
version: '1.1.0',
|
||||
description: 'Add 10 to count',
|
||||
applies: 'nouns',
|
||||
transform: (m) => typeof bagOf(m).count === 'number' ? withBag(m, { count: (bagOf(m).count as number) + 10 }) : null
|
||||
transform: (m) => typeof m.count === 'number' ? { ...m, count: (m.count as number) + 10 } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration1, migration2], async () => {
|
||||
|
|
@ -244,7 +229,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Increment v',
|
||||
applies: 'nouns',
|
||||
transform: (m) => typeof bagOf(m).v === 'number' ? withBag(m, { v: (bagOf(m).v as number) + 1 }) : null
|
||||
transform: (m) => typeof m.v === 'number' ? { ...m, v: (m.v as number) + 1 } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -281,7 +266,7 @@ describe('Migration System', () => {
|
|||
version: '2.0.0',
|
||||
description: 'Add y field to entities with x',
|
||||
applies: 'nouns',
|
||||
transform: (m) => 'x' in bagOf(m) && !('y' in bagOf(m)) ? withBag(m, { y: 2 }) : null
|
||||
transform: (m) => 'x' in m && !('y' in m) ? { ...m, y: 2 } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -305,8 +290,8 @@ describe('Migration System', () => {
|
|||
description: 'Replace original with migrated',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
if (bagOf(m).original === true) {
|
||||
return withBag(m, { original: false, migrated: true })
|
||||
if (m.original === true) {
|
||||
return { ...m, original: false, migrated: true }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -338,7 +323,7 @@ describe('Migration System', () => {
|
|||
version: '4.0.0',
|
||||
description: 'Add field',
|
||||
applies: 'nouns',
|
||||
transform: (m) => 'q' in bagOf(m) && !('r' in bagOf(m)) ? withBag(m, { r: 2 }) : null
|
||||
transform: (m) => 'q' in m && !('r' in m) ? { ...m, r: 2 } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -399,7 +384,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Auto migrate test',
|
||||
applies: 'nouns',
|
||||
transform: (m) => 'legacy' in bagOf(m) ? withBag(m, { legacy: false, upgraded: true }) : null
|
||||
transform: (m) => 'legacy' in m ? { ...m, legacy: false, upgraded: true } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -425,7 +410,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Add y to entities with x',
|
||||
applies: 'nouns',
|
||||
transform: (m) => 'x' in bagOf(m) ? withBag(m, { y: true }) : null
|
||||
transform: (m) => 'x' in m ? { ...m, y: true } : null
|
||||
}
|
||||
|
||||
const progressCalls: any[] = []
|
||||
|
|
@ -459,7 +444,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Increment v on entities that have it',
|
||||
applies: 'nouns',
|
||||
transform: (m) => typeof bagOf(m).v === 'number' ? withBag(m, { v: (bagOf(m).v as number) + 1 }) : null
|
||||
transform: (m) => typeof m.v === 'number' ? { ...m, v: (m.v as number) + 1 } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -492,10 +477,9 @@ describe('Migration System', () => {
|
|||
description: 'Rename strength to intensity',
|
||||
applies: 'verbs',
|
||||
transform: (m) => {
|
||||
const bag = bagOf(m)
|
||||
if ('strength' in bag) {
|
||||
const { strength, ...rest } = bag
|
||||
return { ...m, metadata: { ...rest, intensity: strength } }
|
||||
if ('strength' in m) {
|
||||
const { strength, ...rest } = m
|
||||
return { ...rest, intensity: strength }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -523,7 +507,7 @@ describe('Migration System', () => {
|
|||
version: '1.0.0',
|
||||
description: 'Update tag from old to new',
|
||||
applies: 'both',
|
||||
transform: (m) => bagOf(m).tag === 'old' ? withBag(m, { tag: 'new' }) : null
|
||||
transform: (m) => m.tag === 'old' ? { ...m, tag: 'new' } : null
|
||||
}
|
||||
|
||||
await withMigrations([migration], async () => {
|
||||
|
|
@ -593,11 +577,11 @@ describe('Migration System', () => {
|
|||
description: 'Transform that throws on non-number values',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
if ('value' in bagOf(m)) {
|
||||
if (typeof bagOf(m).value !== 'number') {
|
||||
if ('value' in m) {
|
||||
if (typeof m.value !== 'number') {
|
||||
throw new Error('value must be a number')
|
||||
}
|
||||
return withBag(m, { value: (bagOf(m).value as number) * 10 })
|
||||
return { ...m, value: (m.value as number) * 10 }
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
|
@ -631,7 +615,7 @@ describe('Migration System', () => {
|
|||
description: 'Always throws',
|
||||
applies: 'nouns',
|
||||
transform: (m) => {
|
||||
if ('boom' in bagOf(m)) {
|
||||
if ('boom' in m) {
|
||||
throw new Error('deliberate failure')
|
||||
}
|
||||
return null
|
||||
|
|
|
|||
|
|
@ -56,7 +56,7 @@ describe('find({ orderBy }) sort bug regression', () => {
|
|||
|
||||
const results = await brain.find({
|
||||
type: NounType.Concept,
|
||||
orderBy: 'system.createdAt',
|
||||
orderBy: 'createdAt',
|
||||
order: 'desc',
|
||||
limit: 1
|
||||
})
|
||||
|
|
@ -76,7 +76,7 @@ describe('find({ orderBy }) sort bug regression', () => {
|
|||
|
||||
const results = await brain.find({
|
||||
type: NounType.Concept,
|
||||
orderBy: 'system.createdAt',
|
||||
orderBy: 'createdAt',
|
||||
order: 'asc',
|
||||
limit: 1
|
||||
})
|
||||
|
|
@ -94,7 +94,7 @@ describe('find({ orderBy }) sort bug regression', () => {
|
|||
|
||||
const results = await brain.find({
|
||||
type: NounType.Concept,
|
||||
orderBy: 'system.createdAt',
|
||||
orderBy: 'createdAt',
|
||||
order: 'desc'
|
||||
})
|
||||
|
||||
|
|
@ -115,7 +115,7 @@ describe('find({ orderBy }) sort bug regression', () => {
|
|||
|
||||
const results = await brain.find({
|
||||
type: NounType.Concept,
|
||||
orderBy: 'system.updatedAt',
|
||||
orderBy: 'updatedAt',
|
||||
order: 'desc',
|
||||
limit: 1
|
||||
})
|
||||
|
|
@ -136,7 +136,7 @@ describe('find({ orderBy }) sort bug regression', () => {
|
|||
const id3 = await brain.add({ data: 'third', type: NounType.Concept })
|
||||
|
||||
const results = await brain.find({
|
||||
orderBy: 'system.createdAt',
|
||||
orderBy: 'createdAt',
|
||||
order: 'desc',
|
||||
limit: 2
|
||||
})
|
||||
|
|
@ -215,6 +215,7 @@ describe('resolveEntityField helper', () => {
|
|||
'id',
|
||||
'vector',
|
||||
'connections',
|
||||
'level',
|
||||
'type',
|
||||
'confidence',
|
||||
'weight',
|
||||
|
|
@ -227,9 +228,5 @@ describe('resolveEntityField helper', () => {
|
|||
for (const field of expected) {
|
||||
expect(STANDARD_ENTITY_FIELDS.has(field)).toBe(true)
|
||||
}
|
||||
// `level` is deliberately NOT resolvable: it is HNSW plumbing, and listing
|
||||
// it here shadowed user metadata named `level` in every by-name read
|
||||
// (the reserved-name shadow bug). Plumbing stays out of the resolver.
|
||||
expect(STANDARD_ENTITY_FIELDS.has('level')).toBe(false)
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -1,126 +0,0 @@
|
|||
/**
|
||||
* @module tests/integration/update-write-granularity
|
||||
* @description Write-granularity law for update() (SELF-ENGINE-RESTART-GRIND,
|
||||
* 2026-07-29): a metadata-only update must NEVER rewrite the noun record —
|
||||
* the record carries the full vector, so an unconditional save turns every
|
||||
* metadata touch into a whole-vector rewrite + fsync. Under a read-heavy
|
||||
* consumer sweep bumping per-entity stats this amplified into disk saturation
|
||||
* on a production deployment. Laws:
|
||||
* (1) metadata-only update() → zero saveNoun calls (metadata leg only);
|
||||
* (2) data/vector/type-changing update() → saveNoun runs (the vector leg and
|
||||
* HNSW reindex still happen when the vector side actually changed);
|
||||
* (3) the metadata-only path still lands: merged metadata readable, _rev
|
||||
* bumped, find() by the new field sees the entity.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
|
||||
import { Brainy } from '../../src/brainy.js'
|
||||
import { NounType } from '../../src/types/graphTypes.js'
|
||||
|
||||
const stubEmbedding = async (text: string): Promise<number[]> => {
|
||||
const hash = text.split('').reduce((acc, char) => acc + char.charCodeAt(0), 0)
|
||||
return new Array(384).fill(0).map((_, i) => Math.sin(hash + i))
|
||||
}
|
||||
|
||||
describe('update() write granularity', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
brain = new Brainy({
|
||||
requireSubtype: false,
|
||||
storage: { type: 'memory' as const },
|
||||
embeddingFunction: stubEmbedding
|
||||
})
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('metadata-only update never rewrites the noun record (no vector rewrite)', async () => {
|
||||
const id = await brain.add({
|
||||
data: 'granularity law subject',
|
||||
type: NounType.Concept,
|
||||
metadata: { touched: 0 }
|
||||
})
|
||||
|
||||
const storage = (brain as any).storage
|
||||
const saveNounSpy = vi.spyOn(storage, 'saveNoun')
|
||||
|
||||
await brain.update({ id, metadata: { touched: 1 } })
|
||||
|
||||
expect(saveNounSpy).not.toHaveBeenCalled()
|
||||
saveNounSpy.mockRestore()
|
||||
|
||||
// The metadata leg still landed with full semantics.
|
||||
const after = await brain.get(id, { includeVectors: true })
|
||||
expect(after?.metadata?.touched).toBe(1)
|
||||
expect(after?._rev).toBe(2)
|
||||
expect(Array.isArray(after?.vector) && after!.vector!.length).toBe(384)
|
||||
|
||||
const found = await brain.find({ where: { touched: 1 } })
|
||||
expect(found.some((r: any) => r.id === id)).toBe(true)
|
||||
})
|
||||
|
||||
it('confidence/weight/subtype-only updates also skip the noun record', async () => {
|
||||
const id = await brain.add({
|
||||
data: 'reserved-field touch subject',
|
||||
type: NounType.Concept,
|
||||
metadata: {}
|
||||
})
|
||||
|
||||
const storage = (brain as any).storage
|
||||
const saveNounSpy = vi.spyOn(storage, 'saveNoun')
|
||||
|
||||
await brain.update({ id, confidence: 0.5, weight: 2, subtype: 'note' })
|
||||
|
||||
expect(saveNounSpy).not.toHaveBeenCalled()
|
||||
saveNounSpy.mockRestore()
|
||||
|
||||
const after = await brain.get(id)
|
||||
expect(after?.confidence).toBe(0.5)
|
||||
expect(after?.subtype).toBe('note')
|
||||
})
|
||||
|
||||
it('data-changing update still writes the noun record and reindexes', async () => {
|
||||
const id = await brain.add({
|
||||
data: 'original embedded text',
|
||||
type: NounType.Concept,
|
||||
metadata: {}
|
||||
})
|
||||
|
||||
const before = await brain.get(id, { includeVectors: true })
|
||||
|
||||
const storage = (brain as any).storage
|
||||
const saveNounSpy = vi.spyOn(storage, 'saveNoun')
|
||||
|
||||
await brain.update({ id, data: 'completely different embedded text' })
|
||||
|
||||
expect(saveNounSpy).toHaveBeenCalled()
|
||||
saveNounSpy.mockRestore()
|
||||
|
||||
const after = await brain.get(id, { includeVectors: true })
|
||||
expect(after?.data).toBe('completely different embedded text')
|
||||
expect(after?.vector).not.toEqual(before?.vector)
|
||||
})
|
||||
|
||||
it('explicit-vector update still writes the noun record', async () => {
|
||||
const id = await brain.add({
|
||||
data: 'vector swap subject',
|
||||
type: NounType.Concept,
|
||||
metadata: {}
|
||||
})
|
||||
|
||||
const storage = (brain as any).storage
|
||||
const saveNounSpy = vi.spyOn(storage, 'saveNoun')
|
||||
|
||||
const newVector = new Array(384).fill(0).map((_, i) => Math.cos(i))
|
||||
await brain.update({ id, vector: newVector })
|
||||
|
||||
expect(saveNounSpy).toHaveBeenCalled()
|
||||
saveNounSpy.mockRestore()
|
||||
|
||||
const after = await brain.get(id, { includeVectors: true })
|
||||
expect(after?.vector?.[0]).toBeCloseTo(1) // cos(0)
|
||||
})
|
||||
})
|
||||
|
|
@ -244,10 +244,7 @@ describe('Metadata index cleanup after remove / removeMany', () => {
|
|||
const noConfidenceId = await addEntity({ type: 'thing' })
|
||||
const withConfidenceId = await addEntity({ type: 'thing', confidence: 0.9 })
|
||||
|
||||
// system.confidence — confidence is an engine scalar (an add() param),
|
||||
// never a metadata field; bare 'confidence' now addresses the user's
|
||||
// own metadata bag under the sealed field-addressing law.
|
||||
const results = await brain.find({ where: { 'system.confidence': { exists: true } } })
|
||||
const results = await brain.find({ where: { confidence: { exists: true } } })
|
||||
const ids = results.map(r => r.id)
|
||||
|
||||
expect(ids).toContain(withConfidenceId)
|
||||
|
|
@ -258,8 +255,7 @@ describe('Metadata index cleanup after remove / removeMany', () => {
|
|||
const noWeightId = await addEntity({ type: 'thing' })
|
||||
const withWeightId = await addEntity({ type: 'thing', weight: 0.5 })
|
||||
|
||||
// system.weight — same reasoning as system.confidence above.
|
||||
const results = await brain.find({ where: { 'system.weight': { exists: true } } })
|
||||
const results = await brain.find({ where: { weight: { exists: true } } })
|
||||
const ids = results.map(r => r.id)
|
||||
|
||||
expect(ids).toContain(withWeightId)
|
||||
|
|
@ -273,12 +269,11 @@ describe('Metadata index cleanup after remove / removeMany', () => {
|
|||
const id = await addEntity({ type: 'thing' })
|
||||
await brain.remove(id)
|
||||
|
||||
// Entity must not appear in any confidence query. system.confidence —
|
||||
// same addressing as the two tests above.
|
||||
const existsTrue = await brain.find({ where: { 'system.confidence': { exists: true } } })
|
||||
// Entity must not appear in any confidence query
|
||||
const existsTrue = await brain.find({ where: { confidence: { exists: true } } })
|
||||
expect(existsTrue.map(r => r.id)).not.toContain(id)
|
||||
|
||||
const existsFalse = await brain.find({ where: { 'system.confidence': { exists: false } } })
|
||||
const existsFalse = await brain.find({ where: { confidence: { exists: false } } })
|
||||
expect(existsFalse.map(r => r.id)).not.toContain(id)
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -42,8 +42,7 @@ describe('find({ where, orderBy }) bounds the sort to the page (CTX-BR-FIND-ORDE
|
|||
return real(f, ob, o, topK)
|
||||
}
|
||||
|
||||
// system.createdAt — entity age, not a user metadata field named 'createdAt'.
|
||||
const results = await brain.find({ where: { bucket: 'x' }, orderBy: 'system.createdAt', order: 'desc', limit: 5 })
|
||||
const results = await brain.find({ where: { bucket: 'x' }, orderBy: 'createdAt', order: 'desc', limit: 5 })
|
||||
|
||||
expect(results).toHaveLength(5)
|
||||
// Page-bounded: ~ limit (5) + a small hidden-tier over-fetch — NOT all 50 matches.
|
||||
|
|
|
|||
|
|
@ -245,10 +245,7 @@ describe('rc.8 no-freeze migration deference (isMigrating / stampBrainFormat / b
|
|||
it('the brain-format marker module exports the compiled epoch + data-format constants', () => {
|
||||
// cor imports these from '@soulcraft/brainy/brain-format' (Hook 3) so both
|
||||
// sides share ONE source of truth — no duplicated constant to drift.
|
||||
// Epoch 3: the namespace-law key split (bare user keys · literal
|
||||
// 'system.<field>' scalars, 2026-08-03) — every brain rebuilds onto the
|
||||
// frozen keys at first open. (Epoch 2 same day: `level` indexability.)
|
||||
expect(EXPECTED_INDEX_EPOCH).toBe(3)
|
||||
expect(EXPECTED_INDEX_EPOCH).toBe(1)
|
||||
expect(CURRENT_DATA_FORMAT).toBe('8.0')
|
||||
})
|
||||
})
|
||||
|
|
|
|||
251
tests/unit/brainy/reserved-field-policy.test.ts
Normal file
251
tests/unit/brainy/reserved-field-policy.test.ts
Normal file
|
|
@ -0,0 +1,251 @@
|
|||
/**
|
||||
* @module tests/unit/brainy/reserved-field-policy
|
||||
* @description The 8.0 `reservedFieldPolicy` matrix — what happens when an
|
||||
* untyped (JavaScript) caller smuggles a Brainy-reserved field INSIDE the
|
||||
* `metadata` bag of a write call, past the compile-time guard.
|
||||
*
|
||||
* 8.0 is a clean break with no silent failures. The decided contract:
|
||||
* - `'throw'` (DEFAULT): a reserved key in the bag throws a clear Error naming
|
||||
* the offending key(s) and the correct write path. No remap, no data loss.
|
||||
* - `'warn'`: legacy remap PLUS a one-shot (per method+field, per process)
|
||||
* warning for EVERY reserved key found.
|
||||
* - `'remap'`: the pre-8.0 silent remap, no warning.
|
||||
*
|
||||
* The deep correctness of the remap itself (top-level precedence, system-managed
|
||||
* drops, transact()/with() mirrors, read-side splitting) lives in
|
||||
* tests/unit/brainy/update-reserved-metadata-remap.test.ts (which now runs under
|
||||
* `reservedFieldPolicy: 'remap'`). This file pins the POLICY SELECTION and the
|
||||
* throw/warn behaviors.
|
||||
*
|
||||
* Compile-time callers can't write these shapes at all (see
|
||||
* tests/unit/types/reserved-metadata-keys.test-d.ts); the `as object` widenings
|
||||
* below simulate untyped callers.
|
||||
*/
|
||||
|
||||
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
|
||||
import { Brainy } from '../../../src/index.js'
|
||||
import { NounType, VerbType } from '../../../src/types/graphTypes.js'
|
||||
import { createTestConfig } from '../../helpers/test-factory.js'
|
||||
import { prodLog } from '../../../src/utils/logger.js'
|
||||
|
||||
describe('reservedFieldPolicy', () => {
|
||||
describe("default policy is 'throw'", () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
// No reservedFieldPolicy override → resolves to 'throw'.
|
||||
brain = new Brainy(createTestConfig())
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
it('add() throws naming the offending key and the correct write path', async () => {
|
||||
await expect(
|
||||
brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'x',
|
||||
metadata: { confidence: 0.8 } as object
|
||||
})
|
||||
).rejects.toThrow(/metadata\.confidence is a reserved field/)
|
||||
|
||||
// The error names the right param and the reserved list for discoverability.
|
||||
await expect(
|
||||
brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'x',
|
||||
metadata: { confidence: 0.8 } as object
|
||||
})
|
||||
).rejects.toThrow(/'confidence' param.*RESERVED_ENTITY_FIELDS/s)
|
||||
})
|
||||
|
||||
it('add() lists EVERY offending key when several are present', async () => {
|
||||
const err = await brain
|
||||
.add({
|
||||
type: NounType.Person,
|
||||
data: 'multi',
|
||||
metadata: { confidence: 0.5, weight: 0.6, subtype: 'employee' } as object
|
||||
})
|
||||
.catch((e) => e as Error)
|
||||
expect(err).toBeInstanceOf(Error)
|
||||
expect(err.message).toMatch(/confidence/)
|
||||
expect(err.message).toMatch(/weight/)
|
||||
expect(err.message).toMatch(/subtype/)
|
||||
})
|
||||
|
||||
it('update() throws on a reserved key in the patch', async () => {
|
||||
const id = await brain.add({ type: NounType.Concept, subtype: 'general', data: 'y' })
|
||||
await expect(
|
||||
brain.update({ id, metadata: { confidence: 0.3 } as object })
|
||||
).rejects.toThrow(/metadata\.confidence is a reserved field/)
|
||||
})
|
||||
|
||||
it('relate() throws on a reserved key in the bag', async () => {
|
||||
const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' })
|
||||
const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' })
|
||||
await expect(
|
||||
brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.RelatedTo,
|
||||
subtype: 'colleague',
|
||||
metadata: { confidence: 0.4 } as object
|
||||
})
|
||||
).rejects.toThrow(/metadata\.confidence is a reserved field.*RESERVED_RELATION_FIELDS/s)
|
||||
})
|
||||
|
||||
it('updateRelation() throws on a reserved key in the patch', async () => {
|
||||
const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' })
|
||||
const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' })
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct'
|
||||
})
|
||||
await expect(
|
||||
brain.updateRelation({ id: relId, metadata: { weight: 0.2 } as object })
|
||||
).rejects.toThrow(/metadata\.weight is a reserved field/)
|
||||
})
|
||||
|
||||
it('transact() add op throws on a reserved key in the bag', async () => {
|
||||
await expect(
|
||||
brain.transact([
|
||||
{
|
||||
op: 'add',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'tx',
|
||||
metadata: { confidence: 0.7 } as object
|
||||
}
|
||||
])
|
||||
).rejects.toThrow(/metadata\.confidence is a reserved field/)
|
||||
})
|
||||
|
||||
it('a custom (non-reserved) key in the bag does NOT throw', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'ok',
|
||||
metadata: { status: 'draft', rating: 4 }
|
||||
})
|
||||
const entity = await brain.get(id)
|
||||
expect(entity?.metadata).toEqual({ status: 'draft', rating: 4 })
|
||||
})
|
||||
})
|
||||
|
||||
describe("'remap' policy remaps silently (no warning)", () => {
|
||||
let brain: Brainy
|
||||
let warnSpy: ReturnType<typeof vi.spyOn>
|
||||
|
||||
beforeEach(async () => {
|
||||
warnSpy = vi.spyOn(prodLog, 'warn').mockImplementation(() => {})
|
||||
brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' }))
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
warnSpy.mockRestore()
|
||||
})
|
||||
|
||||
it('lifts user-mutable reserved fields to top-level without warning', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Person,
|
||||
data: 'remap lift',
|
||||
metadata: { confidence: 0.8, weight: 0.6, subtype: 'employee', dept: 'eng' } as object
|
||||
})
|
||||
const entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.8)
|
||||
expect(entity?.weight).toBe(0.6)
|
||||
expect(entity?.subtype).toBe('employee')
|
||||
expect(entity?.metadata).toEqual({ dept: 'eng' })
|
||||
// 'remap' is silent about reserved fields (unrelated storage logs may fire,
|
||||
// so assert specifically that no reserved-field warning was emitted).
|
||||
const reservedWarned = warnSpy.mock.calls.some((c) =>
|
||||
String(c[0]).includes('reserved field')
|
||||
)
|
||||
expect(reservedWarned).toBe(false)
|
||||
})
|
||||
|
||||
it('preserves _originalId on natural-key ids through the remap path', async () => {
|
||||
// A speculative view applies the same normalization and maps a natural-key
|
||||
// id to a stable UUID, preserving the caller's original string.
|
||||
const base = await brain.now()
|
||||
const speculative = await base.with([
|
||||
{
|
||||
op: 'add',
|
||||
id: 'remap-spec-entity',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'spec',
|
||||
metadata: { confidence: 0.65, custom: 'spec' } as object
|
||||
}
|
||||
])
|
||||
const entity = await speculative.get('remap-spec-entity')
|
||||
expect(entity?.confidence).toBe(0.65)
|
||||
expect(entity?.metadata).toEqual({ custom: 'spec', _originalId: 'remap-spec-entity' })
|
||||
await speculative.release()
|
||||
await base.release()
|
||||
})
|
||||
})
|
||||
|
||||
describe("'warn' policy remaps AND warns once per key", () => {
|
||||
let brain: Brainy
|
||||
let warnSpy: ReturnType<typeof vi.spyOn>
|
||||
|
||||
beforeEach(async () => {
|
||||
warnSpy = vi.spyOn(prodLog, 'warn').mockImplementation(() => {})
|
||||
brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'warn' }))
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
warnSpy.mockRestore()
|
||||
})
|
||||
|
||||
it('remaps the value (same as remap) and emits a warning naming the field', async () => {
|
||||
// Use a method+field combo unique to this test so the per-process one-shot
|
||||
// registry has not already consumed it.
|
||||
const id = await brain.add({
|
||||
type: NounType.Person,
|
||||
data: 'warn lift',
|
||||
// weight is user-mutable → remapped; this is the only 'warn'-policy
|
||||
// add({ weight }) in the suite, so the one-shot warning fires here.
|
||||
metadata: { weight: 0.42, dept: 'eng' } as object
|
||||
})
|
||||
const entity = await brain.get(id)
|
||||
// Value is honored (remap still happens under 'warn').
|
||||
expect(entity?.weight).toBe(0.42)
|
||||
expect(entity?.metadata).toEqual({ dept: 'eng' })
|
||||
// And a warning was emitted naming the reserved field.
|
||||
expect(warnSpy).toHaveBeenCalled()
|
||||
const warned = warnSpy.mock.calls.some((c) =>
|
||||
String(c[0]).includes("'weight'")
|
||||
)
|
||||
expect(warned).toBe(true)
|
||||
})
|
||||
|
||||
it('warns for system-managed keys too (closes the historical gap)', async () => {
|
||||
// Pre-8.0 only system-managed fields warned; 'warn' warns for every key.
|
||||
// 'createdBy' (system-managed on update) is unique to this test.
|
||||
const id = await brain.add({ type: NounType.Concept, subtype: 'general', data: 'sys' })
|
||||
warnSpy.mockClear()
|
||||
await brain.update({ id, metadata: { createdBy: 'nope', keep: 'me' } as object })
|
||||
const entity = await brain.get(id)
|
||||
// System-managed key dropped; custom field merged.
|
||||
expect((entity?.metadata as Record<string, unknown>)?.createdBy).toBeUndefined()
|
||||
expect((entity?.metadata as Record<string, unknown>)?.keep).toBe('me')
|
||||
// A warning was emitted for the dropped system-managed key.
|
||||
const warned = warnSpy.mock.calls.some((c) =>
|
||||
String(c[0]).includes("'createdBy'")
|
||||
)
|
||||
expect(warned).toBe(true)
|
||||
})
|
||||
})
|
||||
})
|
||||
403
tests/unit/brainy/update-reserved-metadata-remap.test.ts
Normal file
403
tests/unit/brainy/update-reserved-metadata-remap.test.ts
Normal file
|
|
@ -0,0 +1,403 @@
|
|||
/**
|
||||
* @module tests/unit/brainy/update-reserved-metadata-remap
|
||||
* @description Regression tests for the reserved-field metadata-bag trap,
|
||||
* ported from the 7.x fix and extended to the full 8.0 contract.
|
||||
*
|
||||
* History: `add({metadata: {confidence}})` lifted reserved fields to their
|
||||
* canonical top-level location, but `update({metadata: {confidence}})`
|
||||
* silently dropped the same shape — the patch value survived the merge and
|
||||
* was then clobbered by the preserve-existing spread. A production
|
||||
* consumer's confidence-evolution writes no-oped for weeks before being
|
||||
* caught by reading values back.
|
||||
*
|
||||
* These tests pin the LEGACY REMAP behavior, which in 8.0 is opt-in via
|
||||
* `reservedFieldPolicy: 'remap'` (the default is `'throw'` — see the policy
|
||||
* matrix in tests/unit/brainy/reserved-field-policy.test.ts). The brain in
|
||||
* every test below is constructed with `reservedFieldPolicy: 'remap'` so these
|
||||
* deep correctness assertions about the remap path stay exercised.
|
||||
*
|
||||
* Remap contract under test (every write path, entities AND relationships):
|
||||
* - user-mutable reserved fields (`confidence`, `weight`, `subtype` — plus
|
||||
* `service`/`createdBy` at add()/relate() time) remap from the metadata
|
||||
* bag to their dedicated top-level param, with top-level winning when both
|
||||
* are present;
|
||||
* - system-managed reserved fields (`createdAt`, `_rev`, `noun`/`verb`,
|
||||
* `data`, …) are dropped from the bag;
|
||||
* - the same normalization applies to `transact()` operations and `with()`
|
||||
* speculative views;
|
||||
* - reads NEVER echo a reserved field inside `metadata`.
|
||||
*
|
||||
* TypeScript callers can't write these shapes at all (compile-time guard on
|
||||
* the metadata param types — see tests/unit/types/reserved-metadata-keys.test-d.ts);
|
||||
* these tests simulate untyped (JavaScript) callers, hence the `as object`
|
||||
* widenings on the metadata literals.
|
||||
*/
|
||||
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
||||
import { Brainy } from '../../../src/index.js'
|
||||
import { NounType, VerbType } from '../../../src/types/graphTypes.js'
|
||||
import { createTestConfig } from '../../helpers/test-factory.js'
|
||||
|
||||
describe('reserved-field metadata remap (8.0 legacy remap path)', () => {
|
||||
let brain: Brainy
|
||||
|
||||
beforeEach(async () => {
|
||||
// The remap path is opt-in in 8.0 (default policy is 'throw').
|
||||
brain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' }))
|
||||
await brain.init()
|
||||
})
|
||||
|
||||
afterEach(async () => {
|
||||
await brain.close()
|
||||
})
|
||||
|
||||
describe('update() — the ported 7.x regression', () => {
|
||||
it('remaps metadata.confidence to the top-level field (the production repro)', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'x',
|
||||
metadata: { confidence: 0.8 } as object
|
||||
})
|
||||
|
||||
// Top-level write works (always did)
|
||||
await brain.update({ id, confidence: 0.42 })
|
||||
let entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.42)
|
||||
|
||||
// Metadata-patch write — silently dropped pre-fix, remapped now
|
||||
await brain.update({ id, metadata: { confidence: 0.33 } as object })
|
||||
entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.33)
|
||||
// The reserved key must not linger inside the metadata bag
|
||||
expect((entity?.metadata as Record<string, unknown>)?.confidence).toBeUndefined()
|
||||
})
|
||||
|
||||
it('remaps metadata.weight and metadata.subtype the same way', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'y',
|
||||
metadata: {}
|
||||
})
|
||||
|
||||
await brain.update({ id, metadata: { weight: 0.7, subtype: 'specialized' } as object })
|
||||
const entity = await brain.get(id)
|
||||
expect(entity?.weight).toBe(0.7)
|
||||
expect(entity?.subtype).toBe('specialized')
|
||||
expect((entity?.metadata as Record<string, unknown>)?.weight).toBeUndefined()
|
||||
expect((entity?.metadata as Record<string, unknown>)?.subtype).toBeUndefined()
|
||||
})
|
||||
|
||||
it('top-level param wins when both top-level and metadata-patch carry the field', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'z',
|
||||
metadata: { confidence: 0.5 } as object
|
||||
})
|
||||
|
||||
await brain.update({ id, confidence: 0.9, metadata: { confidence: 0.1 } as object })
|
||||
const entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.9)
|
||||
})
|
||||
|
||||
it('drops system-managed fields from patches without corrupting the entity', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'w',
|
||||
metadata: { keep: 'me' }
|
||||
})
|
||||
const before = await brain.get(id)
|
||||
|
||||
await brain.update({
|
||||
id,
|
||||
metadata: { createdAt: 1, _rev: 999, noun: 'organization', other: 'applied' } as object
|
||||
})
|
||||
const after = await brain.get(id)
|
||||
|
||||
expect(after?.createdAt).toBe(before?.createdAt) // immutable
|
||||
expect(after?.type).toBe('concept') // noun patch ignored
|
||||
expect(after?._rev).toBe((before?._rev ?? 1) + 1) // _rev patch ignored; normal bump applied
|
||||
expect((after?.metadata as Record<string, unknown>)?.other).toBe('applied') // custom fields still merge
|
||||
expect((after?.metadata as Record<string, unknown>)?.keep).toBe('me')
|
||||
expect((after?.metadata as Record<string, unknown>)?._rev).toBeUndefined()
|
||||
expect((after?.metadata as Record<string, unknown>)?.createdAt).toBeUndefined()
|
||||
expect((after?.metadata as Record<string, unknown>)?.noun).toBeUndefined()
|
||||
})
|
||||
|
||||
it('custom (non-reserved) metadata patches are unaffected by the remap', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'v',
|
||||
metadata: { status: 'draft' }
|
||||
})
|
||||
|
||||
await brain.update({ id, metadata: { status: 'reviewed', rating: 4.5 } })
|
||||
const entity = await brain.get(id)
|
||||
expect((entity?.metadata as Record<string, unknown>)?.status).toBe('reviewed')
|
||||
expect((entity?.metadata as Record<string, unknown>)?.rating).toBe(4.5)
|
||||
})
|
||||
})
|
||||
|
||||
describe('add() — explicit lift, identical contract', () => {
|
||||
it('lifts confidence/weight/subtype out of the bag to top level', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Person,
|
||||
data: 'lift check',
|
||||
metadata: { confidence: 0.8, weight: 0.6, subtype: 'employee', dept: 'eng' } as object
|
||||
})
|
||||
|
||||
const entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.8)
|
||||
expect(entity?.weight).toBe(0.6)
|
||||
expect(entity?.subtype).toBe('employee')
|
||||
expect(entity?.metadata).toEqual({ dept: 'eng' })
|
||||
})
|
||||
|
||||
it('lifts service (settable at add time) and lets the top-level param win', async () => {
|
||||
const lifted = await brain.add({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'service lift',
|
||||
metadata: { service: 'orders' } as object
|
||||
})
|
||||
expect((await brain.get(lifted))?.service).toBe('orders')
|
||||
|
||||
const topLevelWins = await brain.add({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'service precedence',
|
||||
service: 'billing',
|
||||
metadata: { service: 'orders' } as object
|
||||
})
|
||||
const entity = await brain.get(topLevelWins)
|
||||
expect(entity?.service).toBe('billing')
|
||||
expect((entity?.metadata as Record<string, unknown>)?.service).toBeUndefined()
|
||||
})
|
||||
|
||||
it('a remapped subtype satisfies subtype enforcement like a top-level one', async () => {
|
||||
brain.requireSubtype(NounType.Document)
|
||||
|
||||
// Top-level missing, but the bag carries it — must not throw.
|
||||
const id = await brain.add({
|
||||
type: NounType.Document,
|
||||
data: 'enforcement via remap',
|
||||
metadata: { subtype: 'invoice' } as object
|
||||
})
|
||||
expect((await brain.get(id))?.subtype).toBe('invoice')
|
||||
|
||||
// Neither place carries it — must throw.
|
||||
await expect(
|
||||
brain.add({ type: NounType.Document, data: 'no subtype anywhere' })
|
||||
).rejects.toThrow(/subtype/)
|
||||
})
|
||||
})
|
||||
|
||||
describe('transact() — same remap on add and update ops', () => {
|
||||
it('normalizes reserved fields in transact add + update ops', async () => {
|
||||
const db1 = await brain.transact([
|
||||
{
|
||||
op: 'add',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'tx',
|
||||
metadata: { confidence: 0.7, custom: 'a' } as object
|
||||
}
|
||||
])
|
||||
const id = db1.receipt!.ids[0]
|
||||
|
||||
let entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.7)
|
||||
expect(entity?.metadata).toEqual({ custom: 'a' })
|
||||
|
||||
await brain.transact([
|
||||
{ op: 'update', id, metadata: { confidence: 0.25, custom: 'b' } as object }
|
||||
])
|
||||
entity = await brain.get(id)
|
||||
expect(entity?.confidence).toBe(0.25)
|
||||
expect(entity?.metadata).toEqual({ custom: 'b' })
|
||||
expect((entity?.metadata as Record<string, unknown>)?.confidence).toBeUndefined()
|
||||
})
|
||||
|
||||
it('historical asOf() reads surface reserved fields ONLY top-level', async () => {
|
||||
const db1 = await brain.transact([
|
||||
{
|
||||
op: 'add',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'historical',
|
||||
metadata: { confidence: 0.9, custom: 'past' } as object
|
||||
}
|
||||
])
|
||||
const id = db1.receipt!.ids[0]
|
||||
|
||||
// Move the world forward so generation db1 is historical.
|
||||
await brain.transact([{ op: 'update', id, confidence: 0.1, metadata: { custom: 'now' } }])
|
||||
|
||||
const past = await brain.asOf(db1.generation)
|
||||
const historical = await past.get(id)
|
||||
expect(historical?.confidence).toBe(0.9)
|
||||
expect(historical?.metadata).toEqual({ custom: 'past' })
|
||||
await past.release()
|
||||
})
|
||||
|
||||
it('with() speculative views apply the same normalization', async () => {
|
||||
const base = await brain.now()
|
||||
const speculative = await base.with([
|
||||
{
|
||||
op: 'add',
|
||||
id: 'spec-entity',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'spec',
|
||||
metadata: { confidence: 0.65, custom: 'spec' } as object
|
||||
}
|
||||
])
|
||||
|
||||
const entity = await speculative.get('spec-entity')
|
||||
expect(entity?.confidence).toBe(0.65)
|
||||
// 8.0 id normalization: a natural-key id is mapped to a stable UUID and
|
||||
// the caller's original string is preserved under _originalId — surfaced
|
||||
// here exactly as the durable transact()/add() paths do.
|
||||
expect(entity?.metadata).toEqual({ custom: 'spec', _originalId: 'spec-entity' })
|
||||
await speculative.release()
|
||||
await base.release()
|
||||
})
|
||||
})
|
||||
|
||||
describe('read paths never echo reserved fields inside metadata', () => {
|
||||
it('find() (storage pagination path) returns custom-only metadata with reserved fields top-level', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'pagination echo check',
|
||||
confidence: 0.8,
|
||||
weight: 0.6,
|
||||
metadata: { dept: 'eng' }
|
||||
})
|
||||
|
||||
// No query/filter → served by the direct storage pagination path
|
||||
// (getNounsWithPagination), which historically echoed the full flat
|
||||
// record (noun/subtype/createdAt/… inside metadata).
|
||||
const results = await brain.find({ limit: 50 })
|
||||
const result = results.find((r) => r.id === id)
|
||||
expect(result).toBeDefined()
|
||||
expect(result?.entity.metadata).toEqual({ dept: 'eng' })
|
||||
expect(result?.entity.type).toBe(NounType.Person)
|
||||
expect(result?.entity.subtype).toBe('employee')
|
||||
expect(result?.entity.confidence).toBe(0.8)
|
||||
expect(result?.entity.weight).toBe(0.6)
|
||||
expect(typeof result?.entity.createdAt).toBe('number')
|
||||
expect(result?.entity._rev).toBe(1)
|
||||
})
|
||||
|
||||
it('related() by target surfaces reserved fields top-level, custom-only metadata', async () => {
|
||||
const a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'src' })
|
||||
const b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'tgt' })
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
confidence: 0.9,
|
||||
weight: 0.5,
|
||||
service: 'orders',
|
||||
metadata: { note: 'target path' }
|
||||
})
|
||||
|
||||
const relations = await brain.related({ to: b })
|
||||
const rel = relations.find((r) => r.id === relId)
|
||||
expect(rel).toBeDefined()
|
||||
expect(rel?.metadata).toEqual({ note: 'target path' })
|
||||
expect(rel?.subtype).toBe('direct')
|
||||
expect(rel?.confidence).toBe(0.9)
|
||||
expect(rel?.weight).toBe(0.5)
|
||||
expect(rel?.service).toBe('orders')
|
||||
expect(typeof rel?.createdAt).toBe('number')
|
||||
})
|
||||
})
|
||||
|
||||
describe('relationships — relate() / updateRelation() mirror', () => {
|
||||
let a: string
|
||||
let b: string
|
||||
|
||||
beforeEach(async () => {
|
||||
a = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'A' })
|
||||
b = await brain.add({ type: NounType.Person, subtype: 'employee', data: 'B' })
|
||||
})
|
||||
|
||||
it('relate() persists the top-level confidence and service params', async () => {
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
confidence: 0.77,
|
||||
service: 'orders'
|
||||
})
|
||||
|
||||
const relations = await brain.related({ from: a })
|
||||
const rel = relations.find((r) => r.id === relId)
|
||||
expect(rel?.confidence).toBe(0.77)
|
||||
expect(rel?.service).toBe('orders')
|
||||
})
|
||||
|
||||
it('relate() remaps reserved fields out of the metadata bag', async () => {
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.RelatedTo,
|
||||
subtype: 'colleague',
|
||||
metadata: { confidence: 0.4, weight: 0.3, role: 'peer' } as object
|
||||
})
|
||||
|
||||
const relations = await brain.related({ from: a })
|
||||
const rel = relations.find((r) => r.id === relId)
|
||||
expect(rel?.confidence).toBe(0.4)
|
||||
expect(rel?.weight).toBe(0.3)
|
||||
expect(rel?.metadata).toEqual({ role: 'peer' })
|
||||
})
|
||||
|
||||
it('relation.metadata never echoes the verb type key', async () => {
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.RelatedTo,
|
||||
subtype: 'colleague',
|
||||
metadata: { note: 'no echo' }
|
||||
})
|
||||
|
||||
const relations = await brain.related({ from: a })
|
||||
const rel = relations.find((r) => r.id === relId)
|
||||
expect(rel?.type).toBe(VerbType.RelatedTo)
|
||||
expect((rel?.metadata as Record<string, unknown>)?.verb).toBeUndefined()
|
||||
expect(rel?.metadata).toEqual({ note: 'no echo' })
|
||||
})
|
||||
|
||||
it('updateRelation() remaps the user-mutable trio and preserves service', async () => {
|
||||
const relId = await brain.relate({
|
||||
from: a,
|
||||
to: b,
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
service: 'orders',
|
||||
metadata: { keep: 'me' }
|
||||
})
|
||||
|
||||
await brain.updateRelation({
|
||||
id: relId,
|
||||
metadata: { confidence: 0.55, subtype: 'dotted-line', extra: 'applied' } as object
|
||||
})
|
||||
|
||||
const relations = await brain.related({ from: a })
|
||||
const rel = relations.find((r) => r.id === relId)
|
||||
expect(rel?.confidence).toBe(0.55)
|
||||
expect(rel?.subtype).toBe('dotted-line')
|
||||
expect(rel?.service).toBe('orders') // fixed at relate() time, never erased by updates
|
||||
expect(rel?.metadata).toEqual({ keep: 'me', extra: 'applied' })
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
@ -198,47 +198,60 @@ describe('visibility (8.0 reserved field)', () => {
|
|||
expect(entity?.visibility).toBeUndefined()
|
||||
})
|
||||
|
||||
it('metadata.visibility is the USER’s field (field-addressing law) — stored verbatim, never lifted to the engine tier', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
data: 'y',
|
||||
metadata: { visibility: 'internal', tag: 't' } as object
|
||||
})
|
||||
const entity = await brain.get(id)
|
||||
// The user's field lives in the bag, verbatim…
|
||||
expect((entity?.metadata as Record<string, unknown>)?.visibility).toBe('internal')
|
||||
expect((entity?.metadata as Record<string, unknown>)?.tag).toBe('t')
|
||||
// …and the ENGINE tier is untouched: absent === public, so the entity
|
||||
// stays visible on default reads (the engine tier is set only via the
|
||||
// dedicated visibility param and reads at system.visibility).
|
||||
expect(entity?.visibility).toBeUndefined()
|
||||
const visible = await brain.find({ type: NounType.Concept, limit: 20 })
|
||||
expect(visible.map((r) => r.id)).toContain(id)
|
||||
it('an untyped caller passing visibility inside metadata is normalized under reservedFieldPolicy:"remap" (lifted to top-level)', async () => {
|
||||
// Simulate a JavaScript caller smuggling the reserved key past the compile-time guard.
|
||||
// The legacy remap behavior is now opt-in (8.0 default is 'throw').
|
||||
const remapBrain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' }))
|
||||
await remapBrain.init()
|
||||
try {
|
||||
const id = await remapBrain.add({
|
||||
type: NounType.Concept,
|
||||
data: 'y',
|
||||
metadata: { visibility: 'internal', tag: 't' } as object
|
||||
})
|
||||
const entity = await remapBrain.get(id)
|
||||
// Lifted to the top-level field…
|
||||
expect(entity?.visibility).toBe('internal')
|
||||
// …and stripped from the metadata bag.
|
||||
expect((entity?.metadata as Record<string, unknown>)?.visibility).toBeUndefined()
|
||||
expect((entity?.metadata as Record<string, unknown>)?.tag).toBe('t')
|
||||
// It is excluded from the default count, exactly like a top-level internal write.
|
||||
expect(await remapBrain.getNounCount()).toBe(0)
|
||||
} finally {
|
||||
await remapBrain.close()
|
||||
}
|
||||
})
|
||||
|
||||
it('a user field valued "system" cannot smuggle the Brainy-only tier — it is just user data', async () => {
|
||||
const id = await brain.add({
|
||||
type: NounType.Concept,
|
||||
data: 'z',
|
||||
metadata: { visibility: 'system' } as object
|
||||
})
|
||||
const entity = await brain.get(id)
|
||||
// Engine tier unaffected → entity stays public (counted, visible);
|
||||
// the string 'system' is ordinary user data in the bag.
|
||||
expect(entity?.visibility).toBeUndefined()
|
||||
expect((entity?.metadata as Record<string, unknown>)?.visibility).toBe('system')
|
||||
const found = await brain.find({ type: NounType.Concept, limit: 10 })
|
||||
expect(found.map((r) => r.id)).toContain(id)
|
||||
it('a "system" value smuggled through metadata is dropped under reservedFieldPolicy:"remap", not honored', async () => {
|
||||
// 'system' is Brainy-only; an untyped caller must not be able to set it.
|
||||
const remapBrain = new Brainy(createTestConfig({ reservedFieldPolicy: 'remap' }))
|
||||
await remapBrain.init()
|
||||
try {
|
||||
const id = await remapBrain.add({
|
||||
type: NounType.Concept,
|
||||
data: 'z',
|
||||
metadata: { visibility: 'system' } as object
|
||||
})
|
||||
const entity = await remapBrain.get(id)
|
||||
// The smuggled 'system' was dropped → entity stays public (counted, visible).
|
||||
expect(entity?.visibility).toBeUndefined()
|
||||
expect(await remapBrain.getNounCount()).toBe(1)
|
||||
const found = await remapBrain.find({ type: NounType.Concept, limit: 10 })
|
||||
expect(found.map((r) => r.id)).toContain(id)
|
||||
} finally {
|
||||
await remapBrain.close()
|
||||
}
|
||||
})
|
||||
|
||||
it('a forged system.visibility key in metadata refuses loudly at the write door', async () => {
|
||||
it('an untyped caller passing visibility inside metadata throws under the default policy', async () => {
|
||||
// 8.0 default: no silent remap — a reserved key in the bag is a loud error.
|
||||
await expect(
|
||||
brain.add({
|
||||
type: NounType.Concept,
|
||||
data: 'throws',
|
||||
metadata: { 'system.visibility': 'internal' } as object
|
||||
metadata: { visibility: 'internal', tag: 't' } as object
|
||||
})
|
||||
).rejects.toThrow(/system\./)
|
||||
).rejects.toThrow(/visibility.*reserved field/)
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -186,52 +186,4 @@ describe('fact log — round-trip, framing, reconcile, rotation, scan', () => {
|
|||
await log.sync()
|
||||
expect(log.segmentPaths()).toEqual([]) // only a tail exists — nothing sealed
|
||||
})
|
||||
|
||||
describe('scanFacts liveness contract (Stage-2 D1)', () => {
|
||||
it('a wedged store fails LOUDLY within the first-batch bound — never a silent hang', async () => {
|
||||
// Force a sealed segment (tiny rotateBytes) so the scan must READ from
|
||||
// storage, then wedge that read: the exact production shape (a
|
||||
// backlogged brain whose segment read never returned).
|
||||
const mem: any = new MemoryStorage()
|
||||
await mem.init()
|
||||
const wedgeable = new FactLog(mem, { rotateBytes: 1 })
|
||||
await wedgeable.open(0)
|
||||
await wedgeable.append(fact(1))
|
||||
await wedgeable.append(fact(2)) // second append rotates → seg 1 sealed
|
||||
await wedgeable.sync()
|
||||
|
||||
const realRead = mem.readRawBytes.bind(mem)
|
||||
mem.readRawBytes = (p: string) =>
|
||||
p.includes('facts/seg-') ? new Promise(() => {}) : realRead(p) // hangs forever
|
||||
|
||||
const scan = wedgeable.scanFacts({ firstBatchTimeoutMs: 200 })
|
||||
const started = Date.now()
|
||||
await expect(scan.batches().next()).rejects.toThrow(/no first batch within 200ms/)
|
||||
expect(Date.now() - started).toBeLessThan(5_000) // bound held, not a hang
|
||||
})
|
||||
|
||||
it('a healthy scan is unaffected — first batch well inside the bound, all facts delivered', async () => {
|
||||
for (let g = 1; g <= 5; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
const scan = log.scanFacts({ batchSize: 2 })
|
||||
const all: CommitFact[] = []
|
||||
for await (const b of scan.batches()) all.push(...b.facts)
|
||||
expect(all.map((f) => f.generation)).toEqual([1, 2, 3, 4, 5])
|
||||
expect(scan.summary().factsYielded).toBe(5)
|
||||
})
|
||||
|
||||
it('consumer think-time between pulls never counts against the producer', async () => {
|
||||
for (let g = 1; g <= 4; g++) await log.append(fact(g))
|
||||
await log.sync()
|
||||
// Bound tighter than the consumer's pause: only the FIRST pull is
|
||||
// raced, so a slow consumer after batch 1 must not trip the deadline.
|
||||
const gen = log.scanFacts({ batchSize: 2, firstBatchTimeoutMs: 150 }).batches()
|
||||
const first = await gen.next()
|
||||
expect(first.done).toBe(false)
|
||||
await new Promise((r) => setTimeout(r, 400)) // dawdle past the bound
|
||||
const second = await gen.next()
|
||||
expect(second.done).toBe(false)
|
||||
expect((await gen.next()).done).toBe(true)
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -1,145 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/db/fieldAddressing
|
||||
* @description Unit pins for the one field-addressing law (ruled 2026-08-03).
|
||||
* These pin the PURE half of the law — parsing, the ruled maps, plumbing
|
||||
* invisibility, refusal text — including the RELATION map, which cannot be
|
||||
* pinned through the public query API today (related() carries no
|
||||
* field-addressing options): the verb mirror is contract-tested here at the
|
||||
* module level so the two engines cannot drift on it.
|
||||
*/
|
||||
import { describe, it, expect } from 'vitest'
|
||||
import {
|
||||
SYSTEM_ENTITY_SCALARS,
|
||||
SYSTEM_RELATION_SCALARS,
|
||||
PLUMBING_FIELDS,
|
||||
parseFieldAddress,
|
||||
buildUnresolvableMessage,
|
||||
InvalidFieldAddressError
|
||||
} from '../../../src/db/fieldAddressing.js'
|
||||
|
||||
describe('field-addressing law — pure module pins', () => {
|
||||
it('the entity system map is EXACTLY the ruled ten scalars', () => {
|
||||
expect([...SYSTEM_ENTITY_SCALARS].sort()).toEqual(
|
||||
[
|
||||
'confidence',
|
||||
'createdAt',
|
||||
'createdBy',
|
||||
'id',
|
||||
'service',
|
||||
'subtype',
|
||||
'type',
|
||||
'updatedAt',
|
||||
'visibility',
|
||||
'weight'
|
||||
].sort()
|
||||
)
|
||||
})
|
||||
|
||||
it('the relation system map is the ruled verb mirror', () => {
|
||||
expect([...SYSTEM_RELATION_SCALARS].sort()).toEqual(
|
||||
[
|
||||
'verb',
|
||||
'sourceId',
|
||||
'targetId',
|
||||
'confidence',
|
||||
'createdAt',
|
||||
'createdBy',
|
||||
'service',
|
||||
'subtype',
|
||||
'updatedAt',
|
||||
'visibility',
|
||||
'weight'
|
||||
].sort()
|
||||
)
|
||||
})
|
||||
|
||||
it('plumbing is exactly the ruled five, and none of it leaks into a system map', () => {
|
||||
expect([...PLUMBING_FIELDS].sort()).toEqual(
|
||||
['_rev', 'connections', 'data', 'level', 'vector'].sort()
|
||||
)
|
||||
for (const field of PLUMBING_FIELDS) {
|
||||
expect(SYSTEM_ENTITY_SCALARS.has(field)).toBe(false)
|
||||
expect(SYSTEM_RELATION_SCALARS.has(field)).toBe(false)
|
||||
}
|
||||
})
|
||||
|
||||
it('bare names address user metadata — even when the name matches a system scalar', () => {
|
||||
expect(parseFieldAddress('level', 'entity')).toEqual({
|
||||
scope: 'metadata',
|
||||
field: 'level',
|
||||
raw: 'level'
|
||||
})
|
||||
expect(parseFieldAddress('confidence', 'entity').scope).toBe('metadata')
|
||||
expect(parseFieldAddress('createdAt', 'entity').scope).toBe('metadata')
|
||||
expect(parseFieldAddress('verb', 'relation').scope).toBe('metadata')
|
||||
})
|
||||
|
||||
it('metadata.-prefix is the explicit spelling of the bare form', () => {
|
||||
expect(parseFieldAddress('metadata.level', 'entity')).toEqual({
|
||||
scope: 'metadata',
|
||||
field: 'level',
|
||||
raw: 'metadata.level'
|
||||
})
|
||||
})
|
||||
|
||||
it('system.-prefix reaches exactly the map — entity and relation', () => {
|
||||
for (const field of SYSTEM_ENTITY_SCALARS) {
|
||||
expect(parseFieldAddress(`system.${field}`, 'entity')).toEqual({
|
||||
scope: 'system',
|
||||
field,
|
||||
raw: `system.${field}`
|
||||
})
|
||||
}
|
||||
for (const field of SYSTEM_RELATION_SCALARS) {
|
||||
expect(parseFieldAddress(`system.${field}`, 'relation').scope).toBe('system')
|
||||
}
|
||||
// The structural relation members are NOT entity scalars.
|
||||
expect(() => parseFieldAddress('system.verb', 'entity')).toThrow(InvalidFieldAddressError)
|
||||
expect(() => parseFieldAddress('system.sourceId', 'entity')).toThrow(InvalidFieldAddressError)
|
||||
})
|
||||
|
||||
it('plumbing refuses in the system spelling, on both record kinds', () => {
|
||||
for (const field of PLUMBING_FIELDS) {
|
||||
expect(() => parseFieldAddress(`system.${field}`, 'entity')).toThrow(
|
||||
InvalidFieldAddressError
|
||||
)
|
||||
expect(() => parseFieldAddress(`system.${field}`, 'relation')).toThrow(
|
||||
InvalidFieldAddressError
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
it('refusal text carries the whole valid map — the fix lives in the message', () => {
|
||||
try {
|
||||
parseFieldAddress('system.level', 'entity')
|
||||
expect.unreachable('should have thrown')
|
||||
} catch (e) {
|
||||
const msg = (e as Error).message
|
||||
for (const field of SYSTEM_ENTITY_SCALARS) {
|
||||
expect(msg).toContain(`system.${field}`)
|
||||
}
|
||||
expect(msg).toContain('plumbing')
|
||||
}
|
||||
})
|
||||
|
||||
it('malformed addresses refuse: empty name, bare metadata. prefix', () => {
|
||||
expect(() => parseFieldAddress('', 'entity')).toThrow(InvalidFieldAddressError)
|
||||
expect(() => parseFieldAddress('metadata.', 'entity')).toThrow(InvalidFieldAddressError)
|
||||
})
|
||||
|
||||
it('the did-you-mean names BOTH candidates for a system-colliding bare name', () => {
|
||||
const msg = buildUnresolvableMessage('createdAt', 'entity')
|
||||
expect(msg).toContain('system.createdAt')
|
||||
expect(msg).toContain('metadata.createdAt')
|
||||
})
|
||||
|
||||
it('a non-colliding unknown bare name names both spellings — system.<f> explicitly as NOT valid', () => {
|
||||
// Cross-engine pin (cor's suite greps for both spellings in every
|
||||
// refusal): the metadata candidate is the fix; the system spelling is
|
||||
// named but HONESTLY marked invalid, never offered as a candidate.
|
||||
const msg = buildUnresolvableMessage('scoore', 'entity')
|
||||
expect(msg).toContain('metadata.scoore')
|
||||
expect(msg).toContain('system.scoore')
|
||||
expect(msg).toContain('NOT valid')
|
||||
})
|
||||
})
|
||||
|
|
@ -1,150 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/db/generation-segments
|
||||
* @description The generation-segment store (Stage-2 D1+D3 file format).
|
||||
* Laws: (1) fold → read round-trips deltas and records byte-faithfully via
|
||||
* sidecar point-reads; (2) the manifest is the ONLY discovery path — reopen
|
||||
* reads one file, never a listing; (3) a lost/corrupt sidecar rebuilds from
|
||||
* its segment loudly, a damaged SEGMENT fails loudly (never silent wrong
|
||||
* data); (4) D3 reclaim drops whole segments only and bumps compactedBelow;
|
||||
* (5) the packed digest is deterministic across reopen; (6) immutability —
|
||||
* fold refuses overlap with sealed ranges.
|
||||
*/
|
||||
import { describe, it, expect, beforeEach } from 'vitest'
|
||||
import { MemoryStorage } from '../../../src/storage/adapters/memoryStorage.js'
|
||||
import {
|
||||
GenerationSegmentStore,
|
||||
SEGMENTS_PREFIX,
|
||||
type FoldGeneration
|
||||
} from '../../../src/db/generationSegments.js'
|
||||
|
||||
const UUID = (n: number): string => `00000000-0000-4000-8000-${String(n).padStart(12, '0')}`
|
||||
|
||||
const gen = (g: number, recordCount = 2): FoldGeneration => ({
|
||||
generation: g,
|
||||
timestamp: 1_700_000_000_000 + g,
|
||||
delta: { generation: g, nouns: [UUID(g)], verbs: [], bytes: 123 + g },
|
||||
records: Array.from({ length: recordCount }, (_, i) => ({
|
||||
kind: (i % 2 === 0 ? 'noun' : 'verb') as 'noun' | 'verb',
|
||||
id: UUID(g * 100 + i),
|
||||
record: { metadata: { noun: 'document', v: g }, vector: { v: [g, i] } }
|
||||
}))
|
||||
})
|
||||
|
||||
describe('db/GenerationSegmentStore — the D1+D3 packed tier', () => {
|
||||
let storage: MemoryStorage
|
||||
let store: GenerationSegmentStore
|
||||
|
||||
beforeEach(async () => {
|
||||
storage = new MemoryStorage()
|
||||
await storage.init()
|
||||
store = new GenerationSegmentStore(storage as any)
|
||||
await store.open()
|
||||
})
|
||||
|
||||
it('fold → read round-trips deltas and records via sidecar point-reads', async () => {
|
||||
const meta = await store.fold([gen(1), gen(2), gen(3)])
|
||||
expect(meta).toMatchObject({ firstGeneration: 1, lastGeneration: 3, frames: 3 })
|
||||
expect(meta.checksum).toBeGreaterThan(0)
|
||||
|
||||
expect(store.hasGeneration(2)).toBe(true)
|
||||
expect(store.hasGeneration(4)).toBe(false)
|
||||
|
||||
const d2 = await store.readDelta(2)
|
||||
expect(d2?.delta).toEqual({ generation: 2, nouns: [UUID(2)], verbs: [], bytes: 125 })
|
||||
expect(d2?.timestamp).toBe(1_700_000_000_002)
|
||||
|
||||
const records = await store.readRecords(3)
|
||||
expect(records).toHaveLength(2)
|
||||
expect(records![0]).toEqual({
|
||||
kind: 'noun',
|
||||
id: UUID(300),
|
||||
record: { metadata: { noun: 'document', v: 3 }, vector: { v: [3, 0] } }
|
||||
})
|
||||
// Point read by id, both kinds.
|
||||
expect(await store.readRecord(3, 'verb', UUID(301))).toEqual({
|
||||
metadata: { noun: 'document', v: 3 },
|
||||
vector: { v: [3, 1] }
|
||||
})
|
||||
expect(await store.readRecord(3, 'noun', UUID(999))).toBeNull()
|
||||
})
|
||||
|
||||
it('reopen discovers everything from the manifest alone — no listing', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await store.fold([gen(3), gen(4)])
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(reopened.segments()).toHaveLength(2)
|
||||
expect(reopened.hasGeneration(4)).toBe(true)
|
||||
expect((await reopened.readDelta(1))?.timestamp).toBe(1_700_000_000_001)
|
||||
})
|
||||
|
||||
it('a lost sidecar rebuilds from its segment; a damaged segment fails LOUDLY', async () => {
|
||||
const meta = await store.fold([gen(1), gen(2)])
|
||||
const idxPath = `${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.idx`
|
||||
await storage.deleteRawObject(idxPath)
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
// Rebuild path: still serves correct data.
|
||||
expect((await reopened.readRecords(2))!).toHaveLength(2)
|
||||
|
||||
// Now damage the SEGMENT itself: flip a payload byte → CRC mismatch, loud.
|
||||
const segPath = `${SEGMENTS_PREFIX}/${meta.file}`
|
||||
const bytes = (await storage.readRawBytes(segPath))!
|
||||
bytes[bytes.length - 3] ^= 0xff
|
||||
await storage.writeRawBytes(segPath, bytes)
|
||||
const damaged = new GenerationSegmentStore(storage as any)
|
||||
await damaged.open()
|
||||
;(damaged as any).sidecars.clear()
|
||||
await storage.deleteRawObject(idxPath) // force the sequential rebuild over damaged bytes
|
||||
await expect(damaged.readRecords(2)).rejects.toThrow(/CRC mismatch|damaged/)
|
||||
})
|
||||
|
||||
it('D3 reclaim drops whole segments only and bumps compactedBelow', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await store.fold([gen(3), gen(4)])
|
||||
await store.fold([gen(5), gen(6)])
|
||||
|
||||
// Horizon mid-segment-2 (below 4): only segment 1 is FULLY below → drops.
|
||||
const r1 = await store.dropSegmentsBelow(4)
|
||||
expect(r1).toEqual({ dropped: 1, compactedBelow: 3 })
|
||||
expect(store.hasGeneration(1)).toBe(false)
|
||||
expect(store.hasGeneration(3)).toBe(true) // partial segment survives whole
|
||||
|
||||
// Bytes actually gone.
|
||||
expect(await storage.readRawBytes(`${SEGMENTS_PREFIX}/seg-${String(1).padStart(20, '0')}.bgs`)).toBeNull()
|
||||
|
||||
// Horizon past everything: the rest drop; compactedBelow is durable.
|
||||
const r2 = await store.dropSegmentsBelow(7)
|
||||
expect(r2.dropped).toBe(2)
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(reopened.compactedBelow()).toBe(7)
|
||||
expect(reopened.segments()).toHaveLength(0)
|
||||
})
|
||||
|
||||
it('the packed digest is deterministic across reopen and changes with history', async () => {
|
||||
await store.fold([gen(1), gen(2), gen(3)])
|
||||
const atSeal = await store.digestThroughPacked(3)
|
||||
const midSegment = await store.digestThroughPacked(2)
|
||||
expect(atSeal).not.toBeNull()
|
||||
expect(midSegment).not.toBeNull()
|
||||
expect(midSegment).not.toBe(atSeal)
|
||||
|
||||
const reopened = new GenerationSegmentStore(storage as any)
|
||||
await reopened.open()
|
||||
expect(await reopened.digestThroughPacked(3)).toBe(atSeal)
|
||||
expect(await reopened.digestThroughPacked(2)).toBe(midSegment)
|
||||
|
||||
await reopened.fold([gen(4)])
|
||||
expect(await reopened.digestThroughPacked(4)).not.toBe(atSeal)
|
||||
})
|
||||
|
||||
it('sealed segments are immutable — fold refuses overlap, requires ascending input', async () => {
|
||||
await store.fold([gen(1), gen(2)])
|
||||
await expect(store.fold([gen(2), gen(3)])).rejects.toThrow(/overlaps the packed tier/)
|
||||
await expect(store.fold([gen(4), gen(4)])).rejects.toThrow(/strictly ascending/)
|
||||
await expect(store.fold([])).rejects.toThrow(/at least one generation/)
|
||||
})
|
||||
})
|
||||
|
|
@ -32,7 +32,7 @@ function entity(overrides: Partial<Entity> = {}): Entity {
|
|||
}
|
||||
|
||||
describe('db/whereMatcher — resolveEntityField', () => {
|
||||
it('system.<field> resolves the entity scalar; bare/metadata. reads the metadata bag only (sealed 2026-08-03)', () => {
|
||||
it('resolves standard top-level fields', () => {
|
||||
const e = entity({
|
||||
subtype: 'invoice',
|
||||
service: 'billing',
|
||||
|
|
@ -41,32 +41,17 @@ describe('db/whereMatcher — resolveEntityField', () => {
|
|||
_rev: 3,
|
||||
data: 'payload'
|
||||
})
|
||||
|
||||
// system.<field> is the ONLY spelling that reaches an entity scalar.
|
||||
expect(resolveEntityField(e, 'system.id')).toBe('e-1')
|
||||
expect(resolveEntityField(e, 'system.type')).toBe(NounType.Document)
|
||||
expect(resolveEntityField(e, 'system.subtype')).toBe('invoice')
|
||||
expect(resolveEntityField(e, 'system.service')).toBe('billing')
|
||||
expect(resolveEntityField(e, 'system.confidence')).toBe(0.9)
|
||||
expect(resolveEntityField(e, 'system.weight')).toBe(0.5)
|
||||
expect(resolveEntityField(e, 'system.createdAt')).toBe(1000)
|
||||
expect(resolveEntityField(e, 'system.updatedAt')).toBe(2000)
|
||||
|
||||
// Plumbing (_rev, data) is invisible even via system. — not in the
|
||||
// ten-scalar map, so this internal resolver reads it as absent (the typed
|
||||
// refusal for these lives one layer up, at the query-surface parser).
|
||||
expect(resolveEntityField(e, 'system._rev')).toBeUndefined()
|
||||
expect(resolveEntityField(e, 'system.data')).toBeUndefined()
|
||||
|
||||
// Bare names are ALWAYS the user's metadata field — even when they share
|
||||
// a spelling with an engine scalar, or with the now-dead 'noun' alias.
|
||||
// This entity's metadata bag is empty, so every bare name below reads
|
||||
// absent rather than silently falling back to the entity scalar.
|
||||
expect(resolveEntityField(e, 'id')).toBeUndefined()
|
||||
expect(resolveEntityField(e, 'type')).toBeUndefined()
|
||||
expect(resolveEntityField(e, 'noun')).toBeUndefined() // legacy alias is dead
|
||||
expect(resolveEntityField(e, 'subtype')).toBeUndefined()
|
||||
expect(resolveEntityField(e, 'createdAt')).toBeUndefined()
|
||||
expect(resolveEntityField(e, 'id')).toBe('e-1')
|
||||
expect(resolveEntityField(e, 'type')).toBe(NounType.Document)
|
||||
expect(resolveEntityField(e, 'noun')).toBe(NounType.Document) // alias
|
||||
expect(resolveEntityField(e, 'subtype')).toBe('invoice')
|
||||
expect(resolveEntityField(e, 'service')).toBe('billing')
|
||||
expect(resolveEntityField(e, 'confidence')).toBe(0.9)
|
||||
expect(resolveEntityField(e, 'weight')).toBe(0.5)
|
||||
expect(resolveEntityField(e, '_rev')).toBe(3)
|
||||
expect(resolveEntityField(e, 'createdAt')).toBe(1000)
|
||||
expect(resolveEntityField(e, 'updatedAt')).toBe(2000)
|
||||
expect(resolveEntityField(e, 'data')).toBe('payload')
|
||||
})
|
||||
|
||||
it('resolves custom fields from the metadata bag', () => {
|
||||
|
|
|
|||
|
|
@ -30,9 +30,6 @@ function allTestFiles(dir: string, out: string[] = []): string[] {
|
|||
* conscious decision — a NEW orphan not listed here fails the guard below.
|
||||
*/
|
||||
const MANUAL_ONLY = new Set<string>([
|
||||
// Conformance suites run as an explicit gate stage (both engines run them
|
||||
// by direct invocation), never swept into the unit/integration configs.
|
||||
'tests/conformance/collider-fidelity.test.ts',
|
||||
'tests/api/performance-benchmarks.test.ts',
|
||||
'tests/critical-neural-validation.test.ts',
|
||||
'tests/critical-performance-benchmark.test.ts',
|
||||
|
|
@ -41,15 +38,7 @@ const MANUAL_ONLY = new Set<string>([
|
|||
'tests/package-size-limit.test.ts',
|
||||
'tests/performance/graph-scale-performance.test.ts',
|
||||
'tests/performance/triple-intelligence-scale.test.ts',
|
||||
'tests/performance/typeAware.bench.test.ts',
|
||||
// Cross-engine field-addressing conformance suite: pinned bit-for-bit against
|
||||
// the native accelerator's implementation of the SAME contract, and invoked
|
||||
// directly (`npx vitest run tests/conformance/namespace-law.test.ts`), never
|
||||
// swept into the unit/integration gates — a run against a branch where the
|
||||
// resolver hasn't landed yet must SKIP loudly (see the file's own SELF-SKIP
|
||||
// doc), not silently pass/fail as a side effect of which gate happened to
|
||||
// pick it up.
|
||||
'tests/conformance/namespace-law.test.ts'
|
||||
'tests/performance/typeAware.bench.test.ts'
|
||||
])
|
||||
|
||||
function inGate(rel: string): boolean {
|
||||
|
|
|
|||
|
|
@ -1,127 +0,0 @@
|
|||
/**
|
||||
* @module tests/unit/types/nestedBagRecord
|
||||
* @description Unit pins for the v2 (nested-bag) stored-record layer — the
|
||||
* storage half of the field-addressing law. The write door accepts ANY user
|
||||
* metadata name; what makes that lossless on disk is the record shape:
|
||||
* engine fields top-level, the user bag NESTED verbatim, discriminated by
|
||||
* the engine-written format stamp (never by names — names are the user's).
|
||||
* These pins hold the builders, the discriminator, and the shape-aware
|
||||
* split that every read path (live, batch, historical) routes through.
|
||||
*/
|
||||
import { describe, it, expect } from 'vitest'
|
||||
import {
|
||||
buildNounMetadataRecord,
|
||||
buildVerbMetadataRecord,
|
||||
splitNounMetadataRecord,
|
||||
splitVerbMetadataRecord,
|
||||
isNestedBagRecord,
|
||||
METADATA_RECORD_FORMAT_KEY,
|
||||
NESTED_BAG_FORMAT
|
||||
} from '../../../src/types/reservedFields.js'
|
||||
|
||||
const COLLIDER_BAG = {
|
||||
confidence: 'user-confidence',
|
||||
weight: 'user-weight',
|
||||
subtype: 'user-subtype',
|
||||
createdAt: 'user-createdAt',
|
||||
service: 'user-service',
|
||||
data: 'user-data',
|
||||
noun: 'user-noun',
|
||||
_rev: 'user-rev',
|
||||
level: 7,
|
||||
plain: 'control'
|
||||
}
|
||||
|
||||
describe('v2 nested-bag stored records — build / discriminate / split', () => {
|
||||
it('build → split round-trips a fully colliding user bag VERBATIM', () => {
|
||||
const record = buildNounMetadataRecord(
|
||||
{ noun: 'document', confidence: 0.25, createdAt: 111, updatedAt: 222, _rev: 1 },
|
||||
{ ...COLLIDER_BAG }
|
||||
)
|
||||
expect(isNestedBagRecord(record)).toBe(true)
|
||||
expect(record[METADATA_RECORD_FORMAT_KEY]).toBe(NESTED_BAG_FORMAT)
|
||||
|
||||
const { reserved, custom } = splitNounMetadataRecord(record)
|
||||
// The engine half is exactly what the engine wrote…
|
||||
expect(reserved.noun).toBe('document')
|
||||
expect(reserved.confidence).toBe(0.25)
|
||||
expect(reserved._rev).toBe(1)
|
||||
// …and the user bag comes back byte-for-byte, colliders included.
|
||||
expect(custom).toEqual(COLLIDER_BAG)
|
||||
})
|
||||
|
||||
it('the verb mirror round-trips an edge collider bag verbatim', () => {
|
||||
const record = buildVerbMetadataRecord(
|
||||
{ verb: 'relatedTo', weight: 1.0, confidence: 0.5, createdAt: 333 },
|
||||
{ verb: 'user-verb', confidence: 'user-c', tag: 't' }
|
||||
)
|
||||
expect(isNestedBagRecord(record)).toBe(true)
|
||||
const { reserved, custom } = splitVerbMetadataRecord(record)
|
||||
expect(reserved.verb).toBe('relatedTo')
|
||||
expect(reserved.confidence).toBe(0.5)
|
||||
expect(custom).toEqual({ verb: 'user-verb', confidence: 'user-c', tag: 't' })
|
||||
})
|
||||
|
||||
it('a LEGACY flat record (no stamp) splits BY NAME — sound because the pre-law door refused colliders', () => {
|
||||
const legacy = {
|
||||
noun: 'document',
|
||||
confidence: 0.75,
|
||||
createdAt: 111,
|
||||
_rev: 2,
|
||||
legacyField: 'legacy-value'
|
||||
}
|
||||
expect(isNestedBagRecord(legacy)).toBe(false)
|
||||
const { reserved, custom } = splitNounMetadataRecord(legacy)
|
||||
expect(reserved.confidence).toBe(0.75)
|
||||
expect(reserved._rev).toBe(2)
|
||||
expect(custom).toEqual({ legacyField: 'legacy-value' })
|
||||
})
|
||||
|
||||
it('the stamp is the discriminator, never the name: a legacy user OBJECT field named `metadata` does not fake a v2 record', () => {
|
||||
// Pre-law, 'metadata' was never a reserved name — a flat record could
|
||||
// legally carry a user object field spelled exactly 'metadata'. Without
|
||||
// the engine-written stamp it must split as legacy, with that object
|
||||
// preserved as an ordinary user field.
|
||||
const legacyWithMetadataField = {
|
||||
noun: 'document',
|
||||
confidence: 0.5,
|
||||
metadata: { nested: 'user-object' }
|
||||
}
|
||||
expect(isNestedBagRecord(legacyWithMetadataField)).toBe(false)
|
||||
const { reserved, custom } = splitNounMetadataRecord(legacyWithMetadataField)
|
||||
expect(reserved.confidence).toBe(0.5)
|
||||
expect(custom).toEqual({ metadata: { nested: 'user-object' } })
|
||||
})
|
||||
|
||||
it('a malformed stamp (right key, wrong value / non-object bag) never discriminates as v2', () => {
|
||||
expect(
|
||||
isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: 999, metadata: {} })
|
||||
).toBe(false)
|
||||
expect(
|
||||
isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, metadata: 'not-a-bag' })
|
||||
).toBe(false)
|
||||
expect(
|
||||
isNestedBagRecord({ [METADATA_RECORD_FORMAT_KEY]: NESTED_BAG_FORMAT, metadata: [1, 2] })
|
||||
).toBe(false)
|
||||
expect(isNestedBagRecord(null)).toBe(false)
|
||||
expect(isNestedBagRecord(undefined)).toBe(false)
|
||||
})
|
||||
|
||||
it('the v2 split never surfaces the stamp or the bag container as fields', () => {
|
||||
const record = buildNounMetadataRecord({ noun: 'document', _rev: 1 }, { a: 1 })
|
||||
const { reserved, custom } = splitNounMetadataRecord(record)
|
||||
expect(METADATA_RECORD_FORMAT_KEY in reserved).toBe(false)
|
||||
expect(METADATA_RECORD_FORMAT_KEY in custom).toBe(false)
|
||||
expect('metadata' in reserved).toBe(false)
|
||||
expect(custom).toEqual({ a: 1 })
|
||||
})
|
||||
|
||||
it('builders copy the bag (no aliasing): later caller mutation cannot reach the record', () => {
|
||||
const bag: Record<string, unknown> = { a: 1 }
|
||||
const record = buildNounMetadataRecord({ noun: 'document' }, bag)
|
||||
bag.a = 999
|
||||
bag.b = 'sneaky'
|
||||
expect((record.metadata as Record<string, unknown>).a).toBe(1)
|
||||
expect('b' in (record.metadata as Record<string, unknown>)).toBe(false)
|
||||
})
|
||||
})
|
||||
265
tests/unit/types/reserved-metadata-keys.test-d.ts
Normal file
265
tests/unit/types/reserved-metadata-keys.test-d.ts
Normal file
|
|
@ -0,0 +1,265 @@
|
|||
/**
|
||||
* @module tests/unit/types/reserved-metadata-keys.test-d
|
||||
* @description Compile-time tests for the reserved-field contract (layer 1 of
|
||||
* three — see src/types/reservedFields.ts): a literal reserved key inside any
|
||||
* `metadata` param is a TypeScript error, while the generic `T` ergonomics
|
||||
* stay intact (typed bags, untyped brains, index-signature shapes, and the
|
||||
* documented exemption for consumers who explicitly declare a reserved key in
|
||||
* their own metadata type).
|
||||
*
|
||||
* Runs under vitest typecheck mode (`test.typecheck` in
|
||||
* tests/configs/vitest.unit.config.ts) — these assertions are validated by
|
||||
* `tsc`, never executed. The runtime half of the contract (the write-path
|
||||
* remap for untyped callers) is pinned by
|
||||
* tests/unit/brainy/update-reserved-metadata-remap.test.ts.
|
||||
*/
|
||||
|
||||
import { describe, it, assertType } from 'vitest'
|
||||
import type {
|
||||
AddParams,
|
||||
UpdateParams,
|
||||
RelateParams,
|
||||
UpdateRelationParams,
|
||||
TxOperation
|
||||
} from '../../../src/index.js'
|
||||
import { NounType, VerbType } from '../../../src/types/graphTypes.js'
|
||||
|
||||
describe('reserved entity keys in metadata are compile errors', () => {
|
||||
it('AddParams (untyped brain) rejects every reserved key but stays open for custom fields', () => {
|
||||
// Custom fields of any shape remain legal — exactly the pre-8.0 latitude.
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
metadata: { dept: 'eng', level: 3, tags: ['a', 'b'], nested: { ok: true } }
|
||||
})
|
||||
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'noun' is reserved (the entity type travels via the top-level 'type' param)
|
||||
metadata: { noun: 'organization' }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param)
|
||||
metadata: { subtype: 'contractor' }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'createdAt' is reserved (system-managed)
|
||||
metadata: { createdAt: Date.now() }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'updatedAt' is reserved (system-managed)
|
||||
metadata: { updatedAt: Date.now() }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param)
|
||||
metadata: { confidence: 0.8 }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'weight' is reserved (use the top-level 'weight' param)
|
||||
metadata: { weight: 0.5 }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'service' is reserved (use the top-level 'service' param)
|
||||
metadata: { service: 'orders' }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'data' is reserved (use the top-level 'data' param)
|
||||
metadata: { data: 'content' }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'createdBy' is reserved (use the top-level 'createdBy' param)
|
||||
metadata: { createdBy: { augmentation: 'importer', version: '1.0' } }
|
||||
})
|
||||
assertType<AddParams>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — '_rev' is reserved (system-managed revision counter)
|
||||
metadata: { _rev: 7 }
|
||||
})
|
||||
})
|
||||
|
||||
it('AddParams<T> (typed brain) rejects reserved keys alongside the declared shape', () => {
|
||||
interface EmployeeMeta {
|
||||
dept: string
|
||||
level: number
|
||||
}
|
||||
|
||||
assertType<AddParams<EmployeeMeta>>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
metadata: { dept: 'eng', level: 3 }
|
||||
})
|
||||
|
||||
assertType<AddParams<EmployeeMeta>>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
// @ts-expect-error — 'confidence' is reserved even when T declares other fields
|
||||
metadata: { dept: 'eng', level: 3, confidence: 0.8 }
|
||||
})
|
||||
})
|
||||
|
||||
it('documented exemptions: T-declared reserved keys and index-signature shapes stay assignable', () => {
|
||||
// A consumer who *explicitly* types a reserved key into their metadata
|
||||
// shape keeps a working (if unwise) type — the guard exempts keyof T.
|
||||
interface LegacyMeta {
|
||||
confidence: number
|
||||
note: string
|
||||
}
|
||||
assertType<AddParams<LegacyMeta>>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
metadata: { confidence: 0.8, note: 'declared by the consumer type' }
|
||||
})
|
||||
|
||||
// Index-signature metadata types (keyof T = string) remain fully open.
|
||||
assertType<AddParams<Record<string, unknown>>>({
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
data: 'x',
|
||||
metadata: { anything: 'goes', confidence: 0.8 }
|
||||
})
|
||||
})
|
||||
|
||||
it('UpdateParams patch rejects reserved keys but accepts partial custom patches', () => {
|
||||
interface EmployeeMeta {
|
||||
dept: string
|
||||
level: number
|
||||
}
|
||||
|
||||
// Partial patch of the declared shape is legal.
|
||||
assertType<UpdateParams<EmployeeMeta>>({ id: 'e1', metadata: { dept: 'sales' } })
|
||||
// Untyped patch with custom fields is legal.
|
||||
assertType<UpdateParams>({ id: 'e1', metadata: { status: 'reviewed', rating: 4.5 } })
|
||||
|
||||
// @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param)
|
||||
assertType<UpdateParams>({ id: 'e1', metadata: { confidence: 0.33 } })
|
||||
// @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param)
|
||||
assertType<UpdateParams>({ id: 'e1', metadata: { subtype: 'specialized' } })
|
||||
// @ts-expect-error — '_rev' is reserved (pass 'ifRev' for optimistic concurrency)
|
||||
assertType<UpdateParams>({ id: 'e1', metadata: { _rev: 3 } })
|
||||
// @ts-expect-error — 'confidence' is reserved even when T declares other fields
|
||||
assertType<UpdateParams<EmployeeMeta>>({ id: 'e1', metadata: { confidence: 0.1 } })
|
||||
})
|
||||
})
|
||||
|
||||
describe('reserved relationship keys in metadata are compile errors', () => {
|
||||
it('RelateParams rejects reserved keys but stays open for custom edge fields', () => {
|
||||
assertType<RelateParams>({
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
metadata: { role: 'peer', since: 2024 }
|
||||
})
|
||||
|
||||
assertType<RelateParams>({
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
// @ts-expect-error — 'verb' is reserved (the relationship type travels via the top-level 'type' param)
|
||||
metadata: { verb: 'relatedTo' }
|
||||
})
|
||||
assertType<RelateParams>({
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
// @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param)
|
||||
metadata: { confidence: 0.9 }
|
||||
})
|
||||
assertType<RelateParams>({
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
// @ts-expect-error — 'weight' is reserved (use the top-level 'weight' param)
|
||||
metadata: { weight: 0.4 }
|
||||
})
|
||||
assertType<RelateParams>({
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.ReportsTo,
|
||||
subtype: 'direct',
|
||||
// @ts-expect-error — 'service' is reserved (use the top-level 'service' param)
|
||||
metadata: { service: 'orders' }
|
||||
})
|
||||
})
|
||||
|
||||
it('UpdateRelationParams patch rejects reserved keys', () => {
|
||||
assertType<UpdateRelationParams>({ id: 'r1', metadata: { note: 'fine' } })
|
||||
|
||||
// @ts-expect-error — 'confidence' is reserved (use the top-level 'confidence' param)
|
||||
assertType<UpdateRelationParams>({ id: 'r1', metadata: { confidence: 0.5 } })
|
||||
// @ts-expect-error — 'subtype' is reserved (use the top-level 'subtype' param)
|
||||
assertType<UpdateRelationParams>({ id: 'r1', metadata: { subtype: 'dotted-line' } })
|
||||
// @ts-expect-error — 'createdAt' is reserved (system-managed)
|
||||
assertType<UpdateRelationParams>({ id: 'r1', metadata: { createdAt: 1 } })
|
||||
})
|
||||
})
|
||||
|
||||
describe('transact() operations inherit the same guard', () => {
|
||||
it('TxOperation add/update/relate metadata rejects reserved keys', () => {
|
||||
assertType<TxOperation>({
|
||||
op: 'add',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'tx',
|
||||
metadata: { custom: 'a' }
|
||||
})
|
||||
assertType<TxOperation>({
|
||||
op: 'add',
|
||||
type: NounType.Concept,
|
||||
subtype: 'general',
|
||||
data: 'tx',
|
||||
// @ts-expect-error — 'confidence' is reserved on transact add ops too
|
||||
metadata: { confidence: 0.7 }
|
||||
})
|
||||
assertType<TxOperation>({
|
||||
op: 'update',
|
||||
id: 'e1',
|
||||
// @ts-expect-error — 'weight' is reserved on transact update ops too
|
||||
metadata: { weight: 0.2 }
|
||||
})
|
||||
assertType<TxOperation>({
|
||||
op: 'relate',
|
||||
from: 'a',
|
||||
to: 'b',
|
||||
type: VerbType.RelatedTo,
|
||||
subtype: 'colleague',
|
||||
// @ts-expect-error — 'verb' is reserved on transact relate ops too
|
||||
metadata: { verb: 'contains' }
|
||||
})
|
||||
})
|
||||
})
|
||||
|
|
@ -56,15 +56,11 @@ describe('Zero-Config Parameter Validation', () => {
|
|||
})).toThrow('cannot specify both query and vector')
|
||||
})
|
||||
|
||||
it('should refuse cursor outright — even paired with offset — as an unimplemented option', () => {
|
||||
// cursor is now a typed, unconditional refusal (UnsupportedFindOptionError):
|
||||
// it used to be accepted-and-ignored, only conflicting when offset was also
|
||||
// given. Accepted-and-ignored died as a class — cursor refuses on its own,
|
||||
// so pairing it with offset refuses too, but with the SAME message.
|
||||
it('should reject both cursor and offset', () => {
|
||||
expect(() => validateFindParams({
|
||||
cursor: 'abc123',
|
||||
offset: 10
|
||||
})).toThrow("find() option 'cursor' is not implemented")
|
||||
})).toThrow('cannot use both cursor and offset pagination')
|
||||
})
|
||||
|
||||
it('should validate vector dimensions', () => {
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue