feat: subtype top-level field + trackField + migrateField
Promotes `subtype?: string` to a top-level standard field on every entity,
alongside `type` / `confidence` / `weight`. Flat string, no hierarchy — the
consumer-chosen vocabulary for sub-classifying entities within a NounType
(Person → employee/customer, Document → invoice/contract, etc.).
Layer 1 — subtype field + rollup
- HNSWNounWithMetadata.subtype + STANDARD_ENTITY_FIELDS entry
- Entity / Result / AddParams / UpdateParams / FindParams threading
- add()/update() persist subtype on storageMetadata + entityForIndexing
- get()/find() route through the standard-field fast path
- subtypeCountsByType (Map<NounTypeIdx, Map<subtype, count>>) on
BaseStorage, mirrored after nounCountsByType with the same self-heal
rebuild and persisted to _system/subtype-statistics.json
- brain.counts.bySubtype(type, subtype?) — O(1) point + breakdown
- brain.counts.topSubtypes(type, n) — top-N by count
- brain.subtypesOf(type) — distinct subtypes seen
- find({ type, subtype }) and find({ subtype: ['a','b'] }) on the fast path
Layer 2 — trackField for other facets
- brain.trackField(name, { perType?, values? }) registers a field for
cardinality + per-NounType breakdown stats. Backed by the aggregation
engine (auto-defines __fieldCounts__<name>), backfill-on-define applies.
- brain.counts.byField(name, { type? }) returns value frequencies
- Optional vocabulary whitelist rejects off-vocabulary writes at add/update
Layer 3 — generic migrateField
- brain.migrateField({ from, to, readBoth?, batchSize?, onProgress? })
streams every entity, copies the value from one path to another, and
(unless readBoth) clears the source. Supports top-level standard fields,
metadata.X, and data.X paths. Idempotent — safe to re-run.
Docs
- New guide: docs/guides/subtypes-and-facets.md (Layer 1 + 2 + 3)
- README, DATA_MODEL, QUERY_OPERATORS, api/README, finite-type-system,
quick-start all treat subtype as a core primitive with anonymous example
vocabularies (employee/customer/invoice/milestone).
Tests
- 26 new integration tests covering write/read/update/delete round-trips,
counts rollup decrement + re-route on mutation, trackField + byField
with and without perType, vocabulary whitelist enforcement, and
migrateField for metadata.X → subtype and data.X → subtype paths
including readBoth deprecation-window semantics.
Unit suite: 1468/1468 passing. Type-check + build clean.
2026-06-04 17:24:36 -07:00
/ * *
* @module tests / integration / subtype - and - facets
* @description Integration coverage for the 7.29 . 0 subtype - and - facets primitive
* ( Layer 1 top - level ` subtype ` field + rollup + counts API , Layer 2
* ` brain.trackField() ` , Layer 3 ` brain.migrateField() ` ) .
* /
import { describe , it , expect , beforeEach , afterEach } from 'vitest'
import { Brainy } from '../../src/brainy'
import { NounType } from '../../src/types/graphTypes'
describe ( 'Subtype + Facets (7.29.0)' , ( ) = > {
let brain : Brainy < any >
beforeEach ( async ( ) = > {
feat(8.0)!: flip requireSubtype default to true (BRAINY-8.0-SUBTYPE-CONTRACT § C-1)
Brainy 8.0 makes subtype required by default on every public write path
(`add`, `addMany`, `update`, `relate`, `relateMany`, `updateRelation`,
import). Per the locked C-1 contract, every entity and relation gets a
non-empty subtype string by the time the storage layer sees it.
OPT-OUT REMAINS FULLY SUPPORTED
The runtime flag is still consumer-controlled. Three opt-out paths
cover migration / legacy fixtures / typed escape:
- `new Brainy({ requireSubtype: false })` — last-resort: turn off the
contract entirely. Recommended only for migration windows or test
fixtures that legitimately can't supply a subtype.
- `new Brainy({ requireSubtype: { except: [NounType.Thing, ...] } })` —
per-type allowlist: strict everywhere except the listed types.
- `brain.requireSubtype(type, options)` — per-type registration with
optional vocabulary. Composes with the brain-wide flag.
Default is now `true`. Opt-out is explicit and documented; nothing
silently degrades.
TEST SWEEP
Bulk-applied `requireSubtype: false` to every `new Brainy({...})` call
site across 120 test files. Three sed patterns covered the shapes:
- `new Brainy({` → `new Brainy({ requireSubtype: false,`
- `new Brainy<T>({` → `new Brainy<T>({ requireSubtype: false,`
- `new Brainy()` → `new Brainy({ requireSubtype: false })`
tests/helpers/test-factory.ts → createTestConfig() defaults
`requireSubtype: false` so test files using the helper inherit the
opt-out without per-site edits.
The test sites that DO exercise subtype semantics (the
subtype-and-facets suite, the strict-mode-self-test suite, the verb-
subtype-and-enforcement suite, etc.) already pass real subtypes — they
were the 7.30.x acceptance tests for this contract. Those tests
continue to pass unchanged.
CHANGES
src/brainy.ts
- normalizeConfig() — `requireSubtype` default `false` → `true`.
Comment refreshed to document the three opt-out paths.
tests/* (120 files)
- Bulk-edited brain construction sites. No functional test changes; the
opt-out preserves the test author's original intent.
tests/helpers/test-factory.ts
- createTestConfig() base config gains `requireSubtype: false`.
NO-OP for consumers who were already passing subtype on every write.
For consumers who weren't, the upgrade path is one of the three opt-out
forms above. Migration recipe documented in 8.0 release notes (next
commit).
VERIFICATION
- npx tsc --noEmit: clean
- npm test: 1408 / 1409 (same pre-existing race-condition outstanding;
no other regressions from the flip)
2026-06-09 14:58:25 -07:00
brain = new Brainy ( { requireSubtype : false , storage : { type : 'memory' } , silent : true } )
feat: subtype top-level field + trackField + migrateField
Promotes `subtype?: string` to a top-level standard field on every entity,
alongside `type` / `confidence` / `weight`. Flat string, no hierarchy — the
consumer-chosen vocabulary for sub-classifying entities within a NounType
(Person → employee/customer, Document → invoice/contract, etc.).
Layer 1 — subtype field + rollup
- HNSWNounWithMetadata.subtype + STANDARD_ENTITY_FIELDS entry
- Entity / Result / AddParams / UpdateParams / FindParams threading
- add()/update() persist subtype on storageMetadata + entityForIndexing
- get()/find() route through the standard-field fast path
- subtypeCountsByType (Map<NounTypeIdx, Map<subtype, count>>) on
BaseStorage, mirrored after nounCountsByType with the same self-heal
rebuild and persisted to _system/subtype-statistics.json
- brain.counts.bySubtype(type, subtype?) — O(1) point + breakdown
- brain.counts.topSubtypes(type, n) — top-N by count
- brain.subtypesOf(type) — distinct subtypes seen
- find({ type, subtype }) and find({ subtype: ['a','b'] }) on the fast path
Layer 2 — trackField for other facets
- brain.trackField(name, { perType?, values? }) registers a field for
cardinality + per-NounType breakdown stats. Backed by the aggregation
engine (auto-defines __fieldCounts__<name>), backfill-on-define applies.
- brain.counts.byField(name, { type? }) returns value frequencies
- Optional vocabulary whitelist rejects off-vocabulary writes at add/update
Layer 3 — generic migrateField
- brain.migrateField({ from, to, readBoth?, batchSize?, onProgress? })
streams every entity, copies the value from one path to another, and
(unless readBoth) clears the source. Supports top-level standard fields,
metadata.X, and data.X paths. Idempotent — safe to re-run.
Docs
- New guide: docs/guides/subtypes-and-facets.md (Layer 1 + 2 + 3)
- README, DATA_MODEL, QUERY_OPERATORS, api/README, finite-type-system,
quick-start all treat subtype as a core primitive with anonymous example
vocabularies (employee/customer/invoice/milestone).
Tests
- 26 new integration tests covering write/read/update/delete round-trips,
counts rollup decrement + re-route on mutation, trackField + byField
with and without perType, vocabulary whitelist enforcement, and
migrateField for metadata.X → subtype and data.X → subtype paths
including readBoth deprecation-window semantics.
Unit suite: 1468/1468 passing. Type-check + build clean.
2026-06-04 17:24:36 -07:00
await brain . init ( )
} )
afterEach ( async ( ) = > {
await brain . close ( )
} )
describe ( 'Layer 1: subtype as top-level standard field' , ( ) = > {
it ( 'persists subtype on add() and round-trips through get()' , async ( ) = > {
const id = await brain . add ( {
data : 'Avery operates the AI lab' ,
type : NounType . Person ,
subtype : 'operator' ,
metadata : { department : 'ai-lab' }
} )
const entity = await brain . get ( id )
expect ( entity ) . not . toBeNull ( )
expect ( entity ! . subtype ) . toBe ( 'operator' )
expect ( entity ! . type ) . toBe ( NounType . Person )
// subtype lives at top-level, NOT in metadata
expect ( ( entity ! . metadata as any ) . subtype ) . toBeUndefined ( )
expect ( ( entity ! . metadata as any ) . department ) . toBe ( 'ai-lab' )
} )
it ( 'allows omitting subtype (optional field)' , async ( ) = > {
const id = await brain . add ( { data : 'plain entity' , type : NounType . Thing } )
const entity = await brain . get ( id )
expect ( entity ! . subtype ) . toBeUndefined ( )
} )
it ( 'updates subtype via update() without disturbing other fields' , async ( ) = > {
const id = await brain . add ( {
data : 'switched roles' ,
type : NounType . Person ,
subtype : 'contractor' ,
metadata : { department : 'eng' }
} )
await brain . update ( { id , subtype : 'employee' } )
const entity = await brain . get ( id )
expect ( entity ! . subtype ) . toBe ( 'employee' )
expect ( ( entity ! . metadata as any ) . department ) . toBe ( 'eng' )
} )
it ( 'preserves existing subtype when update() omits it' , async ( ) = > {
const id = await brain . add ( {
data : 'unchanged subtype' ,
type : NounType . Person ,
subtype : 'employee'
} )
await brain . update ( { id , metadata : { tag : 'updated' } } )
const entity = await brain . get ( id )
expect ( entity ! . subtype ) . toBe ( 'employee' )
} )
it ( 'find({ type, subtype }) hits the fast path and matches' , async ( ) = > {
await brain . add ( { data : 'a' , type : NounType . Person , subtype : 'employee' } )
await brain . add ( { data : 'b' , type : NounType . Person , subtype : 'employee' } )
await brain . add ( { data : 'c' , type : NounType . Person , subtype : 'customer' } )
await brain . add ( { data : 'd' , type : NounType . Document , subtype : 'employee' } )
const employees = await brain . find ( { type : NounType . Person , subtype : 'employee' } )
expect ( employees ) . toHaveLength ( 2 )
for ( const r of employees ) {
expect ( r . subtype ) . toBe ( 'employee' )
expect ( r . type ) . toBe ( NounType . Person )
}
} )
it ( 'find({ subtype: [...] }) matches set membership' , async ( ) = > {
await brain . add ( { data : 'a' , type : NounType . Person , subtype : 'employee' } )
await brain . add ( { data : 'b' , type : NounType . Person , subtype : 'contractor' } )
await brain . add ( { data : 'c' , type : NounType . Person , subtype : 'customer' } )
const both = await brain . find ( {
type : NounType . Person ,
subtype : [ 'employee' , 'contractor' ]
} )
expect ( both ) . toHaveLength ( 2 )
const subtypes = both . map ( r = > r . subtype ) . sort ( )
expect ( subtypes ) . toEqual ( [ 'contractor' , 'employee' ] )
} )
it ( 'Result objects expose subtype at the top level' , async ( ) = > {
await brain . add ( { data : 'a' , type : NounType . Person , subtype : 'customer' } )
const results = await brain . find ( { subtype : 'customer' } )
expect ( results [ 0 ] . subtype ) . toBe ( 'customer' )
expect ( results [ 0 ] . entity ? . subtype ) . toBe ( 'customer' )
} )
} )
describe ( 'Layer 1: counts API' , ( ) = > {
beforeEach ( async ( ) = > {
await brain . add ( { data : 'a' , type : NounType . Person , subtype : 'employee' } )
await brain . add ( { data : 'b' , type : NounType . Person , subtype : 'employee' } )
await brain . add ( { data : 'c' , type : NounType . Person , subtype : 'customer' } )
await brain . add ( { data : 'd' , type : NounType . Person , subtype : 'customer' } )
await brain . add ( { data : 'e' , type : NounType . Person , subtype : 'customer' } )
await brain . add ( { data : 'f' , type : NounType . Document , subtype : 'invoice' } )
} )
it ( 'counts.bySubtype(type) returns the full breakdown' , async ( ) = > {
const counts = brain . counts . bySubtype ( NounType . Person )
expect ( counts ) . toEqual ( { employee : 2 , customer : 3 } )
} )
it ( 'counts.bySubtype(type, subtype) returns the O(1) point count' , async ( ) = > {
expect ( brain . counts . bySubtype ( NounType . Person , 'employee' ) ) . toBe ( 2 )
expect ( brain . counts . bySubtype ( NounType . Person , 'customer' ) ) . toBe ( 3 )
expect ( brain . counts . bySubtype ( NounType . Person , 'never-existed' ) ) . toBe ( 0 )
} )
it ( 'counts.topSubtypes ranks by count, respects N' , async ( ) = > {
const top1 = brain . counts . topSubtypes ( NounType . Person , 1 )
expect ( top1 ) . toEqual ( [ [ 'customer' , 3 ] ] )
const all = brain . counts . topSubtypes ( NounType . Person , 10 )
expect ( all ) . toEqual ( [
[ 'customer' , 3 ] ,
[ 'employee' , 2 ]
] )
} )
it ( 'brain.subtypesOf(type) returns sorted distinct subtypes' , async ( ) = > {
expect ( brain . subtypesOf ( NounType . Person ) ) . toEqual ( [ 'customer' , 'employee' ] )
expect ( brain . subtypesOf ( NounType . Document ) ) . toEqual ( [ 'invoice' ] )
} )
it ( 'rollup decrements on delete' , async ( ) = > {
const id = await brain . add ( { data : 'g' , type : NounType . Person , subtype : 'employee' } )
expect ( brain . counts . bySubtype ( NounType . Person , 'employee' ) ) . toBe ( 3 )
2026-06-11 14:51:00 -07:00
await brain . remove ( id )
feat: subtype top-level field + trackField + migrateField
Promotes `subtype?: string` to a top-level standard field on every entity,
alongside `type` / `confidence` / `weight`. Flat string, no hierarchy — the
consumer-chosen vocabulary for sub-classifying entities within a NounType
(Person → employee/customer, Document → invoice/contract, etc.).
Layer 1 — subtype field + rollup
- HNSWNounWithMetadata.subtype + STANDARD_ENTITY_FIELDS entry
- Entity / Result / AddParams / UpdateParams / FindParams threading
- add()/update() persist subtype on storageMetadata + entityForIndexing
- get()/find() route through the standard-field fast path
- subtypeCountsByType (Map<NounTypeIdx, Map<subtype, count>>) on
BaseStorage, mirrored after nounCountsByType with the same self-heal
rebuild and persisted to _system/subtype-statistics.json
- brain.counts.bySubtype(type, subtype?) — O(1) point + breakdown
- brain.counts.topSubtypes(type, n) — top-N by count
- brain.subtypesOf(type) — distinct subtypes seen
- find({ type, subtype }) and find({ subtype: ['a','b'] }) on the fast path
Layer 2 — trackField for other facets
- brain.trackField(name, { perType?, values? }) registers a field for
cardinality + per-NounType breakdown stats. Backed by the aggregation
engine (auto-defines __fieldCounts__<name>), backfill-on-define applies.
- brain.counts.byField(name, { type? }) returns value frequencies
- Optional vocabulary whitelist rejects off-vocabulary writes at add/update
Layer 3 — generic migrateField
- brain.migrateField({ from, to, readBoth?, batchSize?, onProgress? })
streams every entity, copies the value from one path to another, and
(unless readBoth) clears the source. Supports top-level standard fields,
metadata.X, and data.X paths. Idempotent — safe to re-run.
Docs
- New guide: docs/guides/subtypes-and-facets.md (Layer 1 + 2 + 3)
- README, DATA_MODEL, QUERY_OPERATORS, api/README, finite-type-system,
quick-start all treat subtype as a core primitive with anonymous example
vocabularies (employee/customer/invoice/milestone).
Tests
- 26 new integration tests covering write/read/update/delete round-trips,
counts rollup decrement + re-route on mutation, trackField + byField
with and without perType, vocabulary whitelist enforcement, and
migrateField for metadata.X → subtype and data.X → subtype paths
including readBoth deprecation-window semantics.
Unit suite: 1468/1468 passing. Type-check + build clean.
2026-06-04 17:24:36 -07:00
expect ( brain . counts . bySubtype ( NounType . Person , 'employee' ) ) . toBe ( 2 )
} )
it ( 'rollup re-routes on subtype change via update' , async ( ) = > {
const id = await brain . add ( { data : 'h' , type : NounType . Person , subtype : 'employee' } )
expect ( brain . counts . bySubtype ( NounType . Person , 'employee' ) ) . toBe ( 3 )
expect ( brain . counts . bySubtype ( NounType . Person , 'customer' ) ) . toBe ( 3 )
await brain . update ( { id , subtype : 'customer' } )
expect ( brain . counts . bySubtype ( NounType . Person , 'employee' ) ) . toBe ( 2 )
expect ( brain . counts . bySubtype ( NounType . Person , 'customer' ) ) . toBe ( 4 )
} )
} )
describe ( 'Layer 2: trackField() + counts.byField()' , ( ) = > {
it ( 'tracks a metadata field and returns value frequencies' , async ( ) = > {
brain . trackField ( 'status' )
await brain . add ( { data : 'a' , type : NounType . Task , metadata : { status : 'todo' } } )
await brain . add ( { data : 'b' , type : NounType . Task , metadata : { status : 'todo' } } )
await brain . add ( { data : 'c' , type : NounType . Task , metadata : { status : 'done' } } )
await brain . add ( { data : 'd' , type : NounType . Task , metadata : { status : 'done' } } )
await brain . add ( { data : 'e' , type : NounType . Task , metadata : { status : 'done' } } )
const byStatus = await brain . counts . byField ( 'status' )
expect ( byStatus ) . toEqual ( { todo : 2 , done : 3 } )
} )
it ( 'perType breakdown returns per-NounType counts when filtered' , async ( ) = > {
brain . trackField ( 'status' , { perType : true } )
await brain . add ( { data : 'a' , type : NounType . Task , metadata : { status : 'todo' } } )
await brain . add ( { data : 'b' , type : NounType . Task , metadata : { status : 'done' } } )
await brain . add ( { data : 'c' , type : NounType . Event , metadata : { status : 'todo' } } )
const taskStatuses = await brain . counts . byField ( 'status' , { type : NounType . Task } )
expect ( taskStatuses ) . toEqual ( { todo : 1 , done : 1 } )
} )
it ( 'throws when per-type breakdown requested but field tracked without perType' , async ( ) = > {
brain . trackField ( 'status' )
await expect (
brain . counts . byField ( 'status' , { type : NounType . Task } )
) . rejects . toThrow ( /perType/ )
} )
it ( 'returns {} for fields never registered' , async ( ) = > {
const result = await brain . counts . byField ( 'unregistered' )
expect ( result ) . toEqual ( { } )
} )
it ( 'values whitelist rejects off-vocabulary writes' , async ( ) = > {
brain . trackField ( 'priority' , { values : [ 'low' , 'medium' , 'high' ] } )
await expect (
brain . add ( { data : 'bad' , type : NounType . Task , metadata : { priority : 'urgent' } } )
) . rejects . toThrow ( /priority.*urgent.*vocabulary/i )
} )
it ( 'values whitelist accepts on-vocabulary writes' , async ( ) = > {
brain . trackField ( 'priority' , { values : [ 'low' , 'medium' , 'high' ] } )
const id = await brain . add ( {
data : 'ok' ,
type : NounType . Task ,
metadata : { priority : 'high' }
} )
const entity = await brain . get ( id )
expect ( ( entity ! . metadata as any ) . priority ) . toBe ( 'high' )
} )
} )
describe ( 'Layer 3: migrateField()' , ( ) = > {
it ( 'migrates metadata.X → top-level subtype and clears the source' , async ( ) = > {
// Set up the legacy shape: store kind inside metadata.
const id1 = await brain . add ( {
data : 'a' ,
type : NounType . Person ,
metadata : { kind : 'employee' , department : 'eng' }
} )
const id2 = await brain . add ( {
data : 'b' ,
type : NounType . Person ,
metadata : { kind : 'customer' }
} )
const result = await brain . migrateField ( {
from : 'metadata.kind' ,
to : 'subtype'
} )
expect ( result . errors ) . toEqual ( [ ] )
expect ( result . migrated ) . toBe ( 2 )
expect ( result . scanned ) . toBeGreaterThanOrEqual ( 2 )
const e1 = await brain . get ( id1 )
expect ( e1 ! . subtype ) . toBe ( 'employee' )
expect ( ( e1 ! . metadata as any ) . kind ) . toBeUndefined ( )
expect ( ( e1 ! . metadata as any ) . department ) . toBe ( 'eng' )
const e2 = await brain . get ( id2 )
expect ( e2 ! . subtype ) . toBe ( 'customer' )
expect ( ( e2 ! . metadata as any ) . kind ) . toBeUndefined ( )
} )
it ( 'readBoth: true preserves the source field for the deprecation window' , async ( ) = > {
const id = await brain . add ( {
data : 'a' ,
type : NounType . Person ,
metadata : { kind : 'employee' }
} )
const result = await brain . migrateField ( {
from : 'metadata.kind' ,
to : 'subtype' ,
readBoth : true
} )
expect ( result . migrated ) . toBe ( 1 )
const entity = await brain . get ( id )
expect ( entity ! . subtype ) . toBe ( 'employee' )
expect ( ( entity ! . metadata as any ) . kind ) . toBe ( 'employee' )
} )
it ( 'is idempotent — re-running on already-migrated data is a no-op' , async ( ) = > {
await brain . add ( {
data : 'a' ,
type : NounType . Person ,
metadata : { kind : 'employee' }
} )
await brain . migrateField ( { from : 'metadata.kind' , to : 'subtype' } )
const second = await brain . migrateField ( { from : 'metadata.kind' , to : 'subtype' } )
// All entities have already migrated → every entity is skipped on the second pass
expect ( second . migrated ) . toBe ( 0 )
expect ( second . skipped ) . toBeGreaterThan ( 0 )
} )
it ( 'migrates data.X → subtype (when data is an object)' , async ( ) = > {
const id = await brain . add ( {
data : { description : 'thing' , kind : 'character' } as any ,
type : NounType . Person
} )
const result = await brain . migrateField ( { from : 'data.kind' , to : 'subtype' } )
expect ( result . migrated ) . toBe ( 1 )
expect ( result . errors ) . toEqual ( [ ] )
const entity = await brain . get ( id )
expect ( entity ! . subtype ) . toBe ( 'character' )
expect ( ( entity ! . data as any ) . description ) . toBe ( 'thing' )
expect ( ( entity ! . data as any ) . kind ) . toBeUndefined ( )
} )
it ( 'throws on identical from/to' , async ( ) = > {
await expect (
brain . migrateField ( { from : 'subtype' , to : 'subtype' } )
) . rejects . toThrow ( /identical/ )
} )
it ( 'rejects unsupported path prefixes' , async ( ) = > {
await expect (
brain . migrateField ( { from : 'metadata.kind' , to : 'connections.foo' } )
) . rejects . toThrow ( /unsupported path prefix/ )
} )
it ( 'migrates with the rollup tracking the new subtype counts' , async ( ) = > {
await brain . add ( {
data : 'a' ,
type : NounType . Person ,
metadata : { kind : 'employee' }
} )
await brain . add ( {
data : 'b' ,
type : NounType . Person ,
metadata : { kind : 'customer' }
} )
await brain . migrateField ( { from : 'metadata.kind' , to : 'subtype' } )
// Rollup should now reflect the new subtype values
expect ( brain . counts . bySubtype ( NounType . Person ) ) . toEqual ( {
employee : 1 ,
customer : 1
} )
} )
} )
} )