feat: subtype top-level field + trackField + migrateField
Promotes `subtype?: string` to a top-level standard field on every entity,
alongside `type` / `confidence` / `weight`. Flat string, no hierarchy — the
consumer-chosen vocabulary for sub-classifying entities within a NounType
(Person → employee/customer, Document → invoice/contract, etc.).
Layer 1 — subtype field + rollup
- HNSWNounWithMetadata.subtype + STANDARD_ENTITY_FIELDS entry
- Entity / Result / AddParams / UpdateParams / FindParams threading
- add()/update() persist subtype on storageMetadata + entityForIndexing
- get()/find() route through the standard-field fast path
- subtypeCountsByType (Map<NounTypeIdx, Map<subtype, count>>) on
BaseStorage, mirrored after nounCountsByType with the same self-heal
rebuild and persisted to _system/subtype-statistics.json
- brain.counts.bySubtype(type, subtype?) — O(1) point + breakdown
- brain.counts.topSubtypes(type, n) — top-N by count
- brain.subtypesOf(type) — distinct subtypes seen
- find({ type, subtype }) and find({ subtype: ['a','b'] }) on the fast path
Layer 2 — trackField for other facets
- brain.trackField(name, { perType?, values? }) registers a field for
cardinality + per-NounType breakdown stats. Backed by the aggregation
engine (auto-defines __fieldCounts__<name>), backfill-on-define applies.
- brain.counts.byField(name, { type? }) returns value frequencies
- Optional vocabulary whitelist rejects off-vocabulary writes at add/update
Layer 3 — generic migrateField
- brain.migrateField({ from, to, readBoth?, batchSize?, onProgress? })
streams every entity, copies the value from one path to another, and
(unless readBoth) clears the source. Supports top-level standard fields,
metadata.X, and data.X paths. Idempotent — safe to re-run.
Docs
- New guide: docs/guides/subtypes-and-facets.md (Layer 1 + 2 + 3)
- README, DATA_MODEL, QUERY_OPERATORS, api/README, finite-type-system,
quick-start all treat subtype as a core primitive with anonymous example
vocabularies (employee/customer/invoice/milestone).
Tests
- 26 new integration tests covering write/read/update/delete round-trips,
counts rollup decrement + re-route on mutation, trackField + byField
with and without perType, vocabulary whitelist enforcement, and
migrateField for metadata.X → subtype and data.X → subtype paths
including readBoth deprecation-window semantics.
Unit suite: 1468/1468 passing. Type-check + build clean.
This commit is contained in:
parent
e47fea0917
commit
2cdf70ee0f
14 changed files with 1698 additions and 21 deletions
|
|
@ -186,6 +186,7 @@ When you add an entity, Brainy stores these standard fields in the metadata obje
|
|||
| Field | Set By | Description |
|
||||
|-------|--------|-------------|
|
||||
| `noun` | System | Entity type (NounType enum value) |
|
||||
| `subtype` | User | Per-NounType sub-classification (e.g. `'employee'`, `'invoice'`, `'milestone'`). Flat string, no hierarchy. Indexed on the fast path and rolled into per-NounType statistics. |
|
||||
| `data` | System | The raw `data` value (stored opaquely) |
|
||||
| `createdAt` | System | Creation timestamp |
|
||||
| `updatedAt` | System | Last update timestamp |
|
||||
|
|
@ -196,6 +197,30 @@ When you add an entity, Brainy stores these standard fields in the metadata obje
|
|||
|
||||
On read, these standard fields are extracted to top-level Entity properties. The `metadata` field on the returned Entity contains **only your custom fields**.
|
||||
|
||||
### Subtype — sub-classification within a NounType
|
||||
|
||||
`type` (NounType) is a stable 42-value enum. `subtype` is the consumer-chosen string vocabulary *within* a type:
|
||||
|
||||
```typescript
|
||||
// A Person who is an employee:
|
||||
await brain.add({
|
||||
data: 'Avery Brooks — runs the AI lab',
|
||||
type: NounType.Person,
|
||||
subtype: 'employee',
|
||||
metadata: { department: 'ai-lab' }
|
||||
})
|
||||
|
||||
// A Document that is an invoice:
|
||||
await brain.add({
|
||||
data: 'INV-2026-001',
|
||||
type: NounType.Document,
|
||||
subtype: 'invoice',
|
||||
metadata: { amount: 1500 }
|
||||
})
|
||||
```
|
||||
|
||||
`subtype` lives at the **top level** — NOT inside `metadata`, NOT inside `data`. That's how `find({ type, subtype })` routes through the standard-field fast path (column-store hit) instead of the metadata fallback. See **[Subtypes & Facets](./guides/subtypes-and-facets.md)** for the full guide including `trackField()` and `migrateField()`.
|
||||
|
||||
---
|
||||
|
||||
## See Also
|
||||
|
|
|
|||
|
|
@ -194,6 +194,26 @@ brain.find({ where: { noun: NounType.Person } })
|
|||
brain.find({ type: [NounType.Person, NounType.Agent] })
|
||||
```
|
||||
|
||||
### Filter by subtype
|
||||
|
||||
`subtype` is a top-level standard field — takes the column-store fast path, not the metadata fallback. Pair with `type` for the typical "Person who is an employee" query:
|
||||
|
||||
```typescript
|
||||
// Equality on subtype:
|
||||
brain.find({ type: NounType.Person, subtype: 'employee' })
|
||||
|
||||
// Set membership:
|
||||
brain.find({ type: NounType.Person, subtype: ['employee', 'contractor'] })
|
||||
|
||||
// Operator-form predicates use `where`:
|
||||
brain.find({
|
||||
type: NounType.Person,
|
||||
where: { subtype: { exists: true } }
|
||||
})
|
||||
```
|
||||
|
||||
See the **[Subtypes & Facets guide](./guides/subtypes-and-facets.md)** for the full surface.
|
||||
|
||||
### Combine semantic search with filters
|
||||
|
||||
```typescript
|
||||
|
|
|
|||
|
|
@ -109,11 +109,12 @@ Add a single entity to the database.
|
|||
|
||||
```typescript
|
||||
const id = await brain.add({
|
||||
data: 'JavaScript is a programming language', // Text or pre-computed vector
|
||||
type: NounType.Concept, // Required: Entity type
|
||||
metadata: { // Optional metadata
|
||||
category: 'programming',
|
||||
year: 1995
|
||||
data: 'JavaScript is a programming language', // Text or pre-computed vector
|
||||
type: NounType.Concept, // Required: Entity type
|
||||
subtype: 'language', // Optional: sub-classification
|
||||
metadata: { // Optional: queryable fields
|
||||
category: 'programming',
|
||||
year: 1995
|
||||
}
|
||||
})
|
||||
```
|
||||
|
|
@ -121,6 +122,7 @@ const id = await brain.add({
|
|||
**Parameters:**
|
||||
- `data`: `string | number[]` - Content to embed (text auto-embeds) or pre-computed vector
|
||||
- `type`: `NounType` - Entity type (required)
|
||||
- `subtype?`: `string` - Per-product sub-classification within the NounType (top-level standard field, indexed on the fast path). See [Subtypes & Facets](../guides/subtypes-and-facets.md).
|
||||
- `metadata?`: `object` - Structured queryable fields (indexed by MetadataIndex, used in `where` filters)
|
||||
- `id?`: `string` - Custom ID (auto-generated UUID if not provided)
|
||||
- `vector?`: `number[]` - Pre-computed vector (skips auto-embedding)
|
||||
|
|
@ -158,15 +160,20 @@ Update an existing entity.
|
|||
```typescript
|
||||
await brain.update({
|
||||
id: entityId,
|
||||
data: 'Updated content', // Optional: new data
|
||||
metadata: { updated: true } // Optional: new metadata (merges)
|
||||
data: 'Updated content', // Optional: new data
|
||||
subtype: 'archived', // Optional: change sub-classification
|
||||
metadata: { updated: true } // Optional: new metadata (merges)
|
||||
})
|
||||
```
|
||||
|
||||
**Parameters:**
|
||||
- `id`: `string` - Entity ID
|
||||
- `data?`: `string | number[]` - New data/vector
|
||||
- `metadata?`: `object` - Metadata to merge
|
||||
- `type?`: `NounType` - Change entity type
|
||||
- `subtype?`: `string` - Change subtype (omit to preserve existing)
|
||||
- `metadata?`: `object` - Metadata to merge (or replace with `merge: false`)
|
||||
- `confidence?`: `number` - Update classification confidence
|
||||
- `weight?`: `number` - Update entity importance
|
||||
|
||||
**Returns:** `Promise<void>`
|
||||
|
||||
|
|
@ -221,6 +228,7 @@ const results = await brain.find({
|
|||
**FindParams:**
|
||||
- `query?`: `string` - Text for semantic + hybrid search (searches `data` via HNSW + text index)
|
||||
- `type?`: `NounType | NounType[]` - Filter by entity type(s). Alias for `where.noun`.
|
||||
- `subtype?`: `string | string[]` - Filter by sub-classification (top-level standard field, fast path). Single string for equality, array for set membership.
|
||||
- `where?`: `object` - Metadata filters. See **[Query Operators](../QUERY_OPERATORS.md)** for all operators.
|
||||
- `connected?`: `object` - Graph traversal options
|
||||
- `to?`: `string` - Target entity ID
|
||||
|
|
@ -1928,6 +1936,86 @@ const count = await brain.getVerbCount()
|
|||
|
||||
---
|
||||
|
||||
### Subtype & facet APIs
|
||||
|
||||
Full guide: **[Subtypes & Facets](../guides/subtypes-and-facets.md)**.
|
||||
|
||||
#### `counts.bySubtype(type, subtype?)` → `Record<string, number> | number`
|
||||
|
||||
O(1) subtype counts for a NounType (backed by the persisted rollup).
|
||||
|
||||
```typescript
|
||||
brain.counts.bySubtype(NounType.Person)
|
||||
// → { employee: 12, customer: 847, vendor: 34 }
|
||||
|
||||
brain.counts.bySubtype(NounType.Person, 'employee')
|
||||
// → 12
|
||||
```
|
||||
|
||||
#### `counts.topSubtypes(type, n=10)` → `Array<[subtype, count]>`
|
||||
|
||||
Top N subtypes ranked by count.
|
||||
|
||||
```typescript
|
||||
brain.counts.topSubtypes(NounType.Person, 3)
|
||||
// → [['customer', 847], ['employee', 12], ['vendor', 34]]
|
||||
```
|
||||
|
||||
#### `subtypesOf(type)` → `string[]`
|
||||
|
||||
Sorted distinct subtypes seen for a NounType.
|
||||
|
||||
```typescript
|
||||
brain.subtypesOf(NounType.Person)
|
||||
// → ['customer', 'employee', 'vendor']
|
||||
```
|
||||
|
||||
#### `trackField(name, options?)` → `void`
|
||||
|
||||
Register a metadata field for cardinality + per-NounType breakdown stats. With `values: [...]`, validates against the whitelist on `add()`/`update()`.
|
||||
|
||||
```typescript
|
||||
brain.trackField('status') // basic
|
||||
brain.trackField('status', { perType: true }) // with per-NounType breakdown
|
||||
brain.trackField('priority', { values: ['low', 'med', 'high'] }) // strict vocabulary
|
||||
```
|
||||
|
||||
#### `counts.byField(name, options?)` → `Promise<Record<string, number>>`
|
||||
|
||||
Counts by value for a tracked field. Requires `perType: true` registration if filtering by NounType.
|
||||
|
||||
```typescript
|
||||
await brain.counts.byField('status')
|
||||
// → { todo: 12, doing: 3, done: 47 }
|
||||
|
||||
await brain.counts.byField('status', { type: NounType.Task })
|
||||
// → { todo: 8, doing: 2, done: 30 }
|
||||
```
|
||||
|
||||
#### `migrateField(options)` → `Promise<MigrationSummary>`
|
||||
|
||||
Stream-and-rewrite a field across the brain. Supports `metadata.X`, `data.X`, and top-level paths. Idempotent.
|
||||
|
||||
```typescript
|
||||
// One-shot rewrite
|
||||
await brain.migrateField({ from: 'metadata.kind', to: 'subtype' })
|
||||
|
||||
// Deprecation window — keep source field readable
|
||||
await brain.migrateField({ from: 'data.kind', to: 'subtype', readBoth: true })
|
||||
|
||||
// With progress reporting
|
||||
await brain.migrateField({
|
||||
from: 'metadata.kind',
|
||||
to: 'subtype',
|
||||
batchSize: 500,
|
||||
onProgress: ({ scanned, migrated }) => console.log(`${scanned} / ${migrated}`)
|
||||
})
|
||||
```
|
||||
|
||||
Returns `{ scanned: number, migrated: number, skipped: number, errors: Array<{id, error}> }`.
|
||||
|
||||
---
|
||||
|
||||
### `embed(data)` → `Promise<number[]>` ✨
|
||||
|
||||
Generate embedding vector from text or data.
|
||||
|
|
|
|||
|
|
@ -449,6 +449,26 @@ brain.registerNounType('chemical_compound', {
|
|||
})
|
||||
```
|
||||
|
||||
### 1a. Subtypes — sub-classification without hierarchy
|
||||
|
||||
The 42-type taxonomy is intentionally coarse. Per-product vocabulary fits on the **`subtype`** axis — a top-level standard string field on every entity. Flat by design — no hierarchy, no parent chain, no recursive resolution. That preserves the Uint32Array-backed O(1) type stats while giving consumers a place to put `'employee'` / `'customer'` / `'invoice'` / `'milestone'` without burning a slot in the global enum.
|
||||
|
||||
```typescript
|
||||
// Same NounType, different subtypes:
|
||||
await brain.add({ type: NounType.Person, subtype: 'employee' })
|
||||
await brain.add({ type: NounType.Person, subtype: 'customer' })
|
||||
await brain.add({ type: NounType.Document, subtype: 'invoice' })
|
||||
|
||||
// Fast path — column-store hit, not metadata fallback:
|
||||
await brain.find({ type: NounType.Person, subtype: 'employee' })
|
||||
|
||||
// Per-NounType-per-subtype counts maintained incrementally:
|
||||
brain.counts.bySubtype(NounType.Person)
|
||||
// → { employee: 12, customer: 847 }
|
||||
```
|
||||
|
||||
Subtype has its own statistics rollup (`_system/subtype-statistics.json`) maintained alongside `nounCountsByType`, so per-subtype counts stay O(1) at billion scale. Full guide: **[Subtypes & Facets](../guides/subtypes-and-facets.md)**.
|
||||
|
||||
### 2. Semantic not Structural
|
||||
|
||||
```typescript
|
||||
|
|
|
|||
|
|
@ -39,16 +39,20 @@ That's it. Brainy auto-configures storage, loads the embedding model, and builds
|
|||
const reactId: string = await brain.add({
|
||||
data: 'React is a JavaScript library for building user interfaces',
|
||||
type: NounType.Concept,
|
||||
subtype: 'library', // Sub-classification within Concept
|
||||
metadata: { category: 'frontend', year: 2013 }
|
||||
})
|
||||
|
||||
const nextId: string = await brain.add({
|
||||
data: 'Next.js framework for React with server-side rendering',
|
||||
type: NounType.Concept,
|
||||
subtype: 'framework',
|
||||
metadata: { category: 'framework', year: 2016 }
|
||||
})
|
||||
```
|
||||
|
||||
`type` is one of Brainy's 42 stable NounTypes. `subtype` is your free-form sub-classification within that type — flat string, no hierarchy, indexed on the fast path. See **[Subtypes & Facets](./subtypes-and-facets.md)** for the full guide.
|
||||
|
||||
## 4. Create Relationships
|
||||
|
||||
```typescript
|
||||
|
|
|
|||
285
docs/guides/subtypes-and-facets.md
Normal file
285
docs/guides/subtypes-and-facets.md
Normal file
|
|
@ -0,0 +1,285 @@
|
|||
---
|
||||
title: Subtypes & Facets
|
||||
slug: guides/subtypes-and-facets
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 7
|
||||
description: Use the top-level `subtype` field to sub-classify entities within a NounType, track other metadata facets like status or role, and migrate field names without downtime.
|
||||
next:
|
||||
- guides/aggregation
|
||||
- api/reference
|
||||
---
|
||||
|
||||
# Subtypes & Facets
|
||||
|
||||
> Sub-classify entities within a NounType, track arbitrary metadata facets, and migrate field names without downtime.
|
||||
|
||||
## Why this exists
|
||||
|
||||
Brainy's `NounType` (Person, Document, Event, Concept, Task, …) is a stable, flat 42-type taxonomy. It's deliberately coarse — every product needs a way to further classify entities *within* a type. Is this `Person` an employee or a customer? Is this `Document` an invoice or a contract? Is this `Event` a meeting or a milestone?
|
||||
|
||||
Three layers solve this:
|
||||
|
||||
| Layer | Use it for | Shape |
|
||||
|---|---|---|
|
||||
| **`subtype`** | The primary sub-classification of an entity, one value per entity | Top-level standard field |
|
||||
| **`trackField()`** | Other facets you want to count or filter on (`status`, `source`, `role`, `paradigm`) | Registered metadata field |
|
||||
| **`migrateField()`** | Renaming or restructuring fields across an existing dataset | One-shot stream-and-rewrite |
|
||||
|
||||
## Layer 1 — `subtype`
|
||||
|
||||
`subtype` is a top-level standard field on every entity, alongside `type` / `confidence` / `weight`. Flat string, no hierarchy. The vocabulary is your choice — Brainy stores and counts, never validates.
|
||||
|
||||
### Write
|
||||
|
||||
```typescript
|
||||
import { Brainy, NounType } from '@soulcraft/brainy'
|
||||
|
||||
const brain = new Brainy()
|
||||
await brain.init()
|
||||
|
||||
await brain.add({
|
||||
data: 'Avery Brooks — runs the AI lab',
|
||||
type: NounType.Person,
|
||||
subtype: 'employee', // top-level write param
|
||||
metadata: { department: 'ai-lab' }
|
||||
})
|
||||
```
|
||||
|
||||
### Read
|
||||
|
||||
```typescript
|
||||
// Top-level filter — standard-field fast path, no `where` wrapper:
|
||||
const employees = await brain.find({ type: NounType.Person, subtype: 'employee' })
|
||||
|
||||
// Set membership:
|
||||
const internal = await brain.find({
|
||||
type: NounType.Person,
|
||||
subtype: ['employee', 'contractor']
|
||||
})
|
||||
|
||||
// Operator-form predicates use `where`:
|
||||
const typed = await brain.find({
|
||||
type: NounType.Person,
|
||||
where: { subtype: { exists: true } }
|
||||
})
|
||||
```
|
||||
|
||||
### What it looks like
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "01HZK3M7TW...",
|
||||
"type": "person",
|
||||
"subtype": "employee",
|
||||
"data": "Avery Brooks — runs the AI lab",
|
||||
"metadata": { "department": "ai-lab" }
|
||||
}
|
||||
```
|
||||
|
||||
`subtype` lives at the **top level** — NOT inside `metadata`, NOT inside `data`. That's what makes the fast path possible: queries on `subtype` hit the column-store index directly, never the metadata fallback.
|
||||
|
||||
### Counts (O(1))
|
||||
|
||||
```typescript
|
||||
// All subtypes for a NounType
|
||||
brain.counts.bySubtype(NounType.Person)
|
||||
// → { employee: 12, customer: 847, vendor: 34 }
|
||||
|
||||
// Point count
|
||||
brain.counts.bySubtype(NounType.Person, 'employee')
|
||||
// → 12
|
||||
|
||||
// Top N
|
||||
brain.counts.topSubtypes(NounType.Person, 3)
|
||||
// → [['customer', 847], ['employee', 12], ['vendor', 34]]
|
||||
|
||||
// Distinct subtypes for a NounType
|
||||
brain.subtypesOf(NounType.Person)
|
||||
// → ['customer', 'employee', 'vendor']
|
||||
```
|
||||
|
||||
These are O(1) lookups backed by `_system/subtype-statistics.json` — no scan, no storage round-trip. The rollup is incrementally maintained as entities are added, updated, and deleted.
|
||||
|
||||
### Aggregation
|
||||
|
||||
`subtype` is a first-class group-by dimension:
|
||||
|
||||
```typescript
|
||||
const rows = await brain.find({
|
||||
aggregate: {
|
||||
groupBy: ['type', 'subtype'],
|
||||
metrics: { count: { op: 'COUNT' } }
|
||||
}
|
||||
})
|
||||
// [
|
||||
// { groupKey: { type: 'person', subtype: 'customer' }, count: 847 },
|
||||
// { groupKey: { type: 'person', subtype: 'employee' }, count: 12 },
|
||||
// { groupKey: { type: 'document', subtype: 'invoice' }, count: 2103 },
|
||||
// ...
|
||||
// ]
|
||||
```
|
||||
|
||||
## Layer 2 — `trackField()`
|
||||
|
||||
Some metadata fields aren't *the* sub-classification of an entity but are still worth counting: `status`, `source`, `role`, `paradigm`. Promoting each one to top-level would clutter the contract. `trackField()` registers a metadata field for cardinality + per-NounType breakdown stats without the contract growth.
|
||||
|
||||
### Register and query
|
||||
|
||||
```typescript
|
||||
// Track a single facet
|
||||
brain.trackField('status')
|
||||
|
||||
await brain.add({ data: 'Ship subtype', type: NounType.Task, metadata: { status: 'todo' } })
|
||||
await brain.add({ data: 'Write docs', type: NounType.Task, metadata: { status: 'done' } })
|
||||
|
||||
await brain.counts.byField('status')
|
||||
// → { todo: 1, done: 1 }
|
||||
```
|
||||
|
||||
### Per-NounType breakdown
|
||||
|
||||
```typescript
|
||||
brain.trackField('status', { perType: true })
|
||||
|
||||
await brain.counts.byField('status', { type: NounType.Task })
|
||||
// → { todo: 1, done: 1 }
|
||||
```
|
||||
|
||||
### Vocabulary whitelist (opt-in validation)
|
||||
|
||||
```typescript
|
||||
brain.trackField('priority', { values: ['low', 'medium', 'high'] })
|
||||
|
||||
// Throws — 'urgent' isn't in the vocabulary:
|
||||
await brain.add({
|
||||
data: 'Fix bug',
|
||||
type: NounType.Task,
|
||||
metadata: { priority: 'urgent' }
|
||||
})
|
||||
```
|
||||
|
||||
`trackField` piggybacks on the existing aggregation engine (see the [Aggregation guide](./aggregation.md) for the underlying mechanism). Backfill-on-define means the first call to `counts.byField()` scans existing entities; subsequent calls are O(groups).
|
||||
|
||||
### subtype vs trackField — when to use which
|
||||
|
||||
- **`subtype`** when there's one primary sub-classification per entity. Limit yourself to one per NounType. Examples: Person→employee/customer/vendor; Document→invoice/contract/policy.
|
||||
- **`trackField`** for anything else you want to count or filter on. No limit on how many you register. Examples: status, source, role, paradigm.
|
||||
|
||||
## Layer 3 — `migrateField()`
|
||||
|
||||
Use this when you need to rename or restructure a field across an entire dataset — for example, moving a `metadata.kind` convention up to the top-level `subtype` standard field.
|
||||
|
||||
### One-shot rewrite
|
||||
|
||||
```typescript
|
||||
// Starting state: every entity has metadata.kind
|
||||
const result = await brain.migrateField({
|
||||
from: 'metadata.kind',
|
||||
to: 'subtype'
|
||||
})
|
||||
|
||||
console.log(result)
|
||||
// {
|
||||
// scanned: 1500,
|
||||
// migrated: 1500,
|
||||
// skipped: 0,
|
||||
// errors: []
|
||||
// }
|
||||
```
|
||||
|
||||
After this returns, every entity has `subtype` populated from the old `metadata.kind` value, and `metadata.kind` is cleared.
|
||||
|
||||
### Deprecation window — keep both fields readable
|
||||
|
||||
When you can't coordinate all readers and the migration in a single deploy, use `readBoth: true` to preserve the source field alongside the new one:
|
||||
|
||||
```typescript
|
||||
// Phase 1: dual-populate (existing readers still work against metadata.kind):
|
||||
await brain.migrateField({
|
||||
from: 'metadata.kind',
|
||||
to: 'subtype',
|
||||
readBoth: true
|
||||
})
|
||||
|
||||
// ... readers migrate to query subtype at their own pace ...
|
||||
|
||||
// Phase 2: clear the source field when ready:
|
||||
await brain.migrateField({ from: 'metadata.kind', to: 'subtype' })
|
||||
```
|
||||
|
||||
### Supported paths
|
||||
|
||||
| Path form | Refers to |
|
||||
|---|---|
|
||||
| `'subtype'`, `'type'`, `'confidence'` | Top-level standard fields |
|
||||
| `'metadata.X'` | A key under `entity.metadata` |
|
||||
| `'data.X'` | A key under `entity.data` (when `data` is an object) |
|
||||
| `'X'` (bare, non-standard) | Shorthand for `metadata.X` |
|
||||
|
||||
### Idempotent
|
||||
|
||||
`migrateField` is safe to re-run. Entities where the source is absent, or where the destination already holds the same value, are skipped. This makes it safe to use in a deploy-once-then-cleanup workflow.
|
||||
|
||||
### Progress reporting
|
||||
|
||||
```typescript
|
||||
await brain.migrateField({
|
||||
from: 'metadata.kind',
|
||||
to: 'subtype',
|
||||
batchSize: 500,
|
||||
onProgress: ({ scanned, migrated }) => {
|
||||
console.log(`${scanned} scanned, ${migrated} migrated`)
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
## Putting it together
|
||||
|
||||
A realistic adoption sequence for a brain that started without these primitives:
|
||||
|
||||
```typescript
|
||||
import { Brainy, NounType } from '@soulcraft/brainy'
|
||||
|
||||
const brain = new Brainy({ storage: { type: 'filesystem', options: { path: './brain-data' } } })
|
||||
await brain.init()
|
||||
|
||||
// 1. Migrate any existing metadata.kind convention to the new top-level subtype
|
||||
await brain.migrateField({ from: 'metadata.kind', to: 'subtype', readBoth: true })
|
||||
|
||||
// 2. Register the other facets you want counted
|
||||
brain.trackField('status', { perType: true })
|
||||
brain.trackField('source')
|
||||
|
||||
// 3. Use subtype on every new write
|
||||
await brain.add({
|
||||
data: 'Quarterly review',
|
||||
type: NounType.Event,
|
||||
subtype: 'milestone',
|
||||
metadata: { status: 'todo', source: 'planning-session' }
|
||||
})
|
||||
|
||||
// 4. Query the breakdowns
|
||||
brain.counts.bySubtype(NounType.Event)
|
||||
// → { milestone: 14, meeting: 203, deadline: 7 }
|
||||
|
||||
await brain.counts.byField('status', { type: NounType.Event })
|
||||
// → { todo: 12, done: 212 }
|
||||
|
||||
// 5. Once all readers are on the new field, drop the source:
|
||||
await brain.migrateField({ from: 'metadata.kind', to: 'subtype' })
|
||||
```
|
||||
|
||||
## Reference
|
||||
|
||||
- `brain.add({ ..., subtype: 'value' })` — write the field
|
||||
- `brain.update({ id, subtype: 'value' })` — change the field
|
||||
- `brain.find({ type, subtype })` — filter (fast path)
|
||||
- `brain.find({ subtype: ['a', 'b'] })` — set membership
|
||||
- `brain.counts.bySubtype(type, subtype?)` — O(1) counts
|
||||
- `brain.counts.topSubtypes(type, n?)` — top N by count
|
||||
- `brain.subtypesOf(type)` — distinct subtype list
|
||||
- `brain.trackField(name, { perType?, values? })` — register a facet
|
||||
- `brain.counts.byField(name, { type? })` — facet counts
|
||||
- `brain.migrateField({ from, to, readBoth?, batchSize?, onProgress? })` — rewrite a field
|
||||
Loading…
Add table
Add a link
Reference in a new issue