feat(find): field projection — fields resolve from the column store, not the record

A list view that shows a title and a slug hydrates the whole record for every
row, document bodies included, and discards almost all of it. find/get({ fields })
names what is wanted; the column store serves it; the canonical record is opened
only for fields the index cannot supply.

The provider grows an optional getScalarsForIds(ids, fields) door, batched: it
walks each column ONCE and picks out every requested id, rather than re-walking
per row. The column store grows the primitive that was missing — valuesForIds —
because every other read door there answers which entities have a value, and a
projection asks the opposite.

It reads the COLUMN store, never the sparse index: the column keeps raw values,
the sparse index keeps a bucketed form built for range queries, and a projection
served from the latter would return a value that differs from the record's. A
field the column cannot serve is omitted rather than approximated — omission
costs a read, a wrong value is a wrong answer nobody can see.

Two laws the pins hold: fields absent is byte-identical to today, and a missing
field is simply absent rather than an error — so this path deliberately avoids
the strict address resolver, whose UnresolvableFieldError is right for orderBy
and wrong here.

related() takes no fields: a Relation carries from/to as ids and hydrates no
record, so the param would be decorative.
This commit is contained in:
David Snelling 2026-09-02 13:36:33 -07:00
parent 6597c146f7
commit ad0f493f7a
7 changed files with 639 additions and 7 deletions

View file

@ -4193,6 +4193,16 @@ export class Brainy<T = any> implements BrainyInterface<T> {
}
// Route to metadata-only or full entity based on options
// A PROJECTED get goes through the same seam every list page uses, so a
// detail read of two scalars costs an index read rather than a record read.
// It is checked before `includeVectors` because the two are incompatible by
// construction: a projection returns the named fields, and a vector is not
// one of them unless it was named.
if (options?.fields !== undefined && options.fields.length > 0) {
const page = await this.hydratePage([id], options.fields)
return page.get(id) ?? null
}
const includeVectors = options?.includeVectors ?? false // Default: metadata-only (fast)
if (includeVectors) {
@ -4239,6 +4249,170 @@ export class Brainy<T = any> implements BrainyInterface<T> {
* const children = childIds.map(id => childrenMap.get(id)).filter(Boolean)
* ```
*/
/**
* **The projection seam** hydrate a page of ids under an optional `fields`
* projection, opening the canonical record only when the index cannot serve
* what was asked for.
*
* Without a projection this is exactly `batchGet`, byte for byte: the whole
* point is that `fields` absent changes nothing.
*
* With one, the order is: ask the index for the named scalars in a single
* batched door; see which requested fields it actually served; and read
* records ONLY if something is still missing and only to fill those fields.
* A page whose every requested field is index-served performs zero canonical
* reads, which is the whole reason the door exists.
*
* `guardFields` are fetched ALONGSIDE the projection and trimmed off before
* the caller sees them. find()'s index-integrity guard re-validates every row
* against its own predicate, and it reads the entity to do so so a row
* projected down to `title` would fail a `where: { kind }` it genuinely
* matches, and the whole page would vanish. The fields a filter names are
* fields the index can serve by definition, so carrying them costs nothing
* and keeps the guard honest.
*
* A field nothing can supply is simply absent from the row. That is the
* permissive law: a projection asks "these, if you have them", and an
* optional field must not turn a list into an exception. It deliberately does
* NOT route through the strict address resolver, which throws
* `UnresolvableFieldError` for an unknown key that strictness is right for
* `orderBy`, where a typo silently changes the order, and wrong here, where
* the honest answer is "this row does not have that".
*
* @param ids - Canonical ids for the page.
* @param fields - The projection, or undefined for the full record.
* @returns `id → entity`, projected when `fields` was given.
*/
/**
* The index keys find()'s integrity guard reads when it re-validates a row.
*
* The guard calls `entityMatchesFind(entity, params)`, so a projected entity
* must still carry whatever the params constrain otherwise a row that
* genuinely matches is dropped for lacking the evidence. These are fetched
* with the projection and trimmed off before the caller sees them.
*
* @param params - The find params.
* @returns Index keys to carry through hydration.
*/
private guardFieldsFor(params: FindParams<T>): string[] {
const keys: string[] = []
if (params.where && typeof params.where === 'object') {
// Top-level where keys only: nested `anyOf`/`allOf` branches are carried
// by their own keys when the guard walks them, and a filter whose
// evidence is missing keeps the row (the guard's own catch) rather than
// dropping it.
for (const key of Object.keys(params.where as Record<string, unknown>)) {
if (key === 'anyOf' || key === 'allOf' || key === 'not') continue
keys.push(key)
}
}
if (params.type !== undefined) keys.push('system.type')
if (params.subtype !== undefined) keys.push('system.subtype')
if (params.service !== undefined) keys.push('system.service')
if (params.excludeVFS === true) keys.push('vfsType', 'isVFSEntity')
return keys
}
private async hydratePage(
ids: string[],
fields?: readonly string[],
guardFields: readonly string[] = []
): Promise<Map<string, Entity<T>>> {
if (fields === undefined || fields.length === 0) return this.batchGet(ids)
const wanted = [...new Set([...fields, ...guardFields])]
const provider = this.metadataIndex as unknown as MetadataIndexProvider
let served = new Map<string, Record<string, unknown>>()
if (typeof provider.getScalarsForIds === 'function') {
served = await provider.getScalarsForIds(ids, wanted)
}
// Which ids still owe a field? Only those cost a record read, and a page
// that owes nothing costs none at all.
const owing: string[] = []
for (const id of ids) {
const row = served.get(id)
if (row === undefined || wanted.some((f) => !(f in row))) owing.push(id)
}
// The records are read for the OWED fields only; everything the index
// already served is used as-is, so a body field pulls its own record and
// no more than that.
const records = owing.length > 0 ? await this.batchGet(owing) : new Map<string, Entity<T>>()
const out = new Map<string, Entity<T>>()
for (const id of ids) {
const fromIndex = served.get(id)
const record = records.get(id)
// An id neither the index nor storage knows is not a row.
if (fromIndex === undefined && record === undefined) continue
out.set(id, this.projectEntity(id, wanted, fromIndex, record))
}
return out
}
/**
* Build one projected entity: `id`, plus exactly the requested fields that
* something could supply.
*
* Values come from the index first and the record second, and they must agree
* the index only reports what it can serve exactly, so a field it served is
* the record's value. A field neither has is omitted rather than set to
* `undefined`: absent and present-and-undefined are different answers, and a
* caller checking `'slug' in row.metadata` deserves the true one.
*
* @param id - The entity id, always present on the result.
* @param fields - The requested index keys.
* @param fromIndex - What the index served for this id, if anything.
* @param record - The canonical entity, if one had to be read.
* @returns The projected entity.
*/
private projectEntity(
id: string,
fields: readonly string[],
fromIndex: Record<string, unknown> | undefined,
record: Entity<T> | undefined
): Entity<T> {
const projected: Record<string, unknown> = { id }
const metadata: Record<string, unknown> = {}
let sawMetadata = false
for (const field of fields) {
let value: unknown
let found = false
if (fromIndex !== undefined && field in fromIndex) {
value = fromIndex[field]
found = true
} else if (record !== undefined) {
if (field.startsWith('system.')) {
const inner = field.slice('system.'.length)
const bag = record as unknown as Record<string, unknown>
if (inner in bag && bag[inner] !== undefined) {
value = bag[inner]
found = true
}
} else {
const bag = (record.metadata ?? {}) as Record<string, unknown>
if (field in bag) {
value = bag[field]
found = true
}
}
}
if (!found) continue
if (field.startsWith('system.')) {
projected[field.slice('system.'.length)] = value
} else {
metadata[field] = value
sawMetadata = true
}
}
if (sawMetadata) projected.metadata = metadata
return projected as unknown as Entity<T>
}
async batchGet(ids: string[], options?: GetOptions): Promise<Map<string, Entity<T>>> {
// Canonical read (see get): resolves by id from storage, no derived index.
await this.ensureInitialized({ needs: [] })
@ -8037,7 +8211,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
// Batch-load entities for 10x faster cloud storage performance
// GCS: 10 entities = 1×50ms vs 10×50ms = 500ms (10x faster)
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
for (const id of pageIds) {
const entity = entitiesMap.get(id)
if (entity) {
@ -8074,7 +8248,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
if (hiddenIds.size > 0) allUuids = allUuids.filter((id) => !hiddenIds.has(id))
const pageIds = allUuids.slice(offset, offset + limit)
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
for (const id of pageIds) {
const entity = entitiesMap.get(id)
if (entity) {
@ -8102,7 +8276,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
const pageIds = filteredIds.slice(offset, offset + limit)
// Batch-load entities for 10x faster cloud storage performance
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
for (const id of pageIds) {
const entity = entitiesMap.get(id)
if (entity) {
@ -8337,7 +8511,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
// Batch-load entities for current page - O(page_size) instead of O(total_results)
// GCS: 10 entities = 1×50ms vs 10×50ms = 500ms (10x faster)
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
for (const id of pageIds) {
const entity = entitiesMap.get(id)
if (entity) {
@ -8365,7 +8539,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
// Batch-load entities for paginated results (10x faster on GCS)
const sortedResults: Result<T>[] = []
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
for (const id of pageIds) {
const entity = entitiesMap.get(id)
if (entity) {
@ -8470,6 +8644,28 @@ export class Brainy<T = any> implements BrainyInterface<T> {
})
}
// PROJECTION TRIM — applied once, here, AFTER the integrity guard, so every
// find() path is trimmed uniformly and the guard still saw the evidence it
// needs. Hydration carried the guard's fields alongside the projection;
// this removes them, leaving exactly what the caller named.
//
// Rows that reached here from a path the seam does not hydrate (a vector or
// text leg builds its own entities) are trimmed from what they already
// hold, so the ANSWER is the same everywhere — only the cost differs, and
// only on the paths that still read a record.
if (params.fields !== undefined && params.fields.length > 0 && result.length > 0) {
const named = [...new Set(params.fields)]
result = result.map((r) => {
const projected = this.projectEntity(
r.id,
named,
undefined,
r.entity as unknown as Entity<T>
)
return { ...r, entity: projected } as typeof r
})
}
// includeVectors — opt-in vector hydration. Default (false) keeps the perf
// contract: every result path above builds entities via the metadata-only
// fast path, so `entity.vector` is the empty stub. When requested, fetch the
@ -16842,7 +17038,7 @@ export class Brainy<T = any> implements BrainyInterface<T> {
ordered = valued.map((v) => v.id)
}
const pageIds = ordered.slice(offset, offset + limit)
const entitiesMap = await this.batchGet(pageIds)
const entitiesMap = await this.hydratePage(pageIds, params.fields, this.guardFieldsFor(params))
const results: Result<T>[] = []
for (const id of pageIds) {
const entity = entitiesMap.get(id)

View file

@ -292,6 +292,58 @@ export class ColumnStore implements ColumnStoreProvider {
return result
}
/**
* Read this column's value for each of `entityIntIds` the per-id read
* behind `find({ fields })`.
*
* Every other read door here answers "which entities have this value". A
* projection asks the opposite "what value does this entity have" and
* without it a projection has to go to the canonical record for a field the
* column is already holding.
*
* The column is walked ONCE and the wanted ids are picked out as they pass,
* so the cost is O(column) per field rather than O(ids x column). Later
* sources win: the tail buffer holds writes newer than any segment, and
* within the segments a later one supersedes an earlier, exactly as `filter`
* treats them.
*
* Values are EXACT this store keeps raw values, not the bucketed form the
* sparse index uses for range queries which is what makes it safe to
* project from. Deleted entities are skipped; an id with no value in this
* column is simply absent from the result.
*
* @param field - Field name to read.
* @param entityIntIds - Entity integer ids to read values for.
* @returns `entityIntId -> value` for the ids this column holds.
*/
async valuesForIds(
field: string,
entityIntIds: Iterable<number>
): Promise<Map<number, number | string>> {
const wanted = new Set<number>(entityIntIds)
const out = new Map<number, number | string>()
if (wanted.size === 0 || !this.hasField(field)) return out
const deleted = this.deletedEntities.get(field)
const take = (entry: { value: number | string; entityIntId: number }): void => {
if (!wanted.has(entry.entityIntId)) return
if (deleted && deleted.has(entry.entityIntId)) return
out.set(entry.entityIntId, entry.value)
}
// Segments oldest -> newest, then the tail: a later write overwrites an
// earlier one for the same id.
const cursors = await this.getSegmentCursors(field)
for (const cursor of cursors) {
for (const entry of cursor.iterateForward()) take(entry)
}
const tailCursor = this.getTailBufferCursor(field)
if (tailCursor) {
for (const entry of tailCursor.iterateForward()) take(entry)
}
return out
}
/**
* Range filter: find entities where field is within the bounds.
*

View file

@ -2,7 +2,7 @@
* 🧠 BRAINY EMBEDDED PATTERNS
*
* AUTO-GENERATED - DO NOT EDIT
* Generated: 2025-09-29T10:10:00-07:00
* Generated: 2026-08-27T09:18:45-07:00
* Patterns: 220
* Coverage: 94-98% of all queries
*

View file

@ -495,6 +495,45 @@ export interface MetadataIndexProvider {
query: string,
ids: readonly string[]
): Promise<Array<{ id: string; matchCount: number }>>
/**
* @description OPTIONAL: read named SCALAR fields for many ids at once, from
* the index's own value storage, WITHOUT touching the canonical record.
*
* This is the door behind `find/get/related({ fields })`. A list view that
* needs a title and a slug currently hydrates the whole record for every row
* document bodies included and then discards almost all of it. Serving
* the named scalars from the index turns that into an index read.
*
* ## The contract, and the one rule that makes it safe
*
* **Return only what you can serve EXACTLY, and say what you served.** The
* answer is a per-id map of the fields this index actually resolved; the
* caller diffs it against what was requested and reads the canonical record
* for the remainder. An implementation must therefore OMIT a field rather
* than approximate it and omission costs only a record read, while a wrong
* value is a wrong answer nobody can see.
*
* That rule is not hypothetical. This engine's own index buckets
* `system.createdAt` and `system.updatedAt` to the minute for range queries,
* so it cannot serve them exactly and omits them. An engine whose column
* store holds raw values can serve the same fields so the two answer
* differently in COST and identically in CONTENT, which is the only
* difference a projection door is allowed to have.
*
* A field absent from an entity is simply absent from that entity's map. It
* is never an error, and never a `null` standing in for one: absent and
* present-and-null are different answers.
*
* @param ids - Canonical entity ids to read.
* @param fields - Index KEYS (bare = user metadata, `system.*` = engine
* scalar), already address-resolved by the caller.
* @returns `id → { field: value }` for the fields this index served exactly.
* Ids with nothing to serve may be omitted entirely.
*/
getScalarsForIds?(
ids: readonly string[],
fields: readonly string[]
): Promise<Map<string, Record<string, unknown>>>
getSortedIdsForFilter(filter: any, orderBy: string, order?: 'asc' | 'desc', topK?: number): Promise<string[]>
getFilterValues(field: string): Promise<string[]>
getFilterFields(): Promise<string[]>

View file

@ -561,6 +561,33 @@ export interface UpdateRelationParams<T = any> {
* refusal with the fix in hand beats a silent behavior flip.
*/
export interface FindParams<T = any> {
/**
* **Field projection** return only these fields on each row, instead of the
* whole record.
*
* A list view that shows a title and a slug does not need the document body,
* yet without a projection every row hydrates its full record and throws
* almost all of it away. Naming the fields lets them be served from the index
* itself: a scalar the index holds exactly is read from the index, and the
* canonical record is opened ONLY when a requested field cannot be.
*
* Field names follow the one addressing law: a bare name is the user's
* metadata (`'title'`), and `system.*` is an engine scalar
* (`'system.createdAt'`).
*
* - **Absent** the full record, exactly as before.
* - A requested field the entity does not carry is simply **absent** from the
* row. It is never an error a projection asks "give me these if you have
* them", so an optional field must not turn a list into a failure.
* - Every returned row carries `id` (and, on `find`, `score`) regardless: a
* row you cannot identify is not a row.
*
* @example
* // A list page: two user fields and one engine scalar, no document bodies.
* await brain.find({ where: { kind: 'post' }, fields: ['title', 'slug', 'system.createdAt'], limit: 50 })
*/
fields?: readonly string[]
// Vector Intelligence
/** Natural language or semantic search query (embedded and matched via HNSW + text index) */
query?: string
@ -789,6 +816,12 @@ export interface SimilarParams<T = any> {
* Added string ID shorthand syntax
*/
export interface RelatedParams {
// NOTE: `fields` is deliberately NOT offered here. A Relation carries `from`
// and `to` as IDS and hydrates no entity record, so there is nothing for a
// projection to trim — the param would be decorative. Projecting the
// ENDPOINTS would be a new capability (related() hydrating entities), not a
// projection of an existing one, and it belongs in its own decision.
/**
* Filter by source entity ID
*
@ -1414,6 +1447,33 @@ export interface ImportResult {
*
*/
export interface GetOptions {
/**
* **Field projection** return only these fields on each row, instead of the
* whole record.
*
* A list view that shows a title and a slug does not need the document body,
* yet without a projection every row hydrates its full record and throws
* almost all of it away. Naming the fields lets them be served from the index
* itself: a scalar the index holds exactly is read from the index, and the
* canonical record is opened ONLY when a requested field cannot be.
*
* Field names follow the one addressing law: a bare name is the user's
* metadata (`'title'`), and `system.*` is an engine scalar
* (`'system.createdAt'`).
*
* - **Absent** the full record, exactly as before.
* - A requested field the entity does not carry is simply **absent** from the
* row. It is never an error a projection asks "give me these if you have
* them", so an optional field must not turn a list into a failure.
* - Every returned row carries `id` (and, on `find`, `score`) regardless: a
* row you cannot identify is not a row.
*
* @example
* // A list page: two user fields and one engine scalar, no document bodies.
* await brain.find({ where: { kind: 'post' }, fields: ['title', 'slug', 'system.createdAt'], limit: 50 })
*/
fields?: readonly string[]
/**
* Include 384-dimensional vector embeddings in the response
*

View file

@ -2805,6 +2805,67 @@ export class MetadataIndexManager implements MetadataIndexProvider {
return order === 'asc' ? comparison : -comparison
}
/**
* Read named scalar fields for many ids from the COLUMN STORE, without
* touching the canonical record the `find({ fields })` door.
*
* ## Why the column store and not the sparse index
*
* The column store keeps RAW values; the sparse index keeps a normalized,
* bucketed form built for range queries `system.createdAt` is indexed at
* minute precision there. A projection served from the sparse index would
* hand back a value that differs from the record's, which is a wrong answer
* nobody can see. So this door reads the column store, and a field the
* column store does not hold is OMITTED rather than approximated.
*
* ## Why batched
*
* `getFieldValueForEntity` answers one (id, field) pair by walking the
* field's storage; called per row it re-walks the same column for every id.
* This walks each column ONCE and picks out every requested id as it passes:
* O(fields x column) instead of O(ids x fields x column).
*
* Omission is always safe it costs the caller a record read. The caller
* diffs what it asked for against what came back and reads records for the
* remainder, so an index that can serve nothing is slow, never wrong.
*
* @param ids - Canonical entity ids.
* @param fields - Index keys (bare = user metadata, `system.*` = engine scalar).
* @returns `id -> { field: value }` for exactly the pairs this index served.
*/
async getScalarsForIds(
ids: readonly string[],
fields: readonly string[]
): Promise<Map<string, Record<string, unknown>>> {
const out = new Map<string, Record<string, unknown>>()
if (ids.length === 0 || fields.length === 0) return out
// int -> id, so a column hit resolves back to the caller's id. An id the
// mapper does not know cannot be in any column, so it is simply absent.
const idByInt = new Map<number, string>()
for (const id of ids) {
const intId = this.idMapper.getInt(id)
if (intId !== undefined) idByInt.set(intId, id)
}
if (idByInt.size === 0) return out
for (const field of fields) {
if (!this.columnStore.hasField(field)) continue
const values = await this.columnStore.valuesForIds(field, idByInt.keys())
for (const [intId, value] of values) {
const id = idByInt.get(intId)
if (id === undefined) continue
let row = out.get(id)
if (row === undefined) {
row = {}
out.set(id, row)
}
row[field] = value
}
}
return out
}
async getFieldValueForEntity(entityId: string, field: string): Promise<any> {
// `field` arrives as a FROZEN INDEX KEY (bare = user metadata;
// 'system.<field>' = engine scalar). Storage fallbacks read the matching