2026-04-09 16:26:38 -07:00
/ * *
* @module orderby - sort - bug
* @description Regression tests for the ` find({ orderBy: ... }) ` sort bug .
*
* Bug : ` getFieldValueForEntity ` read ` noun.metadata[field] ` for timestamp
* fields , but ` getNoun() ` destructures standard fields to the top level .
* Every entity returned ` undefined ` for the sort key , stable sort preserved
* insertion order , and ` order: 'desc' ` behaved like ` order: 'asc' ` .
*
* Fix : centralized ` resolveEntityField ` helper + ` BUCKETED_INDEX_FIELDS `
* set in coreTypes . ts , used by getFieldValueForEntity .
*
fix: recalibrate find({ limit }) cap + two-tier enforcement + caller location
Brainy 7.30.0 introduced a memory-derived synchronous cap on `find({ limit })`
to prevent OOM. The cap was sound in intent but ~4x too conservative in
calibration: assumed 100 KB per result while typical entity footprint is 7-10 KB
(384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
free-memory box the cap derived to 9000 — breaking common safety-cap patterns
like `find({ type, where, limit: 10_000 })` that typically return 10-500
entities. Surfaced as a runtime regression with cascading 500s degrading
production dashboards.
Three concurrent fixes:
A. RECALIBRATE THE FORMULA
- src/utils/paramValidation.ts:175,196,212 — the three memory-derived priorities
(reservedQueryMemory / containerMemory / freeMemory) all divided by
100 * 1024 * 1024 (100 KB per result, ~10-15x over conservative). Replaced
with a new MAX_LIMIT_KB_PER_RESULT = 25 constant that matches observed
entity size.
- Result: 4 GB container cap goes 10_000 → 40_000; 2 GB cap goes 5_000 →
20_000; 900 MB free-memory cap goes 9_000 → ~36_000. 100k hard ceiling
unchanged. `maxQueryLimit` / `reservedQueryMemory` constructor overrides
unchanged in behavior.
B. TWO-TIER ENFORCEMENT (warn-then-throw)
- Below cap (limit <= maxLimit): silent pass, unchanged.
- Soft tier (maxLimit < limit <= 2 * maxLimit): NEW — one-time warning per
call site (dedup keyed on caller stack frame + limit value), query
proceeds. Pre-7.30.2 code that relied on the cap silently allowing typical
safety-cap limits keeps working; the warning teaches the recipe so consumers
can fix it intentionally.
- Hard tier (limit > 2 * maxLimit): throw with the same teaching message
format. Real OOM territory; the cap stops being a recommendation and becomes
a guardrail.
- The 2x soft margin absorbs typical safety-cap patterns (limit: 10_000
against a 9 K-cap box) without disabling OOM protection. Real OOM territory
on a JS in-memory brain is hundreds of thousands of results, not 10x the
safety cap.
C. IMPROVED ERROR / WARNING MESSAGE
- Same shape as the 7.30.1 enforcement-error messages: state the problem,
name the three escape valves (maxQueryLimit / reservedQueryMemory /
pagination), include caller location, link to docs.
- Extracted findCallerLocation() helper from brainy.ts to a new
src/utils/callerLocation.ts so both the subtype enforcement (7.30.1) and
the limit enforcement (7.30.2) share one implementation without circular
imports.
DOCS
- New docs/guides/find-limits.md (public: true) — full reference: why the cap
exists, the four memory sources the auto-config considers, the three escape
valves with when-to-use-which guidance, and an explicit "pagination is the
future-proof pattern" callout (8.0 may tighten the cap further; pagination
keeps working unchanged).
- docs/api/README.md find() entry gets a one-paragraph `limit` tip + pointer
to the new guide.
- RELEASES.md v7.30.2 entry.
TESTS
- New tests/integration/find-limits.test.ts (9 tests): below-cap silent pass;
soft-tier warns once per call site (dedup verified by exercising same vs.
different source lines via wrapper closures); soft-tier message format
(names all three escape valves + docs link); soft-tier message includes
caller location; hard-tier throws; hard-tier message format same as
soft-tier; consumer maxQueryLimit override raises the cap and shifts both
tiers accordingly; pre-7.30.2 regression scenario explicitly covered.
- tests/unit/utils/memoryLimits.test.ts — 4 tests updated for the recalibrated
cap values (hardcoded expected numbers bumped 4x to match new 25 KB/result
assumption).
- tests/unit/utils/paramValidation.test.ts — auto-limit test extended to cover
the three-tier semantics (below-cap pass / soft-tier silent / hard-tier
throw).
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-
enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.
CORTEX COMPATIBILITY
- Zero Cortex changes required. Every change is JS-side: formula recalibration
runs in ValidationConfig.constructor(), two-tier enforcement runs in
validateFindParams(), both fire before any storage / index / Cortex call.
- The new guide notes that Brainy 8.0's Datomic-style Db.find() may tighten
per-call limits to keep snapshot semantics cheap; pagination remains the
pattern that's guaranteed to keep working.
REPO-WIDE CLEANUP
Brainy is the only Soulcraft project that is open source. This commit also
scrubs closed-source product names and product-specific class/field references
from every tracked file in the repo (src/, docs/, tests/, RELEASES.md,
CHANGELOG.md). Consumer-reported bugs, regression scenarios, and release
notes now refer to "a consumer", "a downstream application", "a production
deployment", or "an internal report" — never to the named product. Two
product-named test files renamed to neutral diagnostic names. CLAUDE.md gains
a project-level guard rule documenting the policy and an example list of the
identifiers that may not appear in tracked code.
Verification
- npx tsc --noEmit: clean
- npm test: 1468 / 1468 unit
- All four integration subtype + verb + strict + find-limits suites: 78/78
- npm run build: clean
- Closed-source product reference audit: clean
2026-06-08 12:34:05 -07:00
* Reported by a consumer 2026 - 04 - 09 .
2026-04-09 16:26:38 -07:00
*
* NOTE : These tests cover FILTERED sort , which is the only supported path .
* Unfiltered ` find({ orderBy }) ` is explicitly rejected until the dedicated
* time - ordered segment index ships ( Track 2 ) .
* /
import { describe , it , expect , beforeEach , afterEach } from 'vitest'
import { Brainy } from '../../src/brainy'
import {
resolveEntityField ,
STANDARD_ENTITY_FIELDS ,
type HNSWNounWithMetadata
} from '../../src/coreTypes'
import { NounType } from '../../src/types/graphTypes'
describe ( 'find({ orderBy }) sort bug regression' , ( ) = > {
let brain : Brainy < any >
beforeEach ( async ( ) = > {
brain = new Brainy ( { storage : { type : 'memory' } , silent : true } )
await brain . init ( )
} )
afterEach ( async ( ) = > {
await brain . close ( )
} )
/ * *
fix: recalibrate find({ limit }) cap + two-tier enforcement + caller location
Brainy 7.30.0 introduced a memory-derived synchronous cap on `find({ limit })`
to prevent OOM. The cap was sound in intent but ~4x too conservative in
calibration: assumed 100 KB per result while typical entity footprint is 7-10 KB
(384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
free-memory box the cap derived to 9000 — breaking common safety-cap patterns
like `find({ type, where, limit: 10_000 })` that typically return 10-500
entities. Surfaced as a runtime regression with cascading 500s degrading
production dashboards.
Three concurrent fixes:
A. RECALIBRATE THE FORMULA
- src/utils/paramValidation.ts:175,196,212 — the three memory-derived priorities
(reservedQueryMemory / containerMemory / freeMemory) all divided by
100 * 1024 * 1024 (100 KB per result, ~10-15x over conservative). Replaced
with a new MAX_LIMIT_KB_PER_RESULT = 25 constant that matches observed
entity size.
- Result: 4 GB container cap goes 10_000 → 40_000; 2 GB cap goes 5_000 →
20_000; 900 MB free-memory cap goes 9_000 → ~36_000. 100k hard ceiling
unchanged. `maxQueryLimit` / `reservedQueryMemory` constructor overrides
unchanged in behavior.
B. TWO-TIER ENFORCEMENT (warn-then-throw)
- Below cap (limit <= maxLimit): silent pass, unchanged.
- Soft tier (maxLimit < limit <= 2 * maxLimit): NEW — one-time warning per
call site (dedup keyed on caller stack frame + limit value), query
proceeds. Pre-7.30.2 code that relied on the cap silently allowing typical
safety-cap limits keeps working; the warning teaches the recipe so consumers
can fix it intentionally.
- Hard tier (limit > 2 * maxLimit): throw with the same teaching message
format. Real OOM territory; the cap stops being a recommendation and becomes
a guardrail.
- The 2x soft margin absorbs typical safety-cap patterns (limit: 10_000
against a 9 K-cap box) without disabling OOM protection. Real OOM territory
on a JS in-memory brain is hundreds of thousands of results, not 10x the
safety cap.
C. IMPROVED ERROR / WARNING MESSAGE
- Same shape as the 7.30.1 enforcement-error messages: state the problem,
name the three escape valves (maxQueryLimit / reservedQueryMemory /
pagination), include caller location, link to docs.
- Extracted findCallerLocation() helper from brainy.ts to a new
src/utils/callerLocation.ts so both the subtype enforcement (7.30.1) and
the limit enforcement (7.30.2) share one implementation without circular
imports.
DOCS
- New docs/guides/find-limits.md (public: true) — full reference: why the cap
exists, the four memory sources the auto-config considers, the three escape
valves with when-to-use-which guidance, and an explicit "pagination is the
future-proof pattern" callout (8.0 may tighten the cap further; pagination
keeps working unchanged).
- docs/api/README.md find() entry gets a one-paragraph `limit` tip + pointer
to the new guide.
- RELEASES.md v7.30.2 entry.
TESTS
- New tests/integration/find-limits.test.ts (9 tests): below-cap silent pass;
soft-tier warns once per call site (dedup verified by exercising same vs.
different source lines via wrapper closures); soft-tier message format
(names all three escape valves + docs link); soft-tier message includes
caller location; hard-tier throws; hard-tier message format same as
soft-tier; consumer maxQueryLimit override raises the cap and shifts both
tiers accordingly; pre-7.30.2 regression scenario explicitly covered.
- tests/unit/utils/memoryLimits.test.ts — 4 tests updated for the recalibrated
cap values (hardcoded expected numbers bumped 4x to match new 25 KB/result
assumption).
- tests/unit/utils/paramValidation.test.ts — auto-limit test extended to cover
the three-tier semantics (below-cap pass / soft-tier silent / hard-tier
throw).
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-
enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.
CORTEX COMPATIBILITY
- Zero Cortex changes required. Every change is JS-side: formula recalibration
runs in ValidationConfig.constructor(), two-tier enforcement runs in
validateFindParams(), both fire before any storage / index / Cortex call.
- The new guide notes that Brainy 8.0's Datomic-style Db.find() may tighten
per-call limits to keep snapshot semantics cheap; pagination remains the
pattern that's guaranteed to keep working.
REPO-WIDE CLEANUP
Brainy is the only Soulcraft project that is open source. This commit also
scrubs closed-source product names and product-specific class/field references
from every tracked file in the repo (src/, docs/, tests/, RELEASES.md,
CHANGELOG.md). Consumer-reported bugs, regression scenarios, and release
notes now refer to "a consumer", "a downstream application", "a production
deployment", or "an internal report" — never to the named product. Two
product-named test files renamed to neutral diagnostic names. CLAUDE.md gains
a project-level guard rule documenting the policy and an example list of the
identifiers that may not appear in tracked code.
Verification
- npx tsc --noEmit: clean
- npm test: 1468 / 1468 unit
- All four integration subtype + verb + strict + find-limits suites: 78/78
- npm run build: clean
- Closed-source product reference audit: clean
2026-06-08 12:34:05 -07:00
* Real consumer use case : filtered sort of chat sessions .
2026-04-09 16:26:38 -07:00
* Before the fix , this returned the OLDEST entity instead of the newest
* because getFieldValueForEntity was reading createdAt from the wrong
* location on the entity .
* /
it ( 'orderBy createdAt desc with filter returns newest matching entity' , async ( ) = > {
// Add 4 entities 20ms apart so createdAt values are distinct.
const id1 = await brain . add ( { data : 'first' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'second' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'third' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
const id4 = await brain . add ( { data : 'fourth' , type : NounType . Concept } )
const results = await brain . find ( {
type : NounType . Concept ,
orderBy : 'createdAt' ,
order : 'desc' ,
limit : 1
} )
expect ( results ) . toHaveLength ( 1 )
// Must be the last-inserted entity, not the first.
expect ( results [ 0 ] . id ) . toBe ( id4 )
expect ( results [ 0 ] . id ) . not . toBe ( id1 )
} )
it ( 'orderBy createdAt asc with filter returns oldest matching entity' , async ( ) = > {
const id1 = await brain . add ( { data : 'first' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'second' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'third' , type : NounType . Concept } )
const results = await brain . find ( {
type : NounType . Concept ,
orderBy : 'createdAt' ,
order : 'asc' ,
limit : 1
} )
expect ( results ) . toHaveLength ( 1 )
expect ( results [ 0 ] . id ) . toBe ( id1 )
} )
it ( 'orderBy createdAt desc with filter returns all matching entities in newest-first order' , async ( ) = > {
const ids : string [ ] = [ ]
for ( let i = 0 ; i < 5 ; i ++ ) {
ids . push ( await brain . add ( { data : ` item- ${ i } ` , type : NounType . Concept } ) )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
}
const results = await brain . find ( {
type : NounType . Concept ,
orderBy : 'createdAt' ,
order : 'desc'
} )
expect ( results ) . toHaveLength ( 5 )
expect ( results . map ( ( r ) = > r . id ) ) . toEqual ( [ . . . ids ] . reverse ( ) )
} )
it ( 'orderBy updatedAt desc with filter returns most-recently-updated matching entity' , async ( ) = > {
const id1 = await brain . add ( { data : 'first' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'second' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'third' , type : NounType . Concept } )
// Touch id1 so it becomes the most-recently-updated.
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . update ( { id : id1 , data : 'first-updated' } )
const results = await brain . find ( {
type : NounType . Concept ,
orderBy : 'updatedAt' ,
order : 'desc' ,
limit : 1
} )
expect ( results ) . toHaveLength ( 1 )
expect ( results [ 0 ] . id ) . toBe ( id1 )
} )
/ * *
feat: unified column store for filtering + sorting at billion scale
Adds a per-field sorted column store (Lucene doc values + roaring bitmap
architecture) that replaces the MetadataIndex sparse index internals for
both filtering and sorting. One system for all field types with exact
precision — no bucketing, no per-entity storage reads.
Column store: binary .cidx segment format, in-memory tail buffers,
LSM-style compaction, k-way merge sort, multi-value support (__words__).
All queries (filter, range, sort, filtered sort) route through the column
store when data is available, falling back to sparse index otherwise.
Key unlocks:
- find({ orderBy: 'createdAt' }) works WITHOUT a filter (previously threw)
- find({ orderBy: 'metadata.price' }) works for custom numeric fields
- Exact timestamp precision (no 1-minute bucketing)
- O(K log S) sort independent of total entity count
New files: src/indexes/columnStore/ (types, format, tail buffer, cursor,
manifest, coordinator — ~700 lines). 101 new unit tests covering binary
format round-trips, CRC validation, sort, filter, range, deletion,
multi-segment merge, persistence, and multi-value (words) fields.
Deleted: metadataIndex-automatic-bucketing.test.ts (bucketing behavior
eliminated by exact-precision column store). Sparse index write path
removed from addToIndex/removeFromIndex. Sparse index legacy code still
present as dead code pending cleanup in next commit.
2026-04-10 11:22:19 -07:00
* Unfiltered sort now works via the unified column store .
* This was the Track 2 motivating use case — previously threw an error .
2026-04-09 16:26:38 -07:00
* /
feat: unified column store for filtering + sorting at billion scale
Adds a per-field sorted column store (Lucene doc values + roaring bitmap
architecture) that replaces the MetadataIndex sparse index internals for
both filtering and sorting. One system for all field types with exact
precision — no bucketing, no per-entity storage reads.
Column store: binary .cidx segment format, in-memory tail buffers,
LSM-style compaction, k-way merge sort, multi-value support (__words__).
All queries (filter, range, sort, filtered sort) route through the column
store when data is available, falling back to sparse index otherwise.
Key unlocks:
- find({ orderBy: 'createdAt' }) works WITHOUT a filter (previously threw)
- find({ orderBy: 'metadata.price' }) works for custom numeric fields
- Exact timestamp precision (no 1-minute bucketing)
- O(K log S) sort independent of total entity count
New files: src/indexes/columnStore/ (types, format, tail buffer, cursor,
manifest, coordinator — ~700 lines). 101 new unit tests covering binary
format round-trips, CRC validation, sort, filter, range, deletion,
multi-segment merge, persistence, and multi-value (words) fields.
Deleted: metadataIndex-automatic-bucketing.test.ts (bucketing behavior
eliminated by exact-precision column store). Sparse index write path
removed from addToIndex/removeFromIndex. Sparse index legacy code still
present as dead code pending cleanup in next commit.
2026-04-10 11:22:19 -07:00
it ( 'orderBy without filter works via column store' , async ( ) = > {
const id1 = await brain . add ( { data : 'first' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
await brain . add ( { data : 'second' , type : NounType . Concept } )
await new Promise ( ( r ) = > setTimeout ( r , 20 ) )
const id3 = await brain . add ( { data : 'third' , type : NounType . Concept } )
const results = await brain . find ( {
orderBy : 'createdAt' ,
order : 'desc' ,
limit : 2
} )
2026-04-09 16:26:38 -07:00
feat: unified column store for filtering + sorting at billion scale
Adds a per-field sorted column store (Lucene doc values + roaring bitmap
architecture) that replaces the MetadataIndex sparse index internals for
both filtering and sorting. One system for all field types with exact
precision — no bucketing, no per-entity storage reads.
Column store: binary .cidx segment format, in-memory tail buffers,
LSM-style compaction, k-way merge sort, multi-value support (__words__).
All queries (filter, range, sort, filtered sort) route through the column
store when data is available, falling back to sparse index otherwise.
Key unlocks:
- find({ orderBy: 'createdAt' }) works WITHOUT a filter (previously threw)
- find({ orderBy: 'metadata.price' }) works for custom numeric fields
- Exact timestamp precision (no 1-minute bucketing)
- O(K log S) sort independent of total entity count
New files: src/indexes/columnStore/ (types, format, tail buffer, cursor,
manifest, coordinator — ~700 lines). 101 new unit tests covering binary
format round-trips, CRC validation, sort, filter, range, deletion,
multi-segment merge, persistence, and multi-value (words) fields.
Deleted: metadataIndex-automatic-bucketing.test.ts (bucketing behavior
eliminated by exact-precision column store). Sparse index write path
removed from addToIndex/removeFromIndex. Sparse index legacy code still
present as dead code pending cleanup in next commit.
2026-04-10 11:22:19 -07:00
expect ( results . length ) . toBeGreaterThanOrEqual ( 2 )
// Newest should be first (desc order)
expect ( results [ 0 ] . id ) . toBe ( id3 )
2026-04-09 16:26:38 -07:00
} )
} )
describe ( 'resolveEntityField helper' , ( ) = > {
const entity : HNSWNounWithMetadata = {
id : 'abc' ,
vector : [ 0.1 , 0.2 ] ,
connections : new Map ( ) ,
level : 0 ,
type : NounType . Concept ,
createdAt : 1700000000000 ,
updatedAt : 1700000060000 ,
confidence : 0.9 ,
weight : 1 ,
service : 'test' ,
data : { title : 'Hello' } ,
metadata : {
customTag : 'green' ,
priority : 5 ,
modified : 1700000120000 // VFS custom field, lives in metadata
}
}
it ( 'reads standard fields from top level' , ( ) = > {
expect ( resolveEntityField ( entity , 'createdAt' ) ) . toBe ( 1700000000000 )
expect ( resolveEntityField ( entity , 'updatedAt' ) ) . toBe ( 1700000060000 )
expect ( resolveEntityField ( entity , 'type' ) ) . toBe ( NounType . Concept )
expect ( resolveEntityField ( entity , 'confidence' ) ) . toBe ( 0.9 )
expect ( resolveEntityField ( entity , 'weight' ) ) . toBe ( 1 )
expect ( resolveEntityField ( entity , 'service' ) ) . toBe ( 'test' )
expect ( resolveEntityField ( entity , 'id' ) ) . toBe ( 'abc' )
} )
it ( 'reads custom fields from metadata' , ( ) = > {
expect ( resolveEntityField ( entity , 'customTag' ) ) . toBe ( 'green' )
expect ( resolveEntityField ( entity , 'priority' ) ) . toBe ( 5 )
} )
it ( 'reads VFS custom fields (modified, accessed) from metadata' , ( ) = > {
// VFS stores `modified` as a custom field, not top-level.
expect ( resolveEntityField ( entity , 'modified' ) ) . toBe ( 1700000120000 )
} )
it ( 'returns undefined for unknown fields' , ( ) = > {
expect ( resolveEntityField ( entity , 'nonexistent' ) ) . toBeUndefined ( )
} )
it ( 'returns undefined for custom fields when metadata is absent' , ( ) = > {
const noMetadata : HNSWNounWithMetadata = { . . . entity , metadata : undefined }
expect ( resolveEntityField ( noMetadata , 'customTag' ) ) . toBeUndefined ( )
} )
it ( 'does not look in metadata for standard fields' , ( ) = > {
// If a standard field is absent top-level, resolver returns undefined
// rather than silently falling through to metadata. This prevents
// misuse from masking bugs.
const withShadowedField : HNSWNounWithMetadata = {
. . . entity ,
// @ts-expect-error intentionally clobbering for the test
createdAt : undefined ,
metadata : { createdAt : 9999 }
}
expect ( resolveEntityField ( withShadowedField , 'createdAt' ) ) . toBeUndefined ( )
} )
it ( 'STANDARD_ENTITY_FIELDS covers every declared top-level field' , ( ) = > {
// Guards against the resolver and the interface drifting out of sync.
const expected = [
'id' ,
'vector' ,
'connections' ,
'level' ,
'type' ,
'confidence' ,
'weight' ,
'createdAt' ,
'updatedAt' ,
'service' ,
'createdBy' ,
'data'
]
for ( const field of expected ) {
expect ( STANDARD_ENTITY_FIELDS . has ( field ) ) . toBe ( true )
}
} )
} )