Brainy 7.30.0 introduced a memory-derived synchronous cap on `find({ limit })`
to prevent OOM. The cap was sound in intent but ~4x too conservative in
calibration: assumed 100 KB per result while typical entity footprint is 7-10 KB
(384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
free-memory box the cap derived to 9000 — breaking common safety-cap patterns
like `find({ type, where, limit: 10_000 })` that typically return 10-500
entities. Surfaced as a runtime regression with cascading 500s degrading
production dashboards.
Three concurrent fixes:
A. RECALIBRATE THE FORMULA
- src/utils/paramValidation.ts:175,196,212 — the three memory-derived priorities
(reservedQueryMemory / containerMemory / freeMemory) all divided by
100 * 1024 * 1024 (100 KB per result, ~10-15x over conservative). Replaced
with a new MAX_LIMIT_KB_PER_RESULT = 25 constant that matches observed
entity size.
- Result: 4 GB container cap goes 10_000 → 40_000; 2 GB cap goes 5_000 →
20_000; 900 MB free-memory cap goes 9_000 → ~36_000. 100k hard ceiling
unchanged. `maxQueryLimit` / `reservedQueryMemory` constructor overrides
unchanged in behavior.
B. TWO-TIER ENFORCEMENT (warn-then-throw)
- Below cap (limit <= maxLimit): silent pass, unchanged.
- Soft tier (maxLimit < limit <= 2 * maxLimit): NEW — one-time warning per
call site (dedup keyed on caller stack frame + limit value), query
proceeds. Pre-7.30.2 code that relied on the cap silently allowing typical
safety-cap limits keeps working; the warning teaches the recipe so consumers
can fix it intentionally.
- Hard tier (limit > 2 * maxLimit): throw with the same teaching message
format. Real OOM territory; the cap stops being a recommendation and becomes
a guardrail.
- The 2x soft margin absorbs typical safety-cap patterns (limit: 10_000
against a 9 K-cap box) without disabling OOM protection. Real OOM territory
on a JS in-memory brain is hundreds of thousands of results, not 10x the
safety cap.
C. IMPROVED ERROR / WARNING MESSAGE
- Same shape as the 7.30.1 enforcement-error messages: state the problem,
name the three escape valves (maxQueryLimit / reservedQueryMemory /
pagination), include caller location, link to docs.
- Extracted findCallerLocation() helper from brainy.ts to a new
src/utils/callerLocation.ts so both the subtype enforcement (7.30.1) and
the limit enforcement (7.30.2) share one implementation without circular
imports.
DOCS
- New docs/guides/find-limits.md (public: true) — full reference: why the cap
exists, the four memory sources the auto-config considers, the three escape
valves with when-to-use-which guidance, and an explicit "pagination is the
future-proof pattern" callout (8.0 may tighten the cap further; pagination
keeps working unchanged).
- docs/api/README.md find() entry gets a one-paragraph `limit` tip + pointer
to the new guide.
- RELEASES.md v7.30.2 entry.
TESTS
- New tests/integration/find-limits.test.ts (9 tests): below-cap silent pass;
soft-tier warns once per call site (dedup verified by exercising same vs.
different source lines via wrapper closures); soft-tier message format
(names all three escape valves + docs link); soft-tier message includes
caller location; hard-tier throws; hard-tier message format same as
soft-tier; consumer maxQueryLimit override raises the cap and shifts both
tiers accordingly; pre-7.30.2 regression scenario explicitly covered.
- tests/unit/utils/memoryLimits.test.ts — 4 tests updated for the recalibrated
cap values (hardcoded expected numbers bumped 4x to match new 25 KB/result
assumption).
- tests/unit/utils/paramValidation.test.ts — auto-limit test extended to cover
the three-tier semantics (below-cap pass / soft-tier silent / hard-tier
throw).
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-
enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.
CORTEX COMPATIBILITY
- Zero Cortex changes required. Every change is JS-side: formula recalibration
runs in ValidationConfig.constructor(), two-tier enforcement runs in
validateFindParams(), both fire before any storage / index / Cortex call.
- The new guide notes that Brainy 8.0's Datomic-style Db.find() may tighten
per-call limits to keep snapshot semantics cheap; pagination remains the
pattern that's guaranteed to keep working.
REPO-WIDE CLEANUP
Brainy is the only Soulcraft project that is open source. This commit also
scrubs closed-source product names and product-specific class/field references
from every tracked file in the repo (src/, docs/, tests/, RELEASES.md,
CHANGELOG.md). Consumer-reported bugs, regression scenarios, and release
notes now refer to "a consumer", "a downstream application", "a production
deployment", or "an internal report" — never to the named product. Two
product-named test files renamed to neutral diagnostic names. CLAUDE.md gains
a project-level guard rule documenting the policy and an example list of the
identifiers that may not appear in tracked code.
Verification
- npx tsc --noEmit: clean
- npm test: 1468 / 1468 unit
- All four integration subtype + verb + strict + find-limits suites: 78/78
- npm run build: clean
- Closed-source product reference audit: clean
232 lines
7.8 KiB
TypeScript
232 lines
7.8 KiB
TypeScript
/**
|
|
* @module orderby-sort-bug
|
|
* @description Regression tests for the `find({ orderBy: ... })` sort bug.
|
|
*
|
|
* Bug: `getFieldValueForEntity` read `noun.metadata[field]` for timestamp
|
|
* fields, but `getNoun()` destructures standard fields to the top level.
|
|
* Every entity returned `undefined` for the sort key, stable sort preserved
|
|
* insertion order, and `order: 'desc'` behaved like `order: 'asc'`.
|
|
*
|
|
* Fix: centralized `resolveEntityField` helper + `BUCKETED_INDEX_FIELDS`
|
|
* set in coreTypes.ts, used by getFieldValueForEntity.
|
|
*
|
|
* Reported by a consumer 2026-04-09.
|
|
*
|
|
* NOTE: These tests cover FILTERED sort, which is the only supported path.
|
|
* Unfiltered `find({ orderBy })` is explicitly rejected until the dedicated
|
|
* time-ordered segment index ships (Track 2).
|
|
*/
|
|
|
|
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
|
|
import { Brainy } from '../../src/brainy'
|
|
import {
|
|
resolveEntityField,
|
|
STANDARD_ENTITY_FIELDS,
|
|
type HNSWNounWithMetadata
|
|
} from '../../src/coreTypes'
|
|
import { NounType } from '../../src/types/graphTypes'
|
|
|
|
describe('find({ orderBy }) sort bug regression', () => {
|
|
let brain: Brainy<any>
|
|
|
|
beforeEach(async () => {
|
|
brain = new Brainy({ storage: { type: 'memory' }, silent: true })
|
|
await brain.init()
|
|
})
|
|
|
|
afterEach(async () => {
|
|
await brain.close()
|
|
})
|
|
|
|
/**
|
|
* Real consumer use case: filtered sort of chat sessions.
|
|
* Before the fix, this returned the OLDEST entity instead of the newest
|
|
* because getFieldValueForEntity was reading createdAt from the wrong
|
|
* location on the entity.
|
|
*/
|
|
it('orderBy createdAt desc with filter returns newest matching entity', async () => {
|
|
// Add 4 entities 20ms apart so createdAt values are distinct.
|
|
const id1 = await brain.add({ data: 'first', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'second', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'third', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
const id4 = await brain.add({ data: 'fourth', type: NounType.Concept })
|
|
|
|
const results = await brain.find({
|
|
type: NounType.Concept,
|
|
orderBy: 'createdAt',
|
|
order: 'desc',
|
|
limit: 1
|
|
})
|
|
|
|
expect(results).toHaveLength(1)
|
|
// Must be the last-inserted entity, not the first.
|
|
expect(results[0].id).toBe(id4)
|
|
expect(results[0].id).not.toBe(id1)
|
|
})
|
|
|
|
it('orderBy createdAt asc with filter returns oldest matching entity', async () => {
|
|
const id1 = await brain.add({ data: 'first', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'second', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'third', type: NounType.Concept })
|
|
|
|
const results = await brain.find({
|
|
type: NounType.Concept,
|
|
orderBy: 'createdAt',
|
|
order: 'asc',
|
|
limit: 1
|
|
})
|
|
|
|
expect(results).toHaveLength(1)
|
|
expect(results[0].id).toBe(id1)
|
|
})
|
|
|
|
it('orderBy createdAt desc with filter returns all matching entities in newest-first order', async () => {
|
|
const ids: string[] = []
|
|
for (let i = 0; i < 5; i++) {
|
|
ids.push(await brain.add({ data: `item-${i}`, type: NounType.Concept }))
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
}
|
|
|
|
const results = await brain.find({
|
|
type: NounType.Concept,
|
|
orderBy: 'createdAt',
|
|
order: 'desc'
|
|
})
|
|
|
|
expect(results).toHaveLength(5)
|
|
expect(results.map((r) => r.id)).toEqual([...ids].reverse())
|
|
})
|
|
|
|
it('orderBy updatedAt desc with filter returns most-recently-updated matching entity', async () => {
|
|
const id1 = await brain.add({ data: 'first', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'second', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'third', type: NounType.Concept })
|
|
|
|
// Touch id1 so it becomes the most-recently-updated.
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.update({ id: id1, data: 'first-updated' })
|
|
|
|
const results = await brain.find({
|
|
type: NounType.Concept,
|
|
orderBy: 'updatedAt',
|
|
order: 'desc',
|
|
limit: 1
|
|
})
|
|
|
|
expect(results).toHaveLength(1)
|
|
expect(results[0].id).toBe(id1)
|
|
})
|
|
|
|
/**
|
|
* Unfiltered sort now works via the unified column store.
|
|
* This was the Track 2 motivating use case — previously threw an error.
|
|
*/
|
|
it('orderBy without filter works via column store', async () => {
|
|
const id1 = await brain.add({ data: 'first', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
await brain.add({ data: 'second', type: NounType.Concept })
|
|
await new Promise((r) => setTimeout(r, 20))
|
|
const id3 = await brain.add({ data: 'third', type: NounType.Concept })
|
|
|
|
const results = await brain.find({
|
|
orderBy: 'createdAt',
|
|
order: 'desc',
|
|
limit: 2
|
|
})
|
|
|
|
expect(results.length).toBeGreaterThanOrEqual(2)
|
|
// Newest should be first (desc order)
|
|
expect(results[0].id).toBe(id3)
|
|
})
|
|
})
|
|
|
|
describe('resolveEntityField helper', () => {
|
|
const entity: HNSWNounWithMetadata = {
|
|
id: 'abc',
|
|
vector: [0.1, 0.2],
|
|
connections: new Map(),
|
|
level: 0,
|
|
type: NounType.Concept,
|
|
createdAt: 1700000000000,
|
|
updatedAt: 1700000060000,
|
|
confidence: 0.9,
|
|
weight: 1,
|
|
service: 'test',
|
|
data: { title: 'Hello' },
|
|
metadata: {
|
|
customTag: 'green',
|
|
priority: 5,
|
|
modified: 1700000120000 // VFS custom field, lives in metadata
|
|
}
|
|
}
|
|
|
|
it('reads standard fields from top level', () => {
|
|
expect(resolveEntityField(entity, 'createdAt')).toBe(1700000000000)
|
|
expect(resolveEntityField(entity, 'updatedAt')).toBe(1700000060000)
|
|
expect(resolveEntityField(entity, 'type')).toBe(NounType.Concept)
|
|
expect(resolveEntityField(entity, 'confidence')).toBe(0.9)
|
|
expect(resolveEntityField(entity, 'weight')).toBe(1)
|
|
expect(resolveEntityField(entity, 'service')).toBe('test')
|
|
expect(resolveEntityField(entity, 'id')).toBe('abc')
|
|
})
|
|
|
|
it('reads custom fields from metadata', () => {
|
|
expect(resolveEntityField(entity, 'customTag')).toBe('green')
|
|
expect(resolveEntityField(entity, 'priority')).toBe(5)
|
|
})
|
|
|
|
it('reads VFS custom fields (modified, accessed) from metadata', () => {
|
|
// VFS stores `modified` as a custom field, not top-level.
|
|
expect(resolveEntityField(entity, 'modified')).toBe(1700000120000)
|
|
})
|
|
|
|
it('returns undefined for unknown fields', () => {
|
|
expect(resolveEntityField(entity, 'nonexistent')).toBeUndefined()
|
|
})
|
|
|
|
it('returns undefined for custom fields when metadata is absent', () => {
|
|
const noMetadata: HNSWNounWithMetadata = { ...entity, metadata: undefined }
|
|
expect(resolveEntityField(noMetadata, 'customTag')).toBeUndefined()
|
|
})
|
|
|
|
it('does not look in metadata for standard fields', () => {
|
|
// If a standard field is absent top-level, resolver returns undefined
|
|
// rather than silently falling through to metadata. This prevents
|
|
// misuse from masking bugs.
|
|
const withShadowedField: HNSWNounWithMetadata = {
|
|
...entity,
|
|
// @ts-expect-error intentionally clobbering for the test
|
|
createdAt: undefined,
|
|
metadata: { createdAt: 9999 }
|
|
}
|
|
expect(resolveEntityField(withShadowedField, 'createdAt')).toBeUndefined()
|
|
})
|
|
|
|
it('STANDARD_ENTITY_FIELDS covers every declared top-level field', () => {
|
|
// Guards against the resolver and the interface drifting out of sync.
|
|
const expected = [
|
|
'id',
|
|
'vector',
|
|
'connections',
|
|
'level',
|
|
'type',
|
|
'confidence',
|
|
'weight',
|
|
'createdAt',
|
|
'updatedAt',
|
|
'service',
|
|
'createdBy',
|
|
'data'
|
|
]
|
|
for (const field of expected) {
|
|
expect(STANDARD_ENTITY_FIELDS.has(field)).toBe(true)
|
|
}
|
|
})
|
|
})
|