brainy/tests/unit/vfs-multi-instance-diagnostic.test.ts
David Snelling 9e307e457f fix: recalibrate find({ limit }) cap + two-tier enforcement + caller location
Brainy 7.30.0 introduced a memory-derived synchronous cap on `find({ limit })`
to prevent OOM. The cap was sound in intent but ~4x too conservative in
calibration: assumed 100 KB per result while typical entity footprint is 7-10 KB
(384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
free-memory box the cap derived to 9000 — breaking common safety-cap patterns
like `find({ type, where, limit: 10_000 })` that typically return 10-500
entities. Surfaced as a runtime regression with cascading 500s degrading
production dashboards.

Three concurrent fixes:

A. RECALIBRATE THE FORMULA
- src/utils/paramValidation.ts:175,196,212 — the three memory-derived priorities
  (reservedQueryMemory / containerMemory / freeMemory) all divided by
  100 * 1024 * 1024 (100 KB per result, ~10-15x over conservative). Replaced
  with a new MAX_LIMIT_KB_PER_RESULT = 25 constant that matches observed
  entity size.
- Result: 4 GB container cap goes 10_000 → 40_000; 2 GB cap goes 5_000 →
  20_000; 900 MB free-memory cap goes 9_000 → ~36_000. 100k hard ceiling
  unchanged. `maxQueryLimit` / `reservedQueryMemory` constructor overrides
  unchanged in behavior.

B. TWO-TIER ENFORCEMENT (warn-then-throw)
- Below cap (limit <= maxLimit): silent pass, unchanged.
- Soft tier (maxLimit < limit <= 2 * maxLimit): NEW — one-time warning per
  call site (dedup keyed on caller stack frame + limit value), query
  proceeds. Pre-7.30.2 code that relied on the cap silently allowing typical
  safety-cap limits keeps working; the warning teaches the recipe so consumers
  can fix it intentionally.
- Hard tier (limit > 2 * maxLimit): throw with the same teaching message
  format. Real OOM territory; the cap stops being a recommendation and becomes
  a guardrail.
- The 2x soft margin absorbs typical safety-cap patterns (limit: 10_000
  against a 9 K-cap box) without disabling OOM protection. Real OOM territory
  on a JS in-memory brain is hundreds of thousands of results, not 10x the
  safety cap.

C. IMPROVED ERROR / WARNING MESSAGE
- Same shape as the 7.30.1 enforcement-error messages: state the problem,
  name the three escape valves (maxQueryLimit / reservedQueryMemory /
  pagination), include caller location, link to docs.
- Extracted findCallerLocation() helper from brainy.ts to a new
  src/utils/callerLocation.ts so both the subtype enforcement (7.30.1) and
  the limit enforcement (7.30.2) share one implementation without circular
  imports.

DOCS
- New docs/guides/find-limits.md (public: true) — full reference: why the cap
  exists, the four memory sources the auto-config considers, the three escape
  valves with when-to-use-which guidance, and an explicit "pagination is the
  future-proof pattern" callout (8.0 may tighten the cap further; pagination
  keeps working unchanged).
- docs/api/README.md find() entry gets a one-paragraph `limit` tip + pointer
  to the new guide.
- RELEASES.md v7.30.2 entry.

TESTS
- New tests/integration/find-limits.test.ts (9 tests): below-cap silent pass;
  soft-tier warns once per call site (dedup verified by exercising same vs.
  different source lines via wrapper closures); soft-tier message format
  (names all three escape valves + docs link); soft-tier message includes
  caller location; hard-tier throws; hard-tier message format same as
  soft-tier; consumer maxQueryLimit override raises the cap and shifts both
  tiers accordingly; pre-7.30.2 regression scenario explicitly covered.
- tests/unit/utils/memoryLimits.test.ts — 4 tests updated for the recalibrated
  cap values (hardcoded expected numbers bumped 4x to match new 25 KB/result
  assumption).
- tests/unit/utils/paramValidation.test.ts — auto-limit test extended to cover
  the three-tier semantics (below-cap pass / soft-tier silent / hard-tier
  throw).
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-
  enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.

CORTEX COMPATIBILITY
- Zero Cortex changes required. Every change is JS-side: formula recalibration
  runs in ValidationConfig.constructor(), two-tier enforcement runs in
  validateFindParams(), both fire before any storage / index / Cortex call.
- The new guide notes that Brainy 8.0's Datomic-style Db.find() may tighten
  per-call limits to keep snapshot semantics cheap; pagination remains the
  pattern that's guaranteed to keep working.

REPO-WIDE CLEANUP
Brainy is the only Soulcraft project that is open source. This commit also
scrubs closed-source product names and product-specific class/field references
from every tracked file in the repo (src/, docs/, tests/, RELEASES.md,
CHANGELOG.md). Consumer-reported bugs, regression scenarios, and release
notes now refer to "a consumer", "a downstream application", "a production
deployment", or "an internal report" — never to the named product. Two
product-named test files renamed to neutral diagnostic names. CLAUDE.md gains
a project-level guard rule documenting the policy and an example list of the
identifiers that may not appear in tracked code.

Verification
- npx tsc --noEmit: clean
- npm test: 1468 / 1468 unit
- All four integration subtype + verb + strict + find-limits suites: 78/78
- npm run build: clean
- Closed-source product reference audit: clean
2026-06-08 12:49:43 -07:00

171 lines
7.1 KiB
TypeScript
Raw Permalink Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

/**
* VFS Multi-instance Diagnostic Test
*
* Tests to verify VFS import behavior and identify if VFS creates only wrappers or also graph entities
*/
import { describe, it, expect, beforeEach } from 'vitest'
import { Brainy, NounType } from '../../src/index.js'
describe('VFS Multi-instance Diagnostic', () => {
let brain: Brainy
beforeEach(async () => {
brain = new Brainy({
storage: { type: 'memory' }
})
await brain.init()
})
it('should verify VFS creates document wrappers AND allows entity filtering', async () => {
console.log('\n🔬 VFS Multi-instance Diagnostic Test\n')
console.log('='.repeat(70))
// Step 1: Add entities directly (control group)
console.log('\n1⃣ Adding entities directly (without VFS)...\n')
await brain.add({ data: 'Person 1', type: NounType.Person, metadata: { name: 'Person 1' } })
await brain.add({ data: 'Person 2', type: NounType.Person, metadata: { name: 'Person 2' } })
await brain.add({ data: 'Location 1', type: NounType.Location, metadata: { name: 'Location 1' } })
const beforeVfs = await brain.find({ limit: 100 })
console.log(` Total entities: ${beforeVfs.length}`)
const peopleBefore = await brain.find({ type: NounType.Person, limit: 100 })
console.log(` Person filter: ${peopleBefore.length} (expected: 2)`)
expect(peopleBefore.length).toBe(2)
console.log(' ✅ Type filtering works on direct entities\n')
// Step 2: Use VFS to create files
console.log('2⃣ Creating VFS files...\n')
const vfs = brain.vfs
await vfs.init()
await vfs.mkdir('/test', { recursive: true })
// Create a VFS file with entity data
const personData = {
id: 'ent_person_test',
name: 'John Smith',
type: 'person',
metadata: { source: 'test' }
}
await vfs.writeFile('/test/john.json', Buffer.from(JSON.stringify(personData, null, 2)))
console.log(' Created VFS file: /test/john.json')
// Step 3: Check what entities exist now
console.log('\n3⃣ Analyzing entities after VFS...\n')
const afterVfs = await brain.find({ limit: 100 })
console.log(` Total entities: ${afterVfs.length}`)
// Count by type
const typeCounts: Record<string, number> = {}
for (const result of afterVfs) {
const type = result.type || 'unknown'
typeCounts[type] = (typeCounts[type] || 0) + 1
}
console.log('\n Entity type breakdown:')
for (const [type, count] of Object.entries(typeCounts)) {
console.log(` - ${type}: ${count}`)
}
// Count VFS wrappers vs regular entities
const vfsWrappers = afterVfs.filter(e => e.metadata?.vfsType === 'file')
const regularEntities = afterVfs.filter(e => !e.metadata?.vfsType)
console.log(`\n VFS wrappers: ${vfsWrappers.length}`)
console.log(` Regular entities: ${regularEntities.length}`)
// Step 4: Test type filtering after VFS
console.log('\n4⃣ Testing type filtering after VFS...\n')
const peopleAfter = await brain.find({ type: NounType.Person, limit: 100 })
console.log(` Person filter: ${peopleAfter.length} (expected: 2 - same as before)`)
const documents = await brain.find({ type: NounType.Document, limit: 100 })
console.log(` Document filter: ${documents.length} (expected: ${vfsWrappers.length})`)
// Step 5: Analyze VFS wrapper structure
console.log('\n5⃣ Analyzing VFS wrapper structure...\n')
const wrapper = vfsWrappers[0]
if (wrapper) {
console.log(' VFS Wrapper Entity:')
console.log(` - ID: ${wrapper.id}`)
console.log(` - Type: ${wrapper.type}`)
console.log(` - VFS Type: ${wrapper.metadata?.vfsType}`)
console.log(` - Path: ${wrapper.metadata?.path}`)
console.log(` - Has rawData: ${!!wrapper.metadata?.rawData}`)
if (wrapper.metadata?.rawData) {
const decoded = Buffer.from(wrapper.metadata.rawData, 'base64').toString()
const embedded = JSON.parse(decoded)
console.log(`\n Embedded Entity Data:`)
console.log(` - Name: ${embedded.name}`)
console.log(` - Type: ${embedded.type}`)
console.log(`\n 🔍 KEY FINDING:`)
console.log(` Wrapper type: "${wrapper.type}"`)
console.log(` Embedded type: "${embedded.type}"`)
console.log(` Filtering by type="${embedded.type}" searches wrapper type, not embedded!`)
}
}
// Step 6: Diagnosis
console.log('\n' + '='.repeat(70))
console.log('📋 DIAGNOSIS\n')
if (peopleAfter.length === peopleBefore.length) {
console.log('✅ VFS does NOT create duplicate graph entities')
console.log('✅ VFS only creates document wrappers')
console.log('✅ Type filtering works on original entities, ignores VFS wrappers')
console.log('\nThis means:')
console.log(' - VFS files are type="document" wrappers')
console.log(' - Original entities keep their types')
console.log(' - filter({ type: "person" }) returns original entities only')
} else {
console.log('❌ Unexpected behavior - VFS may have created additional entities')
}
console.log('\n' + '='.repeat(70) + '\n')
// Assertions
expect(peopleAfter.length).toBe(2) // Should still be 2, VFS doesn't create person entities
expect(documents.length).toBeGreaterThan(0) // VFS creates document wrappers
expect(vfsWrappers.length).toBeGreaterThan(0) // Should have VFS wrappers
})
it('should verify import creates BOTH VFS wrappers AND graph entities', async () => {
// This test would require creating a test Excel file and running import
// For now, we'll document the expected behavior based on code analysis
console.log('\n📚 Expected Import Behavior (from code analysis):\n')
console.log('When you run brain.import("file.xlsx", { vfsPath: "/imports" }):')
console.log('\n1. ImportCoordinator.execute() calls:')
console.log(' a) vfsGenerator.generate() - creates VFS file wrappers')
console.log(' - Each entity → JSON file in VFS')
console.log(' - Wrapper entity with type="document"')
console.log(' - Entity data stored in metadata.rawData (base64)')
console.log('')
console.log(' b) createGraphEntities() - creates graph entities')
console.log(' - Each entity → graph entity with proper type')
console.log(' - type="person", "location", "concept", etc.')
console.log(' - metadata.vfsPath points to VFS file')
console.log('')
console.log('2. Result: Database contains BOTH:')
console.log(' - VFS wrappers (type="document", vfsType="file")')
console.log(' - Graph entities (type="person", etc., vfsPath set)')
console.log('')
console.log('3. Type filtering:')
console.log(' - filter({ type: "person" }) → returns graph entities')
console.log(' - filter({ type: "document" }) → returns VFS wrappers')
console.log('')
console.log('If a consumer gets 0 results, likely causes:')
console.log(' ❌ Only VFS wrappers created (createEntities: false)')
console.log(' ❌ Import not completing before query')
console.log(' ❌ Querying different Brainy instance')
console.log('')
})
})