feat: COW always-on architecture + cloud storage clear() fix (v5.11.0)
Major architectural improvements and critical bug fixes:
## COW Always-On Architecture
- Removed cowEnabled flag from BaseStorage (COW cannot be disabled)
- Eliminated marker file system (checkClearMarker, createClearMarker)
- Simplified all code paths to assume COW is always enabled
- COW automatically re-initializes after clear() operations
## Critical Bug Fix: Cloud Storage clear()
- Fixed GCS clear() using correct paths (branches/ instead of entities/nouns/)
- Fixed S3Compatible clear() path structure
- Fixed R2 clear() implementation
- Fixed Azure, FileSystem, OPFS, Memory clear() COW flag handling
- clear() now deletes: branches/, _cow/, _system/
- Result: Cloud buckets can now be fully cleared (previously impossible)
## Container Memory Detection
- Auto-detect Docker/K8s/Cloud Run memory limits (cgroup v1/v2)
- Smart memory allocation (75% graph data, 25% query operations)
- Environment variable support (CLOUD_RUN_MEMORY, MEMORY_LIMIT)
- Production-grade containerized deployment support
## CommitLog streamHistory Feature
- Added streamable commit history with pagination
- Efficient memory usage for large commit histories
- Support for branch filtering and time ranges
## Comprehensive Storage Documentation
- Complete v5.11.0 file structure reference
- Detailed path construction algorithms
- 8 common storage scenarios with examples
- Type-first storage, sharding, COW architecture explained
- Public docs: docs/architecture/data-storage-architecture.md (1063 lines)
## Files Modified (14 files)
- All 8 storage adapters (GCS, S3, R2, Azure, FS, OPFS, Memory, Historical)
- BaseStorage core architecture
- CommitLog with streaming
- Brainy memory configuration
- Parameter validation with container detection
- Storage architecture documentation
## Breaking Changes
NONE - COW was already enabled by default. This removes the ability to disable it.
## Migration
No action required. Upgrade and clear() will work correctly on cloud storage.
## Impact
- Users can now clear cloud storage buckets completely
- No more corrupted buckets after clear() operations
- Container deployments automatically optimize memory allocation
- COW is mandatory and always enabled (safer, simpler)
v5.11.0 - Production ready
2025-11-18 13:44:02 -08:00
/ * *
* Integration tests for Memory Enhancements ( v5 . 11.0 )
*
* End - to - end tests verifying :
* - streamHistory ( ) works in production scenarios
* - Container detection works correctly
* - Memory limits are applied correctly
* - getMemoryStats ( ) provides accurate information
* - All features work together seamlessly
* /
import { describe , it , expect , beforeEach , afterEach } from 'vitest'
import { Brainy } from '../../src/brainy.js'
import { mkdtempSync , rmSync } from 'fs'
import { tmpdir } from 'os'
import { join } from 'path'
describe ( 'Memory Enhancements Integration (v5.11.0)' , ( ) = > {
let testDir : string
let originalEnv : Record < string , string | undefined >
beforeEach ( ( ) = > {
testDir = mkdtempSync ( join ( tmpdir ( ) , 'brainy-memory-integration-' ) )
originalEnv = {
CLOUD_RUN_MEMORY : process.env.CLOUD_RUN_MEMORY ,
MEMORY_LIMIT : process.env.MEMORY_LIMIT
}
} )
afterEach ( ( ) = > {
rmSync ( testDir , { recursive : true , force : true } )
Object . keys ( originalEnv ) . forEach ( key = > {
if ( originalEnv [ key ] === undefined ) {
delete process . env [ key ]
} else {
process . env [ key ] = originalEnv [ key ]
}
} )
} )
fix: recalibrate find({ limit }) cap + two-tier enforcement + caller location
Brainy 7.30.0 introduced a memory-derived synchronous cap on `find({ limit })`
to prevent OOM. The cap was sound in intent but ~4x too conservative in
calibration: assumed 100 KB per result while typical entity footprint is 7-10 KB
(384-dim float32 vector ≈ 1.5 KB + standard fields + metadata). On a 900 MB
free-memory box the cap derived to 9000 — breaking common safety-cap patterns
like `find({ type, where, limit: 10_000 })` that typically return 10-500
entities. Surfaced as a runtime regression with cascading 500s degrading
production dashboards.
Three concurrent fixes:
A. RECALIBRATE THE FORMULA
- src/utils/paramValidation.ts:175,196,212 — the three memory-derived priorities
(reservedQueryMemory / containerMemory / freeMemory) all divided by
100 * 1024 * 1024 (100 KB per result, ~10-15x over conservative). Replaced
with a new MAX_LIMIT_KB_PER_RESULT = 25 constant that matches observed
entity size.
- Result: 4 GB container cap goes 10_000 → 40_000; 2 GB cap goes 5_000 →
20_000; 900 MB free-memory cap goes 9_000 → ~36_000. 100k hard ceiling
unchanged. `maxQueryLimit` / `reservedQueryMemory` constructor overrides
unchanged in behavior.
B. TWO-TIER ENFORCEMENT (warn-then-throw)
- Below cap (limit <= maxLimit): silent pass, unchanged.
- Soft tier (maxLimit < limit <= 2 * maxLimit): NEW — one-time warning per
call site (dedup keyed on caller stack frame + limit value), query
proceeds. Pre-7.30.2 code that relied on the cap silently allowing typical
safety-cap limits keeps working; the warning teaches the recipe so consumers
can fix it intentionally.
- Hard tier (limit > 2 * maxLimit): throw with the same teaching message
format. Real OOM territory; the cap stops being a recommendation and becomes
a guardrail.
- The 2x soft margin absorbs typical safety-cap patterns (limit: 10_000
against a 9 K-cap box) without disabling OOM protection. Real OOM territory
on a JS in-memory brain is hundreds of thousands of results, not 10x the
safety cap.
C. IMPROVED ERROR / WARNING MESSAGE
- Same shape as the 7.30.1 enforcement-error messages: state the problem,
name the three escape valves (maxQueryLimit / reservedQueryMemory /
pagination), include caller location, link to docs.
- Extracted findCallerLocation() helper from brainy.ts to a new
src/utils/callerLocation.ts so both the subtype enforcement (7.30.1) and
the limit enforcement (7.30.2) share one implementation without circular
imports.
DOCS
- New docs/guides/find-limits.md (public: true) — full reference: why the cap
exists, the four memory sources the auto-config considers, the three escape
valves with when-to-use-which guidance, and an explicit "pagination is the
future-proof pattern" callout (8.0 may tighten the cap further; pagination
keeps working unchanged).
- docs/api/README.md find() entry gets a one-paragraph `limit` tip + pointer
to the new guide.
- RELEASES.md v7.30.2 entry.
TESTS
- New tests/integration/find-limits.test.ts (9 tests): below-cap silent pass;
soft-tier warns once per call site (dedup verified by exercising same vs.
different source lines via wrapper closures); soft-tier message format
(names all three escape valves + docs link); soft-tier message includes
caller location; hard-tier throws; hard-tier message format same as
soft-tier; consumer maxQueryLimit override raises the cap and shifts both
tiers accordingly; pre-7.30.2 regression scenario explicitly covered.
- tests/unit/utils/memoryLimits.test.ts — 4 tests updated for the recalibrated
cap values (hardcoded expected numbers bumped 4x to match new 25 KB/result
assumption).
- tests/unit/utils/paramValidation.test.ts — auto-limit test extended to cover
the three-tier semantics (below-cap pass / soft-tier silent / hard-tier
throw).
- Existing suites unchanged: subtype-and-facets 26/26, verb-subtype-and-
enforcement 30/30, strict-mode-self-test 13/13. Unit 1468/1468.
CORTEX COMPATIBILITY
- Zero Cortex changes required. Every change is JS-side: formula recalibration
runs in ValidationConfig.constructor(), two-tier enforcement runs in
validateFindParams(), both fire before any storage / index / Cortex call.
- The new guide notes that Brainy 8.0's Datomic-style Db.find() may tighten
per-call limits to keep snapshot semantics cheap; pagination remains the
pattern that's guaranteed to keep working.
REPO-WIDE CLEANUP
Brainy is the only Soulcraft project that is open source. This commit also
scrubs closed-source product names and product-specific class/field references
from every tracked file in the repo (src/, docs/, tests/, RELEASES.md,
CHANGELOG.md). Consumer-reported bugs, regression scenarios, and release
notes now refer to "a consumer", "a downstream application", "a production
deployment", or "an internal report" — never to the named product. Two
product-named test files renamed to neutral diagnostic names. CLAUDE.md gains
a project-level guard rule documenting the policy and an example list of the
identifiers that may not appear in tracked code.
Verification
- npx tsc --noEmit: clean
- npm test: 1468 / 1468 unit
- All four integration subtype + verb + strict + find-limits suites: 78/78
- npm run build: clean
- Closed-source product reference audit: clean
2026-06-08 12:34:05 -07:00
describe ( 'Production Workflow: Consumer Snapshot Timeline' , ( ) = > {
feat: COW always-on architecture + cloud storage clear() fix (v5.11.0)
Major architectural improvements and critical bug fixes:
## COW Always-On Architecture
- Removed cowEnabled flag from BaseStorage (COW cannot be disabled)
- Eliminated marker file system (checkClearMarker, createClearMarker)
- Simplified all code paths to assume COW is always enabled
- COW automatically re-initializes after clear() operations
## Critical Bug Fix: Cloud Storage clear()
- Fixed GCS clear() using correct paths (branches/ instead of entities/nouns/)
- Fixed S3Compatible clear() path structure
- Fixed R2 clear() implementation
- Fixed Azure, FileSystem, OPFS, Memory clear() COW flag handling
- clear() now deletes: branches/, _cow/, _system/
- Result: Cloud buckets can now be fully cleared (previously impossible)
## Container Memory Detection
- Auto-detect Docker/K8s/Cloud Run memory limits (cgroup v1/v2)
- Smart memory allocation (75% graph data, 25% query operations)
- Environment variable support (CLOUD_RUN_MEMORY, MEMORY_LIMIT)
- Production-grade containerized deployment support
## CommitLog streamHistory Feature
- Added streamable commit history with pagination
- Efficient memory usage for large commit histories
- Support for branch filtering and time ranges
## Comprehensive Storage Documentation
- Complete v5.11.0 file structure reference
- Detailed path construction algorithms
- 8 common storage scenarios with examples
- Type-first storage, sharding, COW architecture explained
- Public docs: docs/architecture/data-storage-architecture.md (1063 lines)
## Files Modified (14 files)
- All 8 storage adapters (GCS, S3, R2, Azure, FS, OPFS, Memory, Historical)
- BaseStorage core architecture
- CommitLog with streaming
- Brainy memory configuration
- Parameter validation with container detection
- Storage architecture documentation
## Breaking Changes
NONE - COW was already enabled by default. This removes the ability to disable it.
## Migration
No action required. Upgrade and clear() will work correctly on cloud storage.
## Impact
- Users can now clear cloud storage buckets completely
- No more corrupted buckets after clear() operations
- Container deployments automatically optimize memory allocation
- COW is mandatory and always enabled (safer, simpler)
v5.11.0 - Production ready
2025-11-18 13:44:02 -08:00
it ( 'should stream 1000 snapshots efficiently' , async ( ) = > {
const brain = new Brainy ( {
storage : {
type : 'filesystem' ,
options : { path : testDir }
} ,
silent : true
} )
await brain . init ( )
// Create a realistic workflow: user creates entities and saves snapshots
for ( let i = 0 ; i < 100 ; i ++ ) {
// Add 10 entities
for ( let j = 0 ; j < 10 ; j ++ ) {
await brain . add ( {
type : 'document' ,
data : ` Chapter ${ i } , Section ${ j } `
} )
}
// Save snapshot
await brain . commit ( { message : ` Version ${ i } ` , captureState : true } )
}
// Stream history efficiently
const snapshots : string [ ] = [ ]
const startTime = Date . now ( )
for await ( const commit of brain . streamHistory ( { limit : 100 } ) ) {
snapshots . push ( commit . message )
}
const duration = Date . now ( ) - startTime
expect ( snapshots . length ) . toBe ( 100 )
expect ( duration ) . toBeLessThan ( 5000 ) // Should complete in < 5s
await brain . close ( )
} )
it ( 'should handle large snapshot with filtering' , async ( ) = > {
const brain = new Brainy ( {
storage : {
type : 'filesystem' ,
options : { path : testDir }
} ,
silent : true
} )
await brain . init ( )
// Create snapshots by different authors
for ( let i = 0 ; i < 50 ; i ++ ) {
await brain . add ( { type : 'document' , data : ` Content ${ i } ` } )
await brain . commit ( {
message : ` Snapshot ${ i } ` ,
author : i % 3 === 0 ? 'alice' : i % 3 === 1 ? 'bob' : 'charlie' ,
captureState : true
} )
}
// Stream only alice's snapshots
const aliceSnapshots : any [ ] = [ ]
for await ( const commit of brain . streamHistory ( { author : 'alice' , limit : 100 } ) ) {
aliceSnapshots . push ( commit )
}
expect ( aliceSnapshots . length ) . toBeGreaterThan ( 0 )
aliceSnapshots . forEach ( commit = > {
expect ( commit . author ) . toBe ( 'alice' )
} )
await brain . close ( )
} )
} )
describe ( 'Production Workflow: Cloud Run Deployment' , ( ) = > {
it ( 'should detect 4GB Cloud Run container and set optimal limits' , async ( ) = > {
process . env . CLOUD_RUN_MEMORY = '4Gi'
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
// Verify container detected
expect ( stats . memory . containerLimit ) . toBe ( 4 * 1024 * 1024 * 1024 )
// Verify optimal limits
expect ( stats . limits . basis ) . toBe ( 'containerMemory' )
expect ( stats . limits . maxQueryLimit ) . toBe ( 10000 ) // 25% of 4GB
// Verify can query with these limits
for ( let i = 0 ; i < 100 ; i ++ ) {
await brain . add ( { type : 'document' , data : ` Entity ${ i } ` } )
}
const results = await brain . find ( { limit : 10000 } ) // Should not throw
expect ( results . length ) . toBe ( 100 )
await brain . close ( )
} )
it ( 'should allow manual override for power users in containers' , async ( ) = > {
process . env . CLOUD_RUN_MEMORY = '4Gi'
const brain = new Brainy ( {
storage : { type : 'memory' } ,
maxQueryLimit : 50000 , // Override auto-detection
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
// Container detected but override used
expect ( stats . memory . containerLimit ) . toBe ( 4 * 1024 * 1024 * 1024 )
expect ( stats . limits . basis ) . toBe ( 'override' )
expect ( stats . limits . maxQueryLimit ) . toBe ( 50000 )
await brain . close ( )
} )
it ( 'should use reservedQueryMemory for fine-grained control' , async ( ) = > {
const brain = new Brainy ( {
storage : { type : 'memory' } ,
reservedQueryMemory : 2 * 1024 * 1024 * 1024 , // Reserve 2GB
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
expect ( stats . limits . basis ) . toBe ( 'reservedMemory' )
expect ( stats . limits . maxQueryLimit ) . toBe ( 20000 ) // 2GB / 100MB * 1000
await brain . close ( )
} )
} )
describe ( 'Combined Features: Stream + Memory Management' , ( ) = > {
it ( 'should stream large histories within memory limits' , async ( ) = > {
process . env . CLOUD_RUN_MEMORY = '2Gi'
const brain = new Brainy ( {
storage : {
type : 'filesystem' ,
options : { path : testDir }
} ,
silent : true
} )
await brain . init ( )
// Create many snapshots
for ( let i = 0 ; i < 200 ; i ++ ) {
await brain . add ( { type : 'document' , data : ` Data ${ i } ` } )
if ( i % 5 === 0 ) {
await brain . commit ( { message : ` Checkpoint ${ i / 5 } ` , captureState : true } )
}
}
// Verify memory limits
const stats = brain . getMemoryStats ( )
expect ( stats . limits . maxQueryLimit ) . toBe ( 5000 ) // 2GB * 0.25
// Stream all snapshots (should be ~40)
const heapBefore = process . memoryUsage ( ) . heapUsed
let count = 0
for await ( const commit of brain . streamHistory ( { limit : 100 } ) ) {
count ++
}
const heapAfter = process . memoryUsage ( ) . heapUsed
const heapGrowth = heapAfter - heapBefore
expect ( count ) . toBeGreaterThan ( 0 )
expect ( heapGrowth ) . toBeLessThan ( 20 * 1024 * 1024 ) // < 20MB growth
await brain . close ( )
} )
} )
describe ( 'Memory Stats API' , ( ) = > {
it ( 'should provide actionable recommendations' , async ( ) = > {
process . env . CLOUD_RUN_MEMORY = '4Gi'
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
// Should have recommendations array
expect ( stats . recommendations ) . toBeDefined ( )
expect ( Array . isArray ( stats . recommendations ) ) . toBe ( true )
// Should not have error recommendations for well-configured system
const errorRecommendations = stats . recommendations ? . filter ( r = >
r . toLowerCase ( ) . includes ( 'error' ) || r . toLowerCase ( ) . includes ( 'warning' )
)
expect ( errorRecommendations ? . length || 0 ) . toBe ( 0 )
await brain . close ( )
} )
it ( 'should help debug low limits in production' , async ( ) = > {
// Simulate production issue: container detected but low free memory
process . env . CLOUD_RUN_MEMORY = '4Gi'
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
// User can see why limits are what they are
expect ( stats . limits . basis ) . toBeDefined ( )
expect ( stats . memory . containerLimit ) . toBeDefined ( )
expect ( stats . config ) . toBeDefined ( )
// Can debug: "Why is my limit only 10k?"
// Answer: basis='containerMemory', container=4GB, 4GB * 0.25 = 1GB -> 10k
await brain . close ( )
} )
} )
describe ( 'Zero-Config Behavior' , ( ) = > {
it ( 'should work optimally with no configuration' , async ( ) = > {
process . env . CLOUD_RUN_MEMORY = '4Gi'
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
// Zero config, optimal behavior
const stats = brain . getMemoryStats ( )
expect ( stats . limits . maxQueryLimit ) . toBe ( 10000 )
expect ( stats . limits . basis ) . toBe ( 'containerMemory' )
// Should work without issues
for ( let i = 0 ; i < 100 ; i ++ ) {
await brain . add ( { type : 'note' , data : ` Test ${ i } ` } )
}
const results = await brain . find ( { limit : 100 } )
expect ( results . length ) . toBe ( 100 )
await brain . close ( )
} )
it ( 'should work on bare metal without containers' , async ( ) = > {
// No container env vars
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
const stats = brain . getMemoryStats ( )
expect ( stats . memory . containerLimit ) . toBeNull ( )
expect ( stats . limits . basis ) . toBe ( 'freeMemory' )
expect ( stats . limits . maxQueryLimit ) . toBeGreaterThan ( 0 )
await brain . close ( )
} )
} )
describe ( 'Backward Compatibility' , ( ) = > {
it ( 'should not break existing code' , async ( ) = > {
const brain = new Brainy ( {
storage : { type : 'memory' } ,
silent : true
} )
await brain . init ( )
// Existing APIs work
await brain . add ( { type : 'document' , data : 'Test' } )
const results = await brain . find ( { limit : 10 } )
expect ( results . length ) . toBe ( 1 )
await brain . close ( )
} )
it ( 'should keep getHistory() working as before' , async ( ) = > {
const brain = new Brainy ( {
storage : {
type : 'filesystem' ,
options : { path : testDir }
} ,
silent : true
} )
await brain . init ( )
// Create some snapshots
for ( let i = 0 ; i < 10 ; i ++ ) {
await brain . add ( { type : 'note' , data : ` Test ${ i } ` } )
await brain . commit ( { message : ` Snapshot ${ i } ` , captureState : true } )
}
// Old API still works
const history = await brain . getHistory ( { limit : 10 } )
expect ( history . length ) . toBe ( 10 )
expect ( history [ 0 ] ) . toHaveProperty ( 'hash' )
expect ( history [ 0 ] ) . toHaveProperty ( 'message' )
await brain . close ( )
} )
} )
} )