feat: COW always-on architecture + cloud storage clear() fix (v5.11.0)
Major architectural improvements and critical bug fixes: ## COW Always-On Architecture - Removed cowEnabled flag from BaseStorage (COW cannot be disabled) - Eliminated marker file system (checkClearMarker, createClearMarker) - Simplified all code paths to assume COW is always enabled - COW automatically re-initializes after clear() operations ## Critical Bug Fix: Cloud Storage clear() - Fixed GCS clear() using correct paths (branches/ instead of entities/nouns/) - Fixed S3Compatible clear() path structure - Fixed R2 clear() implementation - Fixed Azure, FileSystem, OPFS, Memory clear() COW flag handling - clear() now deletes: branches/, _cow/, _system/ - Result: Cloud buckets can now be fully cleared (previously impossible) ## Container Memory Detection - Auto-detect Docker/K8s/Cloud Run memory limits (cgroup v1/v2) - Smart memory allocation (75% graph data, 25% query operations) - Environment variable support (CLOUD_RUN_MEMORY, MEMORY_LIMIT) - Production-grade containerized deployment support ## CommitLog streamHistory Feature - Added streamable commit history with pagination - Efficient memory usage for large commit histories - Support for branch filtering and time ranges ## Comprehensive Storage Documentation - Complete v5.11.0 file structure reference - Detailed path construction algorithms - 8 common storage scenarios with examples - Type-first storage, sharding, COW architecture explained - Public docs: docs/architecture/data-storage-architecture.md (1063 lines) ## Files Modified (14 files) - All 8 storage adapters (GCS, S3, R2, Azure, FS, OPFS, Memory, Historical) - BaseStorage core architecture - CommitLog with streaming - Brainy memory configuration - Parameter validation with container detection - Storage architecture documentation ## Breaking Changes NONE - COW was already enabled by default. This removes the ability to disable it. ## Migration No action required. Upgrade and clear() will work correctly on cloud storage. ## Impact - Users can now clear cloud storage buckets completely - No more corrupted buckets after clear() operations - Container deployments automatically optimize memory allocation - COW is mandatory and always enabled (safer, simpler) v5.11.0 - Production ready
This commit is contained in:
parent
28160a3052
commit
3e8b9aacc8
17 changed files with 2925 additions and 1207 deletions
207
src/brainy.ts
207
src/brainy.ts
|
|
@ -33,6 +33,7 @@ import { NULL_HASH } from './storage/cow/constants.js'
|
|||
import { createPipeline } from './streaming/pipeline.js'
|
||||
import { configureLogger, LogLevel } from './utils/logger.js'
|
||||
import { TransactionManager } from './transaction/TransactionManager.js'
|
||||
import { ValidationConfig } from './utils/paramValidation.js'
|
||||
import {
|
||||
SaveNounMetadataOperation,
|
||||
SaveNounOperation,
|
||||
|
|
@ -134,6 +135,15 @@ export class Brainy<T = any> implements BrainyInterface<T> {
|
|||
// Normalize configuration with defaults
|
||||
this.config = this.normalizeConfig(config)
|
||||
|
||||
// Configure memory limits (v5.11.0)
|
||||
// This must happen early, before any validation occurs
|
||||
if (this.config.maxQueryLimit !== undefined || this.config.reservedQueryMemory !== undefined) {
|
||||
ValidationConfig.reconfigure({
|
||||
maxQueryLimit: this.config.maxQueryLimit,
|
||||
reservedQueryMemory: this.config.reservedQueryMemory
|
||||
})
|
||||
}
|
||||
|
||||
// Setup core components
|
||||
this.distance = cosineDistance
|
||||
this.embedder = this.setupEmbedder()
|
||||
|
|
@ -3255,6 +3265,85 @@ export class Brainy<T = any> implements BrainyInterface<T> {
|
|||
})
|
||||
}
|
||||
|
||||
/**
|
||||
* Stream commit history (memory-efficient)
|
||||
*
|
||||
* Use this for large commit histories (1000s of snapshots) where memory
|
||||
* efficiency is critical. Yields commits one at a time without accumulating
|
||||
* them in memory.
|
||||
*
|
||||
* For small histories (< 100 commits), use getHistory() for simpler API.
|
||||
*
|
||||
* @param options - History options
|
||||
* @param options.limit - Maximum number of commits to stream
|
||||
* @param options.since - Only include commits after this timestamp
|
||||
* @param options.until - Only include commits before this timestamp
|
||||
* @param options.author - Filter by author name
|
||||
*
|
||||
* @yields Commit metadata in reverse chronological order (newest first)
|
||||
*
|
||||
* @example
|
||||
* ```typescript
|
||||
* // Stream all commits without memory accumulation
|
||||
* for await (const commit of brain.streamHistory({ limit: 10000 })) {
|
||||
* console.log(`${commit.timestamp}: ${commit.message}`)
|
||||
* }
|
||||
*
|
||||
* // Stream with filtering
|
||||
* for await (const commit of brain.streamHistory({
|
||||
* author: 'alice',
|
||||
* since: Date.now() - 86400000 // Last 24 hours
|
||||
* })) {
|
||||
* // Process commit
|
||||
* }
|
||||
* ```
|
||||
*/
|
||||
async *streamHistory(options?: {
|
||||
limit?: number
|
||||
since?: number
|
||||
until?: number
|
||||
author?: string
|
||||
}): AsyncIterableIterator<{
|
||||
hash: string
|
||||
message: string
|
||||
author: string
|
||||
timestamp: number
|
||||
metadata?: Record<string, any>
|
||||
}> {
|
||||
await this.ensureInitialized()
|
||||
|
||||
if (!('commitLog' in this.storage) || !('refManager' in this.storage)) {
|
||||
throw new Error('History streaming requires COW-enabled storage (v5.0.0+)')
|
||||
}
|
||||
|
||||
const commitLog = (this.storage as any).commitLog
|
||||
const currentBranch = await this.getCurrentBranch()
|
||||
const blobStorage = (this.storage as any).blobStorage
|
||||
|
||||
// Stream commits from CommitLog
|
||||
for await (const commit of commitLog.streamHistory(currentBranch, {
|
||||
maxCount: options?.limit,
|
||||
since: options?.since,
|
||||
until: options?.until
|
||||
})) {
|
||||
// Filter by author if specified
|
||||
if (options?.author && commit.author !== options.author) {
|
||||
continue
|
||||
}
|
||||
|
||||
// Map to expected format (compute hash for commit)
|
||||
yield {
|
||||
hash: blobStorage.constructor.hash(
|
||||
Buffer.from(JSON.stringify(commit))
|
||||
),
|
||||
message: commit.message,
|
||||
author: commit.author,
|
||||
timestamp: commit.timestamp,
|
||||
metadata: commit.metadata
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Get total count of nouns - O(1) operation
|
||||
* @returns Promise that resolves to the total number of nouns
|
||||
|
|
@ -3273,6 +3362,119 @@ export class Brainy<T = any> implements BrainyInterface<T> {
|
|||
return this.storage.getVerbCount()
|
||||
}
|
||||
|
||||
/**
|
||||
* Get memory statistics and limits (v5.11.0)
|
||||
*
|
||||
* Returns detailed memory information including:
|
||||
* - Current heap usage
|
||||
* - Container memory limits (if detected)
|
||||
* - Query limits and how they were calculated
|
||||
* - Memory allocation recommendations
|
||||
*
|
||||
* Use this to debug why query limits are low or to understand
|
||||
* memory allocation in production environments.
|
||||
*
|
||||
* @returns Memory statistics and configuration
|
||||
*
|
||||
* @example
|
||||
* ```typescript
|
||||
* const stats = brain.getMemoryStats()
|
||||
* console.log(`Query limit: ${stats.limits.maxQueryLimit}`)
|
||||
* console.log(`Basis: ${stats.limits.basis}`)
|
||||
* console.log(`Free memory: ${Math.round(stats.memory.free / 1024 / 1024)}MB`)
|
||||
* ```
|
||||
*/
|
||||
getMemoryStats(): {
|
||||
memory: {
|
||||
heapUsed: number
|
||||
heapTotal: number
|
||||
external: number
|
||||
rss: number
|
||||
free: number
|
||||
total: number
|
||||
containerLimit: number | null
|
||||
}
|
||||
limits: {
|
||||
maxQueryLimit: number
|
||||
maxQueryLength: number
|
||||
maxVectorDimensions: number
|
||||
basis: 'override' | 'reservedMemory' | 'containerMemory' | 'freeMemory'
|
||||
}
|
||||
config: {
|
||||
maxQueryLimit?: number
|
||||
reservedQueryMemory?: number
|
||||
}
|
||||
recommendations?: string[]
|
||||
} {
|
||||
const config = ValidationConfig.getInstance()
|
||||
const heapStats = process.memoryUsage ? process.memoryUsage() : {
|
||||
heapUsed: 0,
|
||||
heapTotal: 0,
|
||||
external: 0,
|
||||
rss: 0
|
||||
}
|
||||
|
||||
// Get system memory info
|
||||
let freeMemory = 0
|
||||
let totalMemory = 0
|
||||
if (typeof window === 'undefined') {
|
||||
try {
|
||||
const os = require('node:os')
|
||||
freeMemory = os.freemem()
|
||||
totalMemory = os.totalmem()
|
||||
} catch (e) {
|
||||
// OS module not available
|
||||
}
|
||||
}
|
||||
|
||||
const stats = {
|
||||
memory: {
|
||||
heapUsed: heapStats.heapUsed,
|
||||
heapTotal: heapStats.heapTotal,
|
||||
external: heapStats.external,
|
||||
rss: heapStats.rss,
|
||||
free: freeMemory,
|
||||
total: totalMemory,
|
||||
containerLimit: config.detectedContainerLimit
|
||||
},
|
||||
limits: {
|
||||
maxQueryLimit: config.maxLimit,
|
||||
maxQueryLength: config.maxQueryLength,
|
||||
maxVectorDimensions: config.maxVectorDimensions,
|
||||
basis: config.limitBasis
|
||||
},
|
||||
config: {
|
||||
maxQueryLimit: this.config.maxQueryLimit,
|
||||
reservedQueryMemory: this.config.reservedQueryMemory
|
||||
},
|
||||
recommendations: [] as string[]
|
||||
}
|
||||
|
||||
// Generate recommendations based on stats
|
||||
if (stats.limits.basis === 'freeMemory' && stats.memory.containerLimit) {
|
||||
stats.recommendations.push(
|
||||
`Container detected (${Math.round(stats.memory.containerLimit / 1024 / 1024)}MB) but limits based on free memory. ` +
|
||||
`Consider setting reservedQueryMemory config option for better limits.`
|
||||
)
|
||||
}
|
||||
|
||||
if (stats.limits.maxQueryLimit < 5000 && stats.memory.containerLimit && stats.memory.containerLimit > 2 * 1024 * 1024 * 1024) {
|
||||
stats.recommendations.push(
|
||||
`Query limit is low (${stats.limits.maxQueryLimit}) despite ${Math.round(stats.memory.containerLimit / 1024 / 1024 / 1024)}GB container. ` +
|
||||
`Consider: new Brainy({ reservedQueryMemory: 1073741824 }) to reserve 1GB for queries.`
|
||||
)
|
||||
}
|
||||
|
||||
if (stats.limits.basis === 'override') {
|
||||
stats.recommendations.push(
|
||||
`Using explicit maxQueryLimit override (${stats.limits.maxQueryLimit}). ` +
|
||||
`Auto-detection bypassed.`
|
||||
)
|
||||
}
|
||||
|
||||
return stats
|
||||
}
|
||||
|
||||
// ============= SUB-APIS =============
|
||||
|
||||
/**
|
||||
|
|
@ -4791,7 +4993,10 @@ export class Brainy<T = any> implements BrainyInterface<T> {
|
|||
disableMetrics: config?.disableMetrics ?? false,
|
||||
disableAutoOptimize: config?.disableAutoOptimize ?? false,
|
||||
batchWrites: config?.batchWrites ?? true,
|
||||
maxConcurrentOperations: config?.maxConcurrentOperations ?? 10
|
||||
maxConcurrentOperations: config?.maxConcurrentOperations ?? 10,
|
||||
// Memory management options (v5.11.0)
|
||||
maxQueryLimit: config?.maxQueryLimit ?? undefined as any,
|
||||
reservedQueryMemory: config?.reservedQueryMemory ?? undefined as any
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue