The old config-generation subsystem (src/config/ + autoConfiguration.ts) was
superseded during the 8.0 rework and never wired into init(): it emitted settings
for a partitioning subsystem that no longer exists and probed deleted cloud env
vars. The live zero-config path is inline — recall preset → HNSW knobs, storage
auto-detect, auto persistMode, container-memory-aware cache sizing.
The storage progressive-init / cloud-detection cluster was equally dead after the
cloud adapters were dropped: isCloudStorage() is permanently false (no overriders),
scheduleBackgroundInit/runBackgroundInit were never called (the latter an empty
body), initMode was never assigned, and Brainy.isFullyInitialized()/
awaitBackgroundInit() were always-trivial with zero callers. scheduleCountPersist()
collapses to its only-ever-taken immediate write-through path.
Removed:
- src/config/{index,zeroConfig,storageAutoConfig,modelAutoConfig,sharedConfigManager}.ts
- src/utils/autoConfiguration.ts + the inert BrainyZeroConfig export
- Brainy.isFullyInitialized()/awaitBackgroundInit() (+ BrainyInterface decls)
- InitMode type, isCloudStorage/detectCloudEnvironment/resolveInitMode,
scheduleBackgroundInit/runBackgroundInit/ensureValidatedForWrite and their state
- Dead cloud env-var probes (K_SERVICE/K_REVISION/AWS_LAMBDA_FUNCTION_NAME/
FUNCTIONS_TARGET/AZURE_FUNCTIONS_ENVIRONMENT)
Kept (verified live): production-detection logging (environment.ts), container-
memory cache sizing (memoryDetection/paramValidation), on-disk hash bucketing
(sharding.ts).
Docs: scrubbed deleted-subsystem references (JS quantization knobs, cloud/OPFS
adapters, partitioning, old zero-config API) across 14 files; deleted two wholly-
obsolete feature docs (complete-feature-list, v3-features); rewrote
architecture/zero-config for 8.0.
~3,700 LOC removed. Build clean; 1392 unit + 24 db-mvcc green.
5.4 KiB
5.4 KiB
Architecture Overview
Brainy is a multi-dimensional AI database that combines vector similarity, graph relationships, and metadata filtering into a unified query system. This document provides a comprehensive overview of the system architecture.
Core Components
Brainy (Main Entry Point)
The central orchestrator that manages all subsystems:
- 4-Index Architecture: MetadataIndex, vector index, GraphAdjacencyIndex, DeletedItemsIndex (see Index Architecture)
- Storage System: FileSystem and Memory adapters
- Augmentation System: Extensible plugin architecture
- Triple Intelligence: Unified query engine
Triple Intelligence Engine
Brainy's revolutionary feature that unifies three types of search:
- Vector Search: Semantic similarity via the pluggable vector index
- Graph Traversal: Relationship-based queries
- Field Filtering: Precise metadata filtering with O(1) performance
// Single query combining all three intelligence types
const results = await brain.find({
like: "machine learning papers", // Vector similarity
connected: { to: "research-team", depth: 2 }, // Graph traversal
where: { published: { $gte: "2024-01-01" } } // Metadata filtering
})
Storage Architecture
brainy-data/
├── _system/ # System management
│ └── statistics.json
├── nouns/ # Entity data storage
│ └── {uuid}.json
├── metadata/ # Metadata and indexing
│ ├── {uuid}.json
│ ├── __entity_registry__.json
│ └── __metadata_index__*.json
├── verbs/ # Relationship storage
└── locks/ # Concurrent access control
Vector Index
Pluggable vector index (VectorIndexProvider) for efficient nearest-neighbor search. The default JS implementation, JsHnswVectorIndex, uses a hierarchical graph:
- Performance: O(log n) search complexity
- Configurable recall:
fast/balanced/accuratepresets trade recall for latency - Scalable: Handles millions of vectors per process
- Persistent: Serializable to storage
- Swappable: Replace with a native implementation (such as
@soulcraft/cortex) via the plugin system without changing application code
Metadata Index Manager
High-performance field indexing system:
- O(1) Lookups: Inverted index for field→value→IDs mapping
- Query Support: equals, anyOf, allOf, range queries
- Chunked Storage: Supports massive datasets
- Auto-indexing: Automatically maintains indexes on updates
Performance Characteristics
Operation Complexity
- Vector Search: O(log n) via the vector index
- Field Filtering: O(1) via inverted indexes
- Graph Traversal: O(V + E) for breadth-first search
- Add Operation: O(log n) for index insertion
- Update Operation: O(1) for metadata updates
Memory Usage
- Base Memory: ~50MB for core system
- Per Vector: ~1KB (384 dimensions × 4 bytes)
- Index Overhead: ~20% of vector data
- Cache Size: Configurable (default 1000 entries)
Throughput
- Writes: 1000+ ops/second (with batching)
- Reads: 10,000+ ops/second
- Search: 100+ queries/second (varies by complexity)
Augmentation System
Brainy's extensible plugin architecture allows for powerful enhancements:
Core Augmentations
- Entity Registry: High-speed deduplication for streaming data
- Batch Processing: Optimized bulk operations
- Request Deduplicator: Prevents duplicate processing
Creating Custom Augmentations
class CustomAugmentation extends BrainyAugmentation {
async onInit(brain: Brainy): Promise<void> {
// Initialize augmentation
}
async onAdd(item: any, brain: Brainy): Promise<any> {
// Process item before adding
return item
}
}
Caching Strategy
Multi-layered caching for optimal performance:
- Search Cache: LRU cache for query results
- Metadata Cache: Field index caching
- Pattern Cache: NLP pattern matching cache
- Entity Cache: In-memory entity registry
Integration Points
Key Objects for Extensions
brain.index: Access the vector indexbrain.metadataIndex: Access field indexingbrain.graphIndex: Access graph adjacency indexbrain.storage: Access storage layerbrain.augmentations: Access augmentation manager
For detailed information about each index, see Index Architecture.
Event System
brain.on('add', (item) => console.log('Item added:', item))
brain.on('search', (query) => console.log('Search performed:', query))
brain.on('error', (error) => console.error('Error:', error))
Best Practices
When Adding Features
- Check if similar functionality exists
- Consider if it should be an augmentation
- Use existing indexes and caches
- Avoid duplicating functionality
- Follow the established patterns
Performance Optimization
- Use batch operations for bulk data
- Enable appropriate caching
- Choose the right storage adapter
- Configure index parameters for your use case
- Monitor statistics for bottlenecks
Next Steps
- Index Architecture - Deep dive into the 4-index system
- Storage Architecture - Deep dive into storage system
- Triple Intelligence - Advanced query system
- API Reference - Complete API documentation