The old config-generation subsystem (src/config/ + autoConfiguration.ts) was
superseded during the 8.0 rework and never wired into init(): it emitted settings
for a partitioning subsystem that no longer exists and probed deleted cloud env
vars. The live zero-config path is inline — recall preset → HNSW knobs, storage
auto-detect, auto persistMode, container-memory-aware cache sizing.
The storage progressive-init / cloud-detection cluster was equally dead after the
cloud adapters were dropped: isCloudStorage() is permanently false (no overriders),
scheduleBackgroundInit/runBackgroundInit were never called (the latter an empty
body), initMode was never assigned, and Brainy.isFullyInitialized()/
awaitBackgroundInit() were always-trivial with zero callers. scheduleCountPersist()
collapses to its only-ever-taken immediate write-through path.
Removed:
- src/config/{index,zeroConfig,storageAutoConfig,modelAutoConfig,sharedConfigManager}.ts
- src/utils/autoConfiguration.ts + the inert BrainyZeroConfig export
- Brainy.isFullyInitialized()/awaitBackgroundInit() (+ BrainyInterface decls)
- InitMode type, isCloudStorage/detectCloudEnvironment/resolveInitMode,
scheduleBackgroundInit/runBackgroundInit/ensureValidatedForWrite and their state
- Dead cloud env-var probes (K_SERVICE/K_REVISION/AWS_LAMBDA_FUNCTION_NAME/
FUNCTIONS_TARGET/AZURE_FUNCTIONS_ENVIRONMENT)
Kept (verified live): production-detection logging (environment.ts), container-
memory cache sizing (memoryDetection/paramValidation), on-disk hash bucketing
(sharding.ts).
Docs: scrubbed deleted-subsystem references (JS quantization knobs, cloud/OPFS
adapters, partitioning, old zero-config API) across 14 files; deleted two wholly-
obsolete feature docs (complete-feature-list, v3-features); rewrote
architecture/zero-config for 8.0.
~3,700 LOC removed. Build clean; 1392 unit + 24 db-mvcc green.
150 lines
No EOL
5.4 KiB
Markdown
150 lines
No EOL
5.4 KiB
Markdown
# Architecture Overview
|
||
|
||
Brainy is a multi-dimensional AI database that combines vector similarity, graph relationships, and metadata filtering into a unified query system. This document provides a comprehensive overview of the system architecture.
|
||
|
||
## Core Components
|
||
|
||
### Brainy (Main Entry Point)
|
||
The central orchestrator that manages all subsystems:
|
||
- **4-Index Architecture**: MetadataIndex, vector index, GraphAdjacencyIndex, DeletedItemsIndex (see [Index Architecture](./index-architecture.md))
|
||
- **Storage System**: FileSystem and Memory adapters
|
||
- **Augmentation System**: Extensible plugin architecture
|
||
- **Triple Intelligence**: Unified query engine
|
||
|
||
### Triple Intelligence Engine
|
||
Brainy's revolutionary feature that unifies three types of search:
|
||
- **Vector Search**: Semantic similarity via the pluggable vector index
|
||
- **Graph Traversal**: Relationship-based queries
|
||
- **Field Filtering**: Precise metadata filtering with O(1) performance
|
||
|
||
```typescript
|
||
// Single query combining all three intelligence types
|
||
const results = await brain.find({
|
||
like: "machine learning papers", // Vector similarity
|
||
connected: { to: "research-team", depth: 2 }, // Graph traversal
|
||
where: { published: { $gte: "2024-01-01" } } // Metadata filtering
|
||
})
|
||
```
|
||
|
||
### Storage Architecture
|
||
|
||
```
|
||
brainy-data/
|
||
├── _system/ # System management
|
||
│ └── statistics.json
|
||
├── nouns/ # Entity data storage
|
||
│ └── {uuid}.json
|
||
├── metadata/ # Metadata and indexing
|
||
│ ├── {uuid}.json
|
||
│ ├── __entity_registry__.json
|
||
│ └── __metadata_index__*.json
|
||
├── verbs/ # Relationship storage
|
||
└── locks/ # Concurrent access control
|
||
```
|
||
|
||
### Vector Index
|
||
Pluggable vector index (`VectorIndexProvider`) for efficient nearest-neighbor search. The default JS implementation, `JsHnswVectorIndex`, uses a hierarchical graph:
|
||
- **Performance**: O(log n) search complexity
|
||
- **Configurable recall**: `fast` / `balanced` / `accurate` presets trade recall for latency
|
||
- **Scalable**: Handles millions of vectors per process
|
||
- **Persistent**: Serializable to storage
|
||
- **Swappable**: Replace with a native implementation (such as `@soulcraft/cortex`) via the plugin system without changing application code
|
||
|
||
### Metadata Index Manager
|
||
High-performance field indexing system:
|
||
- **O(1) Lookups**: Inverted index for field→value→IDs mapping
|
||
- **Query Support**: equals, anyOf, allOf, range queries
|
||
- **Chunked Storage**: Supports massive datasets
|
||
- **Auto-indexing**: Automatically maintains indexes on updates
|
||
|
||
## Performance Characteristics
|
||
|
||
### Operation Complexity
|
||
- **Vector Search**: O(log n) via the vector index
|
||
- **Field Filtering**: O(1) via inverted indexes
|
||
- **Graph Traversal**: O(V + E) for breadth-first search
|
||
- **Add Operation**: O(log n) for index insertion
|
||
- **Update Operation**: O(1) for metadata updates
|
||
|
||
### Memory Usage
|
||
- **Base Memory**: ~50MB for core system
|
||
- **Per Vector**: ~1KB (384 dimensions × 4 bytes)
|
||
- **Index Overhead**: ~20% of vector data
|
||
- **Cache Size**: Configurable (default 1000 entries)
|
||
|
||
### Throughput
|
||
- **Writes**: 1000+ ops/second (with batching)
|
||
- **Reads**: 10,000+ ops/second
|
||
- **Search**: 100+ queries/second (varies by complexity)
|
||
|
||
## Augmentation System
|
||
|
||
Brainy's extensible plugin architecture allows for powerful enhancements:
|
||
|
||
### Core Augmentations
|
||
- **Entity Registry**: High-speed deduplication for streaming data
|
||
- **Batch Processing**: Optimized bulk operations
|
||
- **Request Deduplicator**: Prevents duplicate processing
|
||
|
||
### Creating Custom Augmentations
|
||
```typescript
|
||
class CustomAugmentation extends BrainyAugmentation {
|
||
async onInit(brain: Brainy): Promise<void> {
|
||
// Initialize augmentation
|
||
}
|
||
|
||
async onAdd(item: any, brain: Brainy): Promise<any> {
|
||
// Process item before adding
|
||
return item
|
||
}
|
||
}
|
||
```
|
||
|
||
## Caching Strategy
|
||
|
||
Multi-layered caching for optimal performance:
|
||
- **Search Cache**: LRU cache for query results
|
||
- **Metadata Cache**: Field index caching
|
||
- **Pattern Cache**: NLP pattern matching cache
|
||
- **Entity Cache**: In-memory entity registry
|
||
|
||
## Integration Points
|
||
|
||
### Key Objects for Extensions
|
||
- `brain.index`: Access the vector index
|
||
- `brain.metadataIndex`: Access field indexing
|
||
- `brain.graphIndex`: Access graph adjacency index
|
||
- `brain.storage`: Access storage layer
|
||
- `brain.augmentations`: Access augmentation manager
|
||
|
||
For detailed information about each index, see [Index Architecture](./index-architecture.md).
|
||
|
||
### Event System
|
||
```typescript
|
||
brain.on('add', (item) => console.log('Item added:', item))
|
||
brain.on('search', (query) => console.log('Search performed:', query))
|
||
brain.on('error', (error) => console.error('Error:', error))
|
||
```
|
||
|
||
## Best Practices
|
||
|
||
### When Adding Features
|
||
1. Check if similar functionality exists
|
||
2. Consider if it should be an augmentation
|
||
3. Use existing indexes and caches
|
||
4. Avoid duplicating functionality
|
||
5. Follow the established patterns
|
||
|
||
### Performance Optimization
|
||
1. Use batch operations for bulk data
|
||
2. Enable appropriate caching
|
||
3. Choose the right storage adapter
|
||
4. Configure index parameters for your use case
|
||
5. Monitor statistics for bottlenecks
|
||
|
||
## Next Steps
|
||
|
||
- [Index Architecture](./index-architecture.md) - Deep dive into the 4-index system
|
||
- [Storage Architecture](./storage-architecture.md) - Deep dive into storage system
|
||
- [Triple Intelligence](./triple-intelligence.md) - Advanced query system
|
||
- [API Reference](../api/README.md) - Complete API documentation |