brainy/docs/architecture/overview.md
David Snelling bf4a333f9b docs: rename the native provider to @soulcraft/cor across public docs and JSDoc
The 3.0 native engine ships as @soulcraft/cor (brainy 8.x <-> cor 3.x are a
version-matched pair); the old @soulcraft/cortex package stays on the 2.x/7.x
line. Public docs and .d.ts-visible comments now name the correct package.
Historical 2.x contract references (e.g. the 2.3.1 read-side fallback) keep
the old name deliberately.
2026-07-02 15:11:41 -07:00

5.4 KiB
Raw Permalink Blame History

Architecture Overview

Brainy is a multi-dimensional AI database that combines vector similarity, graph relationships, and metadata filtering into a unified query system. This document provides a comprehensive overview of the system architecture.

Core Components

Brainy (Main Entry Point)

The central orchestrator that manages all subsystems:

  • 4-Index Architecture: MetadataIndex, vector index, GraphAdjacencyIndex, DeletedItemsIndex (see Index Architecture)
  • Storage System: FileSystem and Memory adapters
  • Augmentation System: Extensible plugin architecture
  • Triple Intelligence: Unified query engine

Triple Intelligence Engine

Brainy's revolutionary feature that unifies three types of search:

  • Vector Search: Semantic similarity via the pluggable vector index
  • Graph Traversal: Relationship-based queries
  • Field Filtering: Precise metadata filtering with O(1) performance
// Single query combining all three intelligence types
const results = await brain.find({
  like: "machine learning papers",              // Vector similarity
  connected: { to: "research-team", depth: 2 }, // Graph traversal  
  where: { published: { $gte: "2024-01-01" } }  // Metadata filtering
})

Storage Architecture

brainy-data/
├── _system/           # System management
│   └── statistics.json
├── nouns/            # Entity data storage
│   └── {uuid}.json
├── metadata/         # Metadata and indexing
│   ├── {uuid}.json
│   ├── __entity_registry__.json
│   └── __metadata_index__*.json
├── verbs/            # Relationship storage
└── locks/            # Concurrent access control

Vector Index

Pluggable vector index (VectorIndexProvider) for efficient nearest-neighbor search. The default JS implementation, JsHnswVectorIndex, uses a hierarchical graph:

  • Performance: O(log n) search complexity
  • Configurable recall: fast / balanced / accurate presets trade recall for latency
  • Scalable: Handles millions of vectors per process
  • Persistent: Serializable to storage
  • Swappable: Replace with a native implementation (such as @soulcraft/cor) via the plugin system without changing application code

Metadata Index Manager

High-performance field indexing system:

  • O(1) Lookups: Inverted index for field→value→IDs mapping
  • Query Support: equals, anyOf, allOf, range queries
  • Chunked Storage: Supports massive datasets
  • Auto-indexing: Automatically maintains indexes on updates

Performance Characteristics

Operation Complexity

  • Vector Search: O(log n) via the vector index
  • Field Filtering: O(1) via inverted indexes
  • Graph Traversal: O(V + E) for breadth-first search
  • Add Operation: O(log n) for index insertion
  • Update Operation: O(1) for metadata updates

Memory Usage

  • Base Memory: ~50MB for core system
  • Per Vector: ~1KB (384 dimensions × 4 bytes)
  • Index Overhead: ~20% of vector data
  • Cache Size: Configurable (default 1000 entries)

Throughput

  • Writes: 1000+ ops/second (with batching)
  • Reads: 10,000+ ops/second
  • Search: 100+ queries/second (varies by complexity)

Augmentation System

Brainy's extensible plugin architecture allows for powerful enhancements:

Core Augmentations

  • Entity Registry: High-speed deduplication for streaming data
  • Batch Processing: Optimized bulk operations
  • Request Deduplicator: Prevents duplicate processing

Creating Custom Augmentations

class CustomAugmentation extends BrainyAugmentation {
  async onInit(brain: Brainy): Promise<void> {
    // Initialize augmentation
  }
  
  async onAdd(item: any, brain: Brainy): Promise<any> {
    // Process item before adding
    return item
  }
}

Caching Strategy

Multi-layered caching for optimal performance:

  • Search Cache: LRU cache for query results
  • Metadata Cache: Field index caching
  • Pattern Cache: NLP pattern matching cache
  • Entity Cache: In-memory entity registry

Integration Points

Key Objects for Extensions

  • brain.index: Access the vector index
  • brain.metadataIndex: Access field indexing
  • brain.graphIndex: Access graph adjacency index
  • brain.storage: Access storage layer
  • brain.augmentations: Access augmentation manager

For detailed information about each index, see Index Architecture.

Event System

brain.on('add', (item) => console.log('Item added:', item))
brain.on('search', (query) => console.log('Search performed:', query))
brain.on('error', (error) => console.error('Error:', error))

Best Practices

When Adding Features

  1. Check if similar functionality exists
  2. Consider if it should be an augmentation
  3. Use existing indexes and caches
  4. Avoid duplicating functionality
  5. Follow the established patterns

Performance Optimization

  1. Use batch operations for bulk data
  2. Enable appropriate caching
  3. Choose the right storage adapter
  4. Configure index parameters for your use case
  5. Monitor statistics for bottlenecks

Next Steps