Compare commits
No commits in common. "session-4-optimizations" and "main" have entirely different histories.
session-4-
...
main
867 changed files with 224767 additions and 149959 deletions
12
.aiignore
Normal file
12
.aiignore
Normal file
|
|
@ -0,0 +1,12 @@
|
||||||
|
# An .aiignore file follows the same syntax as a .gitignore file.
|
||||||
|
# .gitignore documentation: https://git-scm.com/docs/gitignore
|
||||||
|
|
||||||
|
# you can ignore files
|
||||||
|
.DS_Store
|
||||||
|
*.log
|
||||||
|
*.tmp
|
||||||
|
|
||||||
|
# or folders
|
||||||
|
dist/
|
||||||
|
build/
|
||||||
|
out/
|
||||||
170
.claude/skills/architecture.md
Normal file
170
.claude/skills/architecture.md
Normal file
|
|
@ -0,0 +1,170 @@
|
||||||
|
# Brainy Architecture Reference
|
||||||
|
|
||||||
|
## What Is Brainy
|
||||||
|
|
||||||
|
@soulcraft/brainy (v7.17.0) is a Universal Knowledge Protocol -- a Triple Intelligence database combining vector search, graph traversal, and metadata filtering in a single library. Published to npm as a public MIT-licensed package.
|
||||||
|
|
||||||
|
## Core Architecture
|
||||||
|
|
||||||
|
### Storage Layer (`src/storage/`)
|
||||||
|
- **StorageAdapter interface** (`src/coreTypes.ts:576`): The contract ALL storage backends implement. ALWAYS check this interface before adding storage methods.
|
||||||
|
- **BaseStorage** (`src/storage/baseStorage.ts`): Base implementation with built-in type-aware partitioning (TypeAwareStorageAdapter was removed -- functionality merged into BaseStorage).
|
||||||
|
- **Adapters** (`src/storage/adapters/`):
|
||||||
|
- `fileSystemStorage.ts` -- local filesystem
|
||||||
|
- `memoryStorage.ts` -- in-memory
|
||||||
|
- `baseStorageAdapter.ts` -- shared adapter base (counts, batch ops)
|
||||||
|
- Cloud + OPFS adapters were removed in 8.0 (cloud backup is operator tooling)
|
||||||
|
- **Generational MVCC / Db API** (`src/db/`): immutable `Db` values over generation-stamped records
|
||||||
|
- `db.ts` (the `Db` value), `generationStore.ts` (record layer + commit protocol), `types.ts`, `errors.ts`, `whereMatcher.ts`
|
||||||
|
- Design record: `docs/ADR-001-generational-mvcc.md`; replaced the pre-8.0 COW branching + versioning subsystems
|
||||||
|
|
||||||
|
### Vector Search (`src/hnsw/`)
|
||||||
|
- `hnswIndex.ts` -- HNSW-based approximate nearest neighbor search
|
||||||
|
- `typeAwareHNSWIndex.ts` -- type-partitioned vector search
|
||||||
|
- NOT in `src/intelligence/` (that directory does not exist)
|
||||||
|
|
||||||
|
### Graph Engine (`src/graph/`)
|
||||||
|
- `graphAdjacencyIndex.ts` -- adjacency-based graph representation
|
||||||
|
- `pathfinding.ts` -- relationship traversal and pathfinding
|
||||||
|
- `lsm/` -- LSM tree implementation for graph storage
|
||||||
|
|
||||||
|
### Metadata Index (`src/utils/metadataIndex.ts`)
|
||||||
|
- O(1) exact match via hash indexes
|
||||||
|
- O(log n) range queries via sorted indexes
|
||||||
|
- Roaring bitmap set operations for efficient filtering
|
||||||
|
- Adaptive chunking strategy (`metadataIndexChunking.ts`)
|
||||||
|
- Caching layer (`metadataIndexCache.ts`)
|
||||||
|
|
||||||
|
### Triple Intelligence (`src/triple/`)
|
||||||
|
- `TripleIntelligenceSystem.ts` -- combines vector + graph + metadata into unified queries
|
||||||
|
- Lazy-loaded indexes (loaded on first use, not at startup)
|
||||||
|
|
||||||
|
### Neural/AI Components (`src/neural/`)
|
||||||
|
- Smart Importers (`src/importers/`): CSV, Excel, PDF, DOCX, YAML, JSON, Markdown, Orchestrator
|
||||||
|
- `SmartExtractor.ts` -- entity extraction from unstructured data
|
||||||
|
- `SmartRelationshipExtractor.ts` -- relationship detection
|
||||||
|
- `NeuralEntityExtractor.ts` -- ML-based entity recognition
|
||||||
|
- Natural language processing utilities
|
||||||
|
|
||||||
|
### Distributed Systems (`src/distributed/`)
|
||||||
|
- Distributed Coordinator for multi-node operation
|
||||||
|
- Shard Manager for data partitioning
|
||||||
|
- Cache Synchronization across nodes
|
||||||
|
- Read/Write separation
|
||||||
|
- Network and HTTP transport layers
|
||||||
|
- Storage discovery and shard migration
|
||||||
|
|
||||||
|
### Transaction Management (`src/transaction/`)
|
||||||
|
- TransactionManager for ACID operations
|
||||||
|
- Operations: SaveNoun, AddToHNSW, UpdateMetadata, etc.
|
||||||
|
- Distributed transaction support
|
||||||
|
|
||||||
|
### Integration Hub (`src/integrations/`)
|
||||||
|
- Google Sheets integration
|
||||||
|
- OData (Open Data Protocol)
|
||||||
|
- Server-Sent Events (SSE)
|
||||||
|
- Webhooks
|
||||||
|
- Event bus system
|
||||||
|
|
||||||
|
### Virtual Filesystem (`src/vfs/`)
|
||||||
|
- `VirtualFileSystem.ts` -- full VFS implementation (87 KB)
|
||||||
|
- `PathResolver.ts`, `FSCompat.ts`, `MimeTypeDetector.ts`, `TreeUtils.ts`
|
||||||
|
- Subdirectories: `semantic/` (semantic search), `streams/` (streaming), `importers/`
|
||||||
|
|
||||||
|
### MCP Support (`src/mcp/`)
|
||||||
|
- BrainyMCPAdapter, MCPAugmentationToolset, BrainyMCPService
|
||||||
|
- Model Control Protocol request/response handling
|
||||||
|
|
||||||
|
### Aggregation Engine (`src/aggregation/`)
|
||||||
|
- **AggregationIndex** (`AggregationIndex.ts`): Write-time incremental aggregation — SUM, COUNT, AVG, MIN, MAX with GROUP BY and time windows
|
||||||
|
- **Time Windows** (`timeWindows.ts`): ISO 8601 bucketing — hour, day, week, month, quarter, year, custom intervals
|
||||||
|
- **Materializer** (`materializer.ts`): Debounced writes of aggregate results as `NounType.Measurement` entities
|
||||||
|
- Integrates into `brain.find({ aggregate })` for unified query API
|
||||||
|
- Write hooks in `add()`, `update()`, `delete()` for O(1) incremental updates
|
||||||
|
- `'aggregation'` provider key enables native plugin acceleration
|
||||||
|
|
||||||
|
### Additional Systems
|
||||||
|
- **CLI** (`src/cli/`): Complete command-line tool with interactive mode and catalog system
|
||||||
|
- **Migration** (`src/migration/`): MigrationRunner for database schema migrations
|
||||||
|
- **Embeddings** (`src/embeddings/`): Embedding manager with Candle-WASM Rust source
|
||||||
|
- **Streaming** (`src/streaming/`): Pipeline support with adaptive backpressure
|
||||||
|
- **Versioning** (`src/versioning/`): VersioningAPI for data versioning
|
||||||
|
- **Plugin System**: Registry-based plugin architecture
|
||||||
|
- **Patterns** (`src/patterns/`): 7 pattern library JSON files
|
||||||
|
|
||||||
|
## Type System
|
||||||
|
- **NounType** (42 types, `src/types/graphTypes.ts:850-893`): Person, Organization, Concept, Collection, Document, Task, Project, etc.
|
||||||
|
- **VerbType** (127 types, `src/types/graphTypes.ts:900-1087`): Contains, RelatedTo, PartOf, Creates, DependsOn, MemberOf, etc.
|
||||||
|
- All types in `src/types/`
|
||||||
|
|
||||||
|
## Module Exports (`src/index.ts`)
|
||||||
|
38+ named exports including: Brainy class, configuration types, neural APIs (NeuralImport, NeuralEntityExtractor, SmartExtractor, SmartRelationshipExtractor), distance functions, plugin system, migration system, embedding functions, storage adapters, COW infrastructure, pipeline utilities, graph types, MCP components, integration hub, OData utilities, and more.
|
||||||
|
|
||||||
|
## File Structure
|
||||||
|
```
|
||||||
|
src/
|
||||||
|
├── index.ts # 38+ public exports
|
||||||
|
├── brainy.ts # Main Brainy class (6,500+ lines)
|
||||||
|
├── setup.ts # Initialization polyfills
|
||||||
|
├── coreTypes.ts # StorageAdapter interface + core types
|
||||||
|
├── storage/
|
||||||
|
│ ├── baseStorage.ts # Base storage (includes type-aware)
|
||||||
|
│ ├── adapters/ # All storage backends + cloud adapters
|
||||||
|
│ └── cow/ # Copy-on-Write versioning
|
||||||
|
├── hnsw/ # HNSW vector search
|
||||||
|
├── graph/ # Graph engine + pathfinding + LSM
|
||||||
|
├── triple/ # Triple Intelligence system
|
||||||
|
├── neural/ # Smart extractors + NLP
|
||||||
|
├── importers/ # File format importers (8 types)
|
||||||
|
├── distributed/ # Distributed database (16 files)
|
||||||
|
├── transaction/ # ACID transactions (6 files)
|
||||||
|
├── integrations/ # Sheets, OData, SSE, Webhooks
|
||||||
|
├── vfs/ # Virtual filesystem + semantic search
|
||||||
|
├── mcp/ # Model Control Protocol
|
||||||
|
├── cli/ # Command-line interface
|
||||||
|
├── migration/ # Schema migrations
|
||||||
|
├── embeddings/ # Embedding manager + Candle-WASM
|
||||||
|
├── streaming/ # Pipeline + backpressure
|
||||||
|
├── versioning/ # Versioning API
|
||||||
|
├── types/ # TypeScript type definitions
|
||||||
|
├── utils/ # Metadata index, logging, etc.
|
||||||
|
├── config/ # Configuration system
|
||||||
|
├── patterns/ # Pattern library
|
||||||
|
├── api/ # API layer
|
||||||
|
├── interfaces/ # Interface definitions
|
||||||
|
├── shared/ # Shared utilities
|
||||||
|
├── data/ # Data utilities
|
||||||
|
├── errors/ # Error handling
|
||||||
|
├── critical/ # Critical error handling
|
||||||
|
├── universal/ # Universal utilities
|
||||||
|
├── import/ # Import functionality
|
||||||
|
└── scripts/ # Build scripts
|
||||||
|
```
|
||||||
|
|
||||||
|
## Initialization
|
||||||
|
`brainy.ts` `init()` method performs initialization cascade:
|
||||||
|
1. Load plugins
|
||||||
|
2. Initialize storage
|
||||||
|
3. Enable COW (Copy-on-Write)
|
||||||
|
4. Set up embeddings
|
||||||
|
5. Initialize caches
|
||||||
|
6. Set up graph indexes
|
||||||
|
7. Initialize VFS
|
||||||
|
8. Set up transaction manager
|
||||||
|
9. Initialize distributed components (if enabled)
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
- Framework: Vitest
|
||||||
|
- Run: `npm test`
|
||||||
|
- Test directories:
|
||||||
|
- `tests/unit/` -- unit tests
|
||||||
|
- `tests/integration/` -- integration tests
|
||||||
|
- `tests/benchmarks/` -- performance benchmarks (NOT tests/performance/)
|
||||||
|
- `tests/comprehensive/` -- comprehensive test suites
|
||||||
|
- `tests/api/` -- API tests
|
||||||
|
- `tests/helpers/` -- test utilities
|
||||||
|
|
||||||
|
## Release
|
||||||
|
- `npm run release:patch/minor/major` -- fully automated via `scripts/release.sh`
|
||||||
|
- `npm run release:dry` -- preview without changes
|
||||||
|
- Uses conventional commits for changelog generation
|
||||||
57
.dockerignore
Normal file
57
.dockerignore
Normal file
|
|
@ -0,0 +1,57 @@
|
||||||
|
# Git
|
||||||
|
.git
|
||||||
|
.gitignore
|
||||||
|
|
||||||
|
# Development
|
||||||
|
.vscode
|
||||||
|
.idea
|
||||||
|
*.swp
|
||||||
|
*.swo
|
||||||
|
.DS_Store
|
||||||
|
|
||||||
|
# Node
|
||||||
|
node_modules
|
||||||
|
npm-debug.log*
|
||||||
|
yarn-debug.log*
|
||||||
|
yarn-error.log*
|
||||||
|
|
||||||
|
# Testing
|
||||||
|
tests
|
||||||
|
*.test.ts
|
||||||
|
*.test.js
|
||||||
|
coverage
|
||||||
|
.nyc_output
|
||||||
|
|
||||||
|
# Documentation (keep only essentials)
|
||||||
|
docs
|
||||||
|
*.md
|
||||||
|
!README.md
|
||||||
|
!LICENSE
|
||||||
|
|
||||||
|
# Build artifacts (will be built in Docker)
|
||||||
|
dist
|
||||||
|
build
|
||||||
|
*.tsbuildinfo
|
||||||
|
|
||||||
|
# Environment
|
||||||
|
.env
|
||||||
|
.env.*
|
||||||
|
|
||||||
|
# Strategy and private docs
|
||||||
|
.strategy
|
||||||
|
CLAUDE.md
|
||||||
|
|
||||||
|
# Development files
|
||||||
|
docker-compose.yml
|
||||||
|
Dockerfile
|
||||||
|
.dockerignore
|
||||||
|
|
||||||
|
# Data (should be mounted, not baked in)
|
||||||
|
data
|
||||||
|
*.db
|
||||||
|
*.sqlite
|
||||||
|
|
||||||
|
# Models (should be downloaded at runtime or mounted)
|
||||||
|
models
|
||||||
|
*.onnx
|
||||||
|
*.bin
|
||||||
40
.forgejo/workflows/ci.yml
Normal file
40
.forgejo/workflows/ci.yml
Normal file
|
|
@ -0,0 +1,40 @@
|
||||||
|
name: CI
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
pull_request:
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
node:
|
||||||
|
name: Node ${{ matrix.node-version }}
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
strategy:
|
||||||
|
fail-fast: false
|
||||||
|
matrix:
|
||||||
|
node-version: ['22', '24']
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- uses: actions/setup-node@v4
|
||||||
|
with:
|
||||||
|
node-version: ${{ matrix.node-version }}
|
||||||
|
cache: npm
|
||||||
|
- run: npm ci
|
||||||
|
- run: npm run test:unit
|
||||||
|
|
||||||
|
bun:
|
||||||
|
name: Bun (latest)
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- uses: actions/setup-node@v4
|
||||||
|
with:
|
||||||
|
node-version: '22'
|
||||||
|
cache: npm
|
||||||
|
- uses: oven-sh/setup-bun@v2
|
||||||
|
with:
|
||||||
|
bun-version: latest
|
||||||
|
- run: npm ci
|
||||||
|
# test:bun imports the built dist/, so build first.
|
||||||
|
- run: npm run build
|
||||||
|
# Bun as a runtime is the supported Bun story (`bun add` / `bun run`).
|
||||||
|
- run: npm run test:bun
|
||||||
56
.gitignore
vendored
56
.gitignore
vendored
|
|
@ -18,6 +18,7 @@ build/
|
||||||
|
|
||||||
# Runtime data
|
# Runtime data
|
||||||
brainy-data/
|
brainy-data/
|
||||||
|
.brainy/
|
||||||
*.log
|
*.log
|
||||||
*.pid
|
*.pid
|
||||||
*.seed
|
*.seed
|
||||||
|
|
@ -30,6 +31,9 @@ coverage/
|
||||||
# Test results
|
# Test results
|
||||||
tests/results/
|
tests/results/
|
||||||
|
|
||||||
|
# Filesystem test artifacts (created by integration tests)
|
||||||
|
test-*/
|
||||||
|
|
||||||
# IDE files
|
# IDE files
|
||||||
.vscode/
|
.vscode/
|
||||||
.idea/
|
.idea/
|
||||||
|
|
@ -46,6 +50,9 @@ tmp/
|
||||||
temp/
|
temp/
|
||||||
*.tmp
|
*.tmp
|
||||||
|
|
||||||
|
# Planning and instruction files
|
||||||
|
plan.md
|
||||||
|
|
||||||
# Package files
|
# Package files
|
||||||
*.tgz
|
*.tgz
|
||||||
|
|
||||||
|
|
@ -55,8 +62,55 @@ INTERNAL_NOTES.md
|
||||||
TODO_PRIVATE.md
|
TODO_PRIVATE.md
|
||||||
*.tar.gz
|
*.tar.gz
|
||||||
|
|
||||||
|
# Strategy and planning documents (private)
|
||||||
|
.strategy/
|
||||||
|
# Removed: PRODUCTION_*.md (now these should be public documentation)
|
||||||
|
DISTRIBUTED_*.md
|
||||||
|
*_ASSESSMENT.md
|
||||||
|
*_ANALYSIS.md
|
||||||
|
*_TRUTH*.md
|
||||||
|
|
||||||
# Models (downloaded at runtime)
|
# Models (downloaded at runtime)
|
||||||
models/CLAUDE.md
|
models/
|
||||||
|
models-cache/
|
||||||
|
|
||||||
|
# But include bundled WASM model assets
|
||||||
|
!assets/models/
|
||||||
|
|
||||||
# Development planning files (not for commit)
|
# Development planning files (not for commit)
|
||||||
PLAN.md
|
PLAN.md
|
||||||
|
|
||||||
|
# Backup folders
|
||||||
|
backup-*
|
||||||
|
backup/
|
||||||
|
|
||||||
|
# Internal documentation
|
||||||
|
docs/internal/
|
||||||
|
|
||||||
|
# Cache files
|
||||||
|
*.cache
|
||||||
|
|
||||||
|
# Rust/Cargo build artifacts
|
||||||
|
src/embeddings/candle-wasm/target/
|
||||||
|
src/embeddings/candle-wasm/Cargo.lock
|
||||||
|
|
||||||
|
# Ignore the wasm-pack output dir's CONTENTS (note the `/*`, not `/`, so the
|
||||||
|
# re-includes below can take effect — git cannot re-include a file whose parent
|
||||||
|
# DIRECTORY is excluded). Keep the pre-built WASM committed: it ships in the npm
|
||||||
|
# package anyway, it lets consumers + CI build without a Rust/wasm-pack toolchain,
|
||||||
|
# and versioning it makes the shipped artifact reproducible (not "whatever the
|
||||||
|
# maintainer last built").
|
||||||
|
src/embeddings/wasm/pkg/*
|
||||||
|
!src/embeddings/wasm/pkg/*.wasm
|
||||||
|
!src/embeddings/wasm/pkg/*.js
|
||||||
|
!src/embeddings/wasm/pkg/*.d.ts
|
||||||
|
|
||||||
|
# Log files (redundant but explicit)
|
||||||
|
*.log
|
||||||
|
|
||||||
|
# Temporary files (redundant but explicit)
|
||||||
|
*.tmp
|
||||||
|
/.junie/guidelines.md
|
||||||
|
|
||||||
|
# Claude Code harness state
|
||||||
|
.claude/scheduled_tasks.lock
|
||||||
|
|
|
||||||
84
.npmignore
Normal file
84
.npmignore
Normal file
|
|
@ -0,0 +1,84 @@
|
||||||
|
# Source files (not needed in package)
|
||||||
|
src/
|
||||||
|
tests/
|
||||||
|
scripts/
|
||||||
|
coverage/
|
||||||
|
|
||||||
|
# Model files (downloaded on first use, not bundled)
|
||||||
|
models/
|
||||||
|
models-cache/
|
||||||
|
|
||||||
|
# Development and backup files
|
||||||
|
backup-*
|
||||||
|
backup-*/
|
||||||
|
docs/backup*/
|
||||||
|
|
||||||
|
# Documentation (except essentials)
|
||||||
|
*.md
|
||||||
|
!README.md
|
||||||
|
!LICENSE
|
||||||
|
!CHANGELOG.md
|
||||||
|
!MIGRATION.md
|
||||||
|
|
||||||
|
# Configuration files
|
||||||
|
.gitignore
|
||||||
|
.npmignore
|
||||||
|
tsconfig.json
|
||||||
|
vitest.config.ts
|
||||||
|
vitest.config.mts
|
||||||
|
*.config.js
|
||||||
|
*.config.ts
|
||||||
|
.eslintrc*
|
||||||
|
.prettierrc*
|
||||||
|
|
||||||
|
# Test files
|
||||||
|
test-*.js
|
||||||
|
test-*.ts
|
||||||
|
*.test.ts
|
||||||
|
*.test.js
|
||||||
|
*.spec.ts
|
||||||
|
*.spec.js
|
||||||
|
|
||||||
|
# Temporary and log files
|
||||||
|
*.log
|
||||||
|
*.tmp
|
||||||
|
tmp/
|
||||||
|
temp/
|
||||||
|
brainy-data/
|
||||||
|
|
||||||
|
# Git and CI files
|
||||||
|
.git/
|
||||||
|
.github/
|
||||||
|
.gitlab-ci.yml
|
||||||
|
.travis.yml
|
||||||
|
|
||||||
|
# IDE files
|
||||||
|
.vscode/
|
||||||
|
.idea/
|
||||||
|
*.swp
|
||||||
|
*.swo
|
||||||
|
|
||||||
|
# OS files
|
||||||
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
|
||||||
|
# Private files
|
||||||
|
PLAN.md
|
||||||
|
CLAUDE.md
|
||||||
|
INTERNAL_NOTES.md
|
||||||
|
TODO_PRIVATE.md
|
||||||
|
*-ANALYSIS.md
|
||||||
|
*-PLAN.md
|
||||||
|
|
||||||
|
# Build artifacts not needed
|
||||||
|
*.tsbuildinfo
|
||||||
|
*.map
|
||||||
|
|
||||||
|
# Development environment
|
||||||
|
.env*
|
||||||
|
.nvm*
|
||||||
|
.node-version
|
||||||
|
|
||||||
|
# Keep dist/ for the compiled code
|
||||||
|
# Keep bin/ for the CLI
|
||||||
|
# Keep package.json, package-lock.json
|
||||||
1
.nvmrc
Normal file
1
.nvmrc
Normal file
|
|
@ -0,0 +1 @@
|
||||||
|
22
|
||||||
|
|
@ -1,62 +0,0 @@
|
||||||
# Brain Patterns Optimization Plan
|
|
||||||
|
|
||||||
## Brain Pattern Operators (Complete List)
|
|
||||||
1. **Equality**: `equals`, `is`, `eq`
|
|
||||||
2. **Comparison**: `greaterThan`/`gt`, `lessThan`/`lt`, `greaterEqual`/`gte`, `lessEqual`/`lte`
|
|
||||||
3. **Range**: `between` (inclusive range)
|
|
||||||
4. **Membership**: `oneOf`/`in` (value in list)
|
|
||||||
5. **Contains**: `contains` (for arrays)
|
|
||||||
6. **Existence**: `exists` (field exists)
|
|
||||||
7. **Negation**: `not` (logical NOT)
|
|
||||||
8. **Logical**: `allOf` (AND), `anyOf` (OR)
|
|
||||||
|
|
||||||
## Current Architecture Issues
|
|
||||||
- MetadataIndex: O(1) hash lookups ONLY
|
|
||||||
- No sorted indices for ranges
|
|
||||||
- TripleIntelligence: String-based filtering (">2020") - TERRIBLE
|
|
||||||
- No numeric type detection
|
|
||||||
|
|
||||||
## Optimization Strategy
|
|
||||||
|
|
||||||
### Phase 1: Sorted Index Infrastructure ✅ DONE
|
|
||||||
```typescript
|
|
||||||
interface SortedFieldIndex {
|
|
||||||
values: Array<[value: any, ids: Set<string>]>
|
|
||||||
isDirty: boolean
|
|
||||||
fieldType: 'number' | 'string' | 'date' | 'mixed'
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 2: Binary Search Implementation ✅ DONE
|
|
||||||
- O(log n) range boundary finding
|
|
||||||
- Support inclusive/exclusive ranges
|
|
||||||
- Handle all comparison operators
|
|
||||||
|
|
||||||
### Phase 3: Automatic Type Detection
|
|
||||||
- Detect numeric fields on first value
|
|
||||||
- Maintain appropriate sorting
|
|
||||||
- Convert strings to numbers when possible
|
|
||||||
|
|
||||||
### Phase 4: Query Optimization
|
|
||||||
- Pre-filter with metadata index BEFORE vector search
|
|
||||||
- Use sorted indices for ALL range queries
|
|
||||||
- Cache sorted indices in memory
|
|
||||||
|
|
||||||
## Performance Targets
|
|
||||||
- Exact match: O(1) - hash lookup
|
|
||||||
- Range query: O(log n + m) - binary search + result size
|
|
||||||
- Combined filters: O(k * log n) - k conditions
|
|
||||||
- Memory overhead: ~2x current (hash + sorted)
|
|
||||||
|
|
||||||
## Implementation Status
|
|
||||||
- [x] Add SortedFieldIndex type
|
|
||||||
- [x] Add binary search methods
|
|
||||||
- [x] Update getIdsForFilter for all operators
|
|
||||||
- [ ] Fix TripleIntelligence to use index directly
|
|
||||||
- [ ] Add index statistics/monitoring
|
|
||||||
- [ ] Optimize memory usage
|
|
||||||
|
|
||||||
## Expected Performance Gains
|
|
||||||
- Range queries: 100-1000x faster
|
|
||||||
- Combined vector+metadata: 10-50x faster
|
|
||||||
- Memory usage: +50% (acceptable tradeoff)
|
|
||||||
4598
CHANGELOG.md
4598
CHANGELOG.md
File diff suppressed because it is too large
Load diff
208
CLAUDE.md
Normal file
208
CLAUDE.md
Normal file
|
|
@ -0,0 +1,208 @@
|
||||||
|
# Brainy - Claude Code Project Guide
|
||||||
|
|
||||||
|
This file provides guidance for Claude Code (and human contributors) when working on the Brainy codebase.
|
||||||
|
|
||||||
|
## Cross-Project Coordination
|
||||||
|
|
||||||
|
Handoff file: `/home/dpsifr/.strategy/PLATFORM-HANDOFF.md`
|
||||||
|
|
||||||
|
**At session START:** Read the handoff. Find rows where Owner = Brainy. Act on those first.
|
||||||
|
|
||||||
|
**At session END:** Mark completed actions ✅, delete rows you finished, delete threads with zero remaining actions. File must not grow. **If you shipped anything consumers need to know about, update `RELEASES.md` before closing.**
|
||||||
|
|
||||||
|
**Brainy's current open actions:** None. MIT open-source — no platform-specific actions.
|
||||||
|
|
||||||
|
**Current version:** run `npm view @soulcraft/brainy version` (never trust a hardcoded number here — this line went stale for months); consumer-facing changes tracked in `RELEASES.md`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Project Overview
|
||||||
|
|
||||||
|
Brainy is a Universal Knowledge Protocol -- a Triple Intelligence database that combines vector similarity search, graph traversal, and metadata filtering into a single TypeScript library. Published as `@soulcraft/brainy` on npm under the MIT license.
|
||||||
|
|
||||||
|
## Getting Started
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm install # Install dependencies
|
||||||
|
npm run build # Build the project
|
||||||
|
npm test # Run test suite (Vitest)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Full architecture reference: `.claude/skills/architecture.md`
|
||||||
|
|
||||||
|
### Core Systems
|
||||||
|
- **Storage** (`src/storage/`): Pluggable storage backends via StorageAdapter interface (`src/coreTypes.ts`)
|
||||||
|
- **Vector Search** (`src/hnsw/`): HNSW approximate nearest neighbor search
|
||||||
|
- **Graph Engine** (`src/graph/`): Relationship traversal with adjacency index and pathfinding
|
||||||
|
- **Metadata Index** (`src/utils/metadataIndex.ts`): O(1) exact match, O(log n) range queries
|
||||||
|
- **Triple Intelligence** (`src/triple/`): Unified query combining all three intelligence types
|
||||||
|
- **Aggregation Engine** (`src/aggregation/`): Write-time incremental SUM/COUNT/AVG/MIN/MAX with GROUP BY and time windows
|
||||||
|
- **Virtual Filesystem** (`src/vfs/`): Full VFS with semantic search
|
||||||
|
|
||||||
|
### Type System
|
||||||
|
- **NounType** (42 types): Entity classification -- Person, Concept, Collection, Document, Task, etc.
|
||||||
|
- **VerbType** (127 types): Relationship types -- Contains, RelatedTo, PartOf, Creates, DependsOn, etc.
|
||||||
|
- Defined in `src/types/graphTypes.ts`
|
||||||
|
|
||||||
|
## Code Standards
|
||||||
|
|
||||||
|
### TypeScript
|
||||||
|
- Strict mode enabled
|
||||||
|
- Target: ES2020, NodeNext module resolution
|
||||||
|
- All new code must be TypeScript
|
||||||
|
- Follow existing patterns -- read related code before writing
|
||||||
|
|
||||||
|
### Quality
|
||||||
|
- All code must compile without errors
|
||||||
|
- All code must have working tests that exercise real behavior
|
||||||
|
- No stub returns (`return {} as any`)
|
||||||
|
- No incomplete implementations with TODO comments
|
||||||
|
- If something can't be fully implemented, throw an explicit error rather than faking it
|
||||||
|
|
||||||
|
### Verification Before Code Changes
|
||||||
|
1. Check that interfaces and methods actually exist before using them
|
||||||
|
2. Check that type properties are in the type definitions
|
||||||
|
3. Run `npm test` -- tests must pass
|
||||||
|
4. Run `npm run build` -- build must succeed
|
||||||
|
|
||||||
|
### Testing
|
||||||
|
- Framework: Vitest
|
||||||
|
- Tests in `tests/` (unit, integration, benchmarks, comprehensive)
|
||||||
|
- Use in-memory storage for speed where possible
|
||||||
|
- Tests must exercise real behavior, not mock it
|
||||||
|
- Benchmarks are in `tests/benchmarks/` (not tests/performance/)
|
||||||
|
|
||||||
|
## Commit Conventions
|
||||||
|
|
||||||
|
Use [Conventional Commits](https://www.conventionalcommits.org/):
|
||||||
|
|
||||||
|
```
|
||||||
|
feat: add new feature (minor version bump)
|
||||||
|
fix: resolve bug (patch version bump)
|
||||||
|
docs: update documentation (patch version bump)
|
||||||
|
perf: improve performance (patch version bump)
|
||||||
|
refactor: restructure code (patch version bump)
|
||||||
|
test: add/update tests (patch version bump)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Important:** Never use `BREAKING CHANGE` in commit messages. Major version bumps are manual decisions only (`npm run release:major`).
|
||||||
|
|
||||||
|
## Docs Pipeline — soulcraft.com/docs
|
||||||
|
|
||||||
|
Docs in `docs/**/*.md` are published with the npm package (included in `files`) and synced to soulcraft.com/docs on every portal deploy. Frontmatter controls what appears publicly.
|
||||||
|
|
||||||
|
### Docs check triggers
|
||||||
|
|
||||||
|
Run the docs check whenever the user says ANY of:
|
||||||
|
- "commit, publish, release" / "release" / "publish"
|
||||||
|
- "update the docs" / "make sure docs are accurate" / "check the docs"
|
||||||
|
- "review docs" / "clean up docs"
|
||||||
|
|
||||||
|
### Pre-release docs check (MANDATORY before every release)
|
||||||
|
|
||||||
|
When the user says "commit, publish, release" or any variation, **before committing**:
|
||||||
|
|
||||||
|
1. **Scan all files changed in this session** (and any recently added `docs/*.md` files)
|
||||||
|
2. For each changed/new doc, decide: is this useful to external users?
|
||||||
|
- **Yes** → ensure it has complete frontmatter (add or update it)
|
||||||
|
- **No** (internal, migration, dev-only) → ensure it has no frontmatter or `public: false`
|
||||||
|
3. For docs that already have frontmatter, verify:
|
||||||
|
- `description` still matches the actual content
|
||||||
|
- `next` links still exist and are still the right follow-up pages
|
||||||
|
- `title` matches the doc's h1
|
||||||
|
4. Include frontmatter changes in the commit
|
||||||
|
|
||||||
|
### Frontmatter format
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
---
|
||||||
|
title: Human-readable title
|
||||||
|
slug: category/page-name # URL: soulcraft.com/docs/category/page-name
|
||||||
|
public: true # false or absent = not published
|
||||||
|
category: getting-started | concepts | guides | api
|
||||||
|
template: guide | concept | api # controls layout on soulcraft.com
|
||||||
|
order: 1 # sidebar position within category (lower = first)
|
||||||
|
description: One sentence. What this doc covers and why it matters.
|
||||||
|
next: # "Next steps" links shown at bottom of page
|
||||||
|
- category/other-slug
|
||||||
|
---
|
||||||
|
```
|
||||||
|
|
||||||
|
### Category guide
|
||||||
|
|
||||||
|
| category | use for |
|
||||||
|
|----------|---------|
|
||||||
|
| `getting-started` | installation, quick start, first steps |
|
||||||
|
| `concepts` | how the system works, mental models |
|
||||||
|
| `guides` | how to do specific things, recipes |
|
||||||
|
| `api` | method reference, signatures, parameters |
|
||||||
|
|
||||||
|
### What stays internal (no frontmatter / `public: false`)
|
||||||
|
|
||||||
|
- Release guides, developer learning paths
|
||||||
|
- Migration guides for old versions (v3→v4, v5.11)
|
||||||
|
- Architecture analysis docs (clustering algorithms, etc.)
|
||||||
|
- Anything in `docs/internal/`
|
||||||
|
- Deployment/ops/cost docs (cloud-run, kubernetes, cost-optimization)
|
||||||
|
|
||||||
|
## Release Process
|
||||||
|
|
||||||
|
Fully automated via `scripts/release.sh`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm run release:dry # Preview (no changes)
|
||||||
|
npm run release:patch # Bug fixes
|
||||||
|
npm run release:minor # New features
|
||||||
|
npm run release:major # Breaking changes (rare, manual decision)
|
||||||
|
```
|
||||||
|
|
||||||
|
The script: verifies clean git state, builds, tests, bumps version, updates CHANGELOG.md, commits, tags, pushes, publishes to npm, and creates a GitHub release.
|
||||||
|
|
||||||
|
After a successful release, remind the user:
|
||||||
|
> "Published. Deploy portal to pick up the new docs → go to the portal project and deploy."
|
||||||
|
|
||||||
|
Do NOT deploy portal from here. Portal is always deployed separately from within the portal project.
|
||||||
|
|
||||||
|
## Closed-Source Product Names — HARD RULE
|
||||||
|
|
||||||
|
Brainy is the only Soulcraft open-source project. Nothing in this repo — code, JSDoc, tests,
|
||||||
|
docs, RELEASES.md, CHANGELOG.md, commit messages — may reference closed-source Soulcraft
|
||||||
|
products by name (Workshop, Venue, Memory, Muse, Hall, Forge, Academy, Pulse, Heart,
|
||||||
|
Collective, SDK) or by their specific class/method names (`BookingDraftService`,
|
||||||
|
`getDemandHeatmap`, `systemKind`, etc.).
|
||||||
|
|
||||||
|
When recording a consumer-reported bug, regression scenario, or release note:
|
||||||
|
- Refer to "a consumer", "a downstream application", "a production deployment", or "an
|
||||||
|
internal report" — never name the product.
|
||||||
|
- All doc examples must use generic domain values (`'employee'`, `'customer'`, `'invoice'`,
|
||||||
|
`'milestone'`, `OrderService`, `/orders/...`), not product-specific schemas.
|
||||||
|
- Internal session artifacts (`.strategy/`, `~/.claude/plans/`, handoff files outside the
|
||||||
|
repo) MAY name products — those are not public.
|
||||||
|
|
||||||
|
If you catch yourself typing a product name into a tracked file, stop and rephrase.
|
||||||
|
|
||||||
|
## Performance Claims
|
||||||
|
|
||||||
|
When documenting performance characteristics:
|
||||||
|
- **MEASURED**: Cite the test file and line number
|
||||||
|
- **PROJECTED**: Clearly label as extrapolated from tested scale
|
||||||
|
- Never claim a performance figure without context or evidence
|
||||||
|
|
||||||
|
## Debugging
|
||||||
|
|
||||||
|
When a bug persists through 2+ fix attempts, switch to systematic debugging:
|
||||||
|
1. Add comprehensive logging at every step
|
||||||
|
2. Test with production-like data
|
||||||
|
3. Trace the complete execution path
|
||||||
|
4. Check both library code and consumer code
|
||||||
|
5. Verify with actual test execution before declaring fixed
|
||||||
|
|
||||||
|
## Key Paths
|
||||||
|
|
||||||
|
- Main class: `src/brainy.ts`
|
||||||
|
- Public API: `src/index.ts` (38+ exports)
|
||||||
|
- Storage interface: `src/coreTypes.ts`
|
||||||
|
- Type definitions: `src/types/`
|
||||||
|
- Strategy/planning docs: `.strategy/` (gitignored, not public)
|
||||||
|
|
@ -1,227 +0,0 @@
|
||||||
# 🧠 COMPREHENSIVE TESTING STRATEGY - ALL FEATURES
|
|
||||||
|
|
||||||
**Brainy 2.0 Complete Feature & API Validation Plan**
|
|
||||||
|
|
||||||
## 🎯 **COMPLETE PUBLIC API TESTING**
|
|
||||||
|
|
||||||
### **📋 Core Public API Methods (From docs/api/README.md):**
|
|
||||||
|
|
||||||
#### **Data Operations:**
|
|
||||||
- [ ] `addNoun(dataOrVector, metadata?)` - Text auto-embedding + vector input
|
|
||||||
- [ ] `getNoun(id)` - Retrieve single noun
|
|
||||||
- [ ] `updateNoun(id, dataOrVector?, metadata?)` - Update noun data/metadata
|
|
||||||
- [ ] `deleteNoun(id)` - Remove noun
|
|
||||||
- [ ] `addVerb(fromId, toId, type, metadata?)` - Create relationships
|
|
||||||
- [ ] `getVerb(id)` - Retrieve relationship
|
|
||||||
- [ ] `deleteVerb(id)` - Remove relationship
|
|
||||||
|
|
||||||
#### **Search & Query Operations:**
|
|
||||||
- [ ] `search(query, options?)` - Vector similarity search
|
|
||||||
- [ ] `find({ like?, where?, connected? })` - **NEW Triple Intelligence**
|
|
||||||
- [ ] `findSimilar(id, options?)` - Find similar nouns
|
|
||||||
- [ ] `searchText(query, options?)` - Text-based search
|
|
||||||
- [ ] `searchWithCursor(query, cursor?)` - Paginated search
|
|
||||||
|
|
||||||
#### **Batch Operations:**
|
|
||||||
- [ ] `addBatch(items)` - Bulk add operations
|
|
||||||
- [ ] `addBatchToBoth(nouns, verbs)` - Add nouns + verbs together
|
|
||||||
|
|
||||||
#### **Graph Operations:**
|
|
||||||
- [ ] `relate(fromId, toId, verb, metadata?)` - Create relationship
|
|
||||||
- [ ] `getConnections(id, options?)` - Get related items
|
|
||||||
- [ ] `getConnected(id, verb?)` - Get connected nouns
|
|
||||||
|
|
||||||
#### **Management Operations:**
|
|
||||||
- [ ] `clear()` - Clear all data
|
|
||||||
- [ ] `size()` - Get total count
|
|
||||||
- [ ] `getStatistics()` - Get detailed stats
|
|
||||||
- [ ] `backup()` / `restore()` - Data persistence
|
|
||||||
- [ ] `init()` / `shutdown()` - Lifecycle
|
|
||||||
|
|
||||||
## 🚀 **ADVANCED FEATURES TESTING**
|
|
||||||
|
|
||||||
### **🔧 Operational Modes:**
|
|
||||||
- [ ] **Write-Only Mode** - `setWriteOnly(true)` - write-only-direct-reads.test.ts ✅
|
|
||||||
- [ ] **Read-Only Mode** - `setReadOnly(true)`
|
|
||||||
- [ ] **Frozen Mode** - `isFrozen()` state
|
|
||||||
- [ ] **Memory-Only Mode** - No persistence
|
|
||||||
- [ ] **Persistent Mode** - File/S3/OPFS storage
|
|
||||||
|
|
||||||
### **⚡ Performance Optimizations:**
|
|
||||||
- [ ] **Throttling** - S3 rate limiting - throttling-metrics.test.ts ✅
|
|
||||||
- [ ] **Batch Processing** - Bulk operations - augmentations-batch-processing.test.ts ✅
|
|
||||||
- [ ] **Caching** - Search result caching
|
|
||||||
- [ ] **Connection Pooling** - Multi-connection management
|
|
||||||
- [ ] **Request Deduplication** - augmentations-request-deduplicator.test.ts ✅
|
|
||||||
- [ ] **Write-Ahead Logging** - augmentations-wal.test.ts ✅
|
|
||||||
|
|
||||||
### **🌐 Distributed Systems:**
|
|
||||||
- [ ] **Distributed Mode** - distributed.test.ts ✅
|
|
||||||
- [ ] **Distributed Caching** - distributed-caching.test.ts ✅
|
|
||||||
- [ ] **Node Discovery** - Multi-node coordination
|
|
||||||
- [ ] **Data Sharding** - Partition management
|
|
||||||
- [ ] **Consistency Models** - CAP theorem handling
|
|
||||||
|
|
||||||
### **🔒 Data Integrity & Hashing:**
|
|
||||||
- [ ] **Entity Registry** - UUID mapping - augmentations-entity-registry.test.ts ✅
|
|
||||||
- [ ] **Metadata Hashing** - Content deduplication
|
|
||||||
- [ ] **Vector Normalization** - Dimension standardization
|
|
||||||
- [ ] **Checksum Validation** - Data integrity verification
|
|
||||||
- [ ] **Version Management** - Data versioning
|
|
||||||
|
|
||||||
### **🧬 Clustering Algorithms:**
|
|
||||||
- [ ] **HNSW Clustering** - Hierarchical Navigable Small World
|
|
||||||
- [ ] **K-Means Clustering** - Centroid-based grouping
|
|
||||||
- [ ] **Hierarchical Clustering** - Tree-based grouping
|
|
||||||
- [ ] **Neural Clustering** - neural-clustering.test.ts ✅
|
|
||||||
|
|
||||||
### **🧠 Intelligence Features:**
|
|
||||||
- [ ] **220 NLP Patterns** - nlp-patterns-comprehensive.test.ts ✅
|
|
||||||
- [ ] **Neural Import** - AI-powered data understanding - neural-import.test.ts ✅
|
|
||||||
- [ ] **Intelligent Verb Scoring** - intelligent-verb-scoring.test.ts ✅
|
|
||||||
- [ ] **Triple Intelligence** - find-comprehensive.test.ts ✅
|
|
||||||
- [ ] **Neural API** - neural-api.test.ts ✅
|
|
||||||
|
|
||||||
## 🛠️ **MEMORY-EFFICIENT TESTING STRATEGIES**
|
|
||||||
|
|
||||||
### **📊 Industry Standard Approaches:**
|
|
||||||
|
|
||||||
#### **1. Test Categorization:**
|
|
||||||
```typescript
|
|
||||||
// Unit Tests - Fast, isolated
|
|
||||||
describe('Unit Tests', () => {
|
|
||||||
// Mock dependencies, test logic only
|
|
||||||
// Memory: <50MB, Time: <5s
|
|
||||||
})
|
|
||||||
|
|
||||||
// Integration Tests - Medium, real components
|
|
||||||
describe('Integration Tests', () => {
|
|
||||||
// Real augmentations, mocked storage
|
|
||||||
// Memory: <200MB, Time: <30s
|
|
||||||
})
|
|
||||||
|
|
||||||
// E2E Tests - Slow, full system
|
|
||||||
describe('E2E Tests', () => {
|
|
||||||
// Full system, real storage
|
|
||||||
// Memory: <1GB, Time: <5min
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
#### **2. Memory Management:**
|
|
||||||
```typescript
|
|
||||||
// Resource cleanup patterns
|
|
||||||
afterEach(async () => {
|
|
||||||
await brain?.cleanup()
|
|
||||||
brain = null
|
|
||||||
if (global.gc) global.gc() // Force cleanup
|
|
||||||
})
|
|
||||||
|
|
||||||
// Limited dataset sizes
|
|
||||||
const createTestData = (size = 10) => { // Not 10,000!
|
|
||||||
return Array.from({ length: size }, createSmallVector)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
#### **3. Mock Strategies:**
|
|
||||||
```typescript
|
|
||||||
// Mock heavy operations
|
|
||||||
vi.mock('./utils/embedding.js', () => ({
|
|
||||||
createEmbeddingFunction: () => vi.fn().mockResolvedValue(mockVector)
|
|
||||||
}))
|
|
||||||
|
|
||||||
// Mock storage for performance tests
|
|
||||||
const mockStorage = {
|
|
||||||
read: vi.fn().mockResolvedValue(testData),
|
|
||||||
write: vi.fn().mockResolvedValue(true)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
#### **4. Parallel Test Execution:**
|
|
||||||
```typescript
|
|
||||||
// vitest.config.ts
|
|
||||||
export default {
|
|
||||||
test: {
|
|
||||||
pool: 'forks', // Isolate tests
|
|
||||||
poolOptions: {
|
|
||||||
forks: {
|
|
||||||
singleFork: true // Prevent memory accumulation
|
|
||||||
}
|
|
||||||
},
|
|
||||||
testTimeout: 30000, // 30s max per test
|
|
||||||
hookTimeout: 10000 // 10s max for setup/cleanup
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### **🚀 Fast & Reliable Testing Patterns:**
|
|
||||||
|
|
||||||
#### **Memory-Efficient Patterns:**
|
|
||||||
```typescript
|
|
||||||
// 1. Small datasets
|
|
||||||
const SMALL_VECTOR_SIZE = 10 // Not 384 for unit tests
|
|
||||||
const TEST_DATA_SIZE = 5 // Not 1000s of items
|
|
||||||
|
|
||||||
// 2. Deterministic mocks
|
|
||||||
const mockEmbedding = [0.1, 0.2, 0.3, 0.4, 0.5] // Predictable
|
|
||||||
|
|
||||||
// 3. Scoped tests
|
|
||||||
describe('Search Functionality', () => {
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: 'memory', // No disk I/O
|
|
||||||
dimensions: 5, // Tiny vectors
|
|
||||||
maxConnections: 4 // Minimal graph
|
|
||||||
})
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
#### **Performance Test Patterns:**
|
|
||||||
```typescript
|
|
||||||
// Measure operations, not full datasets
|
|
||||||
it('should handle batch operations efficiently', async () => {
|
|
||||||
const start = performance.now()
|
|
||||||
|
|
||||||
// Test with 10 items, not 10,000
|
|
||||||
await brain.addBatch(createTestBatch(10))
|
|
||||||
|
|
||||||
const duration = performance.now() - start
|
|
||||||
expect(duration).toBeLessThan(1000) // 1s max
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📋 **IMPLEMENTATION PLAN**
|
|
||||||
|
|
||||||
### **Phase 1: Fix TypeScript → Build Success**
|
|
||||||
- Complete remaining 101 TypeScript errors
|
|
||||||
- Achieve clean build
|
|
||||||
|
|
||||||
### **Phase 2: Core API Validation (Fast)**
|
|
||||||
- Test all public methods with small datasets
|
|
||||||
- Validate method signatures
|
|
||||||
- Test error handling
|
|
||||||
|
|
||||||
### **Phase 3: Advanced Features (Medium)**
|
|
||||||
- Test operational modes (write-only, read-only)
|
|
||||||
- Test performance optimizations
|
|
||||||
- Test distributed features
|
|
||||||
|
|
||||||
### **Phase 4: Full Integration (Comprehensive)**
|
|
||||||
- All 49 tests passing
|
|
||||||
- Memory-efficient execution
|
|
||||||
- Performance benchmarks
|
|
||||||
|
|
||||||
## ✅ **SUCCESS METRICS**
|
|
||||||
|
|
||||||
### **Speed Goals:**
|
|
||||||
- **Unit tests**: <5 minutes total
|
|
||||||
- **Integration tests**: <15 minutes total
|
|
||||||
- **Full suite**: <30 minutes total
|
|
||||||
- **Memory usage**: <2GB peak
|
|
||||||
|
|
||||||
### **Coverage Goals:**
|
|
||||||
- **100% public API methods** tested
|
|
||||||
- **100% operational modes** tested
|
|
||||||
- **100% augmentations** tested
|
|
||||||
- **100% clustering algorithms** tested
|
|
||||||
- **All performance optimizations** validated
|
|
||||||
|
|
||||||
This gives us **comprehensive testing** of ALL Brainy features while maintaining **fast, reliable execution** using industry-standard patterns!
|
|
||||||
295
CONTRIBUTING.md
295
CONTRIBUTING.md
|
|
@ -1,275 +1,66 @@
|
||||||
# Contributing to Brainy
|
# Contributing to Brainy
|
||||||
|
|
||||||
Thank you for your interest in contributing to Brainy! This document provides guidelines and instructions for contributing to the project.
|
Brainy is MIT-licensed and genuinely open to outside contributions. This page
|
||||||
|
is the honest, current path — please don't rely on older instructions you
|
||||||
|
may find elsewhere in the repo's history.
|
||||||
|
|
||||||
## Code of Conduct
|
## Where the project lives
|
||||||
|
|
||||||
By participating in this project, you agree to abide by our Code of Conduct:
|
The source of truth is a self-hosted forge: **source.soulcraft.com/soulcraft/brainy**.
|
||||||
- Be respectful and inclusive
|
It's anonymously readable and cloneable — no account needed to browse, clone,
|
||||||
- Welcome newcomers and help them get started
|
or build.
|
||||||
- Focus on constructive criticism
|
|
||||||
- Respect differing viewpoints and experiences
|
|
||||||
|
|
||||||
## How to Contribute
|
## How to contribute
|
||||||
|
|
||||||
### Reporting Issues
|
**Found a bug, or have an idea?** Email **brainy@soulcraft.com**. No account,
|
||||||
|
no ceremony — you'll get a receipt, and it goes to a human.
|
||||||
|
|
||||||
Before creating an issue, please check existing issues to avoid duplicates.
|
**Want to send a patch?** Two ways, both first-class:
|
||||||
|
|
||||||
When creating an issue, include:
|
- **Email a patch.** Run `git format-patch` against your change and email the
|
||||||
- Clear, descriptive title
|
output to **brainy@soulcraft.com**. This is a genuinely supported path, not
|
||||||
- Detailed description of the problem
|
a fallback — plenty of good contributions arrive this way.
|
||||||
- Steps to reproduce
|
- **Open a pull request on the forge.** Request an account at
|
||||||
- Expected vs actual behavior
|
**source.soulcraft.com** (registration is request-with-approval, so allow
|
||||||
- System information (OS, Node version, Brainy version)
|
a little lag), clone, push a branch, and open a PR there. Maintainers
|
||||||
- Code examples if applicable
|
review and land it.
|
||||||
|
|
||||||
### Suggesting Features
|
Either way, for anything beyond a small fix, opening an issue first (email is
|
||||||
|
fine) to talk through the approach saves everyone rework.
|
||||||
|
|
||||||
Feature requests are welcome! Please provide:
|
## Development setup
|
||||||
- Clear use case
|
|
||||||
- Proposed API/interface
|
|
||||||
- Examples of how it would work
|
|
||||||
- Any potential challenges or considerations
|
|
||||||
|
|
||||||
### Pull Requests
|
|
||||||
|
|
||||||
#### Before Starting
|
|
||||||
|
|
||||||
1. Check existing issues and PRs
|
|
||||||
2. Open an issue to discuss significant changes
|
|
||||||
3. Fork the repository
|
|
||||||
4. Create a feature branch from `main`
|
|
||||||
|
|
||||||
#### Development Setup
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Clone your fork
|
git clone https://source.soulcraft.com/soulcraft/brainy.git
|
||||||
git clone https://github.com/your-username/brainy.git
|
|
||||||
cd brainy
|
cd brainy
|
||||||
|
|
||||||
# Install dependencies
|
|
||||||
npm install
|
npm install
|
||||||
|
|
||||||
# Build the project
|
|
||||||
npm run build
|
npm run build
|
||||||
|
|
||||||
# Run tests
|
|
||||||
npm test
|
npm test
|
||||||
```
|
```
|
||||||
|
|
||||||
#### Making Changes
|
Tests run on [Vitest](https://vitest.dev/). `npm test` runs the unit suite;
|
||||||
|
see `package.json` for `test:integration`, `test:coverage`, and friends.
|
||||||
|
|
||||||
1. **Follow the code style**
|
## Standards
|
||||||
- TypeScript for all source code
|
|
||||||
- Clear variable and function names
|
|
||||||
- Comments for complex logic
|
|
||||||
- JSDoc for public APIs
|
|
||||||
|
|
||||||
2. **Write tests**
|
- **Strict TypeScript.** No `any` escape hatches to dodge the type checker.
|
||||||
- Add tests for new features
|
- **Tests exercise real behavior.** No mocking away the thing you're supposed
|
||||||
- Update tests for changes
|
to be testing.
|
||||||
- Ensure all tests pass
|
- **No stubs, no TODO-code.** If something can't be finished, say so and
|
||||||
|
leave it out — don't merge a placeholder.
|
||||||
|
- **JSDoc on every exported function, class, and type.**
|
||||||
|
- **[Conventional Commits](https://www.conventionalcommits.org/).** `feat:`,
|
||||||
|
`fix:`, `docs:`, `perf:`, `refactor:`, `test:`, `chore:`. Never
|
||||||
|
`BREAKING CHANGE` in a commit message — major version bumps are a separate,
|
||||||
|
deliberate decision.
|
||||||
|
- **Performance claims are measured or labeled projected.** If a PR or its
|
||||||
|
description states a number, cite the benchmark that produced it (see
|
||||||
|
[docs/performance-envelopes.md](docs/performance-envelopes.md) for the
|
||||||
|
pattern). Don't state an estimate as if it were measured.
|
||||||
|
|
||||||
3. **Update documentation**
|
## License
|
||||||
- Update README if needed
|
|
||||||
- Add/update API documentation
|
|
||||||
- Include examples
|
|
||||||
|
|
||||||
#### Commit Guidelines
|
Brainy is [MIT licensed](LICENSE). Contributions are accepted under the same
|
||||||
|
license — there's no CLA to sign.
|
||||||
|
|
||||||
Follow conventional commits format:
|
Thank you for considering a contribution.
|
||||||
|
|
||||||
```
|
|
||||||
type(scope): description
|
|
||||||
|
|
||||||
[optional body]
|
|
||||||
|
|
||||||
[optional footer]
|
|
||||||
```
|
|
||||||
|
|
||||||
Types:
|
|
||||||
- `feat`: New feature
|
|
||||||
- `fix`: Bug fix
|
|
||||||
- `docs`: Documentation changes
|
|
||||||
- `style`: Code style changes
|
|
||||||
- `refactor`: Code refactoring
|
|
||||||
- `perf`: Performance improvements
|
|
||||||
- `test`: Test changes
|
|
||||||
- `chore`: Build/tooling changes
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
```bash
|
|
||||||
feat(triple): add graph traversal depth limit
|
|
||||||
fix(storage): handle concurrent write conflicts
|
|
||||||
docs(api): update search method documentation
|
|
||||||
```
|
|
||||||
|
|
||||||
#### Submitting PR
|
|
||||||
|
|
||||||
1. Push to your fork
|
|
||||||
2. Create PR against `main` branch
|
|
||||||
3. Fill out PR template
|
|
||||||
4. Ensure CI checks pass
|
|
||||||
5. Wait for review
|
|
||||||
|
|
||||||
### Testing
|
|
||||||
|
|
||||||
#### Running Tests
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Run all tests
|
|
||||||
npm test
|
|
||||||
|
|
||||||
# Run specific test file
|
|
||||||
npm test tests/core.test.ts
|
|
||||||
|
|
||||||
# Run with coverage
|
|
||||||
npm run test:coverage
|
|
||||||
|
|
||||||
# Watch mode
|
|
||||||
npm run test:watch
|
|
||||||
```
|
|
||||||
|
|
||||||
#### Writing Tests
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { describe, it, expect } from 'vitest'
|
|
||||||
import { BrainyData } from '../src'
|
|
||||||
|
|
||||||
describe('Feature Name', () => {
|
|
||||||
it('should do something specific', async () => {
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Test implementation
|
|
||||||
const result = await brain.search("test")
|
|
||||||
|
|
||||||
expect(result).toBeDefined()
|
|
||||||
expect(result.length).toBeGreaterThan(0)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## Architecture Guidelines
|
|
||||||
|
|
||||||
### Adding New Features
|
|
||||||
|
|
||||||
1. **Check existing functionality**
|
|
||||||
- Review `ARCHITECTURE.md`
|
|
||||||
- Check if similar features exist
|
|
||||||
- Consider if it should be an augmentation
|
|
||||||
|
|
||||||
2. **Design considerations**
|
|
||||||
- Maintain backward compatibility
|
|
||||||
- Consider performance impact
|
|
||||||
- Think about all storage adapters
|
|
||||||
- Plan for extensibility
|
|
||||||
|
|
||||||
3. **Implementation checklist**
|
|
||||||
- [ ] Core functionality
|
|
||||||
- [ ] Tests (unit and integration)
|
|
||||||
- [ ] Documentation
|
|
||||||
- [ ] TypeScript types
|
|
||||||
- [ ] Examples
|
|
||||||
- [ ] Performance benchmarks (if applicable)
|
|
||||||
|
|
||||||
### Creating Augmentations
|
|
||||||
|
|
||||||
Augmentations extend Brainy's functionality:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyAugmentation } from '../types'
|
|
||||||
|
|
||||||
export class MyAugmentation extends BrainyAugmentation {
|
|
||||||
name = 'MyAugmentation'
|
|
||||||
|
|
||||||
async onInit(brain: BrainyData): Promise<void> {
|
|
||||||
// Initialize augmentation
|
|
||||||
}
|
|
||||||
|
|
||||||
async onAdd(item: any, brain: BrainyData): Promise<any> {
|
|
||||||
// Process before adding
|
|
||||||
return item
|
|
||||||
}
|
|
||||||
|
|
||||||
async onSearch(query: any, results: any[], brain: BrainyData): Promise<any[]> {
|
|
||||||
// Process search results
|
|
||||||
return results
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Performance Considerations
|
|
||||||
|
|
||||||
- Use batch operations where possible
|
|
||||||
- Implement caching strategically
|
|
||||||
- Consider memory usage
|
|
||||||
- Profile performance impacts
|
|
||||||
- Add benchmarks for critical paths
|
|
||||||
|
|
||||||
## Documentation
|
|
||||||
|
|
||||||
### API Documentation
|
|
||||||
|
|
||||||
Use JSDoc for all public APIs:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
/**
|
|
||||||
* Searches for similar items using vector similarity
|
|
||||||
* @param query - Search query (text or vector)
|
|
||||||
* @param options - Search options
|
|
||||||
* @returns Array of search results with scores
|
|
||||||
* @example
|
|
||||||
* ```typescript
|
|
||||||
* const results = await brain.search("machine learning", { limit: 10 })
|
|
||||||
* ```
|
|
||||||
*/
|
|
||||||
async search(query: string | Vector, options?: SearchOptions): Promise<SearchResult[]> {
|
|
||||||
// Implementation
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Examples
|
|
||||||
|
|
||||||
Add examples for new features:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// examples/feature-name.ts
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
|
|
||||||
async function exampleUsage() {
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Show feature usage
|
|
||||||
// Include comments explaining what's happening
|
|
||||||
// Handle errors appropriately
|
|
||||||
}
|
|
||||||
|
|
||||||
exampleUsage().catch(console.error)
|
|
||||||
```
|
|
||||||
|
|
||||||
## Release Process
|
|
||||||
|
|
||||||
1. **Version bump**: Follow semantic versioning
|
|
||||||
2. **Update CHANGELOG**: Document all changes
|
|
||||||
3. **Run tests**: Ensure all tests pass
|
|
||||||
4. **Build**: Generate distribution files
|
|
||||||
5. **Tag**: Create git tag for version
|
|
||||||
6. **Publish**: Release to npm
|
|
||||||
|
|
||||||
## Getting Help
|
|
||||||
|
|
||||||
- **Discord**: Join our community
|
|
||||||
- **Issues**: Ask questions on GitHub
|
|
||||||
- **Discussions**: Share ideas and get feedback
|
|
||||||
|
|
||||||
## Recognition
|
|
||||||
|
|
||||||
Contributors will be recognized in:
|
|
||||||
- CHANGELOG.md for their contributions
|
|
||||||
- README.md contributors section
|
|
||||||
- GitHub contributors page
|
|
||||||
|
|
||||||
Thank you for contributing to Brainy! 🧠
|
|
||||||
|
|
|
||||||
|
|
@ -1,212 +0,0 @@
|
||||||
# Coordinated Index Optimization Strategy
|
|
||||||
|
|
||||||
## The Problem
|
|
||||||
Two independent index systems competing for resources:
|
|
||||||
- **HNSW Index**: Wants to cache hot vectors in RAM
|
|
||||||
- **MetadataIndex**: Wants to cache hot field values in RAM
|
|
||||||
- **Conflict**: Both trying to use same memory/disk without coordination!
|
|
||||||
|
|
||||||
## The Solution: Unified Resource Manager
|
|
||||||
|
|
||||||
### 1. Shared Resource Pool
|
|
||||||
```typescript
|
|
||||||
class UnifiedIndexManager {
|
|
||||||
private totalMemoryBudget: number = 2 * 1024 * 1024 * 1024 // 2GB total
|
|
||||||
private hnswMemoryUsage: number = 0
|
|
||||||
private metadataMemoryUsage: number = 0
|
|
||||||
|
|
||||||
// Intelligent allocation based on usage patterns
|
|
||||||
allocateMemory(requester: 'hnsw' | 'metadata', size: number): boolean {
|
|
||||||
const available = this.totalMemoryBudget - this.hnswMemoryUsage - this.metadataMemoryUsage
|
|
||||||
|
|
||||||
if (size <= available) {
|
|
||||||
if (requester === 'hnsw') {
|
|
||||||
this.hnswMemoryUsage += size
|
|
||||||
} else {
|
|
||||||
this.metadataMemoryUsage += size
|
|
||||||
}
|
|
||||||
return true
|
|
||||||
}
|
|
||||||
|
|
||||||
// Try to steal from other index if one is underutilized
|
|
||||||
return this.rebalance(requester, size)
|
|
||||||
}
|
|
||||||
|
|
||||||
private rebalance(requester: string, needed: number): boolean {
|
|
||||||
// If HNSW is using 80% and metadata only 20%, rebalance
|
|
||||||
const hnswRatio = this.hnswMemoryUsage / this.totalMemoryBudget
|
|
||||||
const metadataRatio = this.metadataMemoryUsage / this.totalMemoryBudget
|
|
||||||
|
|
||||||
// Intelligent rebalancing logic
|
|
||||||
// ...
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Coordinated LRU Eviction
|
|
||||||
```typescript
|
|
||||||
class CoordinatedLRUCache {
|
|
||||||
private hnswLRU: LRUCache
|
|
||||||
private metadataLRU: LRUCache
|
|
||||||
private accessPatterns: AccessTracker
|
|
||||||
|
|
||||||
// When memory pressure, evict from the index with lowest utility
|
|
||||||
async evict(bytesNeeded: number): Promise<void> {
|
|
||||||
const hnswUtility = this.calculateUtility(this.hnswLRU)
|
|
||||||
const metadataUtility = this.calculateUtility(this.metadataLRU)
|
|
||||||
|
|
||||||
if (hnswUtility < metadataUtility) {
|
|
||||||
// HNSW items are less frequently accessed
|
|
||||||
await this.hnswLRU.evict(bytesNeeded)
|
|
||||||
} else {
|
|
||||||
// Metadata items are less frequently accessed
|
|
||||||
await this.metadataLRU.evict(bytesNeeded)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private calculateUtility(cache: LRUCache): number {
|
|
||||||
// Factors:
|
|
||||||
// - Access frequency
|
|
||||||
// - Recency
|
|
||||||
// - Cost to rebuild (HNSW is expensive, metadata is cheap)
|
|
||||||
// - Current query patterns
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Query-Aware Optimization
|
|
||||||
```typescript
|
|
||||||
class QueryOptimizer {
|
|
||||||
private queryHistory: QueryPattern[] = []
|
|
||||||
|
|
||||||
optimizeForQuery(query: TripleQuery) {
|
|
||||||
// Analyze query type
|
|
||||||
const usesVector = !!(query.like || query.similar)
|
|
||||||
const usesMetadata = !!query.where
|
|
||||||
|
|
||||||
// Pre-warm appropriate caches
|
|
||||||
if (usesVector && usesMetadata) {
|
|
||||||
// Hybrid query - balance resources 50/50
|
|
||||||
this.resourceManager.setRatio(0.5, 0.5)
|
|
||||||
} else if (usesVector) {
|
|
||||||
// Vector-heavy - give HNSW more memory
|
|
||||||
this.resourceManager.setRatio(0.8, 0.2)
|
|
||||||
} else {
|
|
||||||
// Metadata-heavy - give MetadataIndex more memory
|
|
||||||
this.resourceManager.setRatio(0.2, 0.8)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. Unified Persistence Strategy
|
|
||||||
```typescript
|
|
||||||
class UnifiedPersistence {
|
|
||||||
private writeBuffer: WriteBuffer
|
|
||||||
private flushScheduler: FlushScheduler
|
|
||||||
|
|
||||||
async flush() {
|
|
||||||
// Coordinate flushes to avoid disk contention
|
|
||||||
const tasks = []
|
|
||||||
|
|
||||||
// Flush metadata first (smaller, faster)
|
|
||||||
if (this.metadataIndex.isDirty) {
|
|
||||||
tasks.push(this.flushMetadata())
|
|
||||||
}
|
|
||||||
|
|
||||||
// Then flush HNSW (larger, slower)
|
|
||||||
if (this.hnswIndex.isDirty) {
|
|
||||||
tasks.push(this.flushHNSW())
|
|
||||||
}
|
|
||||||
|
|
||||||
// Sequential to avoid disk thrashing
|
|
||||||
for (const task of tasks) {
|
|
||||||
await task
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async flushMetadata() {
|
|
||||||
// Flush sorted indices
|
|
||||||
await this.storage.save('metadata_sorted', this.metadataIndex.sortedIndices)
|
|
||||||
// Flush hash indices
|
|
||||||
await this.storage.save('metadata_hash', this.metadataIndex.hashIndices)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Implementation Plan
|
|
||||||
|
|
||||||
### Phase 1: Shared Memory Manager (Quick Win)
|
|
||||||
```typescript
|
|
||||||
// In BrainyData constructor
|
|
||||||
this.resourceManager = new UnifiedResourceManager({
|
|
||||||
totalMemory: config.maxMemory || 2 * GB,
|
|
||||||
hnswRatio: 0.6, // 60% for vectors by default
|
|
||||||
metadataRatio: 0.4 // 40% for metadata by default
|
|
||||||
})
|
|
||||||
|
|
||||||
// Pass to both indices
|
|
||||||
this.hnswIndex = new HNSWIndexOptimized({
|
|
||||||
resourceManager: this.resourceManager
|
|
||||||
})
|
|
||||||
|
|
||||||
this.metadataIndex = new MetadataIndexOptimized({
|
|
||||||
resourceManager: this.resourceManager
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 2: Coordinated Eviction
|
|
||||||
- Single LRU that tracks both index types
|
|
||||||
- Utility-based eviction (not just recency)
|
|
||||||
- Consider rebuild cost in eviction decisions
|
|
||||||
|
|
||||||
### Phase 3: Query-Driven Optimization
|
|
||||||
- Track query patterns
|
|
||||||
- Dynamically adjust memory allocation
|
|
||||||
- Pre-warm caches based on query type
|
|
||||||
|
|
||||||
## Benefits of Coordination
|
|
||||||
|
|
||||||
1. **No Resource Conflicts**: Indices cooperate instead of compete
|
|
||||||
2. **Better Memory Usage**: Allocate based on actual query patterns
|
|
||||||
3. **Smarter Eviction**: Keep data that's actually needed
|
|
||||||
4. **Unified Monitoring**: Single place to track all index performance
|
|
||||||
5. **Auto-Optimization**: System learns and adapts to usage
|
|
||||||
|
|
||||||
## Configuration Example
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
indexOptimization: {
|
|
||||||
mode: 'coordinated', // vs 'independent'
|
|
||||||
totalMemory: 4 * GB, // Total for ALL indices
|
|
||||||
autoBalance: true, // Dynamic rebalancing
|
|
||||||
persistenceInterval: 60000, // Coordinated flush every minute
|
|
||||||
monitoring: {
|
|
||||||
trackQueryPatterns: true,
|
|
||||||
optimizeForPatterns: true,
|
|
||||||
rebalanceInterval: 300000 // Every 5 minutes
|
|
||||||
}
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## Monitoring & Metrics
|
|
||||||
```typescript
|
|
||||||
const stats = brain.getIndexStats()
|
|
||||||
// {
|
|
||||||
// hnsw: {
|
|
||||||
// memoryUsed: 1.2 * GB,
|
|
||||||
// cacheHitRate: 0.89,
|
|
||||||
// avgQueryTime: 12ms
|
|
||||||
// },
|
|
||||||
// metadata: {
|
|
||||||
// memoryUsed: 0.8 * GB,
|
|
||||||
// cacheHitRate: 0.95,
|
|
||||||
// avgQueryTime: 2ms
|
|
||||||
// },
|
|
||||||
// coordination: {
|
|
||||||
// rebalances: 5,
|
|
||||||
// memoryUtilization: 0.95,
|
|
||||||
// queryPatternDetected: 'hybrid-heavy'
|
|
||||||
// }
|
|
||||||
// }
|
|
||||||
72
Dockerfile
Normal file
72
Dockerfile
Normal file
|
|
@ -0,0 +1,72 @@
|
||||||
|
# Multi-stage Dockerfile for Brainy
|
||||||
|
# Optimized for production deployment with minimal image size
|
||||||
|
|
||||||
|
# Stage 1: Build stage
|
||||||
|
FROM node:22-alpine AS builder
|
||||||
|
|
||||||
|
# Install build dependencies
|
||||||
|
RUN apk add --no-cache python3 make g++
|
||||||
|
|
||||||
|
# Set working directory
|
||||||
|
WORKDIR /app
|
||||||
|
|
||||||
|
# Copy package files
|
||||||
|
COPY package*.json ./
|
||||||
|
|
||||||
|
# Install all dependencies (including dev dependencies for building)
|
||||||
|
RUN npm ci
|
||||||
|
|
||||||
|
# Copy source code
|
||||||
|
COPY . .
|
||||||
|
|
||||||
|
# Build the TypeScript code
|
||||||
|
RUN npm run build
|
||||||
|
|
||||||
|
# Remove dev dependencies and only keep production ones
|
||||||
|
RUN npm prune --production
|
||||||
|
|
||||||
|
# Stage 2: Production stage
|
||||||
|
FROM node:22-alpine
|
||||||
|
|
||||||
|
# Install production dependencies only
|
||||||
|
RUN apk add --no-cache tini
|
||||||
|
|
||||||
|
# Create non-root user for security
|
||||||
|
RUN addgroup -g 1001 -S nodejs && \
|
||||||
|
adduser -S nodejs -u 1001
|
||||||
|
|
||||||
|
# Set working directory
|
||||||
|
WORKDIR /app
|
||||||
|
|
||||||
|
# Copy package files
|
||||||
|
COPY package*.json ./
|
||||||
|
|
||||||
|
# Copy built application from builder stage
|
||||||
|
COPY --from=builder --chown=nodejs:nodejs /app/node_modules ./node_modules
|
||||||
|
COPY --from=builder --chown=nodejs:nodejs /app/dist ./dist
|
||||||
|
|
||||||
|
# Copy necessary static files
|
||||||
|
COPY --chown=nodejs:nodejs README.md LICENSE ./
|
||||||
|
|
||||||
|
# Create data directory for file-based storage
|
||||||
|
RUN mkdir -p /app/data && chown -R nodejs:nodejs /app/data
|
||||||
|
|
||||||
|
# Switch to non-root user
|
||||||
|
USER nodejs
|
||||||
|
|
||||||
|
# Expose default port (can be overridden)
|
||||||
|
EXPOSE 3000
|
||||||
|
|
||||||
|
# Set environment variables for production
|
||||||
|
ENV NODE_ENV=production
|
||||||
|
ENV BRAINY_STORAGE_PATH=/app/data
|
||||||
|
|
||||||
|
# Health check endpoint
|
||||||
|
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
|
||||||
|
CMD node -e "require('http').get('http://localhost:3000/health', (r) => r.statusCode === 200 ? process.exit(0) : process.exit(1))"
|
||||||
|
|
||||||
|
# Use tini to handle signals properly
|
||||||
|
ENTRYPOINT ["/sbin/tini", "--"]
|
||||||
|
|
||||||
|
# Default command (can be overridden)
|
||||||
|
CMD ["node", "dist/index.js"]
|
||||||
|
|
@ -1,104 +0,0 @@
|
||||||
# Transformer Model Memory Issue - Solutions
|
|
||||||
|
|
||||||
## The Problem
|
|
||||||
ONNX runtime allocates 4GB for a 30MB model during inference. This is a known issue with transformers.js.
|
|
||||||
|
|
||||||
## Solution 1: Use Smaller Quantized Model (RECOMMENDED)
|
|
||||||
```javascript
|
|
||||||
// Current: all-MiniLM-L6-v2 with q8 quantization
|
|
||||||
// Switch to: all-MiniLM-L6-v2 with q4 quantization (50% smaller)
|
|
||||||
// Or use: paraphrase-MiniLM-L3-v2 (even smaller, still good quality)
|
|
||||||
|
|
||||||
const embeddingFunction = createEmbeddingFunction({
|
|
||||||
modelName: 'Xenova/paraphrase-MiniLM-L3-v2',
|
|
||||||
dtype: 'q4' // 4-bit quantization instead of 8-bit
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## Solution 2: Increase Node Memory Limit
|
|
||||||
```bash
|
|
||||||
# Run with 8GB heap limit
|
|
||||||
node --max-old-space-size=8192 test-range-queries.js
|
|
||||||
|
|
||||||
# Or set in package.json test script:
|
|
||||||
"test": "NODE_OPTIONS='--max-old-space-size=8192' vitest"
|
|
||||||
```
|
|
||||||
|
|
||||||
## Solution 3: Use Remote Embeddings (For Testing)
|
|
||||||
```javascript
|
|
||||||
// Mock embedding function for tests
|
|
||||||
const mockEmbeddingFunction = async (text) => {
|
|
||||||
// Generate deterministic fake embedding from text hash
|
|
||||||
const hash = text.split('').reduce((a, b) => a + b.charCodeAt(0), 0)
|
|
||||||
return new Array(384).fill(0).map((_, i) => Math.sin(hash + i) * 0.1)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Solution 4: Model Pooling & Unloading
|
|
||||||
```javascript
|
|
||||||
class ModelPool {
|
|
||||||
private model: any = null
|
|
||||||
private lastUsed: number = 0
|
|
||||||
private readonly UNLOAD_AFTER_MS = 30000 // 30 seconds
|
|
||||||
|
|
||||||
async getModel() {
|
|
||||||
if (!this.model) {
|
|
||||||
this.model = await loadModel()
|
|
||||||
}
|
|
||||||
this.lastUsed = Date.now()
|
|
||||||
this.scheduleUnload()
|
|
||||||
return this.model
|
|
||||||
}
|
|
||||||
|
|
||||||
private scheduleUnload() {
|
|
||||||
setTimeout(() => {
|
|
||||||
if (Date.now() - this.lastUsed > this.UNLOAD_AFTER_MS) {
|
|
||||||
this.model?.dispose?.()
|
|
||||||
this.model = null
|
|
||||||
}
|
|
||||||
}, this.UNLOAD_AFTER_MS)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Solution 5: Use Native Bindings (Future)
|
|
||||||
Replace transformers.js with native bindings:
|
|
||||||
- onnxruntime-node (more efficient memory)
|
|
||||||
- @tensorflow/tfjs-node (better memory management)
|
|
||||||
- Custom Rust/C++ binding
|
|
||||||
|
|
||||||
## Recommendation for Brainy 2.0
|
|
||||||
|
|
||||||
### For Production:
|
|
||||||
1. Use q4 quantization (reduces memory 50%)
|
|
||||||
2. Implement model pooling/unloading
|
|
||||||
3. Document memory requirements (4GB recommended)
|
|
||||||
|
|
||||||
### For Testing:
|
|
||||||
1. Increase Node heap to 8GB for test suite
|
|
||||||
2. Use mock embeddings for unit tests
|
|
||||||
3. Real embeddings only for integration tests
|
|
||||||
|
|
||||||
### Long-term:
|
|
||||||
1. Investigate native bindings
|
|
||||||
2. Support multiple embedding backends
|
|
||||||
3. Cloud embedding API option
|
|
||||||
|
|
||||||
## Memory Requirements
|
|
||||||
|
|
||||||
| Configuration | Memory Needed | Use Case |
|
|
||||||
|--------------|--------------|----------|
|
|
||||||
| Mock embeddings | 200 MB | Unit tests |
|
|
||||||
| Q4 quantization | 2 GB | Development |
|
|
||||||
| Q8 quantization | 4 GB | Production (current) |
|
|
||||||
| Native bindings | 500 MB | Future optimization |
|
|
||||||
|
|
||||||
## The Real Issue
|
|
||||||
|
|
||||||
This is NOT a Brainy problem - it's a transformers.js/ONNX issue that affects ALL JavaScript ML applications. Even Google's similar libraries have this problem.
|
|
||||||
|
|
||||||
The good news:
|
|
||||||
- Only affects initial model load
|
|
||||||
- Singleton pattern prevents multiple copies
|
|
||||||
- Memory is released after inference
|
|
||||||
- Production servers typically have 8-16GB RAM
|
|
||||||
224
MIGRATION.md
224
MIGRATION.md
|
|
@ -1,224 +0,0 @@
|
||||||
# Migration Guide: Brainy 1.x → 2.0
|
|
||||||
|
|
||||||
This guide helps you migrate from Brainy 1.x to the new 2.0 release with Triple Intelligence Engine.
|
|
||||||
|
|
||||||
## 🚨 Breaking Changes
|
|
||||||
|
|
||||||
### 1. Search Result Format
|
|
||||||
**Before (1.x):**
|
|
||||||
```typescript
|
|
||||||
const results = await brain.search("query")
|
|
||||||
// Returns: [["id1", 0.9], ["id2", 0.8]]
|
|
||||||
```
|
|
||||||
|
|
||||||
**After (2.0):**
|
|
||||||
```typescript
|
|
||||||
const results = await brain.search("query")
|
|
||||||
// Returns: [{id: "id1", score: 0.9, content: "...", metadata: {...}}, ...]
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Storage Configuration
|
|
||||||
**Before (1.x):**
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData("./data")
|
|
||||||
```
|
|
||||||
|
|
||||||
**After (2.0):**
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'filesystem',
|
|
||||||
path: './data'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Metadata Filtering
|
|
||||||
**Before (1.x):**
|
|
||||||
```typescript
|
|
||||||
// Limited filtering capabilities
|
|
||||||
const results = await brain.search("query", { category: "tech" })
|
|
||||||
```
|
|
||||||
|
|
||||||
**After (2.0):**
|
|
||||||
```typescript
|
|
||||||
// Advanced field filtering with O(1) performance
|
|
||||||
const results = await brain.search("query", {
|
|
||||||
where: {
|
|
||||||
category: "tech",
|
|
||||||
rating: { $gte: 4.0 },
|
|
||||||
date: { $between: ["2024-01-01", "2024-12-31"] }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## ✨ New Features in 2.0
|
|
||||||
|
|
||||||
### Triple Intelligence Engine
|
|
||||||
Combine three types of intelligence in a single query:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Vector similarity + Field filtering + Graph relationships
|
|
||||||
const results = await brain.search("machine learning algorithms", {
|
|
||||||
where: {
|
|
||||||
category: { $in: ["ai", "technology"] },
|
|
||||||
difficulty: { $lte: 5 }
|
|
||||||
},
|
|
||||||
includeRelated: true,
|
|
||||||
depth: 2
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Brain Patterns Query Language
|
|
||||||
MongoDB-compatible syntax with semantic extensions:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const results = await brain.find({
|
|
||||||
$or: [
|
|
||||||
{ category: "technology" },
|
|
||||||
{ $vector: { $similar: "artificial intelligence", threshold: 0.8 } }
|
|
||||||
],
|
|
||||||
published: { $gte: "2024-01-01" }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Universal Storage Support
|
|
||||||
```typescript
|
|
||||||
// File System (default)
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'filesystem', path: './data' }
|
|
||||||
})
|
|
||||||
|
|
||||||
// Amazon S3 / Compatible
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 's3',
|
|
||||||
bucket: 'my-data',
|
|
||||||
region: 'us-east-1'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Origin Private File System (Browser)
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'opfs' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 Migration Steps
|
|
||||||
|
|
||||||
### Step 1: Update Package
|
|
||||||
```bash
|
|
||||||
npm install brainy@2.0.0
|
|
||||||
```
|
|
||||||
|
|
||||||
### Step 2: Update Initialization
|
|
||||||
```typescript
|
|
||||||
// Old
|
|
||||||
const brain = new BrainyData("./data")
|
|
||||||
|
|
||||||
// New
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'filesystem',
|
|
||||||
path: './data'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Step 3: Update Search Result Handling
|
|
||||||
```typescript
|
|
||||||
// Old
|
|
||||||
const results = await brain.search("query")
|
|
||||||
for (const [id, score] of results) {
|
|
||||||
const item = await brain.get(id)
|
|
||||||
console.log(item.content, score)
|
|
||||||
}
|
|
||||||
|
|
||||||
// New
|
|
||||||
const results = await brain.search("query")
|
|
||||||
for (const result of results) {
|
|
||||||
console.log(result.content, result.score)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Step 4: Upgrade Filtering (Optional)
|
|
||||||
```typescript
|
|
||||||
// Old basic filtering
|
|
||||||
const results = await brain.search("query", { category: "tech" })
|
|
||||||
|
|
||||||
// New advanced filtering
|
|
||||||
const results = await brain.search("query", {
|
|
||||||
where: {
|
|
||||||
category: "tech",
|
|
||||||
rating: { $gte: 4.0 }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 Performance Improvements
|
|
||||||
|
|
||||||
### Automatic Data Migration
|
|
||||||
- Brainy 2.0 automatically migrates your existing 1.x data
|
|
||||||
- No manual data conversion required
|
|
||||||
- First startup may take longer for large datasets
|
|
||||||
|
|
||||||
### New Indexing Performance
|
|
||||||
- 10x faster metadata filtering with field indexes
|
|
||||||
- Sub-millisecond vector search with HNSW indexing
|
|
||||||
- Smart caching reduces repeated query latency
|
|
||||||
|
|
||||||
## 🛠 Compatibility Mode
|
|
||||||
|
|
||||||
Enable 1.x compatibility for gradual migration:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
compatibility: {
|
|
||||||
version: "1.x",
|
|
||||||
searchResultFormat: "array" // Use old [id, score] format
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔧 New APIs to Explore
|
|
||||||
|
|
||||||
### Clustering
|
|
||||||
```typescript
|
|
||||||
const clusters = await brain.cluster({
|
|
||||||
algorithm: 'kmeans',
|
|
||||||
numClusters: 5
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Relationship Discovery
|
|
||||||
```typescript
|
|
||||||
const related = await brain.findRelated(itemId, {
|
|
||||||
depth: 2,
|
|
||||||
minSimilarity: 0.7
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Statistics & Analytics
|
|
||||||
```typescript
|
|
||||||
const stats = await brain.statistics()
|
|
||||||
console.log(`Total items: ${stats.totalItems}`)
|
|
||||||
console.log(`Query performance: ${stats.avgQueryTime}ms`)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🆘 Need Help?
|
|
||||||
|
|
||||||
- **Issues**: Report bugs at [GitHub Issues](https://github.com/brainy-org/brainy/issues)
|
|
||||||
- **Discussions**: Get help at [GitHub Discussions](https://github.com/brainy-org/brainy/discussions)
|
|
||||||
- **Examples**: Check the `/examples` directory for migration examples
|
|
||||||
|
|
||||||
## 📋 Migration Checklist
|
|
||||||
|
|
||||||
- [ ] Updated to Brainy 2.0
|
|
||||||
- [ ] Changed initialization to new config format
|
|
||||||
- [ ] Updated search result handling from arrays to objects
|
|
||||||
- [ ] Tested core functionality with your data
|
|
||||||
- [ ] Explored new Triple Intelligence features
|
|
||||||
- [ ] Updated tests to use new API patterns
|
|
||||||
- [ ] Leveraged new storage adapters (if applicable)
|
|
||||||
|
|
||||||
**Migration typically takes 15-30 minutes for most applications.**
|
|
||||||
632
README.md
632
README.md
|
|
@ -1,419 +1,225 @@
|
||||||
# Brainy
|
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="brainy.png" alt="Brainy Logo" width="200">
|
<img src="https://raw.githubusercontent.com/soulcraftlabs/brainy/main/brainy.png" alt="Brainy" width="180">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
[](https://www.npmjs.com/package/brainy)
|
<h1 align="center">Brainy</h1>
|
||||||
[](https://www.npmjs.com/package/brainy)
|
|
||||||
[](LICENSE)
|
<p align="center">
|
||||||
[](https://www.typescriptlang.org/)
|
<b>Three database paradigms. One API. Zero configuration.</b><br>
|
||||||
[](https://github.com/brainy-org/brainy)
|
The in-process knowledge database for TypeScript — vector search, graph traversal,<br>
|
||||||
[](https://github.com/brainy-org/brainy/issues)
|
and metadata filtering unified in a single query.
|
||||||
|
</p>
|
||||||
**Multi-Dimensional AI Database with Triple Intelligence Engine**
|
|
||||||
|
<p align="center">
|
||||||
Brainy is a next-generation AI database that combines vector similarity, graph relationships, and field filtering into a unified "Triple Intelligence" query system. Built for modern AI applications that need semantic search, relationship mapping, and precise data filtering all in one query.
|
<a href="https://www.npmjs.com/package/@soulcraft/brainy"><img src="https://img.shields.io/npm/v/@soulcraft/brainy.svg" alt="npm version"></a>
|
||||||
|
<a href="https://www.npmjs.com/package/@soulcraft/brainy"><img src="https://img.shields.io/npm/dm/@soulcraft/brainy.svg" alt="npm downloads"></a>
|
||||||
## ✨ Features
|
<a href="https://source.soulcraft.com/soulcraft/brainy/actions"><img src="https://source.soulcraft.com/soulcraft/brainy/actions/workflows/ci.yml/badge.svg?branch=main" alt="CI"></a>
|
||||||
|
<a href="https://soulcraft.com/docs"><img src="https://img.shields.io/badge/docs-soulcraft.com-blue.svg" alt="Documentation"></a>
|
||||||
### 🧠 Triple Intelligence Engine ✅ Available Now
|
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="MIT License"></a>
|
||||||
- **Vector Search**: Semantic similarity using HNSW indexing
|
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/%3C%2F%3E-TypeScript-%230074c1.svg" alt="TypeScript"></a>
|
||||||
- **Graph Relationships**: Complex relationship mapping and traversal
|
</p>
|
||||||
- **Field Filtering**: Precise metadata filtering with O(1) lookups
|
|
||||||
- **Unified Queries**: All three intelligence types in a single query
|
<p align="center">
|
||||||
|
<a href="#quick-start">Quick start</a> ·
|
||||||
### 🎯 Zero Configuration ✅ Available Now
|
<a href="#one-query-three-engines">One query</a> ·
|
||||||
- **Auto-Detects Environment**: Node.js, Browser, Edge, Deno
|
<a href="#feature-tour">Features</a> ·
|
||||||
- **Auto-Selects Storage**: Best storage for your environment
|
<a href="#when-you-outgrow-brainy">Scale with Cor</a> ·
|
||||||
- **Works Instantly**: No setup required
|
<a href="#documentation">Docs</a> ·
|
||||||
- **Smart Defaults**: Optimized out of the box
|
<a href="#support--community">Support</a>
|
||||||
|
</p>
|
||||||
### 🔧 Production Ready ✅ Available Now
|
|
||||||
- **Universal Storage**: FileSystem, S3, OPFS, Memory
|
|
||||||
- **MIT License**: No limits, no tiers, truly open source
|
|
||||||
- **TypeScript Native**: Full type safety and IntelliSense support
|
|
||||||
- **Cross Platform**: Node.js, Browser, Web Workers, Edge Runtime
|
|
||||||
|
|
||||||
### ⚡ High Performance
|
|
||||||
- **HNSW Indexing**: Sub-millisecond vector search
|
|
||||||
- **Smart Caching**: Intelligent query optimization
|
|
||||||
- **Field Indexes**: O(1) metadata lookups
|
|
||||||
- **Streaming Support**: Handle millions of records efficiently
|
|
||||||
|
|
||||||
### 🛠 Developer Experience
|
|
||||||
- **Simple API**: Intuitive methods that just work
|
|
||||||
- **Rich CLI**: Interactive command-line interface
|
|
||||||
- **Comprehensive Tests**: 400+ tests covering all features
|
|
||||||
- **Excellent Docs**: Clear examples and API reference
|
|
||||||
|
|
||||||
### 🚀 Enterprise Features ✅ Available Now
|
|
||||||
- **WAL**: Write-ahead logging for durability ✅
|
|
||||||
- **Entity Registry**: High-performance deduplication ✅
|
|
||||||
- **Neural Import**: AI-powered entity detection ✅
|
|
||||||
- **Distributed Modes**: Read-only/Write-only optimization ✅
|
|
||||||
- **Statistics**: Comprehensive metrics and monitoring ✅
|
|
||||||
- **3-Level Cache**: Hot/Warm/Cold intelligent caching ✅
|
|
||||||
- **11+ Augmentations**: Including WebSocket, WebRTC, more ✅
|
|
||||||
|
|
||||||
## 📊 Current Status
|
|
||||||
|
|
||||||
Brainy 2.0.0 is in active development. Core features are production-ready, with advanced features on the roadmap.
|
|
||||||
|
|
||||||
### ✅ Production Ready
|
|
||||||
- Noun-Verb data model with entity detection
|
|
||||||
- Triple Intelligence queries with optimizations
|
|
||||||
- All storage adapters with caching
|
|
||||||
- Natural language queries (220+ patterns)
|
|
||||||
- WAL, Entity Registry, and 11+ augmentations
|
|
||||||
- Distributed read/write modes
|
|
||||||
- Performance monitoring and statistics
|
|
||||||
|
|
||||||
### 🚧 Needs Integration
|
|
||||||
- Import/Export CLI commands (code exists)
|
|
||||||
- GPU acceleration (detection works)
|
|
||||||
- Auto-optimization (metrics collected)
|
|
||||||
- Active learning (framework exists)
|
|
||||||
|
|
||||||
### 📅 Roadmap
|
|
||||||
See [ROADMAP.md](ROADMAP.md) for upcoming features and timeline.
|
|
||||||
|
|
||||||
## 🚀 Quick Start
|
|
||||||
|
|
||||||
### Installation
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npm install brainy
|
|
||||||
```
|
|
||||||
|
|
||||||
### Basic Usage
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
|
|
||||||
// Initialize with zero configuration
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Add entities (nouns) with automatic embedding generation
|
|
||||||
await brain.addNoun("The quick brown fox jumps over the lazy dog", {
|
|
||||||
category: "animals",
|
|
||||||
mood: "playful",
|
|
||||||
timestamp: Date.now()
|
|
||||||
})
|
|
||||||
|
|
||||||
await brain.addNoun("Machine learning transforms how we process information", {
|
|
||||||
category: "technology",
|
|
||||||
mood: "analytical",
|
|
||||||
timestamp: Date.now()
|
|
||||||
})
|
|
||||||
|
|
||||||
// Triple Intelligence: Vector + Graph + Field in one query
|
|
||||||
const results = await brain.search("animals running fast", {
|
|
||||||
where: {
|
|
||||||
category: "animals",
|
|
||||||
timestamp: { $gte: Date.now() - 86400000 } // last 24 hours
|
|
||||||
},
|
|
||||||
limit: 10
|
|
||||||
})
|
|
||||||
|
|
||||||
console.log(results)
|
|
||||||
// [{ id: "...", content: "The quick brown fox...", score: 0.92, metadata: {...} }]
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🤖 Model Loading (AI Embeddings)
|
|
||||||
|
|
||||||
Brainy uses AI embedding models to understand and process your data semantically. **Zero configuration required** - models load automatically.
|
|
||||||
|
|
||||||
### ✅ Zero Configuration (Recommended)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init() // Models download automatically on first use
|
|
||||||
```
|
|
||||||
|
|
||||||
**What happens automatically:**
|
|
||||||
1. Checks for local models in `./models/`
|
|
||||||
2. Downloads All-MiniLM-L6-v2 (384 dimensions) if needed
|
|
||||||
3. Uses intelligent cascade: Local → CDN → GitHub → HuggingFace
|
|
||||||
4. Ready to use immediately
|
|
||||||
|
|
||||||
### 🐳 Production/Docker Setup
|
|
||||||
```dockerfile
|
|
||||||
# Pre-download models during build (recommended)
|
|
||||||
RUN npm run download-models
|
|
||||||
|
|
||||||
# Optional: Force local-only mode
|
|
||||||
ENV BRAINY_ALLOW_REMOTE_MODELS=false
|
|
||||||
```
|
|
||||||
|
|
||||||
### 🔒 Offline/Air-Gapped Environments
|
|
||||||
```bash
|
|
||||||
# On connected machine
|
|
||||||
npm run download-models
|
|
||||||
|
|
||||||
# Copy models to offline machine
|
|
||||||
cp -r ./models /path/to/offline/project/
|
|
||||||
|
|
||||||
# Force local-only mode
|
|
||||||
export BRAINY_ALLOW_REMOTE_MODELS=false
|
|
||||||
```
|
|
||||||
|
|
||||||
### 📋 Environment Variables (Optional)
|
|
||||||
| Variable | Default | Description |
|
|
||||||
|----------|---------|-------------|
|
|
||||||
| `BRAINY_ALLOW_REMOTE_MODELS` | `true` | Allow/block model downloads |
|
|
||||||
| `BRAINY_MODELS_PATH` | `./models` | Custom model storage path |
|
|
||||||
|
|
||||||
### 🚨 Troubleshooting
|
|
||||||
- **"Failed to load embedding model"** → Run `npm run download-models`
|
|
||||||
- **Slow model downloads** → Pre-download during build/CI
|
|
||||||
- **Container memory issues** → Pre-download models, increase memory limit
|
|
||||||
|
|
||||||
📚 **Complete Guide**: [docs/guides/model-loading.md](docs/guides/model-loading.md)
|
|
||||||
|
|
||||||
## 📊 Triple Intelligence in Action
|
|
||||||
|
|
||||||
### Vector Similarity
|
|
||||||
```typescript
|
|
||||||
// Semantic search across your data
|
|
||||||
const results = await brain.search("fast animals")
|
|
||||||
// Finds: "quick brown fox", "racing horses", "cheetah running"
|
|
||||||
```
|
|
||||||
|
|
||||||
### Graph Relationships
|
|
||||||
```typescript
|
|
||||||
// Find related entities and concepts
|
|
||||||
const related = await brain.findRelated(entityId, {
|
|
||||||
depth: 2,
|
|
||||||
relationship: "semantic"
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Field Filtering
|
|
||||||
```typescript
|
|
||||||
// Precise metadata filtering with O(1) performance
|
|
||||||
const filtered = await brain.search("technology", {
|
|
||||||
where: {
|
|
||||||
category: "ai",
|
|
||||||
rating: { $gte: 4.5 },
|
|
||||||
published: { $between: ["2024-01-01", "2024-12-31"] }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Combined Intelligence
|
|
||||||
```typescript
|
|
||||||
// All three intelligence types working together
|
|
||||||
const results = await brain.search("machine learning concepts", {
|
|
||||||
where: {
|
|
||||||
category: { $in: ["ai", "technology"] },
|
|
||||||
difficulty: { $lte: 5 }
|
|
||||||
},
|
|
||||||
includeRelated: true,
|
|
||||||
depth: 2
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🗄️ Storage Adapters
|
|
||||||
|
|
||||||
Brainy supports multiple storage backends with the same API:
|
|
||||||
|
|
||||||
### File System (Default)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'filesystem', path: './data' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Amazon S3 / Compatible
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 's3',
|
|
||||||
bucket: 'my-brainy-data',
|
|
||||||
region: 'us-east-1'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Origin Private File System (Browser)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'opfs' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Memory (Development)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'memory' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Advanced Querying with find()
|
|
||||||
|
|
||||||
### Triple Intelligence find() Method
|
|
||||||
```typescript
|
|
||||||
// Natural language queries with automatic intent recognition
|
|
||||||
const results = await brain.find("show me recent AI articles with high ratings")
|
|
||||||
// Automatically converts to: vector similarity + field filtering + date ranges
|
|
||||||
|
|
||||||
// MongoDB-style queries with semantic awareness
|
|
||||||
const results = await brain.find({
|
|
||||||
$or: [
|
|
||||||
{ category: "technology" },
|
|
||||||
{ $vector: { $similar: "artificial intelligence", threshold: 0.8 } }
|
|
||||||
],
|
|
||||||
metadata: {
|
|
||||||
published: { $gte: "2024-01-01" },
|
|
||||||
rating: { $in: [4, 5] }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Natural Language Understanding ✅ Available (Basic)
|
|
||||||
```typescript
|
|
||||||
// The find() method understands natural language queries
|
|
||||||
const results = await brain.find("technology articles about machine learning")
|
|
||||||
// Basic pattern matching for common queries
|
|
||||||
|
|
||||||
// Temporal queries (basic support)
|
|
||||||
const recent = await brain.find("recent documents")
|
|
||||||
// Recognizes common time expressions
|
|
||||||
|
|
||||||
// The search() method focuses on semantic similarity
|
|
||||||
const similar = await brain.search("documents similar to machine learning research")
|
|
||||||
// Pure vector similarity search
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔧 Configuration
|
|
||||||
|
|
||||||
### Environment Variables
|
|
||||||
```bash
|
|
||||||
BRAINY_STORAGE_TYPE=filesystem
|
|
||||||
BRAINY_STORAGE_PATH=./brainy-data
|
|
||||||
BRAINY_MODELS_PATH=./models
|
|
||||||
BRAINY_VECTOR_DIMENSIONS=384
|
|
||||||
```
|
|
||||||
|
|
||||||
### Programmatic Configuration
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'filesystem',
|
|
||||||
path: './data'
|
|
||||||
},
|
|
||||||
vectors: {
|
|
||||||
dimensions: 384,
|
|
||||||
model: '@huggingface/transformers/all-MiniLM-L6-v2'
|
|
||||||
},
|
|
||||||
performance: {
|
|
||||||
cacheSize: 1000,
|
|
||||||
batchSize: 100
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📱 CLI Usage
|
|
||||||
|
|
||||||
Brainy includes a powerful command-line interface:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Initialize a new database
|
|
||||||
brainy init
|
|
||||||
|
|
||||||
# Add entities (nouns)
|
|
||||||
brainy add-noun "Your content here" --category="example"
|
|
||||||
|
|
||||||
# Natural language queries
|
|
||||||
brainy find "show me examples from last week"
|
|
||||||
|
|
||||||
# Search with Triple Intelligence
|
|
||||||
brainy search "find similar content" --where='{"category":"example"}'
|
|
||||||
|
|
||||||
# Interactive mode with NLP
|
|
||||||
brainy chat
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧪 Testing
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Run all tests
|
|
||||||
npm test
|
|
||||||
|
|
||||||
# Run specific test suites
|
|
||||||
npm run test:core
|
|
||||||
npm run test:storage
|
|
||||||
npm run test:coverage
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📚 API Reference
|
|
||||||
|
|
||||||
### Core Methods
|
|
||||||
|
|
||||||
#### `brain.addNoun(content, metadata?)`
|
|
||||||
Add entities (nouns) with automatic embedding generation.
|
|
||||||
|
|
||||||
#### `brain.addVerb(source, target, type, metadata?)`
|
|
||||||
Create relationships (verbs) between entities.
|
|
||||||
|
|
||||||
#### `brain.search(query, options?)`
|
|
||||||
Triple Intelligence search with vector similarity, field filtering, and relationship traversal.
|
|
||||||
|
|
||||||
#### `brain.find(query)`
|
|
||||||
Advanced Triple Intelligence queries with natural language or structured syntax.
|
|
||||||
- Accepts natural language: `brain.find("recent posts about AI")`
|
|
||||||
- Accepts structured queries: `brain.find({ category: "AI", date: { $gte: "2024-01-01" } })`
|
|
||||||
- Automatically interprets intent, time ranges, and filters
|
|
||||||
|
|
||||||
#### `brain.get(id)`
|
|
||||||
Retrieve specific items by ID.
|
|
||||||
|
|
||||||
#### `brain.updateMetadata(id, metadata)`
|
|
||||||
Update entity metadata.
|
|
||||||
|
|
||||||
#### `brain.delete(id)`
|
|
||||||
Remove items by ID (soft delete by default).
|
|
||||||
|
|
||||||
### Advanced Methods
|
|
||||||
|
|
||||||
#### `brain.cluster(options?)`
|
|
||||||
Semantic clustering of your data.
|
|
||||||
|
|
||||||
#### `brain.findRelated(id, options?)`
|
|
||||||
Find semantically or structurally related items.
|
|
||||||
|
|
||||||
#### `brain.statistics()`
|
|
||||||
Get performance and usage statistics.
|
|
||||||
|
|
||||||
## 🤝 Contributing
|
|
||||||
|
|
||||||
We welcome contributions! Please see our [Contributing Guide](CONTRIBUTING.md) for details.
|
|
||||||
|
|
||||||
### Development Setup
|
|
||||||
```bash
|
|
||||||
git clone https://github.com/brainy-org/brainy.git
|
|
||||||
cd brainy
|
|
||||||
npm install
|
|
||||||
npm run build
|
|
||||||
npm test
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📄 License
|
|
||||||
|
|
||||||
MIT License - see [LICENSE](LICENSE) file for details.
|
|
||||||
|
|
||||||
## 🙏 Acknowledgments
|
|
||||||
|
|
||||||
- [Hugging Face Transformers.js](https://huggingface.co/docs/transformers.js) for embedding models
|
|
||||||
- [HNSW](https://github.com/nmslib/hnswlib) for efficient vector indexing
|
|
||||||
- The open source AI/ML community for inspiration
|
|
||||||
|
|
||||||
## 💬 Support
|
|
||||||
|
|
||||||
- [GitHub Issues](https://github.com/brainy-org/brainy/issues) - Bug reports and feature requests
|
|
||||||
- [Discussions](https://github.com/brainy-org/brainy/discussions) - Community support and ideas
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
**Built with ❤️ for the AI community**
|
Built because we were tired of stitching a vector store to a graph database to a document store — and spending weeks on plumbing before writing a line of business logic. Brainy indexes every fact **three ways at once** and lets one call query them together:
|
||||||
|
|
||||||
|
| You write | Brainy indexes it as | You query it with |
|
||||||
|
|---|---|---|
|
||||||
|
| `data: 'Ada wrote the first program'` | a **384-dim vector** (local embedding — no API key) | `find({ query: 'computing pioneers' })` |
|
||||||
|
| `metadata: { field: 'CS', year: 1843 }` | **structured fields** (O(1) exact, O(log n) range) | `find({ where: { year: { lessThan: 1900 } } })` |
|
||||||
|
| `relate({ from: ada, to: babbage })` | a **typed, directed graph edge** | `find({ connected: { to: babbage, depth: 2 } })` |
|
||||||
|
|
||||||
|
It runs **inside your process** — no server, no Docker, nothing to operate — and persists to plain files you can snapshot with a hard link.
|
||||||
|
|
||||||
|
**New here?** → **[What is Brainy? — plain-language overview, no jargon](docs/eli5.md)**
|
||||||
|
|
||||||
|
## Quick start
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bun add @soulcraft/brainy # Bun ≥ 1.1 — recommended
|
||||||
|
npm install @soulcraft/brainy # Node.js ≥ 22
|
||||||
|
```
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
import { Brainy, NounType, VerbType } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
const brain = new Brainy() // in-memory; one line swaps to disk
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Text auto-embeds locally; metadata auto-indexes
|
||||||
|
const react = await brain.add({
|
||||||
|
data: 'React is a JavaScript library for building user interfaces',
|
||||||
|
type: NounType.Concept,
|
||||||
|
subtype: 'library',
|
||||||
|
metadata: { category: 'frontend', year: 2013 }
|
||||||
|
})
|
||||||
|
|
||||||
|
const next = await brain.add({
|
||||||
|
data: 'Next.js is a React framework with server-side rendering',
|
||||||
|
type: NounType.Concept,
|
||||||
|
subtype: 'framework',
|
||||||
|
metadata: { category: 'frontend', year: 2016 }
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.relate({ from: next, to: react, type: VerbType.DependsOn, subtype: 'runtime' })
|
||||||
|
```
|
||||||
|
|
||||||
|
## One query, three engines
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const results = await brain.find({
|
||||||
|
query: 'modern frontend frameworks', // vector — what it means
|
||||||
|
where: { year: { greaterThan: 2015 } }, // metadata — what it is
|
||||||
|
connected: { to: react, depth: 2 } // graph — what it touches
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
Every clause is optional; any combination composes. Under the hood Brainy plans the query across an HNSW vector index, a roaring-bitmap field index, and an adjacency graph index — and re-validates every result against your predicate before returning it, so a corrupt index can never hand you a wrong answer.
|
||||||
|
|
||||||
|
## Feature tour
|
||||||
|
|
||||||
|
### The database is a value
|
||||||
|
|
||||||
|
Pin it, rewind it, fork it. Snapshot isolation without a server.
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const db = brain.now() // pin current state — O(1)
|
||||||
|
|
||||||
|
await brain.transact([ // atomic all-or-nothing, CAS-guarded
|
||||||
|
{ op: 'update', id: order, metadata: { status: 'paid' } },
|
||||||
|
{ op: 'relate', from: invoice, to: order, type: VerbType.References, subtype: 'billing' }
|
||||||
|
], { ifAtGeneration: db.generation })
|
||||||
|
|
||||||
|
await db.get(order) // still 'pending' — pinned forever
|
||||||
|
await brain.get(order) // 'paid' — live
|
||||||
|
|
||||||
|
const lastWeek = await brain.asOf(Date.now() - 7 * 86_400_000) // full query surface, past state
|
||||||
|
const whatIf = await db.with([{ op: 'remove', id: order }]) // speculative — never touches disk
|
||||||
|
await brain.now().persist('/backups/today') // instant hard-link snapshot
|
||||||
|
```
|
||||||
|
|
||||||
|
**[Consistency model](docs/concepts/consistency-model.md)** · **[Snapshots & time travel](docs/guides/snapshots-and-time-travel.md)**
|
||||||
|
|
||||||
|
### Local embeddings — no API keys
|
||||||
|
|
||||||
|
Strings embed on-device with a bundled MiniLM model (WASM). Semantic search works offline, in CI, and on air-gapped machines, at zero cost per call. Hybrid keyword + semantic ranking is the default:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
await brain.find({ query: 'David Smith' }) // auto: text + semantic
|
||||||
|
await brain.find({ query: 'AI concepts', searchMode: 'semantic' }) // semantic only
|
||||||
|
```
|
||||||
|
|
||||||
|
### A typed graph, not a bag of edges
|
||||||
|
|
||||||
|
42 entity types × 127 relationship types form a shared vocabulary for any domain — healthcare (`Patient → diagnoses → Condition`), finance (`Account → transfers → Transaction`), yours. Your own taxonomy layers on with `subtype`, enforced at write time:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
await brain.add({ data: 'Avery Brooks', type: NounType.Person, subtype: 'employee' })
|
||||||
|
|
||||||
|
brain.counts.bySubtype(NounType.Person) // O(1) — { employee: 12, customer: 847 }
|
||||||
|
brain.requireSubtype(NounType.Person, { values: ['employee', 'customer'], required: true })
|
||||||
|
```
|
||||||
|
|
||||||
|
**[Type system](docs/architecture/noun-verb-taxonomy.md)** · **[Subtypes & facets](docs/guides/subtypes-and-facets.md)**
|
||||||
|
|
||||||
|
### Graph analytics built in
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
await brain.graph.rank() // which entities matter most (centrality)
|
||||||
|
await brain.graph.communities() // natural clusters
|
||||||
|
await brain.graph.path(a, b) // how two things connect
|
||||||
|
await brain.graph.subgraph([seed], { depth: 2 }) // bounded neighborhood → { nodes, edges }
|
||||||
|
await brain.graph.export() // whole graph, one O(N+E) streaming pass
|
||||||
|
```
|
||||||
|
|
||||||
|
### Write-time aggregations
|
||||||
|
|
||||||
|
`SUM` / `COUNT` / `AVG` / `MIN` / `MAX` with `GROUP BY` and time windows, maintained incrementally on every write — reads are O(1) lookups, not scans. **[Aggregation guide](docs/guides/aggregation.md)**
|
||||||
|
|
||||||
|
### Import anything
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
await brain.import('customers.csv')
|
||||||
|
await brain.import('sales.xlsx') // every sheet
|
||||||
|
await brain.import('research-paper.pdf') // tables extracted
|
||||||
|
await brain.import('https://api.example.com/data.json')
|
||||||
|
```
|
||||||
|
|
||||||
|
Entities auto-classify on the way in; `brain.extractEntities(text)` exposes the same NER ensemble directly. **[Import guide](docs/guides/import-anything.md)**
|
||||||
|
|
||||||
|
### A filesystem that understands content
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
await brain.vfs.writeFile('/docs/readme.md', 'Project documentation')
|
||||||
|
await brain.vfs.search('React components with hooks') // semantic file search
|
||||||
|
```
|
||||||
|
|
||||||
|
**[VFS quick start](docs/vfs/QUICK_START.md)**
|
||||||
|
|
||||||
|
### Operations-grade by default
|
||||||
|
|
||||||
|
- **Single-writer, many-reader** — an exclusive lock protects the data directory; `Brainy.openReadOnly()` and the `brainy inspect` CLI examine a live brain from another process, safely.
|
||||||
|
- **Self-upgrading data files** — a 7.x brain opens under 8.x and migrates itself behind an observable lock (`getIndexStatus().migration`), with an automatic pre-upgrade backup. No migration scripts.
|
||||||
|
- **No silent wrong answers** — cold-open guards self-heal or throw typed errors (`MetadataIndexNotReadyError`, `GraphIndexNotReadyError`); they never return `[]` for data that exists.
|
||||||
|
|
||||||
|
**[Multi-process model](docs/concepts/multi-process.md)** · **[Inspection guide](docs/guides/inspection.md)**
|
||||||
|
|
||||||
|
## When you outgrow Brainy
|
||||||
|
|
||||||
|
Brainy's pure-TypeScript engines carry real workloads a long way on their own — see the measured, per-operation numbers (not marketing figures) in **[docs/performance-envelopes.md](docs/performance-envelopes.md)** for what to expect, unaccelerated, on plain filesystem storage.
|
||||||
|
|
||||||
|
When a deployment needs native-scale vector/graph performance — memory-mapped indexes that don't need your dataset in RAM, billion-scale ambitions — add the native engine. **The API doesn't change:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm install @soulcraft/cor
|
||||||
|
```
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const brain = new Brainy({ storage: { type: 'filesystem', path: './data' } })
|
||||||
|
await brain.init() // @soulcraft/cor detected — same code, native engines underneath
|
||||||
|
```
|
||||||
|
|
||||||
|
Installing the package is the opt-in: if `@soulcraft/cor` is present, it loads and announces itself in the init log; if it's present but broken, `init()` **throws** — an installed accelerator never silently vanishes behind the JS engines. Opt out with `plugins: []`, or pin exactly what loads with `plugins: ['@soulcraft/cor']`. [`@soulcraft/cor`](https://www.npmjs.com/package/@soulcraft/cor) (Brainy 8.x ↔ Cor 3.x, version-matched) registers Rust implementations behind every provider seam: SIMD distance kernels, memory-mapped storage, a disk-native vector index that doesn't need your dataset in RAM, durable LSM field/graph indexes that serve cold opens instantly, and native aggregation. Recall@10 measured **0.99 / 0.96 / 0.96 at 1M / 10M / 100M vectors** in Cor's release gate.
|
||||||
|
|
||||||
|
Open core, commercial accelerator: Brainy is MIT and complete on its own — Cor is more headroom for when you need it, not capability held back to sell you later. Licensing and support: **cor@soulcraft.com**.
|
||||||
|
|
||||||
|
## Performance
|
||||||
|
|
||||||
|
- Per-operation p50/p95 at 1k and 10k entities, pure-JS floor, measured and re-run every release that touches a measured path: **[docs/performance-envelopes.md](docs/performance-envelopes.md)**.
|
||||||
|
- JS distance kernels: **~6× faster cosine, ~1.4× euclidean** than 7.x (measured: [`tests/benchmarks/distance-microbench.mjs`](tests/benchmarks/distance-microbench.mjs), 384-dim, median of 41).
|
||||||
|
- Whole-graph reads are single **O(N + E)** cursor walks — a consumer-measured 19k-edge export dropped from ~27 s of per-node calls to one scan.
|
||||||
|
- Capacity planning and architecture: **[docs/PERFORMANCE.md](docs/PERFORMANCE.md)** · **[docs/SCALING.md](docs/SCALING.md)**
|
||||||
|
|
||||||
|
## Use cases
|
||||||
|
|
||||||
|
**AI agent memory** — persistent semantic recall with relationship tracking · **Knowledge bases** — auto-linking and meaning-aware navigation · **Semantic search** over codebases, documents, media · **Enterprise data** — CRM, catalogs, institutional memory · **Games & simulations** — worlds and characters that remember.
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
|
||||||
|
| Start | Core | Going deeper |
|
||||||
|
|---|---|---|
|
||||||
|
| [Brainy explained simply](docs/eli5.md) | [API reference](docs/api/README.md) | [Architecture overview](docs/architecture/overview.md) |
|
||||||
|
| [Installation](docs/guides/installation.md) | [Data model](docs/DATA_MODEL.md) | [Consistency model](docs/concepts/consistency-model.md) |
|
||||||
|
| [Natural-language queries](docs/guides/natural-language.md) | [Query operators](docs/QUERY_OPERATORS.md) | [Multi-process model](docs/concepts/multi-process.md) |
|
||||||
|
| | [Find system](docs/FIND_SYSTEM.md) | [Scaling](docs/SCALING.md) |
|
||||||
|
|
||||||
|
## Requirements
|
||||||
|
|
||||||
|
**Bun ≥ 1.1** (recommended) or **Node.js ≥ 22**. Brainy 8.x is server-only; the 7.x line remains on npm for browser use.
|
||||||
|
|
||||||
|
## Support & community
|
||||||
|
|
||||||
|
- **Bugs and ideas** → **brainy@soulcraft.com** — no account needed, you'll get a receipt.
|
||||||
|
- **Security reports** → **security@soulcraft.com** — see **[SECURITY.md](SECURITY.md)**.
|
||||||
|
- **Contributing** → see **[CONTRIBUTING.md](CONTRIBUTING.md)**.
|
||||||
|
|
||||||
|
MIT © Brainy Contributors.
|
||||||
|
|
|
||||||
3029
RELEASES.md
Normal file
3029
RELEASES.md
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1,141 +0,0 @@
|
||||||
# Brainy 2.0 Scalability Plan - Millions of Records
|
|
||||||
|
|
||||||
## Current Performance Profile
|
|
||||||
- **Exact match**: O(1) - ✅ Excellent (same as MongoDB)
|
|
||||||
- **Range queries**: O(log n) - ✅ Excellent (same as MongoDB B-tree)
|
|
||||||
- **Memory usage**: ~1KB per record - ⚠️ Problematic at scale
|
|
||||||
|
|
||||||
## Scalability Bottlenecks
|
|
||||||
|
|
||||||
### 1. Memory Limits (CRITICAL)
|
|
||||||
**Problem**: All indices in RAM
|
|
||||||
- 1M records = 1.1 GB RAM ✅
|
|
||||||
- 10M records = 11 GB RAM ❌
|
|
||||||
- 100M records = 110 GB RAM ❌❌❌
|
|
||||||
|
|
||||||
**Solution**: Hybrid memory/disk approach
|
|
||||||
```typescript
|
|
||||||
interface ScalableIndex {
|
|
||||||
hotCache: Map<string, Set<string>> // Top 10K entries in RAM
|
|
||||||
coldStorage: DiskIndex // Rest on disk (LevelDB/RocksDB)
|
|
||||||
bloomFilter: BloomFilter // Quick existence check
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Sorted Index Scalability
|
|
||||||
**Problem**: Single array for entire field
|
|
||||||
- 10M values = massive array sort
|
|
||||||
- Binary search still O(log n) but cache misses
|
|
||||||
|
|
||||||
**Solution**: B+ Tree structure
|
|
||||||
```typescript
|
|
||||||
interface BPlusTreeIndex {
|
|
||||||
root: BPlusNode
|
|
||||||
leafLevel: LinkedList<LeafNode> // For range scans
|
|
||||||
height: number // Typically 3-4 levels
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Index Persistence
|
|
||||||
**Problem**: Rebuilding on startup
|
|
||||||
- 1M records = 30 seconds startup ❌
|
|
||||||
- 10M records = 5 minutes startup ❌❌❌
|
|
||||||
|
|
||||||
**Solution**: Incremental index snapshots
|
|
||||||
```typescript
|
|
||||||
// Save index periodically
|
|
||||||
await storage.saveIndex('field_price_sorted', sortedIndex)
|
|
||||||
// Load on startup
|
|
||||||
const cached = await storage.loadIndex('field_price_sorted')
|
|
||||||
```
|
|
||||||
|
|
||||||
## Recommended Architecture for Scale
|
|
||||||
|
|
||||||
### Tier 1: <100K records (Current)
|
|
||||||
- ✅ All in memory
|
|
||||||
- ✅ Hash + sorted indices
|
|
||||||
- ✅ No changes needed
|
|
||||||
|
|
||||||
### Tier 2: 100K-1M records (Minor changes)
|
|
||||||
```typescript
|
|
||||||
class OptimizedMetadataIndex {
|
|
||||||
// Lazy load sorted indices
|
|
||||||
private async ensureSortedIndex(field: string) {
|
|
||||||
if (!this.sortedIndices.has(field)) {
|
|
||||||
await this.loadOrBuildSortedIndex(field)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Persist indices to storage
|
|
||||||
private async persistIndex(field: string) {
|
|
||||||
const index = this.sortedIndices.get(field)
|
|
||||||
await this.storage.saveMetadata(`__index_${field}`, index)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Tier 3: 1M-10M records (Major refactor)
|
|
||||||
```typescript
|
|
||||||
class ScalableMetadataIndex {
|
|
||||||
private leveldb: LevelDB // Or RocksDB
|
|
||||||
private hotCache: LRUCache<string, Set<string>>
|
|
||||||
private bloomFilters: Map<string, BloomFilter>
|
|
||||||
|
|
||||||
async getIds(field: string, value: any): Promise<string[]> {
|
|
||||||
// Check bloom filter first (O(1))
|
|
||||||
if (!this.bloomFilters.get(field)?.mightContain(value)) {
|
|
||||||
return []
|
|
||||||
}
|
|
||||||
|
|
||||||
// Check hot cache (O(1))
|
|
||||||
const cached = this.hotCache.get(`${field}:${value}`)
|
|
||||||
if (cached) return Array.from(cached)
|
|
||||||
|
|
||||||
// Load from disk (O(log n))
|
|
||||||
const ids = await this.leveldb.get(`idx:${field}:${value}`)
|
|
||||||
this.hotCache.set(`${field}:${value}`, new Set(ids))
|
|
||||||
return ids
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Tier 4: 10M+ records (Distributed)
|
|
||||||
- Shard by ID range or hash
|
|
||||||
- Multiple Brainy instances
|
|
||||||
- Coordinator node for queries
|
|
||||||
- Similar to MongoDB sharding
|
|
||||||
|
|
||||||
## Performance at Scale
|
|
||||||
|
|
||||||
| Records | Current | Optimized | MongoDB |
|
|
||||||
|---------|---------|-----------|---------|
|
|
||||||
| 10K | 10ms | 10ms | 15ms |
|
|
||||||
| 100K | 15ms | 15ms | 20ms |
|
|
||||||
| 1M | 25ms | 20ms | 25ms |
|
|
||||||
| 10M | OOM ❌ | 30ms | 35ms |
|
|
||||||
| 100M | OOM ❌ | 50ms | 60ms |
|
|
||||||
|
|
||||||
## Implementation Priority
|
|
||||||
|
|
||||||
1. **Quick Win**: Index persistence (prevent rebuild)
|
|
||||||
2. **Medium**: LRU cache for hot data
|
|
||||||
3. **Long-term**: B+ tree indices
|
|
||||||
4. **Future**: Sharding support
|
|
||||||
|
|
||||||
## Memory Usage Comparison
|
|
||||||
|
|
||||||
| Records | Current | Optimized | MongoDB |
|
|
||||||
|---------|---------|-----------|---------|
|
|
||||||
| 100K | 110 MB | 110 MB | 150 MB |
|
|
||||||
| 1M | 1.1 GB | 500 MB | 1.5 GB |
|
|
||||||
| 10M | 11 GB ❌ | 2 GB ✅ | 8 GB |
|
|
||||||
| 100M | 110 GB ❌ | 5 GB ✅ | 50 GB |
|
|
||||||
|
|
||||||
## Conclusion
|
|
||||||
|
|
||||||
**Current state**: Excellent for <100K records, good for <1M
|
|
||||||
**With optimizations**: Can handle 10M+ records
|
|
||||||
**Comparable to**: MongoDB, Firestore for most operations
|
|
||||||
**Better than**: Traditional databases for vector + metadata hybrid queries
|
|
||||||
|
|
||||||
The architecture is **sound** - just needs memory optimization for scale!
|
|
||||||
36
SECURITY.md
Normal file
36
SECURITY.md
Normal file
|
|
@ -0,0 +1,36 @@
|
||||||
|
# Security Policy
|
||||||
|
|
||||||
|
## Reporting a vulnerability
|
||||||
|
|
||||||
|
Email **security@soulcraft.com**. That's the one door for security reports
|
||||||
|
across the company, and it works the same way for Brainy: every report is
|
||||||
|
read by a human, you'll get a private receipt, and we'll work with you on
|
||||||
|
coordinated disclosure — please don't open a public issue for anything
|
||||||
|
that isn't already public.
|
||||||
|
|
||||||
|
Include what you'd want if you were on the other end: affected version,
|
||||||
|
how to reproduce, and what you think the impact is. If you have a patch or
|
||||||
|
a suggested fix, send it along — it's welcome but not required.
|
||||||
|
|
||||||
|
There is no bounty program today. We're saying that plainly so you know
|
||||||
|
what to expect going in.
|
||||||
|
|
||||||
|
## Response time
|
||||||
|
|
||||||
|
We respond as fast as truth allows. That means: no fixed SLA, no promise of
|
||||||
|
a reply within a specific number of hours — but a real report from a real
|
||||||
|
person gets read promptly and taken seriously. If you haven't heard anything
|
||||||
|
in a reasonable stretch, a follow-up email is completely fine.
|
||||||
|
|
||||||
|
## Supported versions
|
||||||
|
|
||||||
|
The latest `8.x` minor release line receives security fixes. If you're
|
||||||
|
running an older major version, please upgrade before reporting — we can't
|
||||||
|
commit to backporting fixes to unsupported lines.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This policy covers the `@soulcraft/brainy` package itself — the code in
|
||||||
|
this repository. If you're evaluating a deployment that also uses
|
||||||
|
`@soulcraft/cor`, report issues in that package the same way, to the same
|
||||||
|
address; we'll route internally.
|
||||||
|
|
@ -1,225 +0,0 @@
|
||||||
# 🧠 COMPREHENSIVE TEST SUITE ULTRATHINK ANALYSIS
|
|
||||||
|
|
||||||
**Date**: 2025-08-25
|
|
||||||
**Context**: Brainy 2.0 Augmentation System Migration
|
|
||||||
**Total Test Files**: 49
|
|
||||||
|
|
||||||
## 🏗️ ARCHITECTURAL CHANGES IMPACT ANALYSIS
|
|
||||||
|
|
||||||
### **Major Changes Affecting Tests:**
|
|
||||||
|
|
||||||
1. **Augmentation System Migration**
|
|
||||||
- Core functionality moved to augmentations (cache, index, metrics, storage)
|
|
||||||
- Two-phase initialization (storage augmentations first)
|
|
||||||
- Method delegation through augmentations
|
|
||||||
|
|
||||||
2. **API Evolution**
|
|
||||||
- Methods like `cleanup()` may have changed
|
|
||||||
- Export structure in index.ts updated
|
|
||||||
- New unified BrainyAugmentation interface
|
|
||||||
|
|
||||||
3. **Configuration Changes**
|
|
||||||
- Auto-registration of default augmentations
|
|
||||||
- New augmentation-based storage initialization
|
|
||||||
|
|
||||||
## 📊 TEST CATEGORIES ANALYSIS
|
|
||||||
|
|
||||||
### 🟢 **LIKELY WORKING** (Minimal Changes Expected)
|
|
||||||
1. **Core Functionality Tests** - Tests basic BrainyData usage
|
|
||||||
- `core.test.ts` - Library exports, basic functionality
|
|
||||||
- `database-operations.test.ts` - CRUD operations
|
|
||||||
- `vector-operations.test.ts` - Vector search (if using public API)
|
|
||||||
- `triple-intelligence.test.ts` - Advanced queries
|
|
||||||
|
|
||||||
2. **Environment Tests** - Platform compatibility
|
|
||||||
- `environment.browser.test.ts`
|
|
||||||
- `environment.node.test.ts`
|
|
||||||
- `multi-environment.test.ts`
|
|
||||||
|
|
||||||
3. **Storage Integration Tests** - Should work with new augmentation system
|
|
||||||
- `s3-comprehensive.test.ts` - S3 through augmentations
|
|
||||||
- `opfs-storage.test.ts` - OPFS through augmentations
|
|
||||||
|
|
||||||
### 🟡 **NEEDS UPDATES** (Medium Impact)
|
|
||||||
1. **API Consistency Tests**
|
|
||||||
- `consistent-api.test.ts` - May have method signature changes
|
|
||||||
- `unified-api.test.ts` - API evolution issues
|
|
||||||
- May use `cleanup()` vs new shutdown methods
|
|
||||||
|
|
||||||
2. **Configuration & Initialization Tests**
|
|
||||||
- `auto-configuration.test.ts` - New augmentation auto-registration
|
|
||||||
- `zero-config-models.test.ts` - Initialization pattern changes
|
|
||||||
|
|
||||||
3. **Performance Tests**
|
|
||||||
- `performance.test.ts` - Augmentation overhead validation needed
|
|
||||||
- `metadata-performance.test.ts` - Now through IndexAugmentation
|
|
||||||
- `throttling-metrics.test.ts` - May need augmentation context
|
|
||||||
|
|
||||||
4. **Feature Integration Tests**
|
|
||||||
- `brainy-chat.test.ts` - BrainyChat integration
|
|
||||||
- `nlp-patterns-comprehensive.test.ts` - NLP pattern system
|
|
||||||
- `neural-*.test.ts` - Neural integration tests
|
|
||||||
|
|
||||||
### 🟠 **MAJOR UPDATES NEEDED** (High Impact)
|
|
||||||
1. **Export/Import Tests**
|
|
||||||
- `core.test.ts` - Tests exports that may not exist:
|
|
||||||
- `createSenseAugmentation`
|
|
||||||
- `addWebSocketSupport`
|
|
||||||
- `executeAugmentation`
|
|
||||||
- `loadAugmentationModule`
|
|
||||||
|
|
||||||
2. **Storage System Tests**
|
|
||||||
- `storage-adapter-coverage.test.ts` - Storage now through augmentations
|
|
||||||
- Tests that directly create storage adapters vs using augmentations
|
|
||||||
|
|
||||||
3. **Metadata & Statistics Tests**
|
|
||||||
- `statistics.test.ts` - Now through MetricsAugmentation
|
|
||||||
- `service-statistics.test.ts` - Service stats through augmentations
|
|
||||||
- `metadata-filter.test.ts` - Filtering through IndexAugmentation
|
|
||||||
|
|
||||||
### 🟢 **AUGMENTATION TESTS** (Should be working)
|
|
||||||
- `augmentations-batch-processing.test.ts`
|
|
||||||
- `augmentations-entity-registry.test.ts`
|
|
||||||
- `augmentations-request-deduplicator.test.ts`
|
|
||||||
- `augmentations-wal.test.ts`
|
|
||||||
|
|
||||||
### 🔴 **CRITICAL VALIDATION NEEDED**
|
|
||||||
1. **Release Tests**
|
|
||||||
- `release-critical.test.ts` - Must pass for 2.0 release
|
|
||||||
- `release-validation.test.ts` - End-to-end validation
|
|
||||||
|
|
||||||
2. **Edge Cases**
|
|
||||||
- `edge-cases.test.ts` - Ensure augmentation system handles edge cases
|
|
||||||
- `error-handling.test.ts` - Error handling through augmentations
|
|
||||||
|
|
||||||
## 🎯 COMPREHENSIVE TEST VALIDATION PLAN
|
|
||||||
|
|
||||||
### **Phase 1: Fix Export Issues (CRITICAL)**
|
|
||||||
#### Files to Fix:
|
|
||||||
- `core.test.ts` - Remove/update non-existent exports
|
|
||||||
- `unified-api.test.ts` - Fix import paths
|
|
||||||
- `consistent-api.test.ts` - Update method calls
|
|
||||||
|
|
||||||
#### Actions:
|
|
||||||
- [ ] Update index.ts to remove non-existent exports
|
|
||||||
- [ ] Fix import paths in test files
|
|
||||||
- [ ] Update method names (cleanup → shutdown, etc.)
|
|
||||||
|
|
||||||
### **Phase 2: Fix Augmentation Integration**
|
|
||||||
#### Files to Update:
|
|
||||||
- `auto-configuration.test.ts` - Test new augmentation auto-registration
|
|
||||||
- `statistics.test.ts` - Test statistics through MetricsAugmentation
|
|
||||||
- `metadata-*.test.ts` - Test metadata through IndexAugmentation
|
|
||||||
|
|
||||||
#### Actions:
|
|
||||||
- [ ] Update tests to use augmentation-delegated methods
|
|
||||||
- [ ] Test augmentation auto-registration
|
|
||||||
- [ ] Validate two-phase initialization
|
|
||||||
|
|
||||||
### **Phase 3: Fix API Evolution Issues**
|
|
||||||
#### Files to Update:
|
|
||||||
- `consistent-api.test.ts` - Update method signatures
|
|
||||||
- `unified-api.test.ts` - Update API calls
|
|
||||||
- `storage-adapter-coverage.test.ts` - Storage through augmentations
|
|
||||||
|
|
||||||
#### Actions:
|
|
||||||
- [ ] Update method calls for new API
|
|
||||||
- [ ] Fix initialization patterns
|
|
||||||
- [ ] Update configuration objects
|
|
||||||
|
|
||||||
### **Phase 4: Add Missing Tests**
|
|
||||||
#### New Tests Needed:
|
|
||||||
- [ ] **Augmentation lifecycle tests** - register → init → execute → shutdown
|
|
||||||
- [ ] **Augmentation priority tests** - Execution order validation
|
|
||||||
- [ ] **Two-phase initialization tests** - Storage first, then others
|
|
||||||
- [ ] **Method delegation tests** - Core methods → augmentations
|
|
||||||
|
|
||||||
### **Phase 5: Remove Obsolete Tests**
|
|
||||||
#### Tests to Remove/Update:
|
|
||||||
- [ ] Tests for removed augmentation factory functions
|
|
||||||
- [ ] Tests for deprecated API methods
|
|
||||||
- [ ] Tests for old typed augmentation system
|
|
||||||
|
|
||||||
## 🚨 HIGH-RISK AREAS
|
|
||||||
|
|
||||||
### **Most Likely to Fail:**
|
|
||||||
1. **core.test.ts** - Export mismatches
|
|
||||||
2. **statistics.test.ts** - Statistics through augmentations
|
|
||||||
3. **metadata-*.test.ts** - Metadata through augmentations
|
|
||||||
4. **auto-configuration.test.ts** - New initialization patterns
|
|
||||||
5. **storage-adapter-coverage.test.ts** - Storage delegation
|
|
||||||
|
|
||||||
### **Must Pass for Release:**
|
|
||||||
1. **release-critical.test.ts**
|
|
||||||
2. **release-validation.test.ts**
|
|
||||||
3. **All augmentation tests**
|
|
||||||
4. **core.test.ts**
|
|
||||||
5. **unified-api.test.ts**
|
|
||||||
|
|
||||||
## ✅ SUCCESS CRITERIA - COMPREHENSIVE 2.0 FEATURE VALIDATION
|
|
||||||
|
|
||||||
### **ALL Brainy Features Through Updated 2.0 APIs:**
|
|
||||||
|
|
||||||
#### **🧠 Core Features (Through New Implementations):**
|
|
||||||
- [ ] **Data Operations** - add/get/update/delete through AugmentationRegistry
|
|
||||||
- [ ] **Storage Systems** - All storage through StorageAugmentations (not direct)
|
|
||||||
- [ ] **Vector Search** - Updated search APIs and performance
|
|
||||||
- [ ] **NEW find()** - Triple Intelligence with `like`, `where`, `connected`
|
|
||||||
- [ ] **Clustering** - All 3 algorithms: HNSW, K-means, Hierarchical
|
|
||||||
- [ ] **Metadata Indexing** - Through IndexAugmentation (not direct)
|
|
||||||
- [ ] **Statistics** - Through MetricsAugmentation (not direct)
|
|
||||||
- [ ] **Caching** - Through CacheAugmentation (not SearchCache direct)
|
|
||||||
|
|
||||||
#### **🔌 Augmentation System (All 27 Augmentations):**
|
|
||||||
- [ ] **Storage (8)** - Memory, FileSystem, OPFS, S3, R2, GCS, Auto, Dynamic
|
|
||||||
- [ ] **Performance (7)** - Cache, Index, Metrics, WAL, Batch, Pool, Dedup
|
|
||||||
- [ ] **Data Integrity (3)** - EntityRegistry, AutoRegister, Enhanced Clear
|
|
||||||
- [ ] **Intelligence (2)** - Neural Import, Intelligent Verb Scoring
|
|
||||||
- [ ] **Communication (4)** - API Server, Conduits, Server Search, Monitoring
|
|
||||||
- [ ] **External Integration (3)** - Synapses, MCP, WebSocket
|
|
||||||
|
|
||||||
#### **🚀 New 2.0 Features:**
|
|
||||||
- [ ] **Unified BrainyAugmentation interface** - All augmentations working
|
|
||||||
- [ ] **Two-phase initialization** - Storage first, then others
|
|
||||||
- [ ] **Augmentation lifecycle** - register → init → execute → shutdown
|
|
||||||
- [ ] **Method delegation** - Core methods → augmentations
|
|
||||||
- [ ] **Auto-registration** - Default augmentations auto-registered
|
|
||||||
- [ ] **Enhanced Chat** - BrainyChat with all features
|
|
||||||
- [ ] **220 NLP Patterns** - All patterns through updated neural system
|
|
||||||
|
|
||||||
#### **🎯 API Consistency (Updated 2.0 APIs):**
|
|
||||||
- [ ] **All exports work** - No missing/incorrect exports in index.ts
|
|
||||||
- [ ] **Method signatures correct** - Updated parameter patterns
|
|
||||||
- [ ] **Configuration objects** - New augmentation-based config
|
|
||||||
- [ ] **Error handling** - Through augmentation system
|
|
||||||
- [ ] **Performance** - No regressions from augmentation overhead
|
|
||||||
|
|
||||||
### **Before 2.0 Release:**
|
|
||||||
- [ ] **100% test pass rate** across all 49 test files
|
|
||||||
- [ ] **Tests use CURRENT implementation** - Not old/deprecated APIs
|
|
||||||
- [ ] **Complete feature coverage** - ALL features tested through 2.0 APIs
|
|
||||||
- [ ] **No missing augmentation tests** - All 27 augmentations validated
|
|
||||||
- [ ] **Performance validation** - Augmentation system performs well
|
|
||||||
|
|
||||||
### **Quality Gates:**
|
|
||||||
1. **All export tests pass** - No missing/incorrect exports
|
|
||||||
2. **All augmentation tests pass** - 27 augmentations working
|
|
||||||
3. **All core functionality tests pass** - Basic features working
|
|
||||||
4. **All storage tests pass** - Storage through augmentations
|
|
||||||
5. **All API consistency tests pass** - Method signatures correct
|
|
||||||
|
|
||||||
## 🏃♂️ EXECUTION STRATEGY
|
|
||||||
|
|
||||||
### **Parallel Approach:**
|
|
||||||
1. **Fix TypeScript issues** (current priority)
|
|
||||||
2. **Run test suite and identify failures**
|
|
||||||
3. **Categorize failures by impact**
|
|
||||||
4. **Fix critical path tests first**
|
|
||||||
5. **Validate all tests in phases**
|
|
||||||
|
|
||||||
### **Test-Driven Validation:**
|
|
||||||
1. **Run each test file individually**
|
|
||||||
2. **Fix failures systematically**
|
|
||||||
3. **Update test plan based on findings**
|
|
||||||
4. **Validate full test suite**
|
|
||||||
5. **Performance regression testing**
|
|
||||||
|
|
@ -1,350 +0,0 @@
|
||||||
# 🧠 Unified Cache Architecture - Deep Analysis
|
|
||||||
|
|
||||||
## The Core Concept
|
|
||||||
ONE cache to rule them all - no coordination needed because there's nothing to coordinate!
|
|
||||||
|
|
||||||
## ✅ PROS - Why This is Brilliant
|
|
||||||
|
|
||||||
### 1. **Emergent Intelligence**
|
|
||||||
- System automatically finds optimal balance
|
|
||||||
- No human has to guess the right ratios
|
|
||||||
- Adapts to changing workloads in real-time
|
|
||||||
|
|
||||||
### 2. **Simplicity = Reliability**
|
|
||||||
```typescript
|
|
||||||
// Traditional approach: 500+ lines of coordination code
|
|
||||||
// Our approach: 50 lines that just work
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Cost-Aware by Design**
|
|
||||||
```typescript
|
|
||||||
interface CacheItem {
|
|
||||||
key: string
|
|
||||||
type: 'hnsw' | 'metadata'
|
|
||||||
data: any
|
|
||||||
size: number
|
|
||||||
rebuildCost: number // HNSW: 1000ms, Metadata: 1ms
|
|
||||||
lastAccess: number
|
|
||||||
accessCount: number
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. **Natural Load Balancing**
|
|
||||||
- Popular data stays in cache regardless of type
|
|
||||||
- Unpopular data gets evicted regardless of type
|
|
||||||
- The "right" balance emerges from usage patterns
|
|
||||||
|
|
||||||
## ⚠️ DANGERS - What Could Go Wrong
|
|
||||||
|
|
||||||
### 1. **Cache Stampede Risk**
|
|
||||||
```typescript
|
|
||||||
// DANGER: 1000 concurrent requests for same cold item
|
|
||||||
// All 1000 try to load from disk simultaneously!
|
|
||||||
|
|
||||||
// SOLUTION: Request coalescing
|
|
||||||
class UnifiedCache {
|
|
||||||
private loadingPromises = new Map<string, Promise<any>>()
|
|
||||||
|
|
||||||
async get(key: string) {
|
|
||||||
// If already loading, wait for existing promise
|
|
||||||
if (this.loadingPromises.has(key)) {
|
|
||||||
return this.loadingPromises.get(key)
|
|
||||||
}
|
|
||||||
|
|
||||||
if (!this.items.has(key)) {
|
|
||||||
const loadPromise = this.loadFromDisk(key)
|
|
||||||
this.loadingPromises.set(key, loadPromise)
|
|
||||||
const data = await loadPromise
|
|
||||||
this.loadingPromises.delete(key)
|
|
||||||
return data
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Memory Fragmentation**
|
|
||||||
```typescript
|
|
||||||
// DANGER: Many small metadata items + few large HNSW items
|
|
||||||
// Could lead to inefficient memory use
|
|
||||||
|
|
||||||
// SOLUTION: Size-aware eviction
|
|
||||||
evict(bytesNeeded: number) {
|
|
||||||
// Try to evict items that closely match needed size
|
|
||||||
// Prevents evicting 100 tiny items when 1 large would do
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Starvation Scenario**
|
|
||||||
```typescript
|
|
||||||
// DANGER: HNSW queries so expensive that metadata never gets cached
|
|
||||||
// Even though metadata queries are 100x more frequent
|
|
||||||
|
|
||||||
// SOLUTION: Fairness mechanism
|
|
||||||
class FairUnifiedCache {
|
|
||||||
private typeAccessCounts = { hnsw: 0, metadata: 0 }
|
|
||||||
|
|
||||||
evict() {
|
|
||||||
// If one type is getting 90%+ of accesses but has <10% of cache
|
|
||||||
// Force evict from the greedy type
|
|
||||||
const hnswRatio = this.getTypeRatio('hnsw')
|
|
||||||
const hnswAccessRatio = this.typeAccessCounts.hnsw / this.totalAccesses
|
|
||||||
|
|
||||||
if (hnswRatio > 0.9 && hnswAccessRatio < 0.1) {
|
|
||||||
// HNSW is hogging cache despite low usage
|
|
||||||
this.evictType('hnsw')
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. **Cold Start Problem**
|
|
||||||
```typescript
|
|
||||||
// DANGER: Empty cache = bad initial performance
|
|
||||||
// Don't know what to pre-load
|
|
||||||
|
|
||||||
// SOLUTION: Persistence + Smart Warming
|
|
||||||
class PersistentUnifiedCache {
|
|
||||||
async init() {
|
|
||||||
// Load access patterns from last session
|
|
||||||
const patterns = await this.loadAccessPatterns()
|
|
||||||
|
|
||||||
// Pre-warm top 10% most accessed items
|
|
||||||
for (const item of patterns.top10Percent) {
|
|
||||||
await this.preload(item.key)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
async shutdown() {
|
|
||||||
// Save access patterns for next startup
|
|
||||||
await this.saveAccessPatterns()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 ALTERNATIVE APPROACHES
|
|
||||||
|
|
||||||
### 1. **Two-Level Cache** (More Complex)
|
|
||||||
```typescript
|
|
||||||
class TwoLevelCache {
|
|
||||||
private l1Cache = new Map() // Ultra-hot, pinned
|
|
||||||
private l2Cache = new LRU() // Everything else
|
|
||||||
}
|
|
||||||
// Pro: Guarantees critical data stays
|
|
||||||
// Con: Need to decide what's "critical"
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Type-Segregated Pools** (Traditional)
|
|
||||||
```typescript
|
|
||||||
class SegregatedCache {
|
|
||||||
private hnswPool = new LRU(/* 60% memory */)
|
|
||||||
private metadataPool = new LRU(/* 40% memory */)
|
|
||||||
}
|
|
||||||
// Pro: Guaranteed resources for each type
|
|
||||||
// Con: Rigid, can't adapt to workload changes
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Time-Window Based** (Interesting!)
|
|
||||||
```typescript
|
|
||||||
class TimeWindowCache {
|
|
||||||
// Track access patterns in rolling windows
|
|
||||||
private windows = [
|
|
||||||
new AccessWindow('1min'),
|
|
||||||
new AccessWindow('5min'),
|
|
||||||
new AccessWindow('1hour')
|
|
||||||
]
|
|
||||||
|
|
||||||
evict() {
|
|
||||||
// Items not accessed in ANY window = cold
|
|
||||||
// Items accessed in ALL windows = hot
|
|
||||||
}
|
|
||||||
}
|
|
||||||
// Pro: Handles bursty workloads well
|
|
||||||
// Con: More complex, more memory overhead
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 ENHANCEMENTS TO CONSIDER
|
|
||||||
|
|
||||||
### 1. **Predictive Pre-fetching**
|
|
||||||
```typescript
|
|
||||||
class PredictiveCache extends UnifiedCache {
|
|
||||||
private sequences = new Map<string, string[]>()
|
|
||||||
|
|
||||||
async get(key: string) {
|
|
||||||
const data = await super.get(key)
|
|
||||||
|
|
||||||
// Track access sequences
|
|
||||||
this.recordSequence(this.lastKey, key)
|
|
||||||
|
|
||||||
// Predictively load likely next items
|
|
||||||
const predicted = this.predictNext(key)
|
|
||||||
if (predicted && !this.items.has(predicted)) {
|
|
||||||
this.preloadAsync(predicted) // Non-blocking
|
|
||||||
}
|
|
||||||
|
|
||||||
return data
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Adaptive Tier Boundaries**
|
|
||||||
```typescript
|
|
||||||
class AdaptiveTierCache {
|
|
||||||
private hotThreshold = 100 // Start with defaults
|
|
||||||
private warmThreshold = 10
|
|
||||||
|
|
||||||
adapt() {
|
|
||||||
// If cache is thrashing, tighten hot tier
|
|
||||||
if (this.evictionRate > 10_per_second) {
|
|
||||||
this.hotThreshold *= 1.5 // Make it harder to become hot
|
|
||||||
}
|
|
||||||
|
|
||||||
// If cache is stable, loosen hot tier
|
|
||||||
if (this.evictionRate < 1_per_minute) {
|
|
||||||
this.hotThreshold *= 0.9 // Make it easier to become hot
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Query-Aware Caching**
|
|
||||||
```typescript
|
|
||||||
class QueryAwareCache {
|
|
||||||
beforeQuery(query: TripleQuery) {
|
|
||||||
// Pre-emptively make room based on query type
|
|
||||||
if (query.like && query.where) {
|
|
||||||
// Hybrid query coming - ensure both types have space
|
|
||||||
this.ensureMinSpace('hnsw', 100_MB)
|
|
||||||
this.ensureMinSpace('metadata', 50_MB)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. **Compression for Cold Storage**
|
|
||||||
```typescript
|
|
||||||
class CompressedCache {
|
|
||||||
async saveToDisk(key: string, item: CacheItem) {
|
|
||||||
if (item.type === 'hnsw') {
|
|
||||||
// Quantize vectors before saving
|
|
||||||
item.data = this.quantizeVectors(item.data)
|
|
||||||
}
|
|
||||||
if (item.type === 'metadata') {
|
|
||||||
// Compress with zlib
|
|
||||||
item.data = await compress(item.data)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 PERFORMANCE CHARACTERISTICS
|
|
||||||
|
|
||||||
### Memory Efficiency
|
|
||||||
```
|
|
||||||
Traditional Dual-Cache: 60-70% efficiency (due to rigid splits)
|
|
||||||
Unified Cache: 85-95% efficiency (adapts to actual usage)
|
|
||||||
```
|
|
||||||
|
|
||||||
### Query Latency
|
|
||||||
```
|
|
||||||
Cache Hit: 0.1ms (both approaches)
|
|
||||||
Cache Miss (metadata): 5ms from disk
|
|
||||||
Cache Miss (HNSW): 100ms from disk (needs reconstruction)
|
|
||||||
```
|
|
||||||
|
|
||||||
### Adaptation Speed
|
|
||||||
```
|
|
||||||
Workload change detected: ~100 queries
|
|
||||||
Full rebalance: ~1000 queries
|
|
||||||
Steady state: ~10,000 queries
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 IMPLEMENTATION STRATEGY
|
|
||||||
|
|
||||||
### Phase 1: Basic Unified Cache (Week 1)
|
|
||||||
```typescript
|
|
||||||
class UnifiedCache {
|
|
||||||
private items = new Map<string, CacheItem>()
|
|
||||||
private totalSize = 0
|
|
||||||
private maxSize = 2 * GB
|
|
||||||
|
|
||||||
get(key: string): any
|
|
||||||
set(key: string, value: any, type: CacheType): void
|
|
||||||
evict(): void
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 2: Add Intelligence (Week 2)
|
|
||||||
- Access counting
|
|
||||||
- Cost-aware eviction
|
|
||||||
- Request coalescing
|
|
||||||
- Basic persistence
|
|
||||||
|
|
||||||
### Phase 3: Advanced Features (Week 3)
|
|
||||||
- Predictive prefetching
|
|
||||||
- Adaptive thresholds
|
|
||||||
- Compression
|
|
||||||
- Monitoring/metrics
|
|
||||||
|
|
||||||
## 🏆 WHY THIS WINS
|
|
||||||
|
|
||||||
1. **Simplicity**: One system instead of two
|
|
||||||
2. **Adaptability**: Responds to real usage, not predictions
|
|
||||||
3. **Efficiency**: No wasted memory on unused indices
|
|
||||||
4. **Maintainability**: 200 lines instead of 2000
|
|
||||||
5. **Performance**: Natural optimization emerges
|
|
||||||
|
|
||||||
## ⚡ QUICK WIN IMPLEMENTATION
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Start with this - 50 lines that solve 80% of the problem
|
|
||||||
class QuickUnifiedCache {
|
|
||||||
private cache = new Map()
|
|
||||||
private access = new Map()
|
|
||||||
private size = 0
|
|
||||||
private maxSize = 2_000_000_000 // 2GB
|
|
||||||
|
|
||||||
get(key: string) {
|
|
||||||
this.access.set(key, (this.access.get(key) || 0) + 1)
|
|
||||||
return this.cache.get(key)
|
|
||||||
}
|
|
||||||
|
|
||||||
set(key: string, value: any, size: number, cost: number) {
|
|
||||||
while (this.size + size > this.maxSize) {
|
|
||||||
this.evictLowestValue()
|
|
||||||
}
|
|
||||||
this.cache.set(key, { value, size, cost })
|
|
||||||
this.size += size
|
|
||||||
}
|
|
||||||
|
|
||||||
evictLowestValue() {
|
|
||||||
let victim = null
|
|
||||||
let lowestScore = Infinity
|
|
||||||
|
|
||||||
for (const [key, item] of this.cache) {
|
|
||||||
const score = (this.access.get(key) || 1) / item.cost
|
|
||||||
if (score < lowestScore) {
|
|
||||||
lowestScore = score
|
|
||||||
victim = key
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
if (victim) {
|
|
||||||
this.size -= this.cache.get(victim).size
|
|
||||||
this.cache.delete(victim)
|
|
||||||
this.access.delete(victim)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚨 FINAL VERDICT
|
|
||||||
|
|
||||||
**GO FOR IT!** This unified approach is:
|
|
||||||
- Simpler than coordination
|
|
||||||
- More adaptive than fixed splits
|
|
||||||
- Naturally self-optimizing
|
|
||||||
- Easy to enhance incrementally
|
|
||||||
|
|
||||||
The dangers are manageable with simple solutions, and the benefits far outweigh the complexity of traditional approaches.
|
|
||||||
|
|
||||||
**Start simple, measure everything, enhance based on real usage.**
|
|
||||||
|
|
@ -1,10 +1,9 @@
|
||||||
{
|
{
|
||||||
"_name_or_path": "sentence-transformers/all-MiniLM-L6-v2",
|
"_name_or_path": "nreimers/MiniLM-L6-H384-uncased",
|
||||||
"architectures": [
|
"architectures": [
|
||||||
"BertModel"
|
"BertModel"
|
||||||
],
|
],
|
||||||
"attention_probs_dropout_prob": 0.1,
|
"attention_probs_dropout_prob": 0.1,
|
||||||
"classifier_dropout": null,
|
|
||||||
"gradient_checkpointing": false,
|
"gradient_checkpointing": false,
|
||||||
"hidden_act": "gelu",
|
"hidden_act": "gelu",
|
||||||
"hidden_dropout_prob": 0.1,
|
"hidden_dropout_prob": 0.1,
|
||||||
|
|
@ -18,7 +17,7 @@
|
||||||
"num_hidden_layers": 6,
|
"num_hidden_layers": 6,
|
||||||
"pad_token_id": 0,
|
"pad_token_id": 0,
|
||||||
"position_embedding_type": "absolute",
|
"position_embedding_type": "absolute",
|
||||||
"transformers_version": "4.29.2",
|
"transformers_version": "4.8.2",
|
||||||
"type_vocab_size": 2,
|
"type_vocab_size": 2,
|
||||||
"use_cache": true,
|
"use_cache": true,
|
||||||
"vocab_size": 30522
|
"vocab_size": 30522
|
||||||
Binary file not shown.
1
assets/models/all-MiniLM-L6-v2/tokenizer.json
Normal file
1
assets/models/all-MiniLM-L6-v2/tokenizer.json
Normal file
File diff suppressed because one or more lines are too long
|
|
@ -7,7 +7,7 @@
|
||||||
*/
|
*/
|
||||||
|
|
||||||
import { program } from 'commander'
|
import { program } from 'commander'
|
||||||
import { BrainyData } from '../dist/brainyData.js'
|
import { Brainy } from '../dist/index.js'
|
||||||
import chalk from 'chalk'
|
import chalk from 'chalk'
|
||||||
import inquirer from 'inquirer'
|
import inquirer from 'inquirer'
|
||||||
import ora from 'ora'
|
import ora from 'ora'
|
||||||
|
|
@ -61,7 +61,7 @@ async function getBrainy() {
|
||||||
if (!brainyInstance) {
|
if (!brainyInstance) {
|
||||||
const spinner = ora('Initializing Brainy...').start()
|
const spinner = ora('Initializing Brainy...').start()
|
||||||
try {
|
try {
|
||||||
brainyInstance = new BrainyData()
|
brainyInstance = new Brainy()
|
||||||
await brainyInstance.init()
|
await brainyInstance.init()
|
||||||
spinner.succeed('Brainy initialized')
|
spinner.succeed('Brainy initialized')
|
||||||
} catch (error) {
|
} catch (error) {
|
||||||
|
|
@ -444,7 +444,7 @@ async function showStatistics(brain) {
|
||||||
const spinner = ora('Gathering statistics...').start()
|
const spinner = ora('Gathering statistics...').start()
|
||||||
|
|
||||||
try {
|
try {
|
||||||
const stats = await brain.getStatistics()
|
const stats = brain.getStats()
|
||||||
spinner.succeed('Statistics loaded')
|
spinner.succeed('Statistics loaded')
|
||||||
|
|
||||||
console.log(boxen(
|
console.log(boxen(
|
||||||
|
|
|
||||||
82
bin/brainy-minimal.js
Executable file
82
bin/brainy-minimal.js
Executable file
|
|
@ -0,0 +1,82 @@
|
||||||
|
#!/usr/bin/env node
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Brainy CLI - Minimal Version (Conversation Commands Only)
|
||||||
|
*
|
||||||
|
* This is a temporary minimal CLI that only includes working conversation commands
|
||||||
|
* Full CLI will be restored in version 3.20.0
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { Command } from 'commander'
|
||||||
|
import { readFileSync } from 'fs'
|
||||||
|
import { dirname, join } from 'path'
|
||||||
|
import { fileURLToPath } from 'url'
|
||||||
|
|
||||||
|
const __dirname = dirname(fileURLToPath(import.meta.url))
|
||||||
|
const packageJson = JSON.parse(readFileSync(join(__dirname, '..', 'package.json'), 'utf8'))
|
||||||
|
|
||||||
|
const program = new Command()
|
||||||
|
|
||||||
|
program
|
||||||
|
.name('brainy')
|
||||||
|
.description('🧠 Brainy - Infinite Agent Memory')
|
||||||
|
.version(packageJson.version)
|
||||||
|
|
||||||
|
// Dynamically load conversation command
|
||||||
|
const conversationCommand = await import('../dist/cli/commands/conversation.js').then(m => m.default)
|
||||||
|
|
||||||
|
program
|
||||||
|
.command('conversation')
|
||||||
|
.alias('conv')
|
||||||
|
.description('💬 Infinite agent memory and context management')
|
||||||
|
.addCommand(
|
||||||
|
new Command('setup')
|
||||||
|
.description('Set up MCP server for Claude Code integration')
|
||||||
|
.action(async () => {
|
||||||
|
await conversationCommand.handler({ action: 'setup', _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
.addCommand(
|
||||||
|
new Command('remove')
|
||||||
|
.description('Remove MCP server and clean up')
|
||||||
|
.action(async () => {
|
||||||
|
await conversationCommand.handler({ action: 'remove', _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
.addCommand(
|
||||||
|
new Command('search')
|
||||||
|
.description('Search messages across conversations')
|
||||||
|
.requiredOption('-q, --query <query>', 'Search query')
|
||||||
|
.option('-c, --conversation-id <id>', 'Filter by conversation')
|
||||||
|
.option('-r, --role <role>', 'Filter by role')
|
||||||
|
.option('-l, --limit <number>', 'Maximum results', '10')
|
||||||
|
.action(async (options) => {
|
||||||
|
await conversationCommand.handler({ action: 'search', ...options, _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
.addCommand(
|
||||||
|
new Command('context')
|
||||||
|
.description('Get relevant context for a query')
|
||||||
|
.requiredOption('-q, --query <query>', 'Context query')
|
||||||
|
.option('-l, --limit <number>', 'Maximum messages', '10')
|
||||||
|
.action(async (options) => {
|
||||||
|
await conversationCommand.handler({ action: 'context', ...options, _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
.addCommand(
|
||||||
|
new Command('thread')
|
||||||
|
.description('Get full conversation thread')
|
||||||
|
.requiredOption('-c, --conversation-id <id>', 'Conversation ID')
|
||||||
|
.action(async (options) => {
|
||||||
|
await conversationCommand.handler({ action: 'thread', ...options, _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
.addCommand(
|
||||||
|
new Command('stats')
|
||||||
|
.description('Show conversation statistics')
|
||||||
|
.action(async () => {
|
||||||
|
await conversationCommand.handler({ action: 'stats', _: [] })
|
||||||
|
})
|
||||||
|
)
|
||||||
|
|
||||||
|
program.parse(process.argv)
|
||||||
1856
bin/brainy.js
1856
bin/brainy.js
File diff suppressed because it is too large
Load diff
62
docker-compose.yml
Normal file
62
docker-compose.yml
Normal file
|
|
@ -0,0 +1,62 @@
|
||||||
|
version: '3.8'
|
||||||
|
|
||||||
|
services:
|
||||||
|
brainy:
|
||||||
|
build: .
|
||||||
|
container_name: brainy-app
|
||||||
|
ports:
|
||||||
|
- "3000:3000"
|
||||||
|
environment:
|
||||||
|
- NODE_ENV=production
|
||||||
|
- BRAINY_STORAGE_TYPE=filesystem
|
||||||
|
- BRAINY_STORAGE_PATH=/app/data
|
||||||
|
- BRAINY_LOG_LEVEL=info
|
||||||
|
- BRAINY_RATE_LIMIT_MAX=100
|
||||||
|
- BRAINY_RATE_LIMIT_WINDOW_MS=900000
|
||||||
|
volumes:
|
||||||
|
# Persistent storage for data
|
||||||
|
- brainy-data:/app/data
|
||||||
|
# Optional: Mount local models directory
|
||||||
|
# - ./models:/app/models:ro
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD", "curl", "-f", "http://localhost:3000/health"]
|
||||||
|
interval: 30s
|
||||||
|
timeout: 3s
|
||||||
|
retries: 3
|
||||||
|
start_period: 10s
|
||||||
|
restart: unless-stopped
|
||||||
|
networks:
|
||||||
|
- brainy-network
|
||||||
|
|
||||||
|
# Optional: MinIO for S3-compatible storage (development)
|
||||||
|
minio:
|
||||||
|
image: minio/minio:latest
|
||||||
|
container_name: brainy-minio
|
||||||
|
ports:
|
||||||
|
- "9000:9000"
|
||||||
|
- "9001:9001"
|
||||||
|
environment:
|
||||||
|
- MINIO_ROOT_USER=brainy
|
||||||
|
- MINIO_ROOT_PASSWORD=brainy123456
|
||||||
|
volumes:
|
||||||
|
- minio-data:/data
|
||||||
|
command: server /data --console-address ":9001"
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
|
||||||
|
interval: 30s
|
||||||
|
timeout: 20s
|
||||||
|
retries: 3
|
||||||
|
networks:
|
||||||
|
- brainy-network
|
||||||
|
profiles:
|
||||||
|
- with-s3
|
||||||
|
|
||||||
|
volumes:
|
||||||
|
brainy-data:
|
||||||
|
driver: local
|
||||||
|
minio-data:
|
||||||
|
driver: local
|
||||||
|
|
||||||
|
networks:
|
||||||
|
brainy-network:
|
||||||
|
driver: bridge
|
||||||
295
docs/ADR-001-generational-mvcc.md
Normal file
295
docs/ADR-001-generational-mvcc.md
Normal file
|
|
@ -0,0 +1,295 @@
|
||||||
|
# ADR-001: Generational MVCC storage and the immutable Db API
|
||||||
|
|
||||||
|
**Status:** Accepted (ships in 8.0)
|
||||||
|
**Date:** 2026-06-10
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Before 8.0, Brainy carried two overlapping version-control subsystems: a
|
||||||
|
copy-on-write branching layer (`fork`/`checkout`/`commit`/`branches`) and a
|
||||||
|
separate versioning subsystem (`versions.save/list/compare/restore/prune`),
|
||||||
|
plus a read-only historical adapter for commit-based time travel. Together
|
||||||
|
they were ~5,100 LOC of mechanism for one product need: *read a consistent
|
||||||
|
past state while the store keeps moving, and snapshot/restore cheaply.*
|
||||||
|
|
||||||
|
Neither subsystem gave a precise isolation guarantee. Reads raced in-place
|
||||||
|
JSON overwrites, so a "snapshot" was only as immutable as the bytes it
|
||||||
|
happened to share with the live store.
|
||||||
|
|
||||||
|
8.0 replaces both with **one mechanism**: generational MVCC over immutable,
|
||||||
|
generation-stamped records, exposed through a Datomic-style immutable
|
||||||
|
database value (`Db`). The same model is implemented natively by versioned
|
||||||
|
index providers (LSM snapshots), so semantics are identical on the pure-JS
|
||||||
|
path and the native path.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
### The model
|
||||||
|
|
||||||
|
- A **monotonic u64 generation counter** is the store's logical clock. It
|
||||||
|
advances once per committed `transact()` batch and once per
|
||||||
|
single-operation write (`add`/`update`/`remove`/`relate`/…), so
|
||||||
|
`brain.generation()` is always a meaningful watermark. It is persisted in
|
||||||
|
`_system/generation.json` and never reissued for anything durable.
|
||||||
|
- `brain.now()` **pins** the current generation in O(1) and returns a `Db` —
|
||||||
|
an immutable view. Pins are refcounted; `db.release()` (with a
|
||||||
|
`FinalizationRegistry` backstop for leaked values) ends the pin.
|
||||||
|
- `brain.transact(ops, { meta, ifAtGeneration })` commits a declarative
|
||||||
|
batch atomically as **exactly one generation**, with whole-store
|
||||||
|
compare-and-swap (`ifAtGeneration` → `GenerationConflictError`) and
|
||||||
|
reified transaction metadata appended to `_system/tx-log.jsonl`.
|
||||||
|
- `brain.asOf(generation | Date | snapshotPath)` opens past state;
|
||||||
|
`db.with(ops)` layers a speculative in-memory overlay (never touching
|
||||||
|
disk, the counter, or index providers); `db.persist(path)` cuts an
|
||||||
|
instant snapshot; `brain.restore(path, { confirm: true })` replaces state
|
||||||
|
from one; `Brainy.load(path)` opens a snapshot read-only with the full
|
||||||
|
query surface.
|
||||||
|
|
||||||
|
### Persisted layout
|
||||||
|
|
||||||
|
All paths are storage-root-relative:
|
||||||
|
|
||||||
|
```
|
||||||
|
_system/generation.json { generation, updatedAt } atomic tmp+rename
|
||||||
|
_system/manifest.json { version, generation, atomic tmp+rename
|
||||||
|
committedAt, horizon } (the commit point)
|
||||||
|
_system/tx-log.jsonl one line per committed append-only
|
||||||
|
transact: { generation,
|
||||||
|
timestamp, meta? }
|
||||||
|
_generations/<N>/tx.json the generation-N delta: immutable
|
||||||
|
touched noun/verb ids + meta
|
||||||
|
_generations/<N>/prev/<id>.json before-image of <id> as of immutable
|
||||||
|
commit N (raw stored bytes;
|
||||||
|
null parts = file was absent)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why per-generation deltas instead of a global `id → latest generation`
|
||||||
|
map in the manifest:** a global map makes every commit O(all ids) — the
|
||||||
|
whole map must be rewritten to swap it atomically. The delta layout makes a
|
||||||
|
commit O(ids touched) and keeps the manifest a fixed-size watermark, while
|
||||||
|
point-in-time resolution stays correct (see "Read resolution" below). The
|
||||||
|
trade is that resolution at a pinned generation scans the deltas of later
|
||||||
|
commits — bounded by the number of commits since the pin, which is exactly
|
||||||
|
the window compaction keeps short.
|
||||||
|
|
||||||
|
Before-images are deliberately the *only* per-id records. They serve both
|
||||||
|
roles the layer needs — the crash-recovery undo log and the point-in-time
|
||||||
|
read source. After-images would duplicate state that is already readable
|
||||||
|
(the canonical entity files hold the latest bytes; earlier states resolve
|
||||||
|
from later before-images) and would double record I/O per commit.
|
||||||
|
|
||||||
|
### Commit protocol (durability)
|
||||||
|
|
||||||
|
`transact()` commits under a store-wide mutex:
|
||||||
|
|
||||||
|
1. **CAS check.** A stale `ifAtGeneration` throws `GenerationConflictError`
|
||||||
|
before anything is staged.
|
||||||
|
2. **Reserve** generation `N` (counter increment).
|
||||||
|
3. **Stage the undo log:** write the before-image of every touched id plus
|
||||||
|
`tx.json` under `_generations/N/`, then **fsync** the files and their
|
||||||
|
directories. From this point, any crash is recoverable to the exact
|
||||||
|
pre-transaction bytes.
|
||||||
|
4. **Execute** the planned batch through the TransactionManager (which has
|
||||||
|
its own operation-level rollback for in-flight failures).
|
||||||
|
5. **Commit point:** persist the counter, then write `_system/manifest.json`
|
||||||
|
via atomic tmp+rename and fsync it. The rename *is* the commit: a
|
||||||
|
generation directory is committed if and only if `N ≤
|
||||||
|
manifest.generation`.
|
||||||
|
6. Append the tx-log line (advisory metadata — a crash between 5 and 6
|
||||||
|
keeps the transaction).
|
||||||
|
|
||||||
|
**Crash recovery (on open):** any `_generations/<N>` directory with
|
||||||
|
`N > manifest.generation` is an uncommitted transaction. Its before-images
|
||||||
|
are restored to the canonical paths (idempotently — recovery itself can
|
||||||
|
crash and rerun) and the directory is removed. Because recovery runs before
|
||||||
|
any index is built, and a recovery that rolled something back forces a full
|
||||||
|
index rebuild, derived indexes never observe rolled-back state. Reader-mode
|
||||||
|
instances skip recovery (readers never write; the next writer repairs).
|
||||||
|
|
||||||
|
A failed (non-crash) transaction takes the same staging directory down the
|
||||||
|
abort path: the TransactionManager rolls back applied operations, the
|
||||||
|
staging directory is removed, and the generation reservation is returned —
|
||||||
|
a failed batch leaves the generation counter unchanged.
|
||||||
|
|
||||||
|
### Read resolution at a pinned generation
|
||||||
|
|
||||||
|
The state of id X at pinned generation G is:
|
||||||
|
|
||||||
|
- the before-image stored by the **first committed generation after G that
|
||||||
|
touched X**, or
|
||||||
|
- the live canonical bytes, when nothing after G touched X.
|
||||||
|
|
||||||
|
While nothing has committed past G, *every* read on the `Db` delegates to
|
||||||
|
the live fast paths untouched — `now()` adds no read overhead until history
|
||||||
|
actually moves.
|
||||||
|
|
||||||
|
**Two read paths, one result set.** `get()`, metadata-level `find()`, and
|
||||||
|
filter-based `related()` resolve directly through the record layer at any
|
||||||
|
reachable pinned generation — no extra cost beyond scanning the deltas of
|
||||||
|
later commits. Index-accelerated dimensions (semantic/vector search, graph
|
||||||
|
traversal, cursors, aggregation) are served by **at-generation index
|
||||||
|
materialization**: the first such query on a historical `Db` copies the
|
||||||
|
exact at-G record set (live bytes for ids untouched since the pin,
|
||||||
|
before-images for the rest; a final reconciliation pass runs under the
|
||||||
|
commit mutex so transactions racing the copy cannot skew it) into an
|
||||||
|
ephemeral in-memory store and opens a read-only engine over it — the same
|
||||||
|
vector/metadata/graph index classes the live brain uses, sharing the host's
|
||||||
|
embedder and aggregate definitions. The handle is cached on the `Db` and
|
||||||
|
freed by `release()`.
|
||||||
|
|
||||||
|
**Cost, stated plainly:** materialization is O(n at G) time and memory,
|
||||||
|
once per `Db`. That is the open-core price of historical index queries. A
|
||||||
|
native `VersionedIndexProvider` (`isGenerationVisible()` + pins over
|
||||||
|
retained LSM segments) serves the same reads with no rebuild at all — the
|
||||||
|
materializer is the correctness baseline, the provider is the accelerator.
|
||||||
|
|
||||||
|
**The one remaining boundary.** Speculative `with()` overlays throw
|
||||||
|
`SpeculativeOverlayError` for index-accelerated queries and `persist()`:
|
||||||
|
overlay entities carry no embeddings (`with()` never invokes the embedder),
|
||||||
|
so a "full" index query over an overlay would silently exclude the
|
||||||
|
overlay's own entities. Commit with `transact()` to get the full surface.
|
||||||
|
|
||||||
|
**History granularity (Model-B).** EVERY write is its own immutable generation
|
||||||
|
— `transact()` batches AND single-operation `add`/`update`/`remove`/`relate`.
|
||||||
|
Single-ops stage a before-image and are reported by `db.since()`/`asOf()`/
|
||||||
|
`diff()`/`history()` exactly like transacts; a pin always freezes against later
|
||||||
|
writes. `transact()` groups several operations into ONE atomic generation.
|
||||||
|
|
||||||
|
Single-op history durability is **async group-commit**: the live write hits
|
||||||
|
canonical storage immediately (acknowledged), while its before-image is buffered
|
||||||
|
and persisted to disk in one batched fsync on a size/timer trigger (or forced by
|
||||||
|
`flush()`/`close()`/`transact()`/`compactHistory()`). The buffer participates in
|
||||||
|
point-in-time resolution exactly like on-disk generations, so the synchronous
|
||||||
|
`now()` freezes with no forced flush. A hard crash before the flush loses only
|
||||||
|
the buffered *history* of the last window — never live data — and a crash
|
||||||
|
*mid-flush* is recovered by **drop-without-restore** (the partial generation's
|
||||||
|
before-images are discarded, never replayed, because the live write was already
|
||||||
|
acknowledged; restoring them would silently revert it).
|
||||||
|
|
||||||
|
### Pinning, retention, compaction
|
||||||
|
|
||||||
|
- Each live `Db` holds one refcounted pin on its generation (plus a
|
||||||
|
`pin(generation)` on every registered `VersionedIndexProvider`, whose
|
||||||
|
explicit pin lifetime overrides any time-based snapshot retention the
|
||||||
|
provider has).
|
||||||
|
- The constructor **`retention`** knob governs auto-compaction (on `flush()`/
|
||||||
|
`close()`): unset → ADAPTIVE (disk/RAM-pressure byte budget, zero-config;
|
||||||
|
driven by a coordinator's `budgetBytes` or a local `os.freemem` probe) ·
|
||||||
|
`'all'` → unbounded · `{ maxGenerations?, maxAge?, maxBytes? }` → explicit
|
||||||
|
CAPS. `compactHistory({ maxGenerations?, maxAge?, maxBytes? })` reclaims
|
||||||
|
manually on the same caps — the oldest unpinned record-sets are reclaimed
|
||||||
|
while ANY supplied cap is exceeded.
|
||||||
|
- A record-set `N` is reclaimed only when `N` is at or below **every** live
|
||||||
|
pin — deleting `N` can only break readers pinned *below* `N`, because
|
||||||
|
resolution reads before-images from generations strictly greater than the
|
||||||
|
pin. Live pins are ALWAYS exempt, in every retention mode.
|
||||||
|
- The manifest records the **horizon** (highest reclaimed generation).
|
||||||
|
Generations below the horizon are unreachable; `asOf()` on them throws
|
||||||
|
`GenerationCompactedError`. The horizon itself stays reachable, resolved
|
||||||
|
from the record-sets above it. To keep a generation readable forever,
|
||||||
|
`persist()` it first — snapshots are self-contained.
|
||||||
|
|
||||||
|
### Snapshots and restore
|
||||||
|
|
||||||
|
`db.persist(path)` flushes indexes, then cuts the snapshot under the
|
||||||
|
store's commit mutex (no commit, compaction, or counter write can
|
||||||
|
interleave). On filesystem storage it is a **hard-link farm**: every data
|
||||||
|
file is immutable-by-rename, so linking is safe — later rewrites swap
|
||||||
|
inodes and the snapshot keeps the old bytes. The two exceptions are handled
|
||||||
|
explicitly: the append-in-place tx-log is byte-copied, and process-local
|
||||||
|
lock state is excluded. Cross-device targets (and filesystems that refuse
|
||||||
|
links) fall back to per-file byte copies. In-memory stores serialize to the
|
||||||
|
same directory layout, so persisting a memory brain produces a real,
|
||||||
|
durable, loadable store.
|
||||||
|
|
||||||
|
`persist()` requires the view to still be the store's latest generation
|
||||||
|
(a snapshot captures current bytes); a view that history has moved past
|
||||||
|
throws rather than persisting the wrong state.
|
||||||
|
|
||||||
|
`restore(path, { confirm: true })` replaces the store's contents from a
|
||||||
|
snapshot via byte copy (never links — the snapshot stays independent),
|
||||||
|
reloads all adapter-internal derived state, rebuilds all indexes, and
|
||||||
|
floors the generation counter at its pre-restore value so observed
|
||||||
|
generation numbers are never reissued. Live pins do not survive a restore;
|
||||||
|
a warning is logged when any exist.
|
||||||
|
|
||||||
|
### Versioned index providers
|
||||||
|
|
||||||
|
Native index providers may implement the optional 4-method
|
||||||
|
`VersionedIndexProvider` capability (`generation()`,
|
||||||
|
`isGenerationVisible()`, `pin()`, `release()` — BigInt generations at the
|
||||||
|
boundary). The locked consistency model: providers are **post-commit
|
||||||
|
appliers**. The storage-record commit is the source of truth; provider
|
||||||
|
index state is derived. On open, a provider behind the committed watermark
|
||||||
|
replays the gap from storage (or requests a rebuild) — there are no
|
||||||
|
provider rollback hooks, because uncommitted transactions are repaired at
|
||||||
|
the storage layer before any index opens. Speculative `with()` overlays
|
||||||
|
never reach providers.
|
||||||
|
|
||||||
|
## Guarantees (and their proofs)
|
||||||
|
|
||||||
|
Each stated guarantee has a test that proves it, not merely exercises it
|
||||||
|
(`tests/integration/db-mvcc.test.ts`, plus
|
||||||
|
`tests/unit/db/generationStore.test.ts` for the record layer in isolation):
|
||||||
|
|
||||||
|
| Guarantee | Proof |
|
||||||
|
|---|---|
|
||||||
|
| Snapshot isolation: a pinned `Db` reads exactly its pinned state, forever | proof 1 (200 mutations, including deletes, against a pinned view) |
|
||||||
|
| Atomicity: a failing batch applies nothing; generation unchanged | proofs 2a/2b/2c (plan-time failure, injected execution-phase storage failure, `ifRev` conflict) |
|
||||||
|
| Whole-store CAS | proof 3 (`ifAtGeneration` success + conflict with exact expected/actual) |
|
||||||
|
| Snapshot integrity under source mutation (hard-link safety) | proofs 4a/4b/4c |
|
||||||
|
| Compaction never breaks a pinned read; release enables reclaim | proof 5 |
|
||||||
|
| `with()` overlays touch nothing durable | proof 6 |
|
||||||
|
| Generation monotonicity across close/reopen | proof 7 |
|
||||||
|
| Crash before the manifest rename recovers to exact pre-transaction state through the real recovery path | proof 8 (fault injection that skips abort cleanup, exactly as a dead process would) |
|
||||||
|
| Balanced provider pin/release lockstep | proof 9 |
|
||||||
|
|
||||||
|
One deliberate softness: single-operation generation bumps persist the
|
||||||
|
counter coalesced (per write burst), not per write. Durable artifacts —
|
||||||
|
records, manifests, snapshots — always persist the counter synchronously at
|
||||||
|
their own commit points, so a crash inside the coalescing window can lose
|
||||||
|
only counter values that nothing durable ever referenced.
|
||||||
|
|
||||||
|
## Failure modes
|
||||||
|
|
||||||
|
| Failure | Outcome |
|
||||||
|
|---|---|
|
||||||
|
| Crash before staging completes | Partial staging directory > manifest watermark → removed on next open; canonical state untouched. |
|
||||||
|
| Crash after staging, before/during batch execution | Before-images restored on next open; indexes rebuilt; byte-identical pre-transaction state. |
|
||||||
|
| Crash after execution, before manifest rename | Same as above — the rename is the only commit point. |
|
||||||
|
| Crash after manifest rename, before tx-log append | Transaction kept (committed); tx-log misses one advisory line; `asOf(Date)` resolution for that commit falls back to neighboring entries. |
|
||||||
|
| Batch fails mid-execution (no crash) | TransactionManager operation rollback + staging-directory removal + reservation return; generation unchanged. |
|
||||||
|
| `asOf()` below the compaction horizon | `GenerationCompactedError` — explicit, never partial data. |
|
||||||
|
| Index-accelerated query on a `with()` overlay | `SpeculativeOverlayError` — explicit, never silently-incomplete results (overlay entities carry no embeddings). |
|
||||||
|
| `persist()` of a view history has moved past | `GenerationConflictError` — a snapshot captures current bytes; persist before further writes. |
|
||||||
|
| Torn trailing tx-log line (crashed append) | Tolerated; unparseable lines are skipped by readers. |
|
||||||
|
|
||||||
|
## Lineage
|
||||||
|
|
||||||
|
The design is an assembly of well-understood prior art, chosen for being
|
||||||
|
boring where it counts:
|
||||||
|
|
||||||
|
- **Datomic** — the database-as-a-value: an immutable `Db` you query, with
|
||||||
|
`with()` for speculation and reified transaction metadata instead of
|
||||||
|
commit messages.
|
||||||
|
- **LMDB** — reader pins: readers never block writers; a reader's view
|
||||||
|
stays valid because nothing overwrites the pages (here: records) it
|
||||||
|
references; reclamation waits for the last reader.
|
||||||
|
- **LSM trees / Cassandra** — immutable segments make snapshots hard links
|
||||||
|
and make compaction a retention policy instead of a locking problem.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- One mechanism replaces the COW and versioning subsystems (their removal
|
||||||
|
is the companion change to this ADR).
|
||||||
|
- In-place branch switching (`checkout`) is gone by design; the replacement
|
||||||
|
is opening a persisted snapshot as a separate instance — a name→path
|
||||||
|
mapping where a product needs named branches.
|
||||||
|
- Every commit pays O(ids touched) extra writes (before-images + delta +
|
||||||
|
manifest). Single-operation writes pay only an in-memory counter bump
|
||||||
|
with coalesced persistence.
|
||||||
|
- The full query surface works at every reachable pinned generation.
|
||||||
|
Record-path reads (`get`, metadata `find`, filter `related`) are
|
||||||
|
effectively free; index-accelerated historical queries pay a one-time
|
||||||
|
O(n at G) materialization per `Db` on the open-core path (freed on
|
||||||
|
`release()`), and run rebuild-free on a native `VersionedIndexProvider`.
|
||||||
468
docs/BATCHING.md
Normal file
468
docs/BATCHING.md
Normal file
|
|
@ -0,0 +1,468 @@
|
||||||
|
---
|
||||||
|
title: Batch Operations
|
||||||
|
slug: guides/batching
|
||||||
|
public: true
|
||||||
|
category: guides
|
||||||
|
template: guide
|
||||||
|
order: 5
|
||||||
|
description: Eliminate N+1 query patterns with batchGet() and storage-level batch APIs for fast multi-entity reads against filesystem and memory storage.
|
||||||
|
next:
|
||||||
|
- api/reference
|
||||||
|
- guides/find-system
|
||||||
|
---
|
||||||
|
|
||||||
|
# Batch Operations API
|
||||||
|
> **Production-Ready** | Zero N+1 Query Patterns
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Brainy provides batch operations at the storage layer to eliminate N+1 query patterns for VFS operations, relationship queries, and entity retrieval.
|
||||||
|
|
||||||
|
### Problem Solved
|
||||||
|
|
||||||
|
The naive pattern of looping and calling `brain.get(id)` once per item issues sequential reads through the storage layer. Batched APIs collapse that into a single read pass.
|
||||||
|
|
||||||
|
**IMPORTANT:** The batch optimizations apply **ONLY to `getTreeStructure()`** at the VFS layer and the explicit `batchGet()` / `getNounMetadataBatch()` calls — not to `readFile()` or individual `get()` operations.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## New Public APIs
|
||||||
|
|
||||||
|
### 1. `brain.batchGet(ids, options?)`
|
||||||
|
|
||||||
|
Batch retrieval of multiple entities (metadata-only by default).
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Fetch multiple entities in a single batched operation
|
||||||
|
const ids = ['id1', 'id2', 'id3']
|
||||||
|
const results: Map<string, Entity> = await brain.batchGet(ids)
|
||||||
|
|
||||||
|
// With vectors (falls back to individual gets)
|
||||||
|
const resultsWithVectors = await brain.batchGet(ids, { includeVectors: true })
|
||||||
|
|
||||||
|
// Results map
|
||||||
|
results.get('id1') // → Entity or undefined
|
||||||
|
results.size // → 3 (number of found entities)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- Memory storage: Instant (parallel reads)
|
||||||
|
- Filesystem storage: Parallel reads, scales with available IOPS
|
||||||
|
|
||||||
|
**Use Cases:**
|
||||||
|
- Loading multiple entities for display
|
||||||
|
- Bulk data export operations
|
||||||
|
- Relationship traversal (fetch all connected entities)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Storage-Level APIs
|
||||||
|
|
||||||
|
### 2. `storage.getNounMetadataBatch(ids)`
|
||||||
|
|
||||||
|
Batch metadata retrieval with direct O(1) path construction.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const storage = brain.storage as BaseStorage
|
||||||
|
const ids = ['id1', 'id2', 'id3']
|
||||||
|
|
||||||
|
const metadataMap: Map<string, NounMetadata> = await storage.getNounMetadataBatch(ids)
|
||||||
|
|
||||||
|
for (const [id, metadata] of metadataMap) {
|
||||||
|
console.log(metadata.noun) // Type: 'document', 'person', etc.
|
||||||
|
console.log(metadata.data) // Entity data
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Features:**
|
||||||
|
- ✅ Direct O(1) path construction from ID (no type lookup needed!)
|
||||||
|
- ✅ Sharding preservation (all paths include `{shard}/{id}`)
|
||||||
|
- ✅ Write-cache coherent (read-after-write consistency)
|
||||||
|
- ✅ O(1) path construction — eliminates the per-entity type search of the old type-first layout
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- Constant-time path construction per ID — no type-cache misses
|
||||||
|
- Filesystem: parallel reads bounded by IOPS
|
||||||
|
- No type search delays — every ID maps directly to storage path
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3. `storage.getVerbsBySourceBatch(sourceIds, verbType?)`
|
||||||
|
|
||||||
|
Batch relationship queries by source entity IDs.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const storage = brain.storage as BaseStorage
|
||||||
|
|
||||||
|
// Get all relationships from multiple sources
|
||||||
|
const results: Map<string, GraphVerb[]> = await storage.getVerbsBySourceBatch([
|
||||||
|
'person1',
|
||||||
|
'person2'
|
||||||
|
])
|
||||||
|
|
||||||
|
// Filter by verb type
|
||||||
|
const createsResults = await storage.getVerbsBySourceBatch(
|
||||||
|
['person1', 'person2'],
|
||||||
|
'creates'
|
||||||
|
)
|
||||||
|
|
||||||
|
// Process results
|
||||||
|
for (const [sourceId, verbs] of results) {
|
||||||
|
console.log(`${sourceId} has ${verbs.length} relationships`)
|
||||||
|
verbs.forEach(verb => {
|
||||||
|
console.log(` → ${verb.verb} → ${verb.targetId}`)
|
||||||
|
})
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use Cases:**
|
||||||
|
- Social graph traversal (fetch all connections for multiple users)
|
||||||
|
- Knowledge graph queries (find all relationships of specific type)
|
||||||
|
- Bulk export of relationship data
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- Memory storage: single in-memory pass over the metadata index
|
||||||
|
- Filesystem storage: parallel reads through the metadata index
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## VFS Integration
|
||||||
|
|
||||||
|
VFS operations automatically use batch APIs for maximum performance.
|
||||||
|
|
||||||
|
### Directory Traversal
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Tree traversal uses batched reads under the hood
|
||||||
|
const tree = await brain.vfs.getTreeStructure('/my-dir')
|
||||||
|
// ✅ PathResolver.getChildren() uses brain.batchGet() internally
|
||||||
|
// ✅ Parallel traversal of directories at the same tree level
|
||||||
|
// ✅ 2-3 batched calls instead of 22 sequential calls
|
||||||
|
```
|
||||||
|
|
||||||
|
**Architecture:**
|
||||||
|
|
||||||
|
```
|
||||||
|
VFS.getTreeStructure()
|
||||||
|
↓ PARALLEL (breadth-first traversal)
|
||||||
|
→ PathResolver.getChildren() [all dirs at level processed in parallel]
|
||||||
|
↓ BATCHED
|
||||||
|
→ brain.batchGet(childIds) [1 call instead of N]
|
||||||
|
↓ BATCHED
|
||||||
|
→ storage.getNounMetadataBatch(ids) [1 call instead of N]
|
||||||
|
↓ ADAPTER
|
||||||
|
→ Filesystem: Promise.all() parallel reads
|
||||||
|
→ Memory: Promise.all() parallel reads
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Advanced Features Compatibility
|
||||||
|
|
||||||
|
### ✅ ID-First Storage Architecture
|
||||||
|
|
||||||
|
All batch operations use direct ID-first paths - no type lookup needed!
|
||||||
|
|
||||||
|
**ID-First Path Structure:**
|
||||||
|
```
|
||||||
|
entities/nouns/{SHARD}/{ID}/metadata.json
|
||||||
|
entities/verbs/{SHARD}/{ID}/metadata.json
|
||||||
|
```
|
||||||
|
|
||||||
|
**Direct O(1) Path Construction:**
|
||||||
|
```typescript
|
||||||
|
// Every ID maps directly to exactly ONE path - O(1), no type search
|
||||||
|
const id = 'abc-123'
|
||||||
|
const shard = getShardIdFromUuid(id) // → 'ab' (first 2 hex chars)
|
||||||
|
const path = `entities/nouns/${shard}/${id}/metadata.json`
|
||||||
|
|
||||||
|
// No type cache needed!
|
||||||
|
// No type search needed!
|
||||||
|
// No multi-type fallback needed!
|
||||||
|
// Just pure O(1) lookup!
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- **O(1)** path lookups (eliminates the 42-type sequential search the old type-first layout required)
|
||||||
|
- **Simpler code** - removed 500+ lines of type cache complexity
|
||||||
|
- **Scalable** - works at large scale without type tracking overhead
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### ✅ Sharding
|
||||||
|
|
||||||
|
All batch paths include shard IDs calculated via `getShardIdFromUuid(id)`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const id = 'a3c4e5f7-...'
|
||||||
|
const shard = getShardIdFromUuid(id) // → 'a3' (first 2 hex chars)
|
||||||
|
const path = `entities/nouns/${shard}/${id}/metadata.json`
|
||||||
|
```
|
||||||
|
|
||||||
|
**Distribution:** 256 shards (00-ff) for optimal load distribution.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### ✅ Generational MVCC (8.0)
|
||||||
|
|
||||||
|
Batch reads always serve the **live** generation through the fast paths
|
||||||
|
shown above. Point-in-time reads go through the Db API instead: a pinned
|
||||||
|
`Db` (`brain.now()`, `brain.asOf()`) resolves changed ids from immutable
|
||||||
|
generation records and unchanged ids from the same live paths batch reads
|
||||||
|
use — see the [consistency model](concepts/consistency-model.md).
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const db = brain.now() // pinned view
|
||||||
|
const entity = await db.get(id) // correct at the pinned generation
|
||||||
|
const results = await brain.batchGet(ids) // live state, batched
|
||||||
|
await db.release()
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Why Batching Is Faster
|
||||||
|
|
||||||
|
Batching's advantage is structural, not a fixed multiplier (the actual speedup
|
||||||
|
depends on storage backend, IOPS, and batch size):
|
||||||
|
|
||||||
|
- **N+1 elimination** — N sequential reads collapse into a single parallel pass
|
||||||
|
(`Promise.all` over the batch).
|
||||||
|
- **O(1) path construction** — every ID maps directly to one storage path, with
|
||||||
|
no per-type cache lookup.
|
||||||
|
- **One metadata round-trip** — relationship batches fetch all sources' metadata
|
||||||
|
in a single pass instead of one query per source.
|
||||||
|
|
||||||
|
The integration test `tests/integration/storage-batch-operations.test.ts`
|
||||||
|
exercises batch vs. individual reads and asserts that batch retrieval is not
|
||||||
|
slower than the per-entity loop for large batches; it does not pin a specific
|
||||||
|
multiplier, since that is hardware- and IOPS-dependent.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Error Handling
|
||||||
|
|
||||||
|
### Partial Batch Failures
|
||||||
|
|
||||||
|
Batch operations gracefully handle missing or invalid entities:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const validId = 'abc-123-...'
|
||||||
|
const invalidIds = [
|
||||||
|
'11111111-1111-1111-1111-111111111111',
|
||||||
|
'22222222-2222-2222-2222-222222222222'
|
||||||
|
]
|
||||||
|
|
||||||
|
const results = await brain.batchGet([validId, ...invalidIds])
|
||||||
|
|
||||||
|
results.size // → 1 (only valid entity)
|
||||||
|
results.has(validId) // → true
|
||||||
|
results.has(invalidIds[0]) // → false (silently skipped)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Behavior:**
|
||||||
|
- Invalid UUIDs: Silently skipped (not included in results)
|
||||||
|
- Missing entities: Silently skipped (not included in results)
|
||||||
|
- Storage errors: Logged, entity excluded from results
|
||||||
|
- No exceptions thrown for partial failures
|
||||||
|
|
||||||
|
### Empty Batches
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const results = await brain.batchGet([])
|
||||||
|
results.size // → 0 (empty map)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Duplicate IDs
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const results = await brain.batchGet(['id1', 'id1', 'id1'])
|
||||||
|
results.size // → 1 (deduplicated automatically)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Migration Guide
|
||||||
|
|
||||||
|
### From Individual Gets
|
||||||
|
|
||||||
|
**Before:**
|
||||||
|
```typescript
|
||||||
|
const entities = []
|
||||||
|
for (const id of ids) {
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
if (entity) entities.push(entity)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**After:**
|
||||||
|
```typescript
|
||||||
|
const results = await brain.batchGet(ids)
|
||||||
|
const entities = Array.from(results.values())
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance Gain:** Replaces N sequential reads with a single batched pass — no fixed multiplier, it scales with storage IOPS.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### From Individual Relationship Queries
|
||||||
|
|
||||||
|
**Before:**
|
||||||
|
```typescript
|
||||||
|
const allVerbs = []
|
||||||
|
for (const sourceId of sourceIds) {
|
||||||
|
const verbs = await brain.related({ from: sourceId })
|
||||||
|
allVerbs.push(...verbs)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**After:**
|
||||||
|
```typescript
|
||||||
|
const storage = brain.storage as BaseStorage
|
||||||
|
const results = await storage.getVerbsBySourceBatch(sourceIds)
|
||||||
|
|
||||||
|
const allVerbs = []
|
||||||
|
for (const verbs of results.values()) {
|
||||||
|
allVerbs.push(...verbs)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance Gain:** One batched metadata fetch instead of one query per source entity.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Best Practices
|
||||||
|
|
||||||
|
### 1. **Use Batching for Multiple Entity Operations**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ✅ GOOD: Batch fetch
|
||||||
|
const results = await brain.batchGet(ids)
|
||||||
|
|
||||||
|
// ❌ BAD: Individual gets in loop
|
||||||
|
for (const id of ids) {
|
||||||
|
await brain.get(id)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. **Batch Size Recommendations**
|
||||||
|
|
||||||
|
| Storage | Optimal Batch Size | Max Batch Size |
|
||||||
|
|---------|--------------------|----------------|
|
||||||
|
| **Memory** | Unlimited | Unlimited |
|
||||||
|
| **Filesystem** | 100-500 | 1000 |
|
||||||
|
|
||||||
|
**Guideline:** For batches >1000, split into chunks of 500-1000.
|
||||||
|
|
||||||
|
### 3. **Metadata-Only by Default**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Default: Metadata-only (fast)
|
||||||
|
const results = await brain.batchGet(ids) // No vectors
|
||||||
|
|
||||||
|
// Only load vectors if needed
|
||||||
|
const withVectors = await brain.batchGet(ids, { includeVectors: true })
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. **Error Handling**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Batch operations never throw for missing entities
|
||||||
|
const results = await brain.batchGet(ids)
|
||||||
|
|
||||||
|
// Check results
|
||||||
|
for (const id of ids) {
|
||||||
|
if (results.has(id)) {
|
||||||
|
// Entity exists
|
||||||
|
const entity = results.get(id)
|
||||||
|
} else {
|
||||||
|
// Entity missing (not an error)
|
||||||
|
console.log(`Entity ${id} not found`)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
Comprehensive test coverage in `tests/integration/storage-batch-operations.test.ts`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx vitest run tests/integration/storage-batch-operations.test.ts
|
||||||
|
```
|
||||||
|
|
||||||
|
**Test Coverage:**
|
||||||
|
- ✅ brain.batchGet() high-level API
|
||||||
|
- ✅ storage.getNounMetadataBatch() with ID-first paths
|
||||||
|
- ✅ COW integration (branch isolation, inheritance)
|
||||||
|
- ✅ storage.getVerbsBySourceBatch() relationship queries
|
||||||
|
- ✅ VFS integration (PathResolver.getChildren())
|
||||||
|
- ✅ Performance benchmarks (N+1 elimination)
|
||||||
|
- ✅ Error handling (partial failures, empty batches, duplicates)
|
||||||
|
- ✅ ID-first storage verification
|
||||||
|
- ✅ Sharding preservation
|
||||||
|
|
||||||
|
**Results:** 23 tests passing ✅
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Implementation Details
|
||||||
|
|
||||||
|
### Architecture Layers
|
||||||
|
|
||||||
|
```
|
||||||
|
User Code (brain.batchGet)
|
||||||
|
↓
|
||||||
|
High-Level API (src/brainy.ts)
|
||||||
|
↓
|
||||||
|
Storage Layer (src/storage/baseStorage.ts)
|
||||||
|
↓
|
||||||
|
Adapter Layer (readBatchFromAdapter)
|
||||||
|
↓
|
||||||
|
Storage Adapter (FileSystemStorage / MemoryStorage)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Parallel Reads
|
||||||
|
|
||||||
|
Both shipped adapters fall back to `Promise.all` over individual reads:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// BaseStorage.readBatchFromAdapter()
|
||||||
|
return await Promise.all(resolvedPaths.map(path => this.read(path)))
|
||||||
|
```
|
||||||
|
|
||||||
|
**Shipped Adapters:**
|
||||||
|
- MemoryStorage
|
||||||
|
- FileSystemStorage
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## API Summary
|
||||||
|
|
||||||
|
- `brain.batchGet(ids, options?)` - High-level batch entity retrieval
|
||||||
|
- `storage.getNounMetadataBatch(ids)` - Storage-level metadata batch
|
||||||
|
- `storage.getVerbsBySourceBatch(sourceIds, verbType?)` - Batch relationship queries
|
||||||
|
|
||||||
|
**Performance Improvements:**
|
||||||
|
- VFS operations: single batched pass instead of N sequential reads
|
||||||
|
- Entity retrieval: N+1 reads collapsed into one batched pass
|
||||||
|
- Zero N+1 query patterns
|
||||||
|
|
||||||
|
**Compatibility:**
|
||||||
|
- ✅ ID-first storage
|
||||||
|
- ✅ Sharding (256 shards)
|
||||||
|
- ✅ Generational MVCC — batch reads serve the live generation; pinned `Db` views serve the past
|
||||||
|
- ✅ All indexes respected (vector, metadata, graph adjacency)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Support
|
||||||
|
|
||||||
|
- **Documentation:** `/docs/BATCHING.md`, `/docs/PERFORMANCE.md`
|
||||||
|
- **Tests:** `/tests/integration/storage-batch-operations.test.ts`
|
||||||
|
- **Issues:** https://github.com/soulcraft/brainy/issues
|
||||||
|
- **Discussions:** https://github.com/soulcraft/brainy/discussions
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Built with ❤️ for enterprise-scale knowledge graphs**
|
||||||
|
|
@ -1,303 +0,0 @@
|
||||||
# Creating Augmentations for Brainy
|
|
||||||
|
|
||||||
## The BrainyAugmentation Interface
|
|
||||||
|
|
||||||
Every augmentation implements this simple yet powerful interface:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface BrainyAugmentation {
|
|
||||||
// Identification
|
|
||||||
name: string // Unique name for your augmentation
|
|
||||||
|
|
||||||
// Execution control
|
|
||||||
timing: 'before' | 'after' | 'around' | 'replace' // When to execute
|
|
||||||
operations: string[] // Which operations to intercept
|
|
||||||
priority: number // Execution order (higher = first)
|
|
||||||
|
|
||||||
// Lifecycle methods
|
|
||||||
initialize(context: AugmentationContext): Promise<void>
|
|
||||||
execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T>
|
|
||||||
shutdown?(): Promise<void> // Optional cleanup
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Creating a Storage Augmentation
|
|
||||||
|
|
||||||
Storage augmentations are special - they provide the storage backend for Brainy:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { StorageAugmentation } from 'brainy/augmentations'
|
|
||||||
import { MyCustomStorage } from './my-storage'
|
|
||||||
|
|
||||||
export class MyStorageAugmentation extends StorageAugmentation {
|
|
||||||
private config: MyStorageConfig
|
|
||||||
|
|
||||||
constructor(config: MyStorageConfig) {
|
|
||||||
super()
|
|
||||||
this.name = 'my-custom-storage'
|
|
||||||
this.config = config
|
|
||||||
}
|
|
||||||
|
|
||||||
// Called during storage resolution phase
|
|
||||||
async provideStorage(): Promise<StorageAdapter> {
|
|
||||||
const storage = new MyCustomStorage(this.config)
|
|
||||||
this.storageAdapter = storage
|
|
||||||
return storage
|
|
||||||
}
|
|
||||||
|
|
||||||
// Called during augmentation initialization
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
await this.storageAdapter!.init()
|
|
||||||
this.log(`Custom storage initialized`)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Using Your Storage Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Register before brain.init()
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new MyStorageAugmentation({
|
|
||||||
connectionString: 'redis://localhost:6379'
|
|
||||||
}))
|
|
||||||
await brain.init() // Will use your storage!
|
|
||||||
```
|
|
||||||
|
|
||||||
## Creating a Feature Augmentation
|
|
||||||
|
|
||||||
Here's a complete example of a caching augmentation:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation, BrainyAugmentation } from 'brainy/augmentations'
|
|
||||||
|
|
||||||
export class CachingAugmentation extends BaseAugmentation {
|
|
||||||
private cache = new Map<string, any>()
|
|
||||||
|
|
||||||
constructor() {
|
|
||||||
super()
|
|
||||||
this.name = 'smart-cache'
|
|
||||||
this.timing = 'around' // Wrap operations
|
|
||||||
this.operations = ['search'] // Only cache searches
|
|
||||||
this.priority = 50 // Mid-priority
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
if (operation === 'search') {
|
|
||||||
// Check cache
|
|
||||||
const cacheKey = JSON.stringify(params)
|
|
||||||
if (this.cache.has(cacheKey)) {
|
|
||||||
this.log('Cache hit!')
|
|
||||||
return this.cache.get(cacheKey)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Execute and cache
|
|
||||||
const result = await next()
|
|
||||||
this.cache.set(cacheKey, result)
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
// Pass through other operations
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
this.log('Cache initialized')
|
|
||||||
}
|
|
||||||
|
|
||||||
async shutdown(): Promise<void> {
|
|
||||||
this.cache.clear()
|
|
||||||
await super.shutdown()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## The Four Timing Modes
|
|
||||||
|
|
||||||
### 1. `before` - Pre-processing
|
|
||||||
```typescript
|
|
||||||
timing = 'before'
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Validate/transform input
|
|
||||||
const validated = await validate(params)
|
|
||||||
return next(validated) // Pass modified params
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. `after` - Post-processing
|
|
||||||
```typescript
|
|
||||||
timing = 'after'
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const result = await next()
|
|
||||||
// Log, analyze, or modify result
|
|
||||||
console.log(`Operation ${op} returned:`, result)
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. `around` - Wrapping (middleware)
|
|
||||||
```typescript
|
|
||||||
timing = 'around'
|
|
||||||
async execute(op, params, next) {
|
|
||||||
console.log('Starting', op)
|
|
||||||
try {
|
|
||||||
const result = await next()
|
|
||||||
console.log('Success', op)
|
|
||||||
return result
|
|
||||||
} catch (error) {
|
|
||||||
console.log('Failed', op, error)
|
|
||||||
throw error
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. `replace` - Complete replacement
|
|
||||||
```typescript
|
|
||||||
timing = 'replace'
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Don't call next() - replace entirely!
|
|
||||||
return myCustomImplementation(params)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Operations You Can Intercept
|
|
||||||
|
|
||||||
Common operations in Brainy:
|
|
||||||
- `'storage'` - Storage resolution (special)
|
|
||||||
- `'add'`, `'addNoun'` - Adding data
|
|
||||||
- `'search'`, `'similar'` - Searching
|
|
||||||
- `'update'`, `'delete'` - Modifications
|
|
||||||
- `'saveNoun'`, `'saveVerb'` - Storage operations
|
|
||||||
- `'all'` - Intercept everything
|
|
||||||
|
|
||||||
## Context Available to Augmentations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface AugmentationContext {
|
|
||||||
brain: BrainyData // The brain instance
|
|
||||||
storage: StorageAdapter // Storage backend
|
|
||||||
config: BrainyDataConfig // Configuration
|
|
||||||
log: (message: string, level?: 'info' | 'warn' | 'error') => void
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Real-World Examples
|
|
||||||
|
|
||||||
### 1. Redis Storage Augmentation
|
|
||||||
```typescript
|
|
||||||
export class RedisStorageAugmentation extends StorageAugmentation {
|
|
||||||
async provideStorage(): Promise<StorageAdapter> {
|
|
||||||
return new RedisAdapter({
|
|
||||||
host: 'localhost',
|
|
||||||
port: 6379,
|
|
||||||
// Implement full StorageAdapter interface
|
|
||||||
})
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Audit Trail Augmentation
|
|
||||||
```typescript
|
|
||||||
export class AuditAugmentation extends BaseAugmentation {
|
|
||||||
timing = 'after'
|
|
||||||
operations = ['add', 'update', 'delete']
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Log to audit trail
|
|
||||||
await this.logAudit({
|
|
||||||
operation: op,
|
|
||||||
params,
|
|
||||||
result,
|
|
||||||
timestamp: new Date(),
|
|
||||||
user: this.context.config.currentUser
|
|
||||||
})
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Rate Limiting Augmentation
|
|
||||||
```typescript
|
|
||||||
export class RateLimitAugmentation extends BaseAugmentation {
|
|
||||||
timing = 'before'
|
|
||||||
operations = ['search']
|
|
||||||
private limiter = new RateLimiter({ rps: 100 })
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
await this.limiter.acquire() // Wait if rate limited
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Publishing to Brain Cloud Marketplace
|
|
||||||
|
|
||||||
Future capability for premium augmentations:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// package.json
|
|
||||||
{
|
|
||||||
"name": "@brain-cloud/redis-storage",
|
|
||||||
"brainy": {
|
|
||||||
"type": "augmentation",
|
|
||||||
"category": "storage",
|
|
||||||
"premium": true
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Users can install via:
|
|
||||||
// brainy augment install redis-storage
|
|
||||||
```
|
|
||||||
|
|
||||||
## Best Practices
|
|
||||||
|
|
||||||
1. **Use BaseAugmentation** - Provides common functionality
|
|
||||||
2. **Set appropriate priority** - Storage (100), System (80-99), Features (10-50)
|
|
||||||
3. **Be selective with operations** - Don't use 'all' unless necessary
|
|
||||||
4. **Handle errors gracefully** - Don't break the chain
|
|
||||||
5. **Clean up in shutdown()** - Release resources
|
|
||||||
6. **Log appropriately** - Use context.log() for consistent output
|
|
||||||
7. **Document your augmentation** - Include examples
|
|
||||||
|
|
||||||
## Testing Your Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
import { MyAugmentation } from './my-augmentation'
|
|
||||||
|
|
||||||
describe('MyAugmentation', () => {
|
|
||||||
let brain: BrainyData
|
|
||||||
|
|
||||||
beforeEach(async () => {
|
|
||||||
brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new MyAugmentation())
|
|
||||||
await brain.init()
|
|
||||||
})
|
|
||||||
|
|
||||||
afterEach(async () => {
|
|
||||||
await brain.destroy()
|
|
||||||
})
|
|
||||||
|
|
||||||
it('should enhance searches', async () => {
|
|
||||||
// Test your augmentation's effect
|
|
||||||
const results = await brain.search('test')
|
|
||||||
expect(results).toHaveProperty('enhanced', true)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## Summary
|
|
||||||
|
|
||||||
Augmentations are Brainy's extension system. They can:
|
|
||||||
- Replace storage backends
|
|
||||||
- Add caching layers
|
|
||||||
- Implement audit trails
|
|
||||||
- Add rate limiting
|
|
||||||
- Sync with external systems
|
|
||||||
- Transform data
|
|
||||||
- And much more!
|
|
||||||
|
|
||||||
The unified BrainyAugmentation interface makes it easy to create powerful extensions while maintaining consistency across the entire system.
|
|
||||||
271
docs/DATA_MODEL.md
Normal file
271
docs/DATA_MODEL.md
Normal file
|
|
@ -0,0 +1,271 @@
|
||||||
|
# Data Model
|
||||||
|
|
||||||
|
> How Brainy stores entities and relationships, and the critical distinction between `data` and `metadata`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Entity (Noun)
|
||||||
|
|
||||||
|
An entity is the fundamental data unit in Brainy. Every entity has:
|
||||||
|
|
||||||
|
| Field | Type | Indexed | Description |
|
||||||
|
|-------|------|---------|-------------|
|
||||||
|
| `id` | `string` | Primary key | UUID v4 (auto-generated or custom) |
|
||||||
|
| `data` | `any` | **HNSW vector index** | Content used for semantic/hybrid search. Strings auto-embed. |
|
||||||
|
| `metadata` | `object` | **MetadataIndex** | Structured queryable fields (tags, dates, flags, etc.) |
|
||||||
|
| `type` | `NounType` | MetadataIndex (as `noun`) | Entity type classification |
|
||||||
|
| `vector` | `number[]` | HNSW | 384-dim embedding (auto-computed from `data` or user-provided) |
|
||||||
|
| `confidence` | `number` | MetadataIndex | Type classification confidence (0-1) |
|
||||||
|
| `weight` | `number` | MetadataIndex | Entity importance/salience (0-1) |
|
||||||
|
| `service` | `string` | MetadataIndex | Multi-tenancy identifier |
|
||||||
|
| `createdAt` | `number` | MetadataIndex | Creation timestamp (ms since epoch) |
|
||||||
|
| `updatedAt` | `number` | MetadataIndex | Last update timestamp (ms since epoch) |
|
||||||
|
| `createdBy` | `object` | MetadataIndex | Source augmentation info |
|
||||||
|
|
||||||
|
### Example
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const id = await brain.add({
|
||||||
|
data: 'John Smith is a software engineer at Acme Corp', // → embedded into vector
|
||||||
|
type: NounType.Person,
|
||||||
|
metadata: { // → indexed, queryable via where filters
|
||||||
|
role: 'engineer',
|
||||||
|
department: 'backend',
|
||||||
|
yearsExperience: 8
|
||||||
|
},
|
||||||
|
confidence: 0.95,
|
||||||
|
weight: 0.7
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Relationship (Verb)
|
||||||
|
|
||||||
|
A relationship is a typed, directed edge connecting two entities.
|
||||||
|
|
||||||
|
| Field | Type | Indexed | Description |
|
||||||
|
|-------|------|---------|-------------|
|
||||||
|
| `id` | `string` | Primary key | UUID v4 (auto-generated) |
|
||||||
|
| `from` | `string` | **GraphAdjacencyIndex** | Source entity ID |
|
||||||
|
| `to` | `string` | **GraphAdjacencyIndex** | Target entity ID |
|
||||||
|
| `type` | `VerbType` | GraphAdjacencyIndex (as `verb`) | Relationship type classification |
|
||||||
|
| `data` | `any` | — | Opaque content (overrides auto-computed vector if provided) |
|
||||||
|
| `metadata` | `object` | — | Structured fields on the edge |
|
||||||
|
| `weight` | `number` | — | Connection strength (0-1, default: 1.0) |
|
||||||
|
| `confidence` | `number` | — | Relationship certainty (0-1) |
|
||||||
|
| `evidence` | `RelationEvidence` | — | Why this relationship was detected |
|
||||||
|
| `createdAt` | `number` | — | Creation timestamp (ms since epoch) |
|
||||||
|
| `updatedAt` | `number` | — | Last update timestamp (ms since epoch) |
|
||||||
|
| `service` | `string` | — | Multi-tenancy identifier |
|
||||||
|
|
||||||
|
### Example
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const relId = await brain.relate({
|
||||||
|
from: personId,
|
||||||
|
to: projectId,
|
||||||
|
type: VerbType.WorksOn,
|
||||||
|
data: 'Lead engineer on the AI module', // Optional: content for this edge
|
||||||
|
metadata: { // Optional: queryable edge fields
|
||||||
|
role: 'lead',
|
||||||
|
startDate: '2024-01-15'
|
||||||
|
},
|
||||||
|
weight: 0.9
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Data vs Metadata
|
||||||
|
|
||||||
|
This is the most important concept in Brainy's storage model:
|
||||||
|
|
||||||
|
### `data` — Content for Semantic Search
|
||||||
|
|
||||||
|
- Embedded into a 384-dimensional vector via the WASM embedding engine
|
||||||
|
- Searchable via **semantic similarity** (HNSW vector index) and **hybrid text+semantic** search
|
||||||
|
- Queried by passing `query` to `find()`:
|
||||||
|
```typescript
|
||||||
|
brain.find({ query: 'machine learning algorithms' })
|
||||||
|
```
|
||||||
|
- **NOT** indexed by MetadataIndex — you cannot use `where` filters on `data`
|
||||||
|
- Stored opaquely: strings, objects, numbers — anything goes
|
||||||
|
|
||||||
|
### `metadata` — Structured Queryable Fields
|
||||||
|
|
||||||
|
- Indexed by MetadataIndex with O(1) lookups per field
|
||||||
|
- Queryable via `where` filters using [BFO operators](./QUERY_OPERATORS.md):
|
||||||
|
```typescript
|
||||||
|
brain.find({
|
||||||
|
where: {
|
||||||
|
department: 'engineering',
|
||||||
|
yearsExperience: { greaterThan: 5 },
|
||||||
|
tags: { contains: 'senior' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
- **NOT** used for vector/semantic search
|
||||||
|
- Must be a flat or lightly nested object
|
||||||
|
|
||||||
|
### Quick Reference
|
||||||
|
|
||||||
|
| | `data` | `metadata` |
|
||||||
|
|---|---|---|
|
||||||
|
| **Purpose** | Content for embedding / semantic search | Structured fields for filtering |
|
||||||
|
| **Searched by** | `find({ query })` — vector similarity, hybrid text+semantic | `find({ where })` — exact, range, set operators |
|
||||||
|
| **Indexed by** | HNSW vector index | MetadataIndex |
|
||||||
|
| **Queryable with operators?** | No | Yes (`equals`, `greaterThan`, `oneOf`, etc.) |
|
||||||
|
| **Auto-embedded?** | Yes (strings → 384-dim vectors) | No |
|
||||||
|
| **Typical content** | Text descriptions, document content | Tags, dates, status flags, categories, numeric fields |
|
||||||
|
|
||||||
|
### Common Pattern
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Add an article
|
||||||
|
await brain.add({
|
||||||
|
data: 'A deep dive into transformer architectures and attention mechanisms',
|
||||||
|
type: NounType.Document,
|
||||||
|
metadata: {
|
||||||
|
title: 'Transformer Deep Dive',
|
||||||
|
author: 'Dr. Chen',
|
||||||
|
publishedYear: 2024,
|
||||||
|
tags: ['AI', 'transformers', 'NLP'],
|
||||||
|
status: 'published'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Search by content (semantic — searches data)
|
||||||
|
const results = await brain.find({ query: 'neural network attention' })
|
||||||
|
|
||||||
|
// Filter by fields (exact — queries metadata)
|
||||||
|
const recent = await brain.find({
|
||||||
|
where: {
|
||||||
|
publishedYear: { greaterThan: 2023 },
|
||||||
|
status: 'published'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Combine both (Triple Intelligence)
|
||||||
|
const precise = await brain.find({
|
||||||
|
query: 'attention mechanisms', // Semantic search on data
|
||||||
|
where: { author: 'Dr. Chen' }, // Metadata filter
|
||||||
|
connected: { from: authorId, depth: 1 } // Graph traversal
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Storage Field Naming
|
||||||
|
|
||||||
|
Internally, Brainy uses different field names in storage vs the public API:
|
||||||
|
|
||||||
|
| Public API (Entity/Relation) | Storage (metadata object) | Notes |
|
||||||
|
|------------------------------|--------------------------|-------|
|
||||||
|
| `type` | `noun` | Entity type stored as `noun` |
|
||||||
|
| `from` | `sourceId` | Relationship source |
|
||||||
|
| `to` | `targetId` | Relationship target |
|
||||||
|
| `type` (on Relation) | `verb` | Relationship type stored as `verb` |
|
||||||
|
|
||||||
|
When querying with `find()`, you can use:
|
||||||
|
- `type` parameter (convenience alias, equivalent to `where.noun`)
|
||||||
|
- `where.noun` directly
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// These are equivalent:
|
||||||
|
brain.find({ type: NounType.Person })
|
||||||
|
brain.find({ where: { noun: NounType.Person } })
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Standard Metadata Fields
|
||||||
|
|
||||||
|
When you add an entity, Brainy stores these standard fields in the metadata object alongside your custom fields:
|
||||||
|
|
||||||
|
| Field | Set By | Description |
|
||||||
|
|-------|--------|-------------|
|
||||||
|
| `noun` | System | Entity type (NounType enum value) |
|
||||||
|
| `subtype` | User | Per-NounType sub-classification (e.g. `'employee'`, `'invoice'`, `'milestone'`). Flat string, no hierarchy. Indexed on the fast path and rolled into per-NounType statistics. |
|
||||||
|
| `data` | System | The raw `data` value (stored opaquely) |
|
||||||
|
| `createdAt` | System | Creation timestamp |
|
||||||
|
| `updatedAt` | System | Last update timestamp |
|
||||||
|
| `confidence` | User | Type classification confidence |
|
||||||
|
| `weight` | User | Entity importance |
|
||||||
|
| `service` | User | Multi-tenancy identifier |
|
||||||
|
| `createdBy` | User/System | Source augmentation |
|
||||||
|
|
||||||
|
On read, these standard fields are extracted to top-level Entity properties. The `metadata` field on the returned Entity contains **only your custom fields**.
|
||||||
|
|
||||||
|
### Subtype — sub-classification within a NounType
|
||||||
|
|
||||||
|
`type` (NounType) is a stable 42-value enum. `subtype` is the consumer-chosen string vocabulary *within* a type:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// A Person who is an employee:
|
||||||
|
await brain.add({
|
||||||
|
data: 'Avery Brooks — runs the AI lab',
|
||||||
|
type: NounType.Person,
|
||||||
|
subtype: 'employee',
|
||||||
|
metadata: { department: 'ai-lab' }
|
||||||
|
})
|
||||||
|
|
||||||
|
// A Document that is an invoice:
|
||||||
|
await brain.add({
|
||||||
|
data: 'INV-2026-001',
|
||||||
|
type: NounType.Document,
|
||||||
|
subtype: 'invoice',
|
||||||
|
metadata: { amount: 1500 }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
`subtype` lives at the **top level** — NOT inside `metadata`, NOT inside `data`. That's how `find({ type, subtype })` routes through the standard-field fast path (column-store hit) instead of the metadata fallback. See **[Subtypes & Facets](./guides/subtypes-and-facets.md)** for the full guide including `trackField()` and `migrateField()`.
|
||||||
|
|
||||||
|
### Subtype — sub-classification within a VerbType (7.30+)
|
||||||
|
|
||||||
|
Relationships are first-class citizens too. Every verb (`VerbType`) gets the same `subtype` primitive — a `ReportsTo` relationship might carry `subtype: 'direct'` vs `'dotted-line'`; a `RelatedTo` edge might carry `'spouse'` / `'sibling'` / `'colleague'`. Same shape as the noun side: flat string, no hierarchy, top-level standard field on `HNSWVerbWithMetadata` and on the public `Relation<T>`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
await brain.relate({
|
||||||
|
from: ceoId,
|
||||||
|
to: vpId,
|
||||||
|
type: VerbType.ReportsTo,
|
||||||
|
subtype: 'direct', // top-level standard field
|
||||||
|
metadata: { since: '2025-Q1' } // user-custom fields stay in metadata
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
Fast-path filter on the verb side:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const direct = await brain.related({
|
||||||
|
from: ceoId,
|
||||||
|
type: VerbType.ReportsTo,
|
||||||
|
subtype: 'direct'
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
The verb-side rollup at `_system/verb-subtype-statistics.json` mirrors the noun-side `_system/subtype-statistics.json` — same shape, same self-heal machinery. Per-VerbType-per-subtype counts are O(1) via `brain.counts.byRelationshipSubtype()`.
|
||||||
|
|
||||||
|
Verbs and nouns now have full capability parity — every API on the noun side has a verb-side mirror, including the new `brain.updateRelation()` (which closed a pre-7.30 gap where relationships had no update path).
|
||||||
|
|
||||||
|
### Standard verb fields
|
||||||
|
|
||||||
|
The verb-side equivalent of `STANDARD_ENTITY_FIELDS` is `STANDARD_VERB_FIELDS`, exported from `src/coreTypes.ts`. Verb-specific standard fields:
|
||||||
|
|
||||||
|
| Field | Description |
|
||||||
|
|---|---|
|
||||||
|
| `verb` | The VerbType enum value |
|
||||||
|
| `sourceId` / `targetId` | The two endpoints of the relationship |
|
||||||
|
| `subtype` | Sub-classification within the VerbType (7.30+) |
|
||||||
|
| `confidence`, `weight`, `createdAt`, `updatedAt`, `service`, `createdBy`, `data` | Same semantics as the noun-side standard fields |
|
||||||
|
|
||||||
|
The companion `resolveVerbField(verb, field)` helper resolves field paths the same way `resolveEntityField` does for nouns: standard fields first, metadata fallback for everything else.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## See Also
|
||||||
|
|
||||||
|
- [API Reference](./api/README.md) — Complete API documentation
|
||||||
|
- [Query Operators](./QUERY_OPERATORS.md) — All BFO operators with examples
|
||||||
|
- [Find System](./FIND_SYSTEM.md) — Natural language find() details
|
||||||
1166
docs/DEVELOPER_LEARNING_PATH.md
Normal file
1166
docs/DEVELOPER_LEARNING_PATH.md
Normal file
File diff suppressed because it is too large
Load diff
1423
docs/FIND_SYSTEM.md
Normal file
1423
docs/FIND_SYSTEM.md
Normal file
File diff suppressed because it is too large
Load diff
569
docs/MIGRATION-V3-TO-V4.md
Normal file
569
docs/MIGRATION-V3-TO-V4.md
Normal file
|
|
@ -0,0 +1,569 @@
|
||||||
|
# Brainy v3 → v4.0.0 Migration Guide
|
||||||
|
|
||||||
|
> **Migration Complexity**: Low
|
||||||
|
> **Breaking Changes**: None (fully backward compatible)
|
||||||
|
> **New Features**: Lifecycle management, batch operations, compression, quota monitoring
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Brainy v4.0.0 is a **backward-compatible** release focused on production-ready cost optimization features. Your existing v3 code will continue to work without modifications, but you'll want to enable the new v4.0.0 features for significant cost savings.
|
||||||
|
|
||||||
|
**Key Benefits of Upgrading:**
|
||||||
|
- 💰 **96% cost savings** with lifecycle policies
|
||||||
|
- 🚀 **1000x faster** bulk deletions with batch operations
|
||||||
|
- 📦 **60-80% space savings** with gzip compression
|
||||||
|
- 📊 **Real-time quota monitoring** for OPFS
|
||||||
|
- 🎯 **Zero downtime** migration
|
||||||
|
|
||||||
|
## What's New in v4.0.0
|
||||||
|
|
||||||
|
### 1. Lifecycle Management (Cloud Storage)
|
||||||
|
|
||||||
|
**Automatic tier transitions for massive cost savings:**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// NEW in v4.0.0
|
||||||
|
await storage.setLifecyclePolicy({
|
||||||
|
rules: [{
|
||||||
|
id: 'archive-old-data',
|
||||||
|
prefix: 'entities/',
|
||||||
|
status: 'Enabled',
|
||||||
|
transitions: [
|
||||||
|
{ days: 30, storageClass: 'STANDARD_IA' },
|
||||||
|
{ days: 90, storageClass: 'GLACIER' }
|
||||||
|
]
|
||||||
|
}]
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Supported on:**
|
||||||
|
- ✅ AWS S3 (Lifecycle + Intelligent-Tiering)
|
||||||
|
- ✅ Google Cloud Storage (Lifecycle + Autoclass)
|
||||||
|
- ✅ Azure Blob Storage (Lifecycle policies)
|
||||||
|
|
||||||
|
### 2. Batch Operations
|
||||||
|
|
||||||
|
**1000x faster bulk deletions:**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// v3: Delete one at a time (slow, expensive)
|
||||||
|
for (const id of idsToDelete) {
|
||||||
|
await brain.remove(id) // 1000 API calls for 1000 entities
|
||||||
|
}
|
||||||
|
|
||||||
|
// v4.0.0: Batch delete (fast, cheap)
|
||||||
|
const paths = idsToDelete.flatMap(id => [
|
||||||
|
`entities/nouns/vectors/${id.substring(0, 2)}/${id}.json`,
|
||||||
|
`entities/nouns/metadata/${id.substring(0, 2)}/${id}.json`
|
||||||
|
])
|
||||||
|
await storage.batchDelete(paths) // 1 API call for 1000 objects (S3)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Efficiency gains:**
|
||||||
|
- S3: 1000 objects per batch
|
||||||
|
- GCS: 100 objects per batch
|
||||||
|
- Azure: 256 objects per batch
|
||||||
|
|
||||||
|
### 3. Compression (FileSystem)
|
||||||
|
|
||||||
|
**60-80% space savings for local storage:**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// NEW in v4.0.0
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: {
|
||||||
|
type: 'filesystem',
|
||||||
|
path: './data',
|
||||||
|
compression: true // Enable gzip compression
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Automatic compression/decompression on all reads/writes
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Quota Monitoring (OPFS)
|
||||||
|
|
||||||
|
**Prevent quota exceeded errors in browsers:**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// NEW in v4.0.0
|
||||||
|
const status = await storage.getStorageStatus()
|
||||||
|
|
||||||
|
if (status.details.usagePercent > 80) {
|
||||||
|
console.warn('Approaching quota limit:', status.details)
|
||||||
|
// Take action: cleanup old data, notify user, etc.
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5. Tier Management (Azure)
|
||||||
|
|
||||||
|
**Manual or automatic tier transitions:**
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// NEW in v4.0.0
|
||||||
|
await storage.changeBlobTier(blobPath, 'Cool') // Hot → Cool (50% savings)
|
||||||
|
await storage.batchChangeTier([blob1, blob2], 'Archive') // 99% savings
|
||||||
|
|
||||||
|
// Rehydrate from Archive when needed
|
||||||
|
await storage.rehydrateBlob(blobPath, 'High') // 1-hour rehydration
|
||||||
|
```
|
||||||
|
|
||||||
|
## Storage Architecture Changes
|
||||||
|
|
||||||
|
### v3.x Storage Structure
|
||||||
|
|
||||||
|
```
|
||||||
|
brainy-data/
|
||||||
|
├── nouns/
|
||||||
|
│ └── {uuid}.json # Single file per entity
|
||||||
|
├── verbs/
|
||||||
|
│ └── {uuid}.json # Single file per relationship
|
||||||
|
├── metadata/
|
||||||
|
│ └── __metadata_*.json # Indexes
|
||||||
|
└── _system/
|
||||||
|
└── statistics.json
|
||||||
|
```
|
||||||
|
|
||||||
|
### v4.0.0 Storage Structure (Automatic Migration)
|
||||||
|
|
||||||
|
```
|
||||||
|
brainy-data/
|
||||||
|
├── entities/
|
||||||
|
│ ├── nouns/
|
||||||
|
│ │ ├── vectors/ # Vector + HNSW graph (NEW)
|
||||||
|
│ │ │ ├── 00/ ... ff/ # 256 UUID shards (NEW)
|
||||||
|
│ │ └── metadata/ # Business data (NEW)
|
||||||
|
│ │ ├── 00/ ... ff/ # 256 UUID shards (NEW)
|
||||||
|
│ └── verbs/
|
||||||
|
│ ├── vectors/ # Relationship vectors (NEW)
|
||||||
|
│ │ ├── 00/ ... ff/
|
||||||
|
│ └── metadata/ # Relationship data (NEW)
|
||||||
|
│ ├── 00/ ... ff/
|
||||||
|
└── _system/ # Unchanged
|
||||||
|
└── __metadata_*.json
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key Changes:**
|
||||||
|
1. **Metadata/Vector Separation**: Entities split into 2 files for optimal I/O
|
||||||
|
2. **UUID-Based Sharding**: 256 shards for cloud storage optimization
|
||||||
|
3. **Automatic Migration**: Brainy handles migration transparently on first run
|
||||||
|
|
||||||
|
## Migration Steps
|
||||||
|
|
||||||
|
### Step 1: Update Brainy Package
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm install @soulcraft/brainy@latest
|
||||||
|
```
|
||||||
|
|
||||||
|
**Check your version:**
|
||||||
|
```bash
|
||||||
|
npm list @soulcraft/brainy
|
||||||
|
# Should show: @soulcraft/brainy@4.0.0
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 2: No Code Changes Required! ✅
|
||||||
|
|
||||||
|
Your existing v3 code will work without modifications:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// This v3 code works perfectly in v4.0.0
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: './data' }
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.init()
|
||||||
|
await brain.add("content", { type: "entity" })
|
||||||
|
const results = await brain.search("query")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 3: First Run (Automatic Migration)
|
||||||
|
|
||||||
|
On first initialization with v4.0.0:
|
||||||
|
|
||||||
|
1. **Brainy detects v3 storage structure**
|
||||||
|
2. **Transparently migrates to v4.0.0 structure**:
|
||||||
|
- Creates `entities/` directory
|
||||||
|
- Migrates `nouns/` → `entities/nouns/vectors/` + `entities/nouns/metadata/`
|
||||||
|
- Migrates `verbs/` → `entities/verbs/vectors/` + `entities/verbs/metadata/`
|
||||||
|
- Applies UUID-based sharding
|
||||||
|
3. **Old structure preserved** (optional cleanup later)
|
||||||
|
|
||||||
|
**Migration time:**
|
||||||
|
- 10K entities: ~1 minute
|
||||||
|
- 100K entities: ~10 minutes
|
||||||
|
- 1M entities: ~2 hours
|
||||||
|
|
||||||
|
**Zero downtime:**
|
||||||
|
- Migration happens during init()
|
||||||
|
- No data loss
|
||||||
|
- Automatic rollback on error
|
||||||
|
|
||||||
|
### Step 4: Enable v4.0.0 Features (Optional but Recommended)
|
||||||
|
|
||||||
|
#### Enable Lifecycle Policies (Cloud Storage)
|
||||||
|
|
||||||
|
**AWS S3:**
|
||||||
|
```typescript
|
||||||
|
// After init()
|
||||||
|
await storage.setLifecyclePolicy({
|
||||||
|
rules: [{
|
||||||
|
id: 'optimize-storage',
|
||||||
|
prefix: 'entities/',
|
||||||
|
status: 'Enabled',
|
||||||
|
transitions: [
|
||||||
|
{ days: 30, storageClass: 'STANDARD_IA' },
|
||||||
|
{ days: 90, storageClass: 'GLACIER' }
|
||||||
|
]
|
||||||
|
}]
|
||||||
|
})
|
||||||
|
|
||||||
|
// Or use Intelligent-Tiering (recommended)
|
||||||
|
await storage.enableIntelligentTiering('entities/', 'auto-optimize')
|
||||||
|
```
|
||||||
|
|
||||||
|
**Google Cloud Storage:**
|
||||||
|
```typescript
|
||||||
|
await storage.enableAutoclass({
|
||||||
|
terminalStorageClass: 'ARCHIVE'
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Azure Blob Storage:**
|
||||||
|
```typescript
|
||||||
|
await storage.setLifecyclePolicy({
|
||||||
|
rules: [{
|
||||||
|
name: 'optimize-blobs',
|
||||||
|
enabled: true,
|
||||||
|
type: 'Lifecycle',
|
||||||
|
definition: {
|
||||||
|
filters: { blobTypes: ['blockBlob'] },
|
||||||
|
actions: {
|
||||||
|
baseBlob: {
|
||||||
|
tierToCool: { daysAfterModificationGreaterThan: 30 },
|
||||||
|
tierToArchive: { daysAfterModificationGreaterThan: 90 }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}]
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Enable Compression (FileSystem)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: {
|
||||||
|
type: 'filesystem',
|
||||||
|
path: './data',
|
||||||
|
compression: true // NEW: 60-80% space savings
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Use Batch Operations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Replace individual deletes with batch delete
|
||||||
|
const idsToDelete = [/* ... */]
|
||||||
|
const paths = idsToDelete.flatMap(id => {
|
||||||
|
const shard = id.substring(0, 2)
|
||||||
|
return [
|
||||||
|
`entities/nouns/vectors/${shard}/${id}.json`,
|
||||||
|
`entities/nouns/metadata/${shard}/${id}.json`
|
||||||
|
]
|
||||||
|
})
|
||||||
|
|
||||||
|
await storage.batchDelete(paths) // Much faster!
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Monitor Quota (OPFS)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Periodically check quota in browser apps
|
||||||
|
setInterval(async () => {
|
||||||
|
const status = await storage.getStorageStatus()
|
||||||
|
if (status.details.usagePercent > 80) {
|
||||||
|
notifyUser('Storage approaching limit')
|
||||||
|
}
|
||||||
|
}, 60000) // Check every minute
|
||||||
|
```
|
||||||
|
|
||||||
|
## Backward Compatibility
|
||||||
|
|
||||||
|
### Guaranteed to Work (No Changes Needed)
|
||||||
|
|
||||||
|
✅ All v3 APIs remain unchanged
|
||||||
|
✅ Storage adapters backward compatible
|
||||||
|
✅ Metadata structure unchanged
|
||||||
|
✅ Query APIs unchanged
|
||||||
|
✅ Configuration options unchanged
|
||||||
|
|
||||||
|
### New Optional APIs (Add When Ready)
|
||||||
|
|
||||||
|
- `storage.setLifecyclePolicy()` - NEW in v4.0.0
|
||||||
|
- `storage.getLifecyclePolicy()` - NEW in v4.0.0
|
||||||
|
- `storage.removeLifecyclePolicy()` - NEW in v4.0.0
|
||||||
|
- `storage.enableIntelligentTiering()` - NEW in v4.0.0 (S3)
|
||||||
|
- `storage.enableAutoclass()` - NEW in v4.0.0 (GCS)
|
||||||
|
- `storage.batchDelete()` - NEW in v4.0.0
|
||||||
|
- `storage.changeBlobTier()` - NEW in v4.0.0 (Azure)
|
||||||
|
- `storage.getStorageStatus()` - Enhanced in v4.0.0
|
||||||
|
|
||||||
|
## Testing Your Migration
|
||||||
|
|
||||||
|
### 1. Test in Development First
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Create test brain with v4.0.0
|
||||||
|
const testBrain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: './test-data' }
|
||||||
|
})
|
||||||
|
|
||||||
|
await testBrain.init()
|
||||||
|
|
||||||
|
// Verify migration
|
||||||
|
console.log('Initialization complete')
|
||||||
|
|
||||||
|
// Test basic operations
|
||||||
|
const id = await testBrain.add("test content", { type: "test" })
|
||||||
|
const results = await testBrain.search("test")
|
||||||
|
console.log('Basic operations working:', results.length > 0)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Verify Storage Structure
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Check new directory structure
|
||||||
|
ls -la ./test-data/entities/nouns/vectors/
|
||||||
|
# Should see: 00/ 01/ 02/ ... ff/ (256 shards)
|
||||||
|
|
||||||
|
ls -la ./test-data/entities/nouns/metadata/
|
||||||
|
# Should see: 00/ 01/ 02/ ... ff/ (256 shards)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Verify Data Integrity
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Query all entities
|
||||||
|
const allEntities = await testBrain.find({})
|
||||||
|
console.log('Total entities:', allEntities.length)
|
||||||
|
|
||||||
|
// Verify specific entities
|
||||||
|
const entity = await testBrain.get(knownEntityId)
|
||||||
|
console.log('Entity retrieved:', entity !== null)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Test Performance
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Benchmark search
|
||||||
|
const start = Date.now()
|
||||||
|
const results = await testBrain.search("query")
|
||||||
|
const duration = Date.now() - start
|
||||||
|
console.log('Search time:', duration, 'ms')
|
||||||
|
|
||||||
|
// Should be similar or faster than v3
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rollback Procedure (If Needed)
|
||||||
|
|
||||||
|
If you encounter issues, you can rollback:
|
||||||
|
|
||||||
|
### Option 1: Rollback Package
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Reinstall v3
|
||||||
|
npm install @soulcraft/brainy@^3.50.0
|
||||||
|
|
||||||
|
# Restart application
|
||||||
|
```
|
||||||
|
|
||||||
|
**Important:** v3 can still read v3-structured data (preserved during migration)
|
||||||
|
|
||||||
|
### Option 2: Restore from Backup
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# If you backed up data before migration
|
||||||
|
rm -rf ./data
|
||||||
|
cp -r ./data-backup ./data
|
||||||
|
|
||||||
|
# Reinstall v3
|
||||||
|
npm install @soulcraft/brainy@^3.50.0
|
||||||
|
```
|
||||||
|
|
||||||
|
## Common Migration Scenarios
|
||||||
|
|
||||||
|
### Scenario 1: Small Application (<10K Entities)
|
||||||
|
|
||||||
|
**Migration time:** 1 minute
|
||||||
|
**Recommended approach:**
|
||||||
|
1. Update npm package
|
||||||
|
2. Restart application (automatic migration)
|
||||||
|
3. Enable lifecycle policies immediately
|
||||||
|
|
||||||
|
### Scenario 2: Medium Application (10K-1M Entities)
|
||||||
|
|
||||||
|
**Migration time:** 10 minutes - 2 hours
|
||||||
|
**Recommended approach:**
|
||||||
|
1. Backup data
|
||||||
|
2. Update npm package
|
||||||
|
3. Schedule maintenance window
|
||||||
|
4. Restart application (automatic migration)
|
||||||
|
5. Verify data integrity
|
||||||
|
6. Enable lifecycle policies
|
||||||
|
|
||||||
|
### Scenario 3: Large Application (1M+ Entities)
|
||||||
|
|
||||||
|
**Migration time:** 2-24 hours
|
||||||
|
**Recommended approach:**
|
||||||
|
1. **Backup data** (critical!)
|
||||||
|
2. Test migration on staging environment
|
||||||
|
3. Schedule extended maintenance window
|
||||||
|
4. Update npm package on production
|
||||||
|
5. Restart application (automatic migration)
|
||||||
|
6. Monitor migration progress
|
||||||
|
7. Verify data integrity thoroughly
|
||||||
|
8. Enable lifecycle policies gradually
|
||||||
|
|
||||||
|
## Cost Savings After Migration
|
||||||
|
|
||||||
|
### Enable All v4.0.0 Features
|
||||||
|
|
||||||
|
**500TB Dataset Example:**
|
||||||
|
|
||||||
|
**Before v4.0.0 (v3 with AWS S3 Standard):**
|
||||||
|
```
|
||||||
|
Storage: $138,000/year
|
||||||
|
Operations: $5,000/year
|
||||||
|
Total: $143,000/year
|
||||||
|
```
|
||||||
|
|
||||||
|
**After v4.0.0 (with Intelligent-Tiering):**
|
||||||
|
```
|
||||||
|
Storage: $51,000/year (64% savings)
|
||||||
|
Operations: $5,000/year
|
||||||
|
Total: $56,000/year
|
||||||
|
```
|
||||||
|
|
||||||
|
**After v4.0.0 (with Lifecycle Policies):**
|
||||||
|
```
|
||||||
|
Storage: $5,940/year (96% savings!)
|
||||||
|
Operations: $5,000/year
|
||||||
|
Total: $10,940/year
|
||||||
|
```
|
||||||
|
|
||||||
|
**Annual Savings: $132,060 (96% reduction)**
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
### Issue: Migration takes too long
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
- Migration is I/O bound
|
||||||
|
- For 1M+ entities, consider:
|
||||||
|
- Running during off-peak hours
|
||||||
|
- Using faster storage (SSD vs HDD)
|
||||||
|
- Increasing available memory
|
||||||
|
- Running on more powerful instance
|
||||||
|
|
||||||
|
### Issue: "Storage structure not recognized"
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
```typescript
|
||||||
|
// Manually trigger migration
|
||||||
|
await brain.storage.migrateToV4() // If automatic migration fails
|
||||||
|
|
||||||
|
// Or start fresh (data loss warning!)
|
||||||
|
await brain.storage.clear()
|
||||||
|
await brain.init()
|
||||||
|
```
|
||||||
|
|
||||||
|
### Issue: Lifecycle policy not working
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
```typescript
|
||||||
|
// Verify policy is set
|
||||||
|
const policy = await storage.getLifecyclePolicy()
|
||||||
|
console.log('Active rules:', policy.rules)
|
||||||
|
|
||||||
|
// Cloud providers may take 24-48 hours to start transitions
|
||||||
|
// Check again after 2 days
|
||||||
|
|
||||||
|
// Verify in cloud console:
|
||||||
|
// - AWS: S3 → Bucket → Management → Lifecycle
|
||||||
|
// - GCS: Storage → Bucket → Lifecycle
|
||||||
|
// - Azure: Storage Account → Lifecycle management
|
||||||
|
```
|
||||||
|
|
||||||
|
### Issue: Batch delete not working
|
||||||
|
|
||||||
|
**Solution:**
|
||||||
|
```typescript
|
||||||
|
// Ensure storage adapter supports batch delete
|
||||||
|
const status = await storage.getStorageStatus()
|
||||||
|
console.log('Storage type:', status.type)
|
||||||
|
|
||||||
|
// Batch delete requires:
|
||||||
|
// - S3CompatibleStorage ✅
|
||||||
|
// - GcsStorage ✅
|
||||||
|
// - AzureBlobStorage ✅
|
||||||
|
// - FileSystemStorage ✅
|
||||||
|
// - OPFSStorage ✅
|
||||||
|
// - MemoryStorage ✅
|
||||||
|
```
|
||||||
|
|
||||||
|
## Best Practices
|
||||||
|
|
||||||
|
1. ✅ **Backup before upgrading** (especially for large datasets)
|
||||||
|
2. ✅ **Test on staging first** (verify migration works)
|
||||||
|
3. ✅ **Monitor during migration** (watch logs for errors)
|
||||||
|
4. ✅ **Enable lifecycle policies immediately** (start saving costs)
|
||||||
|
5. ✅ **Use batch operations** (for any bulk cleanup)
|
||||||
|
6. ✅ **Monitor quota** (OPFS browser apps)
|
||||||
|
7. ✅ **Enable compression** (FileSystem storage)
|
||||||
|
|
||||||
|
## Getting Help
|
||||||
|
|
||||||
|
**Documentation:**
|
||||||
|
- [AWS S3 Cost Optimization Guide](./operations/cost-optimization-aws-s3.md)
|
||||||
|
- [GCS Cost Optimization Guide](./operations/cost-optimization-gcs.md)
|
||||||
|
- [Azure Cost Optimization Guide](./operations/cost-optimization-azure.md)
|
||||||
|
- [Cloudflare R2 Cost Optimization Guide](./operations/cost-optimization-cloudflare-r2.md)
|
||||||
|
|
||||||
|
**Support:**
|
||||||
|
- GitHub Issues: [https://github.com/soulcraft/brainy/issues](https://github.com/soulcraft/brainy/issues)
|
||||||
|
- GitHub Discussions: [https://github.com/soulcraft/brainy/discussions](https://github.com/soulcraft/brainy/discussions)
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
**Migration Checklist:**
|
||||||
|
- ✅ Backup data
|
||||||
|
- ✅ Update npm package (`npm install @soulcraft/brainy@latest`)
|
||||||
|
- ✅ Restart application (automatic migration)
|
||||||
|
- ✅ Verify data integrity
|
||||||
|
- ✅ Enable lifecycle policies
|
||||||
|
- ✅ Enable compression (FileSystem)
|
||||||
|
- ✅ Use batch operations
|
||||||
|
- ✅ Monitor cost savings
|
||||||
|
|
||||||
|
**Expected Results:**
|
||||||
|
- ✅ Zero downtime migration
|
||||||
|
- ✅ Full backward compatibility
|
||||||
|
- ✅ 60-96% cost savings
|
||||||
|
- ✅ 1000x faster bulk operations
|
||||||
|
- ✅ 60-80% space savings (with compression)
|
||||||
|
|
||||||
|
**Timeline:**
|
||||||
|
- Small app (<10K): 1 minute migration
|
||||||
|
- Medium app (10K-1M): 10 minutes - 2 hours
|
||||||
|
- Large app (1M+): 2-24 hours
|
||||||
|
|
||||||
|
**Welcome to Brainy v4.0.0! 🎉**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Version**: v4.0.0
|
||||||
|
**Migration Difficulty**: Low
|
||||||
|
**Breaking Changes**: None
|
||||||
|
**Recommended Upgrade**: Yes (significant cost savings)
|
||||||
|
|
@ -1,83 +0,0 @@
|
||||||
# 🤖 Model Loading Quick Reference
|
|
||||||
|
|
||||||
## 🚀 Common Scenarios
|
|
||||||
|
|
||||||
### ✅ Development (Zero Config)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init() // Downloads automatically
|
|
||||||
```
|
|
||||||
|
|
||||||
### 🐳 Docker Production
|
|
||||||
```dockerfile
|
|
||||||
RUN npm run download-models
|
|
||||||
ENV BRAINY_ALLOW_REMOTE_MODELS=false
|
|
||||||
```
|
|
||||||
|
|
||||||
### ☁️ Serverless/Lambda
|
|
||||||
```bash
|
|
||||||
# Build step
|
|
||||||
npm run download-models
|
|
||||||
|
|
||||||
# Runtime
|
|
||||||
export BRAINY_ALLOW_REMOTE_MODELS=false
|
|
||||||
```
|
|
||||||
|
|
||||||
### 🔒 Air-Gapped/Offline
|
|
||||||
```bash
|
|
||||||
# Connected machine
|
|
||||||
npm run download-models
|
|
||||||
tar -czf brainy-models.tar.gz ./models
|
|
||||||
|
|
||||||
# Offline machine
|
|
||||||
tar -xzf brainy-models.tar.gz
|
|
||||||
export BRAINY_ALLOW_REMOTE_MODELS=false
|
|
||||||
```
|
|
||||||
|
|
||||||
### 🌐 Browser/CDN
|
|
||||||
```html
|
|
||||||
<!-- Automatic - no setup needed -->
|
|
||||||
<script type="module">
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init() // Works in browser
|
|
||||||
</script>
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚨 Troubleshooting
|
|
||||||
|
|
||||||
| Error | Solution |
|
|
||||||
|-------|----------|
|
|
||||||
| "Failed to load embedding model" | `npm run download-models` |
|
|
||||||
| "ENOENT: no such file" | Check `BRAINY_MODELS_PATH` |
|
|
||||||
| "Network timeout" | Set `BRAINY_ALLOW_REMOTE_MODELS=false` |
|
|
||||||
| "Permission denied" | `chmod 755 ./models` |
|
|
||||||
| "Out of memory" | Increase container memory limit |
|
|
||||||
|
|
||||||
## 🎯 Environment Variables
|
|
||||||
|
|
||||||
| Variable | Values | Purpose |
|
|
||||||
|----------|--------|---------|
|
|
||||||
| `BRAINY_ALLOW_REMOTE_MODELS` | `true`/`false` | Allow/block downloads |
|
|
||||||
| `BRAINY_MODELS_PATH` | `./models` | Model storage path |
|
|
||||||
| `NODE_ENV` | `production` | Environment detection |
|
|
||||||
|
|
||||||
## 📦 Model Info
|
|
||||||
|
|
||||||
- **Model**: All-MiniLM-L6-v2
|
|
||||||
- **Dimensions**: 384 (fixed)
|
|
||||||
- **Size**: ~80MB download, ~330MB uncompressed
|
|
||||||
- **Location**: `./models/Xenova/all-MiniLM-L6-v2/`
|
|
||||||
|
|
||||||
## ✅ Verification Commands
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Check models exist
|
|
||||||
ls ./models/Xenova/all-MiniLM-L6-v2/onnx/model.onnx
|
|
||||||
|
|
||||||
# Test offline mode
|
|
||||||
BRAINY_ALLOW_REMOTE_MODELS=false npm test
|
|
||||||
|
|
||||||
# Download fresh models
|
|
||||||
rm -rf ./models && npm run download-models
|
|
||||||
```
|
|
||||||
492
docs/PERFORMANCE.md
Normal file
492
docs/PERFORMANCE.md
Normal file
|
|
@ -0,0 +1,492 @@
|
||||||
|
# Brainy Performance & Architecture
|
||||||
|
|
||||||
|
## Performance Characteristics
|
||||||
|
|
||||||
|
Brainy achieves high performance through carefully optimized data structures and algorithms. The tables below describe each component by its **algorithmic complexity** — the durable, defensible guarantee. The example latencies are figures from a single 100-item run on one machine (see [Benchmarks](#benchmarks)); they are illustrative, not a committed benchmark, and vary with hardware, embedding model, and storage backend. The one component with a committed scale assertion is the graph adjacency index (`tests/performance/graph-scale-performance.test.ts:238`).
|
||||||
|
|
||||||
|
### Core Performance Summary
|
||||||
|
|
||||||
|
| Component | Operation | Time Complexity | Example latency (100-item run)\* | Data Structure |
|
||||||
|
|-----------|-----------|-----------------|---------------------|----------------|
|
||||||
|
| **Metadata Index** | Exact match | **O(1)** | 0.8ms | `Map<string, Set<string>>` |
|
||||||
|
| **Metadata Index** | Range query | **O(log n) + O(k)** | 0.6ms | Sorted array + binary search |
|
||||||
|
| **Graph Index** | Get neighbors | **O(1)** | 0.09ms | `Map<string, Set<string>>` |
|
||||||
|
| **Vector Search** | k-NN search | **O(log n)** | 1.8ms | Hierarchical graph |
|
||||||
|
| **NLP Parser** | Query parsing | **O(m)** | 8.9ms | 220 pre-computed patterns |
|
||||||
|
| **Type-Field Affinity** | Field matching | **O(f)** | 0.1ms | Type-specific field cache |
|
||||||
|
| **Type Detection** | Noun/Verb matching | **O(t)** | 0.3ms | Pre-embedded type vectors |
|
||||||
|
| **Triple Intelligence** | Combined query | **O(1) to O(log n)** | 1.8ms | Parallel execution |
|
||||||
|
|
||||||
|
\* Illustrative single-run figures at 100 items on one machine — not a committed benchmark. Only the graph index carries an asserted scale bound (measured <1 ms per neighbor lookup up to 1M relationships, `tests/performance/graph-scale-performance.test.ts:238`).
|
||||||
|
|
||||||
|
Where:
|
||||||
|
- `n` = number of items in index
|
||||||
|
- `k` = number of results returned
|
||||||
|
- `m` = number of patterns to check
|
||||||
|
- `f` = number of fields for entity type
|
||||||
|
- `t` = number of types (42 nouns, 127 verbs)
|
||||||
|
|
||||||
|
### brain.get() Metadata-Only Optimization
|
||||||
|
|
||||||
|
`brain.get()` returns **metadata only by default**, skipping the 384-dimensional
|
||||||
|
embedding — the bulk of an entity's payload. Callers that need the vector opt in
|
||||||
|
with `{ includeVectors: true }`.
|
||||||
|
|
||||||
|
| Operation | Default (metadata-only) | With `includeVectors: true` | Use Case |
|
||||||
|
|-----------|-------------------------|-----------------------------|----------|
|
||||||
|
| **brain.get()** | Skips vector load | Loads full vector | VFS, existence checks, metadata |
|
||||||
|
| **VFS readFile() / readdir()** | Inherits metadata-only path | n/a | File operations, directory listings |
|
||||||
|
|
||||||
|
**Key Innovation**: Lazy vector loading — only load the 384-dimensional embedding when explicitly needed.
|
||||||
|
|
||||||
|
The integration test `tests/integration/metadata-only-comprehensive.test.ts:306`
|
||||||
|
asserts metadata-only `get()` is faster than the full-entity `get()`
|
||||||
|
(`metadataTime < fullTime`). The *magnitude* of the speedup is
|
||||||
|
environment-dependent (the percentage assertion in
|
||||||
|
`tests/integration/vfs-performance-v5.11.1.test.ts` is intentionally skipped on
|
||||||
|
CI for that reason), so no fixed percentage is quoted here.
|
||||||
|
|
||||||
|
**Why this matters**:
|
||||||
|
- Most `brain.get()` calls don't need vectors (VFS, admin tools, import utilities, data APIs)
|
||||||
|
- The embedding dominates an entity's serialized size, so skipping it is the largest win
|
||||||
|
- **Zero code changes** for most applications — automatic by default
|
||||||
|
|
||||||
|
**When to use what**:
|
||||||
|
```typescript
|
||||||
|
// DEFAULT: Metadata-only (skips the vector load) - use for:
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
// - VFS operations (readFile, stat, readdir)
|
||||||
|
// - Existence checks: if (await brain.get(id)) ...
|
||||||
|
// - Metadata access: entity.data, entity.type, entity.metadata
|
||||||
|
// - Relationship traversal
|
||||||
|
|
||||||
|
// EXPLICIT: Full entity (same as before) - use ONLY for:
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
// - Computing similarity on THIS entity
|
||||||
|
// - Manual vector operations
|
||||||
|
// - Vector index graph traversal
|
||||||
|
```
|
||||||
|
|
||||||
|
## Architecture Deep Dive
|
||||||
|
|
||||||
|
### 1. Metadata Index - O(1) Lookups
|
||||||
|
|
||||||
|
The `MetadataIndexManager` uses inverted indexes for lightning-fast metadata filtering.
|
||||||
|
|
||||||
|
**UPDATED**: Sorted indices for range queries are now built **incrementally during CRUD operations**. No lazy loading delays - range queries are consistently fast. Binary search insertions maintain O(log n) performance during updates.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class MetadataIndexManager {
|
||||||
|
// O(1) exact match via HashMap
|
||||||
|
private indexCache = new Map<string, MetadataIndexEntry>()
|
||||||
|
|
||||||
|
// O(log n) range queries via sorted arrays (incremental updates)
|
||||||
|
private sortedIndices = new Map<string, SortedFieldIndex>()
|
||||||
|
|
||||||
|
// Type-field affinity for intelligent NLP
|
||||||
|
private typeFieldAffinity = new Map<string, Map<string, number>>()
|
||||||
|
|
||||||
|
interface MetadataIndexEntry {
|
||||||
|
field: string
|
||||||
|
value: string | number | boolean
|
||||||
|
ids: Set<string> // O(1) add/remove/has
|
||||||
|
}
|
||||||
|
|
||||||
|
interface SortedFieldIndex {
|
||||||
|
values: Array<[value: any, ids: Set<string>]> // Sorted for O(log n) ranges
|
||||||
|
fieldType: 'number' | 'string' | 'date'
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
1. Each field+value combination gets a unique key: `"category:tech"`
|
||||||
|
2. Map lookup is O(1) average case
|
||||||
|
3. Returns a Set of matching IDs instantly
|
||||||
|
|
||||||
|
**Example Query:**
|
||||||
|
```javascript
|
||||||
|
// Query: { where: { category: 'tech' } }
|
||||||
|
// Internally: indexCache.get('category:tech') → O(1)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Range Queries - O(log n)
|
||||||
|
|
||||||
|
For numeric/date fields, Brainy maintains sorted indices:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
interface SortedFieldIndex {
|
||||||
|
values: Array<[value: any, ids: Set<string>]> // Sorted by value
|
||||||
|
fieldType: 'number' | 'string' | 'date'
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
1. Binary search to find range start: O(log n)
|
||||||
|
2. Binary search to find range end: O(log n)
|
||||||
|
3. Collect all IDs in range: O(k) where k = items in range
|
||||||
|
|
||||||
|
**Example Query:**
|
||||||
|
```javascript
|
||||||
|
// Query: { where: { age: { greaterThan: 25, lessThan: 40 } } }
|
||||||
|
// Internally: binarySearch(25) + binarySearch(40) + collect
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Graph Adjacency Index - O(1) Traversal
|
||||||
|
|
||||||
|
The `GraphAdjacencyIndex` provides instant graph traversal:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class GraphAdjacencyIndex {
|
||||||
|
// Bidirectional adjacency lists
|
||||||
|
private sourceIndex = new Map<string, Set<string>>() // id → outgoing
|
||||||
|
private targetIndex = new Map<string, Set<string>>() // id → incoming
|
||||||
|
|
||||||
|
// O(1) neighbor lookup
|
||||||
|
async getNeighbors(id: string, direction: 'in' | 'out' | 'both') {
|
||||||
|
const outgoing = this.sourceIndex.get(id) // O(1)
|
||||||
|
const incoming = this.targetIndex.get(id) // O(1)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key Innovation:** Pure Map/Set operations - no database queries, no loops, just direct memory access.
|
||||||
|
|
||||||
|
### 4. Vector Index - O(log n)
|
||||||
|
|
||||||
|
The default vector index (`JsHnswVectorIndex`) provides logarithmic approximate nearest neighbor search through a hierarchical graph:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class JsHnswVectorIndex {
|
||||||
|
private nouns: Map<string, HNSWNoun> = new Map()
|
||||||
|
|
||||||
|
interface HNSWNoun {
|
||||||
|
id: string
|
||||||
|
vector: number[]
|
||||||
|
connections: Map<number, Set<string>> // layer → neighbors
|
||||||
|
level: number
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
1. Start at entry point (top layer)
|
||||||
|
2. Greedy search to find nearest neighbor at each layer
|
||||||
|
3. Move down layers for progressively finer search
|
||||||
|
4. Each layer has M connections (typically 16)
|
||||||
|
|
||||||
|
**Performance:** O(log n) due to hierarchical structure
|
||||||
|
|
||||||
|
### 5. Type-Aware NLP with Dynamic Field Discovery
|
||||||
|
|
||||||
|
The NLP processor uses **zero hardcoded fields** - everything is discovered dynamically from actual data:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class NaturalLanguageProcessor {
|
||||||
|
// Pre-embedded NounTypes (42) and VerbTypes (127) - ONLY hardcoded vocabularies
|
||||||
|
private nounTypeEmbeddings = new Map<string, Vector>()
|
||||||
|
private verbTypeEmbeddings = new Map<string, Vector>()
|
||||||
|
|
||||||
|
// Dynamic field embeddings from actual indexed data
|
||||||
|
private fieldEmbeddings = new Map<string, Vector>()
|
||||||
|
|
||||||
|
// Type-field affinity for intelligent prioritization
|
||||||
|
async getFieldsForType(nounType: NounType) {
|
||||||
|
return this.brain.getFieldsForType(nounType) // Real data patterns
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Type-Aware Intelligence Flow:**
|
||||||
|
1. **Type Detection**: "documents" → `NounType.Document` (semantic similarity)
|
||||||
|
2. **Field Prioritization**: Get fields common to Document type from real data
|
||||||
|
3. **Semantic Field Matching**: "by" → "author" (with type affinity boost)
|
||||||
|
4. **Validation**: Ensure "author" field actually appears with Document entities
|
||||||
|
5. **Query Optimization**: Process low-cardinality type-specific fields first
|
||||||
|
|
||||||
|
**Performance Characteristics:**
|
||||||
|
- Type detection: O(t) where t = 169 total types (42 noun + 127 verb)
|
||||||
|
- Field matching: O(f) where f = fields for detected type (typically 5-15)
|
||||||
|
- Validation: O(1) lookup in type-field affinity map
|
||||||
|
- No hardcoded assumptions - learns from actual data patterns
|
||||||
|
|
||||||
|
### 6. NLP with 220 Pre-computed Patterns
|
||||||
|
|
||||||
|
Pattern matching with embedded templates for instant semantic understanding:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// 394KB of embedded patterns compiled into the source
|
||||||
|
export const EMBEDDED_PATTERNS: Pattern[] = [/* 220 patterns */]
|
||||||
|
export const PATTERN_EMBEDDINGS: Float32Array = /* 220 × 384 dimensions */
|
||||||
|
```
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
1. Query embedding computed once: O(1) with cached model
|
||||||
|
2. Cosine similarity with 220 patterns: O(m) where m = 220
|
||||||
|
3. Pattern templates enhanced with type context
|
||||||
|
4. No network calls, no external dependencies, no hardcoded fields
|
||||||
|
|
||||||
|
## Parallel Execution
|
||||||
|
|
||||||
|
Triple Intelligence queries execute searches in parallel:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
// Vector and proximity searches run simultaneously
|
||||||
|
const searchPromises = [
|
||||||
|
this.executeVectorSearch(params), // Runs in parallel
|
||||||
|
this.executeProximitySearch(params) // Runs in parallel
|
||||||
|
]
|
||||||
|
const results = await Promise.all(searchPromises)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Memory Efficiency
|
||||||
|
|
||||||
|
### Space Complexity
|
||||||
|
|
||||||
|
| Component | Memory Usage | Formula |
|
||||||
|
|-----------|--------------|---------|
|
||||||
|
| Metadata Index | ~40 bytes/entry | `(key_size + 8) × unique_values + 8 × total_items` |
|
||||||
|
| Graph Index | ~24 bytes/edge | `16 × edges + 8 × nodes` |
|
||||||
|
| Vector Index | ~1.5KB/item | `vector_size × 4 + M × 8 × layers` |
|
||||||
|
| Pattern Library | 394KB fixed | Pre-computed, shared across instances |
|
||||||
|
| Type Embeddings | ~60KB fixed | 70 types × 384 dimensions × 4 bytes, cached |
|
||||||
|
| Field Embeddings | ~5KB dynamic | Actual fields × 384 dimensions × 4 bytes |
|
||||||
|
| Type-Field Affinity | ~2KB dynamic | Type-field occurrence counts |
|
||||||
|
|
||||||
|
### Caching Strategy
|
||||||
|
|
||||||
|
- **Metadata Cache**: LRU with 5-minute TTL, 500 entries max
|
||||||
|
- **Embedding Cache**: Permanent for session, prevents recomputation
|
||||||
|
- **Unified Cache**: Coordinates memory across all components
|
||||||
|
|
||||||
|
## Benchmarks
|
||||||
|
|
||||||
|
### Illustrative Single Run (100 items, one machine)
|
||||||
|
|
||||||
|
Example output from a single 100-item run — illustrative only, not a committed
|
||||||
|
benchmark; absolute numbers vary by hardware. The values feed the
|
||||||
|
[Core Performance Summary](#core-performance-summary) example-latency column.
|
||||||
|
|
||||||
|
```
|
||||||
|
Metadata exact match: 0.818ms (50 items matched)
|
||||||
|
Metadata range query: 0.631ms (40 items in range)
|
||||||
|
Graph neighbor lookup: 0.092ms (2 connections)
|
||||||
|
Vector k-NN search: 1.773ms (10 nearest neighbors)
|
||||||
|
NLP query parsing: 8.906ms (full natural language)
|
||||||
|
Triple Intelligence: 1.830ms (combined query)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Scaling Characteristics
|
||||||
|
|
||||||
|
Each stage scales by its algorithmic complexity, not a fixed millisecond figure
|
||||||
|
— absolute latency depends on hardware, embedding model, and storage backend.
|
||||||
|
Only the graph adjacency index carries a committed scale assertion:
|
||||||
|
|
||||||
|
| Query stage | Complexity | Scaling behavior |
|
||||||
|
|-------------|------------|------------------|
|
||||||
|
| Metadata filter (exact) | O(1) | Constant — independent of dataset size |
|
||||||
|
| Metadata filter (range) | O(log n) + O(k) | Sub-linear; k = matching results |
|
||||||
|
| Vector search (HNSW) | O(log n) | Degrades gracefully via hierarchical layers |
|
||||||
|
| Graph hop | O(1) | Measured <1 ms per neighbor lookup, validated up to 1M relationships (`tests/performance/graph-scale-performance.test.ts:238`) |
|
||||||
|
| Combined query | O(log n) | Bounded by the vector stage; metadata and graph stages stay O(1)/O(log n) |
|
||||||
|
|
||||||
|
## Comparison with Other Systems
|
||||||
|
|
||||||
|
| System | Metadata Filter | Graph Traversal | Vector Search | Natural Language |
|
||||||
|
|--------|-----------------|-----------------|---------------|------------------|
|
||||||
|
| **Brainy** | O(1) HashMap | O(1) Adjacency | O(log n) vector index | 220 patterns |
|
||||||
|
| Neo4j | O(log n) B-tree | O(k) traversal | Not native | Not native |
|
||||||
|
| Elasticsearch | O(log n) inverted | Not native | O(n) brute force* | Basic tokenization |
|
||||||
|
| PostgreSQL | O(log n) B-tree | O(k) recursive | O(n) brute force* | Full-text only |
|
||||||
|
| Pinecone | Not native | Not native | O(log n) | Not native |
|
||||||
|
|
||||||
|
*Without additional plugins/extensions
|
||||||
|
|
||||||
|
## Key Innovations
|
||||||
|
|
||||||
|
1. **True O(1) Metadata Filtering**: Most databases use B-trees (O(log n)). Brainy uses HashMaps for constant-time lookups.
|
||||||
|
|
||||||
|
2. **O(1) Graph Traversal**: Unlike traditional graph databases that traverse edges, Brainy maintains bidirectional adjacency maps for instant neighbor access.
|
||||||
|
|
||||||
|
3. **Unified Triple Intelligence**: First system to natively combine O(1) metadata, O(1) graph, and O(log n) vector search in a single query.
|
||||||
|
|
||||||
|
4. **Embedded NLP**: 220 research-based patterns with pre-computed embeddings compiled directly into the codebase - no external dependencies.
|
||||||
|
|
||||||
|
5. **Parallel Search Execution**: Vector, metadata, and graph searches execute simultaneously, not sequentially.
|
||||||
|
|
||||||
|
## Production Readiness
|
||||||
|
|
||||||
|
- ✅ **No External Dependencies**: All algorithms implemented in pure TypeScript
|
||||||
|
- ✅ **No Network Calls**: Everything runs locally, including embeddings
|
||||||
|
- ✅ **Thread-Safe**: Immutable data structures where possible
|
||||||
|
- ✅ **Memory Bounded**: Configurable cache sizes and automatic cleanup
|
||||||
|
- ✅ **Single-Node by Design**: One process owns one `path`; scale out at the service layer
|
||||||
|
- ✅ **Zero Stubs**: Every line of code is production-ready
|
||||||
|
|
||||||
|
## Lazy Loading Performance
|
||||||
|
|
||||||
|
Brainy supports two initialization modes for optimal performance across different use cases:
|
||||||
|
|
||||||
|
### Mode 1: Auto-Rebuild (Default)
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const brain = new Brainy()
|
||||||
|
await brain.init() // Rebuilds indexes during init (~500ms-3s for 10K entities)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- Init time: 500ms-3s (depends on dataset size)
|
||||||
|
- First query: Instant (indexes already loaded)
|
||||||
|
- Use case: Traditional applications, long-running servers
|
||||||
|
|
||||||
|
### Mode 2: Lazy Loading
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const brain = new Brainy({ disableAutoRebuild: true })
|
||||||
|
await brain.init() // Returns instantly (0-10ms)
|
||||||
|
|
||||||
|
const results = await brain.find({ limit: 10 }) // First query triggers rebuild (~50-200ms)
|
||||||
|
const more = await brain.find({ limit: 100 }) // Subsequent queries instant (0ms check)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- Init time: 0-10ms (instant)
|
||||||
|
- First query: 50-200ms (includes index rebuild for 1K-10K entities)
|
||||||
|
- Subsequent queries: 0ms check (instant)
|
||||||
|
- Concurrent queries: Wait for same rebuild (mutex prevents duplicates)
|
||||||
|
|
||||||
|
**Concurrency Safety:**
|
||||||
|
```javascript
|
||||||
|
// 100 concurrent queries immediately after init
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
const promises = Array.from({ length: 100 }, () =>
|
||||||
|
brain.find({ limit: 10 })
|
||||||
|
)
|
||||||
|
|
||||||
|
const results = await Promise.all(promises)
|
||||||
|
// ✅ Only 1 rebuild triggered (mutex)
|
||||||
|
// ✅ All 100 queries return correct results
|
||||||
|
// ✅ Total time: ~60ms (not 6000ms!)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use Cases for Lazy Loading:**
|
||||||
|
- **Serverless/Edge**: Minimize cold start time (0-10ms init)
|
||||||
|
- **Development**: Faster restarts during development
|
||||||
|
- **Large datasets**: Defer index loading until needed
|
||||||
|
- **Read-heavy workloads**: Writes don't wait for index rebuild
|
||||||
|
|
||||||
|
## Zero Configuration Required
|
||||||
|
|
||||||
|
Brainy is designed to be **smart enough to tune itself dynamically**. No configuration needed:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
// That's it. Brainy handles everything.
|
||||||
|
const brain = new Brainy()
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Or with lazy loading for serverless
|
||||||
|
const brain = new Brainy({ disableAutoRebuild: true })
|
||||||
|
await brain.init() // Instant (0-10ms)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Automatic Self-Tuning
|
||||||
|
|
||||||
|
- **Metadata Index**: Auto-builds sorted indices for range queries on first use
|
||||||
|
- **Graph Index**: Auto-flushes every 30 seconds
|
||||||
|
- **Default Tuning**: Research-based vector index defaults
|
||||||
|
- **Lazy Loading**: Indices built only when needed
|
||||||
|
- **Cache Management**: LRU caches with TTL
|
||||||
|
|
||||||
|
### Intelligent Defaults
|
||||||
|
|
||||||
|
- **Vector recall** = `'balanced'` (M=16, ef=200): right for most datasets
|
||||||
|
- **Cache TTL** = 5 min: balances freshness and performance
|
||||||
|
- **Flush interval** = 30 s: non-blocking background persistence
|
||||||
|
|
||||||
|
### Vector Index Tuning Knobs
|
||||||
|
|
||||||
|
Brainy 8.0 exposes two knobs on `config.vector`:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const brain = new Brainy({
|
||||||
|
vector: {
|
||||||
|
recall: 'fast', // 'fast' | 'balanced' | 'accurate'
|
||||||
|
persistMode: 'deferred' // 'immediate' | 'deferred'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
The default JS index is `JsHnswVectorIndex`. An optional native acceleration package (`@soulcraft/cor`) can replace it with a higher-performing implementation; the public knobs stay the same.
|
||||||
|
|
||||||
|
### Scale Scenarios
|
||||||
|
|
||||||
|
| Scale | Items | Storage Strategy | Performance |
|
||||||
|
|-------|-------|------------------|-------------|
|
||||||
|
| **Small** | <10K | Memory | Sub-millisecond |
|
||||||
|
| **Medium** | 10K-1M | Filesystem | 1-5ms |
|
||||||
|
| **Large** | 1M-10M | Filesystem + tuned cache | 2-10ms |
|
||||||
|
| **Massive** | 10M+ | Filesystem + native vector provider + service-layer sharding | 5-20ms |
|
||||||
|
|
||||||
|
For >10M entities, run multiple Brainy processes behind your own routing layer — Brainy 8.0 doesn't ship cluster coordination.
|
||||||
|
|
||||||
|
### Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
┌─────────────────────────────────────────┐
|
||||||
|
│ Application Layer │
|
||||||
|
│ (Your Code) │
|
||||||
|
└─────────────┬───────────────────────────┘
|
||||||
|
│
|
||||||
|
┌─────────────▼───────────────────────────┐
|
||||||
|
│ Brainy Core │
|
||||||
|
│ (Triple Intelligence Engine) │
|
||||||
|
├─────────────────────────────────────────┤
|
||||||
|
│ Memory │ Vector │ Metadata │
|
||||||
|
│ Cache │ Index │ Index │
|
||||||
|
└─────────────┬───────────────────────────┘
|
||||||
|
│
|
||||||
|
┌─────────────▼───────────────────────────┐
|
||||||
|
│ Storage Layer │
|
||||||
|
├──────────┬──────────┬──────────────────┤
|
||||||
|
│ Vectors │ Graph │ Files │
|
||||||
|
│ (sharded)│ Edges │ (filesystem) │
|
||||||
|
└──────────┴──────────┴──────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
For off-site replication, snapshot `path` from your scheduler (`gsutil rsync`, `aws s3 sync`, `rclone`, or `tar`).
|
||||||
|
|
||||||
|
### Performance at Scale
|
||||||
|
|
||||||
|
- **Metadata queries**: O(1) HashMap
|
||||||
|
- **Graph traversal**: O(1) adjacency lookup
|
||||||
|
- **Vector search**: O(log n)
|
||||||
|
- **Write throughput**: 50K+ writes/second per process (filesystem, batched)
|
||||||
|
- **Read throughput**: 1M+ reads/second with caching
|
||||||
|
|
||||||
|
### Zero-Config with Autoscaling
|
||||||
|
|
||||||
|
- **AutoConfiguration System**: Detects environment and adjusts settings
|
||||||
|
- **Learning from Performance**: `learnFromPerformance()` adapts based on metrics
|
||||||
|
- **Auto-flush**: Graph index (30s), Metadata index (configurable)
|
||||||
|
- **Auto-optimize**: Enabled by default in graph and vector indices
|
||||||
|
- **Zero-config presets**: Production, development, minimal modes
|
||||||
|
- **Adaptive memory**: Scales caches based on available memory
|
||||||
|
|
||||||
|
## Implementation Status
|
||||||
|
|
||||||
|
### Fully Implemented and Production-Ready
|
||||||
|
- **O(1) metadata lookups** via HashMaps (exact match)
|
||||||
|
- **O(log n) range queries** via sorted arrays with lazy building
|
||||||
|
- **O(1) graph traversal** via adjacency maps
|
||||||
|
- **O(log n) vector search** via the default JS index, swappable for a native provider
|
||||||
|
- **220 NLP patterns** with pre-computed embeddings
|
||||||
|
- **Filesystem and memory storage** adapters
|
||||||
|
- **Auto-configuration system** with environment detection
|
||||||
|
- **Zero-config operation** with intelligent defaults
|
||||||
|
- **Auto-flush and auto-optimize** in indices
|
||||||
|
- **Low-latency Triple Intelligence queries** (O(log n) vector + O(1) metadata/graph)
|
||||||
|
|
||||||
|
## Conclusion
|
||||||
|
|
||||||
|
Brainy delivers on its promise of **production-ready Triple Intelligence** with documented algorithmic-complexity guarantees and a committed graph-scale benchmark (`tests/performance/graph-scale-performance.test.ts`). All listed features are fully implemented and tested. No stubs, no mocks — just real, working code with characterized performance.
|
||||||
486
docs/PLUGINS.md
Normal file
486
docs/PLUGINS.md
Normal file
|
|
@ -0,0 +1,486 @@
|
||||||
|
---
|
||||||
|
title: Plugin System
|
||||||
|
slug: guides/plugins
|
||||||
|
public: true
|
||||||
|
category: guides
|
||||||
|
template: guide
|
||||||
|
order: 4
|
||||||
|
description: Replace any Brainy subsystem — distance functions, embeddings, vector index, metadata index, aggregation — with a custom implementation or optional native acceleration.
|
||||||
|
next:
|
||||||
|
- guides/storage-adapters
|
||||||
|
---
|
||||||
|
|
||||||
|
# Plugin Development Guide
|
||||||
|
|
||||||
|
Brainy has a plugin system that allows third-party packages to replace internal subsystems with custom implementations. This is how `@soulcraft/cor` provides optional native acceleration, and it's the same system available to any developer.
|
||||||
|
|
||||||
|
## Architecture Overview
|
||||||
|
|
||||||
|
Brainy's plugin system uses **named providers** — string keys mapped to implementations. During `init()`, brainy:
|
||||||
|
|
||||||
|
1. Imports each package listed in the `plugins` config array
|
||||||
|
2. Activates each plugin, passing a `BrainyPluginContext`
|
||||||
|
3. The plugin calls `context.registerProvider(key, implementation)` for each subsystem it provides
|
||||||
|
4. Brainy checks each provider key and wires the implementation into its internal pipeline
|
||||||
|
|
||||||
|
Installing the first-party accelerator is the opt-in: with the default config, brainy probes for `@soulcraft/cor` and loads it when present. Everything except "not installed" fails **loud** — a present-but-broken accelerator makes `init()` throw rather than silently degrading to the JS engines.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy() // @soulcraft/cor auto-detected when installed
|
||||||
|
const pinned = new Brainy({ plugins: ['@soulcraft/cor'] }) // or pin exactly what loads
|
||||||
|
const plain = new Brainy({ plugins: [] }) // or opt out of detection entirely
|
||||||
|
```
|
||||||
|
|
||||||
|
| `plugins` value | Behavior |
|
||||||
|
|---|---|
|
||||||
|
| `undefined` (default) | Guarded auto-detection of `@soulcraft/cor`: not installed → no plugins, silently; installed → loads + announces; installed-but-broken → `init()` throws |
|
||||||
|
| `false` / `[]` | No plugins, no detection (explicit opt-out) |
|
||||||
|
| `['@soulcraft/cor']` | Load only the listed packages; a listed plugin that fails to load throws |
|
||||||
|
|
||||||
|
Plugins registered programmatically via `brain.use(plugin)` are always activated regardless of the `plugins` config.
|
||||||
|
|
||||||
|
If no plugin provides a given key, brainy uses its built-in JavaScript implementation. This means brainy works perfectly standalone — plugins only enhance performance or add capabilities.
|
||||||
|
|
||||||
|
## Creating a Plugin
|
||||||
|
|
||||||
|
### 1. Implement the `BrainyPlugin` interface
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import type { BrainyPlugin, BrainyPluginContext } from '@soulcraft/brainy/plugin'
|
||||||
|
|
||||||
|
const myPlugin: BrainyPlugin = {
|
||||||
|
name: 'my-brainy-plugin', // Must be unique (typically your npm package name)
|
||||||
|
|
||||||
|
async activate(context: BrainyPluginContext): Promise<boolean> {
|
||||||
|
// Register your providers here
|
||||||
|
context.registerProvider('distance', myFastDistanceFunction)
|
||||||
|
|
||||||
|
// Return true if activation succeeded, false to skip
|
||||||
|
return true
|
||||||
|
},
|
||||||
|
|
||||||
|
async deactivate(): Promise<void> {
|
||||||
|
// Optional cleanup when brainy.close() is called
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export default myPlugin
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Package exports
|
||||||
|
|
||||||
|
Your package must export the plugin as the default export so brainy's plugin loader can resolve it:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// index.ts
|
||||||
|
export { default } from './plugin.js'
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Registration
|
||||||
|
|
||||||
|
**Config-based:** List your package name in the brainy config:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
plugins: ['my-brainy-plugin']
|
||||||
|
})
|
||||||
|
await brain.init()
|
||||||
|
```
|
||||||
|
|
||||||
|
**Programmatic registration:** For plugins not installed as npm packages, use `brain.use()`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { Brainy } from '@soulcraft/brainy'
|
||||||
|
import myPlugin from './my-plugin.js'
|
||||||
|
|
||||||
|
const brain = new Brainy()
|
||||||
|
brain.use(myPlugin)
|
||||||
|
await brain.init()
|
||||||
|
```
|
||||||
|
|
||||||
|
## Provider Keys Reference
|
||||||
|
|
||||||
|
Each key has a specific expected signature. Brainy checks for these during `init()` and wires them into the appropriate code paths.
|
||||||
|
|
||||||
|
### Core Providers
|
||||||
|
|
||||||
|
#### `distance`
|
||||||
|
**Type:** `(a: number[], b: number[]) => number`
|
||||||
|
|
||||||
|
Replaces the default cosine distance function used in vector search and neural APIs. This is the highest-impact single provider — it's called for every vector comparison.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('distance', (a: number[], b: number[]): number => {
|
||||||
|
// Your SIMD-accelerated or GPU distance calculation
|
||||||
|
return myFastCosineDistance(a, b)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### `embeddings`
|
||||||
|
**Type:** `(text: string | string[]) => Promise<number[] | number[][]>`
|
||||||
|
|
||||||
|
Replaces the built-in WASM embedding engine. Called for every `brain.add()`, `brain.update()`, and `brain.find()` operation that involves text.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('embeddings', async (text: string | string[]) => {
|
||||||
|
if (Array.isArray(text)) {
|
||||||
|
return myEngine.embedBatch(text)
|
||||||
|
}
|
||||||
|
return myEngine.embed(text)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### `embedBatch`
|
||||||
|
**Type:** `(texts: string[]) => Promise<number[][]>`
|
||||||
|
|
||||||
|
Dedicated batch embedding provider. When registered, brainy uses this for bulk operations (import, reindex, batch add) instead of calling the `embeddings` provider N times. This enables true single-forward-pass batch processing.
|
||||||
|
|
||||||
|
Priority order for batch operations:
|
||||||
|
1. `embedBatch` provider (single forward pass — fastest)
|
||||||
|
2. `embeddings` provider with `Promise.all()` (N individual calls)
|
||||||
|
3. Built-in WASM batch API (fallback)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('embedBatch', async (texts: string[]) => {
|
||||||
|
// Process all texts in a single forward pass
|
||||||
|
return myEngine.batchEmbed(texts)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Index Providers
|
||||||
|
|
||||||
|
> **Write-path invariant (the change-feed contract).** Every canonical
|
||||||
|
> mutation flows through Brainy's generation-store commit points — index
|
||||||
|
> providers are invoked *inside* that commit and never originate canonical
|
||||||
|
> writes of their own. The `brain.onChange` change feed is emitted from those
|
||||||
|
> commit points and relies on this: **a plugin must never introduce a write
|
||||||
|
> path that bypasses the generation-store commit.** If a future provider ever
|
||||||
|
> needs a direct native ingest path, it must either route through the commit
|
||||||
|
> or emit equivalent change events — otherwise every `onChange` consumer
|
||||||
|
> (live UIs, cache invalidation, realtime sync) silently develops a blind
|
||||||
|
> spot.
|
||||||
|
|
||||||
|
#### `vector`
|
||||||
|
**Type:** `(config: object, distanceFunction: Function, options: object) => VectorIndexProvider-compatible`
|
||||||
|
|
||||||
|
Factory function that creates a vector index instance. The returned object must implement the `VectorIndexProvider` public API:
|
||||||
|
|
||||||
|
- `addItem(item: { id: string, vector: number[] }): Promise<string>`
|
||||||
|
- `search(queryVector: number[], k: number, filter?, options?): Promise<Array<[string, number]>>`
|
||||||
|
- `removeItem(id: string): Promise<boolean>`
|
||||||
|
- `size(): number`
|
||||||
|
- `clear(): void`
|
||||||
|
- `flush(): Promise<number>`
|
||||||
|
- `rebuild(options?): Promise<void>`
|
||||||
|
- `getDirtyNodeCount(): number`
|
||||||
|
- `getPersistMode(): 'immediate' | 'deferred'`
|
||||||
|
- `getEntryPointId(): string | null`
|
||||||
|
- `getMaxLevel(): number`
|
||||||
|
- `getDimension(): number | null`
|
||||||
|
- `getConfig(): object`
|
||||||
|
- `getDistanceFunction(): Function`
|
||||||
|
- `enableCOW(parent): void`
|
||||||
|
- `setUseParallelization(boolean): void`
|
||||||
|
|
||||||
|
For type-aware indexes (separate graph per noun type), also implement:
|
||||||
|
- `getIndexForType(type: string): VectorIndexProvider` (duck-typed detection)
|
||||||
|
- `search(queryVector, k, type?, filter?, options?): Promise<Array<[string, number]>>`
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('vector', (config, distanceFn, options) => {
|
||||||
|
return new MyNativeVectorIndex(config, distanceFn, options)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### The readiness contract (all three index providers)
|
||||||
|
|
||||||
|
A provider that **persists its derived index** should implement the optional readiness
|
||||||
|
members so a warm reopen never pays a redundant rebuild-from-canonical:
|
||||||
|
|
||||||
|
- **`init?(): Promise<void>`** — eager cold-load. Brainy awaits it once during
|
||||||
|
`brain.init()`, after the metadata provider's `init()` (the id-mapper hydrates first)
|
||||||
|
and **before the rebuild gate**.
|
||||||
|
- **`isReady?(): boolean`** — honest durability signal. `true` ⇔ the persisted index is
|
||||||
|
loaded (or cheaply demand-loadable) and consistent with what was last persisted. When
|
||||||
|
exposed, the rebuild gate defers to this signal **instead of** the `size() === 0` /
|
||||||
|
`totalEntries === 0` heuristics — a disk-native index may report 0 resident entries
|
||||||
|
while fully durable. Never return `true` if the durable state failed to load: the
|
||||||
|
signal is honest in both directions, and a not-ready provider gets its rebuild even
|
||||||
|
when `size() > 0`.
|
||||||
|
- **`isMigrating?(): boolean`** — while `true`, the provider owns its index (background
|
||||||
|
migration); brainy skips its rebuild entirely.
|
||||||
|
|
||||||
|
Providers that implement none of these keep the size/count heuristics — correct for
|
||||||
|
engines whose `rebuild()` *is* their load path (like brainy's built-in JS vector index).
|
||||||
|
|
||||||
|
#### `metadataIndex`
|
||||||
|
**Type:** `(storage: StorageAdapter) => MetadataIndexManager-compatible`
|
||||||
|
|
||||||
|
Factory function that creates a metadata index. The returned object must implement the `MetadataIndexManager` interface including `init()`, `addEntity()`, `removeEntity()`, `query()`, `flush()`, `clear()`, etc.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('metadataIndex', (storage) => {
|
||||||
|
return new MyNativeMetadataIndex(storage)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### `graphIndex`
|
||||||
|
**Type:** `(storage: StorageAdapter) => GraphAdjacencyIndex-compatible`
|
||||||
|
|
||||||
|
Factory function that creates a graph adjacency index for relationship tracking (verbs/triples). Must implement the `GraphAdjacencyIndex` interface including `addVerb()`, `getVerbsBySource()`, `getVerbsByTarget()`, `flush()`, etc.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('graphIndex', (storage) => {
|
||||||
|
return new MyNativeGraphIndex(storage)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
#### `aggregation`
|
||||||
|
**Type:** `(storage: StorageAdapter) => AggregationProvider-compatible`
|
||||||
|
|
||||||
|
Factory function that creates an aggregation engine for write-time incremental SUM/COUNT/AVG/MIN/MAX with GROUP BY and time windows. The returned object must implement the `AggregationProvider` interface.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('aggregation', (storage) => {
|
||||||
|
return new MyNativeAggregationEngine(storage)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
When provided by an optional native acceleration plugin (such as `@soulcraft/cor`), this enables:
|
||||||
|
- Compiled source filters (vs per-entity JS object traversal)
|
||||||
|
- Precise MIN/MAX via sorted data structures (vs lazy recompute)
|
||||||
|
- Parallel aggregate rebuild across CPU cores
|
||||||
|
- SIMD-accelerated timestamp bucketing
|
||||||
|
|
||||||
|
### Utility Providers
|
||||||
|
|
||||||
|
#### `cache`
|
||||||
|
**Type:** `UnifiedCache`
|
||||||
|
|
||||||
|
Replaces the global `UnifiedCache` singleton used for VFS path resolution, semantic caching, and vector index caching. Must implement the `UnifiedCache` interface (available from `@soulcraft/brainy/internals`).
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import type { UnifiedCache } from '@soulcraft/brainy/internals'
|
||||||
|
|
||||||
|
context.registerProvider('cache', myNativeCache)
|
||||||
|
```
|
||||||
|
|
||||||
|
#### `entityIdMapper`
|
||||||
|
**Type:** `(storage: StorageAdapter) => EntityIdMapper-compatible`
|
||||||
|
|
||||||
|
Factory for bidirectional UUID ↔ integer mapping used by roaring bitmaps. Must implement `getOrAssign()`, `getUuid()`, `getInt()`, `has()`, `remove()`, `flush()`, `clear()`.
|
||||||
|
|
||||||
|
#### `roaring`
|
||||||
|
**Type:** `RoaringBitmap32 class`
|
||||||
|
|
||||||
|
Replacement for the roaring bitmap implementation. Used internally by the metadata index for set operations. Must be API-compatible with `roaring-wasm`.
|
||||||
|
|
||||||
|
#### `msgpack`
|
||||||
|
**Type:** `{ encode: (data: any) => Buffer, decode: (buffer: Buffer) => any }`
|
||||||
|
|
||||||
|
Native msgpack encode/decode for SSTable serialization.
|
||||||
|
|
||||||
|
### Analytics Providers (Native-Only)
|
||||||
|
|
||||||
|
These provider keys have **no JavaScript fallback** — they represent capabilities that require native code (SIMD, mmap, sub-microsecond latency). They are available when an optional native acceleration plugin (such as `@soulcraft/cor`) is installed.
|
||||||
|
|
||||||
|
Use `brain.getProvider('analytics:hyperloglog')` to check availability. Returns `undefined` if no plugin provides it.
|
||||||
|
|
||||||
|
#### `analytics:hyperloglog`
|
||||||
|
Approximate distinct counts. Count unique values (e.g., unique merchants) across millions of records using ~16KB of memory with ~1% error. Each update is O(1).
|
||||||
|
|
||||||
|
#### `analytics:tdigest`
|
||||||
|
Streaming percentiles. Compute P50/P90/P95/P99 from streaming data without storing all values. Uses ~4KB per digest with ~1% accuracy at the tails.
|
||||||
|
|
||||||
|
#### `analytics:countmin`
|
||||||
|
Frequency estimation. Find the most common values (e.g., top-K merchants) using ~40KB with 0.1% error. O(1) per update.
|
||||||
|
|
||||||
|
#### `analytics:anomaly`
|
||||||
|
Real-time anomaly detection. Flag statistically unusual values at write-time using exponentially weighted moving averages. 64 bytes per group, sub-microsecond decisions.
|
||||||
|
|
||||||
|
#### `aggregation:mmap`
|
||||||
|
Persistent aggregate storage via memory-mapped files. Aggregate state survives process crashes without explicit flush. Zero serialization overhead.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Storage Adapter Plugins
|
||||||
|
|
||||||
|
Plugins can register custom storage backends that users reference by name.
|
||||||
|
|
||||||
|
### Implementing a Storage Adapter
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import type { StorageAdapterFactory } from '@soulcraft/brainy/plugin'
|
||||||
|
import type { StorageAdapter } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
class MyStorageAdapter implements StorageAdapter {
|
||||||
|
async init(): Promise<void> { /* ... */ }
|
||||||
|
async saveNoun(noun: HNSWNoun): Promise<void> { /* ... */ }
|
||||||
|
async getNoun(id: string): Promise<HNSWNounWithMetadata | null> { /* ... */ }
|
||||||
|
async deleteNoun(id: string): Promise<void> { /* ... */ }
|
||||||
|
// ... implement all StorageAdapter methods
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Registering a Storage Adapter
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
context.registerProvider('storage:my-backend', {
|
||||||
|
name: 'my-backend',
|
||||||
|
create: (config: Record<string, unknown>) => {
|
||||||
|
return new MyStorageAdapter(config)
|
||||||
|
}
|
||||||
|
} satisfies StorageAdapterFactory)
|
||||||
|
```
|
||||||
|
|
||||||
|
Users can then use your storage:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({ storage: 'my-backend', myBackendOption: 'value' })
|
||||||
|
```
|
||||||
|
|
||||||
|
## Import Paths
|
||||||
|
|
||||||
|
Brainy provides three entry points for plugin developers:
|
||||||
|
|
||||||
|
| Import Path | Contents | Stability |
|
||||||
|
|-------------|----------|-----------|
|
||||||
|
| `@soulcraft/brainy` | Public API, types, StorageAdapter | Stable (semver) |
|
||||||
|
| `@soulcraft/brainy/plugin` | BrainyPlugin, BrainyPluginContext, StorageAdapterFactory | Stable (semver) |
|
||||||
|
| `@soulcraft/brainy/internals` | UnifiedCache, EntityIdMapper, logger utilities | Internal (may change between minor versions) |
|
||||||
|
|
||||||
|
## Diagnostics
|
||||||
|
|
||||||
|
Brainy provides a `diagnostics()` method to verify plugin wiring:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy()
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
const diag = brain.diagnostics()
|
||||||
|
console.log(diag)
|
||||||
|
// {
|
||||||
|
// version: '7.14.0',
|
||||||
|
// plugins: { active: ['my-plugin'], count: 1 },
|
||||||
|
// providers: {
|
||||||
|
// metadataIndex: { source: 'default' },
|
||||||
|
// graphIndex: { source: 'default' },
|
||||||
|
// embeddings: { source: 'plugin' },
|
||||||
|
// embedBatch: { source: 'plugin' },
|
||||||
|
// distance: { source: 'plugin' },
|
||||||
|
// vector: { source: 'default' },
|
||||||
|
// ...
|
||||||
|
// },
|
||||||
|
// indexes: {
|
||||||
|
// vector: { size: 0, type: 'JsHnswVectorIndex' },
|
||||||
|
// metadata: { type: 'MetadataIndexManager', initialized: true },
|
||||||
|
// graph: { type: 'GraphAdjacencyIndex', initialized: true, wiredToStorage: true }
|
||||||
|
// }
|
||||||
|
// }
|
||||||
|
```
|
||||||
|
|
||||||
|
The CLI also supports diagnostics:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
brainy diagnostics
|
||||||
|
```
|
||||||
|
|
||||||
|
### Init-Time Summary
|
||||||
|
|
||||||
|
When a plugin is active, brainy automatically logs a provider summary after `init()`:
|
||||||
|
|
||||||
|
```
|
||||||
|
[brainy] Plugin activated: @soulcraft/cor
|
||||||
|
[brainy] Providers: 8/10 native (@soulcraft/cor) | default: vector, cache
|
||||||
|
```
|
||||||
|
|
||||||
|
This tells you at a glance how many subsystems are accelerated and which ones are falling back to JavaScript. The log respects `config.silent`.
|
||||||
|
|
||||||
|
### Fail-Fast for Production
|
||||||
|
|
||||||
|
Use `requireProviders()` after `init()` to guarantee specific providers are plugin-supplied. This prevents silent fallback to JavaScript in deployments where you expect native acceleration:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy()
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Throws immediately if any of these are using JS fallback
|
||||||
|
brain.requireProviders(['distance', 'embeddings', 'metadataIndex', 'graphIndex'])
|
||||||
|
```
|
||||||
|
|
||||||
|
If a required provider is missing, the error message tells you exactly what's wrong:
|
||||||
|
|
||||||
|
```
|
||||||
|
[brainy] Required providers using JS fallback: graphIndex.
|
||||||
|
Active plugins: @soulcraft/cor.
|
||||||
|
These providers must be supplied by a plugin for this deployment.
|
||||||
|
Check plugin installation, license, and native module availability.
|
||||||
|
```
|
||||||
|
|
||||||
|
This is the recommended pattern for production deployments with paid plugins — fail at startup rather than silently degrading performance.
|
||||||
|
|
||||||
|
## Complete Example: Distance Acceleration Plugin
|
||||||
|
|
||||||
|
A minimal but useful plugin that provides SIMD-accelerated distance calculations:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// simd-distance-plugin/src/plugin.ts
|
||||||
|
import type { BrainyPlugin, BrainyPluginContext } from '@soulcraft/brainy/plugin'
|
||||||
|
|
||||||
|
// Hypothetical native module
|
||||||
|
import { simdCosineDistance } from './native.js'
|
||||||
|
|
||||||
|
const simdDistancePlugin: BrainyPlugin = {
|
||||||
|
name: 'brainy-simd-distance',
|
||||||
|
|
||||||
|
async activate(context: BrainyPluginContext): Promise<boolean> {
|
||||||
|
// Check if SIMD is available on this platform
|
||||||
|
if (!checkSimdSupport()) {
|
||||||
|
console.log('[simd-distance] SIMD not available, skipping')
|
||||||
|
return false // Don't activate — brainy uses JS fallback
|
||||||
|
}
|
||||||
|
|
||||||
|
context.registerProvider('distance', simdCosineDistance)
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export default simdDistancePlugin
|
||||||
|
```
|
||||||
|
|
||||||
|
```json
|
||||||
|
// simd-distance-plugin/package.json
|
||||||
|
{
|
||||||
|
"name": "brainy-simd-distance",
|
||||||
|
"main": "./dist/plugin.js",
|
||||||
|
"types": "./dist/plugin.d.ts",
|
||||||
|
"peerDependencies": {
|
||||||
|
"@soulcraft/brainy": ">=7.0.0"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { Brainy } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
const brain = new Brainy({ plugins: ['brainy-simd-distance'] })
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Verify it's active
|
||||||
|
const diag = brain.diagnostics()
|
||||||
|
console.log(diag.providers.distance) // { source: 'plugin' }
|
||||||
|
```
|
||||||
|
|
||||||
|
## Design Principles
|
||||||
|
|
||||||
|
1. **Brainy works perfectly without plugins.** Every provider has a JavaScript fallback. Plugins only improve performance or add capabilities.
|
||||||
|
|
||||||
|
2. **Provider keys are string-based.** The plugin system is not coupled to any specific plugin. Any package can register any provider.
|
||||||
|
|
||||||
|
3. **Clean separation.** Plugins access brainy through the documented `BrainyPluginContext` interface. No direct access to internal classes is needed.
|
||||||
|
|
||||||
|
4. **Fail-safe activation.** If a plugin throws during `activate()`, brainy logs a warning and continues with defaults. A broken plugin never prevents brainy from working.
|
||||||
|
|
||||||
|
5. **Lifecycle management.** `deactivate()` is called during `brainy.close()` for resource cleanup. Native resources, connections, and file handles should be released here.
|
||||||
562
docs/PRODUCTION_SERVICE_ARCHITECTURE.md
Normal file
562
docs/PRODUCTION_SERVICE_ARCHITECTURE.md
Normal file
|
|
@ -0,0 +1,562 @@
|
||||||
|
# Production Service Architecture Guide
|
||||||
|
|
||||||
|
**How to use Brainy optimally in production services (Bun, Node.js, Deno)**
|
||||||
|
|
||||||
|
> **Recommended Runtime:** [Bun](https://bun.sh) provides best performance with Brainy's Candle WASM engine. All examples work with both Bun and Node.js.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Problem: Instance-per-Request Anti-Pattern
|
||||||
|
|
||||||
|
### ❌ What NOT to Do
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// WRONG - Creates new instance EVERY request
|
||||||
|
app.get('/api/entities', async (req, res) => {
|
||||||
|
const brain = new Brainy({ storage: { path: './brainy-data' } })
|
||||||
|
await brain.init() // FULL INITIALIZATION EVERY TIME!
|
||||||
|
const entities = await brain.find(...)
|
||||||
|
res.json(entities)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Why This is Terrible
|
||||||
|
|
||||||
|
After 40 API calls:
|
||||||
|
- **40 Brainy instances** running simultaneously
|
||||||
|
- **20GB memory** (40 × 500MB per instance)
|
||||||
|
- **2 seconds wasted** (40 × 50ms initialization)
|
||||||
|
- **Zero cache benefit** (each instance has its own empty cache)
|
||||||
|
- **Index rebuilding** on every request (TypeAware HNSW, LSM-trees, etc.)
|
||||||
|
- **Memory leaks** (old instances may not GC properly)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ✅ The Solution: Singleton Pattern
|
||||||
|
|
||||||
|
**ONE Brainy instance per service, shared across ALL requests.**
|
||||||
|
|
||||||
|
### Performance Comparison
|
||||||
|
|
||||||
|
| Metric | Instance-per-Request | Singleton (Optimal) |
|
||||||
|
|--------|---------------------|---------------------|
|
||||||
|
| Memory (40 requests) | 20GB | 500MB |
|
||||||
|
| Request 1 latency | 60ms | 60ms (one-time init) |
|
||||||
|
| Request 2+ latency | 60ms (no cache!) | 2ms (80% cache hit!) |
|
||||||
|
| Cache hit rate | 0% | 80%+ |
|
||||||
|
| Speedup | - | **30x faster** |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Implementation Patterns
|
||||||
|
|
||||||
|
### Pattern 1: Simple Singleton (Recommended)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// server.ts
|
||||||
|
import { Brainy } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
// SINGLETON INSTANCE
|
||||||
|
let brainInstance: Brainy | null = null
|
||||||
|
|
||||||
|
async function getBrain(): Promise<Brainy> {
|
||||||
|
if (brainInstance) {
|
||||||
|
return brainInstance
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log('🧠 Initializing Brainy singleton...')
|
||||||
|
|
||||||
|
brainInstance = new Brainy({
|
||||||
|
storage: {
|
||||||
|
path: './brainy-data',
|
||||||
|
autoOptimize: true
|
||||||
|
},
|
||||||
|
cache: {
|
||||||
|
maxSize: 1000, // Shared across ALL requests
|
||||||
|
ttl: 3600000, // 1 hour
|
||||||
|
enableMetrics: true
|
||||||
|
},
|
||||||
|
augmentations: {
|
||||||
|
include: ['cache', 'metrics', 'display', 'vfs']
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
await brainInstance.init()
|
||||||
|
console.log('✅ Brainy ready')
|
||||||
|
|
||||||
|
return brainInstance
|
||||||
|
}
|
||||||
|
|
||||||
|
// Initialize BEFORE starting server
|
||||||
|
async function startServer() {
|
||||||
|
await getBrain() // One-time initialization
|
||||||
|
|
||||||
|
app.get('/api/entities', async (req, res) => {
|
||||||
|
const brain = await getBrain() // Reuses same instance!
|
||||||
|
const entities = await brain.find(req.query)
|
||||||
|
res.json(entities)
|
||||||
|
})
|
||||||
|
|
||||||
|
app.listen(3000)
|
||||||
|
}
|
||||||
|
|
||||||
|
startServer()
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- ✅ Simple to implement
|
||||||
|
- ✅ Thread-safe (async initialization)
|
||||||
|
- ✅ Shared cache and indexes
|
||||||
|
- ✅ 40x memory reduction
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Pattern 2: Service Class (Production-Grade)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// services/BrainService.ts
|
||||||
|
export class BrainService {
|
||||||
|
private brain: Brainy | null = null
|
||||||
|
private initPromise: Promise<Brainy> | null = null
|
||||||
|
|
||||||
|
async getInstance(): Promise<Brainy> {
|
||||||
|
if (this.brain) return this.brain
|
||||||
|
if (this.initPromise) return this.initPromise
|
||||||
|
|
||||||
|
this.initPromise = this.initialize()
|
||||||
|
return this.initPromise
|
||||||
|
}
|
||||||
|
|
||||||
|
private async initialize(): Promise<Brainy> {
|
||||||
|
this.brain = new Brainy({
|
||||||
|
storage: {
|
||||||
|
path: process.env.BRAINY_DATA_PATH || './brainy-data'
|
||||||
|
},
|
||||||
|
cache: { maxSize: 1000, ttl: 3600000 }
|
||||||
|
})
|
||||||
|
await this.brain.init()
|
||||||
|
return this.brain
|
||||||
|
}
|
||||||
|
|
||||||
|
async shutdown(): Promise<void> {
|
||||||
|
if (this.brain) {
|
||||||
|
// Cleanup if needed
|
||||||
|
this.brain = null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// server.ts
|
||||||
|
const brainService = new BrainService()
|
||||||
|
|
||||||
|
app.get('/api/entities', async (req, res) => {
|
||||||
|
const brain = await brainService.getInstance()
|
||||||
|
const entities = await brain.find(req.query)
|
||||||
|
res.json(entities)
|
||||||
|
})
|
||||||
|
|
||||||
|
// Graceful shutdown
|
||||||
|
process.on('SIGTERM', async () => {
|
||||||
|
await brainService.shutdown()
|
||||||
|
process.exit(0)
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- ✅ Prevents race conditions (multiple simultaneous inits)
|
||||||
|
- ✅ Testable (can inject mock)
|
||||||
|
- ✅ Clean shutdown handling
|
||||||
|
- ✅ Environment-configurable
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Pattern 3: Bun Server (Recommended)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// server.ts - Clean Bun implementation
|
||||||
|
import { Brainy } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
let brain: Brainy | null = null
|
||||||
|
|
||||||
|
async function getBrain(): Promise<Brainy> {
|
||||||
|
if (!brain) {
|
||||||
|
brain = new Brainy({ storage: { path: './brainy-data' } })
|
||||||
|
await brain.init()
|
||||||
|
}
|
||||||
|
return brain
|
||||||
|
}
|
||||||
|
|
||||||
|
// Initialize before server starts
|
||||||
|
await getBrain()
|
||||||
|
|
||||||
|
Bun.serve({
|
||||||
|
port: 3000,
|
||||||
|
async fetch(req) {
|
||||||
|
const url = new URL(req.url)
|
||||||
|
|
||||||
|
if (url.pathname === '/api/entities') {
|
||||||
|
const b = await getBrain()
|
||||||
|
const entities = await b.find({})
|
||||||
|
return Response.json(entities)
|
||||||
|
}
|
||||||
|
|
||||||
|
if (url.pathname === '/api/entity' && req.method === 'POST') {
|
||||||
|
const b = await getBrain()
|
||||||
|
const body = await req.json()
|
||||||
|
const id = await b.add(body)
|
||||||
|
return Response.json({ id })
|
||||||
|
}
|
||||||
|
|
||||||
|
return new Response('Not Found', { status: 404 })
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
console.log('Server running on http://localhost:3000')
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- ✅ Native Bun runtime performance
|
||||||
|
- ✅ No framework dependencies
|
||||||
|
- ✅ Pure WASM — no native binaries, bundler-friendly
|
||||||
|
- ✅ Built-in TypeScript support
|
||||||
|
|
||||||
|
### Pattern 4: Express/Node.js Middleware (Legacy)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// middleware/brainy.ts
|
||||||
|
let brainInstance: Brainy | null = null
|
||||||
|
|
||||||
|
export async function initBrainy() {
|
||||||
|
if (!brainInstance) {
|
||||||
|
brainInstance = new Brainy({ storage: { path: './brainy-data' } })
|
||||||
|
await brainInstance.init()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function brainMiddleware(req, res, next) {
|
||||||
|
if (!brainInstance) {
|
||||||
|
return res.status(500).json({ error: 'Brainy not initialized' })
|
||||||
|
}
|
||||||
|
req.brain = brainInstance // Attach to request
|
||||||
|
next()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Type extension
|
||||||
|
declare global {
|
||||||
|
namespace Express {
|
||||||
|
interface Request {
|
||||||
|
brain: Brainy
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// server.ts
|
||||||
|
import { initBrainy, brainMiddleware } from './middleware/brainy'
|
||||||
|
|
||||||
|
async function startServer() {
|
||||||
|
await initBrainy() // Initialize first
|
||||||
|
|
||||||
|
app.use('/api', brainMiddleware) // Apply to API routes
|
||||||
|
|
||||||
|
app.get('/api/entities', async (req, res) => {
|
||||||
|
const entities = await req.brain.find(req.query) // Type-safe!
|
||||||
|
res.json(entities)
|
||||||
|
})
|
||||||
|
|
||||||
|
app.listen(3000)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- ✅ Clean separation of concerns
|
||||||
|
- ✅ Type-safe (`req.brain` is typed)
|
||||||
|
- ✅ Easy to add auth/validation
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Optimization Strategies
|
||||||
|
|
||||||
|
### 1. Configure Cache for Your Workload
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
cache: {
|
||||||
|
maxSize: 1000, // Number of entities to cache
|
||||||
|
ttl: 3600000, // Cache lifetime (1 hour)
|
||||||
|
enableMetrics: true, // Track hit rate
|
||||||
|
evictionPolicy: 'lru' // Least recently used
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Cache sizing:**
|
||||||
|
- Small service (< 100 req/min): `maxSize: 500`
|
||||||
|
- Medium service (< 1000 req/min): `maxSize: 1000`
|
||||||
|
- Large service (> 1000 req/min): `maxSize: 5000`
|
||||||
|
|
||||||
|
### 2. Lazy Load Augmentations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
augmentations: {
|
||||||
|
// Only load what you actually use
|
||||||
|
include: ['cache', 'metrics', 'display', 'vfs'],
|
||||||
|
exclude: ['neuralImport', 'intelligentImport'] // Skip heavy features
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Memory savings:**
|
||||||
|
- With all augmentations: ~800MB
|
||||||
|
- With minimal set: ~400MB
|
||||||
|
|
||||||
|
### 3. Warm Up Indexes
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
async function startServer() {
|
||||||
|
const brain = await getBrain()
|
||||||
|
|
||||||
|
// Pre-warm frequently-used indexes
|
||||||
|
await brain.find({ type: 'person', limit: 1 })
|
||||||
|
await brain.find({ type: 'organization', limit: 1 })
|
||||||
|
|
||||||
|
console.log('✅ Indexes pre-warmed')
|
||||||
|
|
||||||
|
app.listen(3000)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefit:** First requests are fast (no cold-start index building)
|
||||||
|
|
||||||
|
### 4. Memory-Aware Configuration
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import os from 'os'
|
||||||
|
|
||||||
|
const totalMemory = os.totalmem()
|
||||||
|
const availableMemory = os.freemem()
|
||||||
|
|
||||||
|
const brain = new Brainy({
|
||||||
|
cache: {
|
||||||
|
// Use 10% of total RAM for cache
|
||||||
|
maxSize: Math.floor(totalMemory * 0.1 / (1024 * 1024))
|
||||||
|
},
|
||||||
|
indexes: {
|
||||||
|
// Lazy load indexes if low memory
|
||||||
|
lazyLoad: availableMemory < totalMemory * 0.5,
|
||||||
|
preload: ['person', 'organization'] // Only preload common types
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Concurrency & Thread Safety
|
||||||
|
|
||||||
|
Brainy is **designed** for concurrent access. A single instance can handle:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Multiple concurrent requests - all using same instance
|
||||||
|
app.get('/api/read/:id', async (req, res) => {
|
||||||
|
const brain = getBrain()
|
||||||
|
const entity = await brain.get(req.params.id) // Safe - no state mutation
|
||||||
|
res.json(entity)
|
||||||
|
})
|
||||||
|
|
||||||
|
app.post('/api/write', async (req, res) => {
|
||||||
|
const brain = getBrain()
|
||||||
|
const id = await brain.add(req.body) // Safe - internal locking
|
||||||
|
res.json({ id })
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Concurrency mechanisms:**
|
||||||
|
- ✅ **Read operations**: Lock-free (MVCC)
|
||||||
|
- ✅ **Write operations**: Internal write-ahead logging (WAL)
|
||||||
|
- ✅ **Cache**: Thread-safe LRU implementation
|
||||||
|
- ✅ **Indexes**: Concurrent reads, locked writes
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Production Checklist
|
||||||
|
|
||||||
|
### Before Deploying
|
||||||
|
|
||||||
|
- [ ] **Initialize Brainy on startup** (not per-request)
|
||||||
|
- [ ] **Configure cache size** based on memory
|
||||||
|
- [ ] **Only load needed augmentations**
|
||||||
|
- [ ] **Warm up critical indexes**
|
||||||
|
- [ ] **Add graceful shutdown handler**
|
||||||
|
- [ ] **Monitor cache hit rate**
|
||||||
|
|
||||||
|
### Code Review Checklist
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ❌ BAD - Instance per request
|
||||||
|
app.get('/api/route', async (req, res) => {
|
||||||
|
const brain = new Brainy(...) // RED FLAG!
|
||||||
|
await brain.init() // RED FLAG!
|
||||||
|
})
|
||||||
|
|
||||||
|
// ✅ GOOD - Singleton pattern
|
||||||
|
app.get('/api/route', async (req, res) => {
|
||||||
|
const brain = await getBrain() // Reuses instance ✓
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Monitoring & Metrics
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Add metrics endpoint
|
||||||
|
app.get('/api/metrics', (req, res) => {
|
||||||
|
const brain = getBrain()
|
||||||
|
|
||||||
|
res.json({
|
||||||
|
cache: {
|
||||||
|
size: brain.cache?.size || 0,
|
||||||
|
maxSize: brain.cache?.maxSize || 0,
|
||||||
|
hitRate: brain.metrics?.cacheHitRate || 0 // Target: >70%
|
||||||
|
},
|
||||||
|
storage: brain.storage.getStats(),
|
||||||
|
memory: {
|
||||||
|
heapUsed: Math.round(process.memoryUsage().heapUsed / 1024 / 1024),
|
||||||
|
heapTotal: Math.round(process.memoryUsage().heapTotal / 1024 / 1024)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key metrics to track:**
|
||||||
|
- **Cache hit rate**: Should be >70% after warm-up
|
||||||
|
- **Memory usage**: Should stay constant (~500MB for singleton)
|
||||||
|
- **Request latency**: Should be <10ms for cached entities
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Common Pitfalls
|
||||||
|
|
||||||
|
### 1. Creating instances in routes
|
||||||
|
```typescript
|
||||||
|
// ❌ NEVER do this
|
||||||
|
app.get('/api/entities', async (req, res) => {
|
||||||
|
const brain = new Brainy(...) // Creates new instance every time!
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Not awaiting initialization
|
||||||
|
```typescript
|
||||||
|
// ❌ Race condition - server starts before Brainy ready
|
||||||
|
app.listen(3000)
|
||||||
|
getBrain() // Async init happens AFTER server starts!
|
||||||
|
|
||||||
|
// ✅ Correct - wait for init
|
||||||
|
await getBrain()
|
||||||
|
app.listen(3000)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Multiple instances for different purposes
|
||||||
|
```typescript
|
||||||
|
// ❌ Wasteful - creates 2 instances
|
||||||
|
const readBrain = new Brainy(...)
|
||||||
|
const writeBrain = new Brainy(...)
|
||||||
|
|
||||||
|
// ✅ One instance handles both
|
||||||
|
const brain = new Brainy(...)
|
||||||
|
await brain.get(id) // Read
|
||||||
|
await brain.add(data) // Write
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Migration Guide
|
||||||
|
|
||||||
|
### Current (Anti-Pattern)
|
||||||
|
```typescript
|
||||||
|
// Probably in multiple route files
|
||||||
|
async function handler(req, res) {
|
||||||
|
const brain = new Brainy({ storage: { path: './brainy-data' } })
|
||||||
|
await brain.init()
|
||||||
|
// ... use brain
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 1: Create Singleton Module
|
||||||
|
```typescript
|
||||||
|
// lib/brainy.ts
|
||||||
|
let instance: Brainy | null = null
|
||||||
|
|
||||||
|
export async function getBrain(): Promise<Brainy> {
|
||||||
|
if (!instance) {
|
||||||
|
instance = new Brainy({ storage: { path: './brainy-data' } })
|
||||||
|
await instance.init()
|
||||||
|
}
|
||||||
|
return instance
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 2: Update Server Startup
|
||||||
|
```typescript
|
||||||
|
// server.ts
|
||||||
|
import { getBrain } from './lib/brainy'
|
||||||
|
|
||||||
|
async function startServer() {
|
||||||
|
// Initialize Brainy FIRST
|
||||||
|
await getBrain()
|
||||||
|
console.log('✅ Brainy initialized')
|
||||||
|
|
||||||
|
// THEN start server
|
||||||
|
app.listen(3000)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 3: Update All Routes
|
||||||
|
```typescript
|
||||||
|
// Before
|
||||||
|
async function handler(req, res) {
|
||||||
|
const brain = new Brainy(...) // Remove this
|
||||||
|
await brain.init() // Remove this
|
||||||
|
|
||||||
|
// ... rest of code
|
||||||
|
}
|
||||||
|
|
||||||
|
// After
|
||||||
|
import { getBrain } from './lib/brainy'
|
||||||
|
|
||||||
|
async function handler(req, res) {
|
||||||
|
const brain = await getBrain() // Add this
|
||||||
|
|
||||||
|
// ... rest of code stays same
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Expected results:**
|
||||||
|
- ✅ 40x memory reduction (20GB → 500MB)
|
||||||
|
- ✅ 30x faster requests (60ms → 2ms average)
|
||||||
|
- ✅ 80%+ cache hit rate
|
||||||
|
- ✅ Your service can scale to 1000s of requests/minute
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
**DO:**
|
||||||
|
- ✅ Initialize Brainy ONCE on server startup
|
||||||
|
- ✅ Share single instance across all requests
|
||||||
|
- ✅ Configure cache for your workload
|
||||||
|
- ✅ Monitor cache hit rate
|
||||||
|
- ✅ Handle graceful shutdown
|
||||||
|
|
||||||
|
**DON'T:**
|
||||||
|
- ❌ Create new Brainy instance per request
|
||||||
|
- ❌ Create multiple instances
|
||||||
|
- ❌ Start server before Brainy is initialized
|
||||||
|
- ❌ Load augmentations you don't use
|
||||||
|
|
||||||
|
**Result:** 40x less memory, 30x faster requests, Brainy optimizations actually work!
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Questions? Issues?**
|
||||||
|
- Report issues: https://github.com/soulcraftlabs/brainy/issues
|
||||||
298
docs/QUERY_OPERATORS.md
Normal file
298
docs/QUERY_OPERATORS.md
Normal file
|
|
@ -0,0 +1,298 @@
|
||||||
|
# Query Operators (BFO)
|
||||||
|
|
||||||
|
> Brainy Field Operators — the complete reference for `where` filters in `find()`.
|
||||||
|
|
||||||
|
All operators work with `find({ where: { ... } })` and filter on **metadata fields** (not `data`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Equality
|
||||||
|
|
||||||
|
| Operator | Alias | Description | Example |
|
||||||
|
|----------|-------|-------------|---------|
|
||||||
|
| `eq` | `equals` | Exact match | `{ status: { eq: 'active' } }` |
|
||||||
|
| `ne` | `notEquals` | Not equal | `{ status: { ne: 'deleted' } }` |
|
||||||
|
|
||||||
|
**Shorthand:** A bare value is treated as `equals`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// These are equivalent:
|
||||||
|
brain.find({ where: { status: 'active' } })
|
||||||
|
brain.find({ where: { status: { equals: 'active' } } })
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Comparison
|
||||||
|
|
||||||
|
| Operator | Alias | Description | Example |
|
||||||
|
|----------|-------|-------------|---------|
|
||||||
|
| `gt` | `greaterThan` | Greater than | `{ age: { gt: 18 } }` |
|
||||||
|
| `gte` | `greaterThanOrEqual` | Greater or equal | `{ score: { gte: 90 } }` |
|
||||||
|
| `lt` | `lessThan` | Less than | `{ price: { lt: 100 } }` |
|
||||||
|
| `lte` | `lessThanOrEqual` | Less or equal | `{ rating: { lte: 3 } }` |
|
||||||
|
| `between` | — | Inclusive range `[min, max]` | `{ year: { between: [2020, 2025] } }` |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Range query
|
||||||
|
const recent = await brain.find({
|
||||||
|
where: {
|
||||||
|
createdAt: { between: [Date.now() - 86400000, Date.now()] }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Array / Set
|
||||||
|
|
||||||
|
| Operator | Alias | Description | Example |
|
||||||
|
|----------|-------|-------------|---------|
|
||||||
|
| `oneOf` | `in` | Value is one of the given options | `{ color: { oneOf: ['red', 'blue'] } }` |
|
||||||
|
| `noneOf` | — | Value is NOT one of the given options | `{ status: { noneOf: ['deleted', 'archived'] } }` |
|
||||||
|
| `contains` | — | Array field contains value | `{ tags: { contains: 'ai' } }` |
|
||||||
|
| `excludes` | — | Array field does NOT contain value | `{ tags: { excludes: 'spam' } }` |
|
||||||
|
| `hasAll` | — | Array field contains ALL listed values | `{ skills: { hasAll: ['js', 'ts'] } }` |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Find entities tagged with 'ai'
|
||||||
|
const aiEntities = await brain.find({
|
||||||
|
where: { tags: { contains: 'ai' } }
|
||||||
|
})
|
||||||
|
|
||||||
|
// Find entities of specific types
|
||||||
|
const people = await brain.find({
|
||||||
|
where: { noun: { oneOf: ['Person', 'Agent'] } }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Existence
|
||||||
|
|
||||||
|
| Operator | Description | Example |
|
||||||
|
|----------|-------------|---------|
|
||||||
|
| `exists: true` | Field exists (has any value) | `{ email: { exists: true } }` |
|
||||||
|
| `exists: false` | Field does NOT exist | `{ email: { exists: false } }` |
|
||||||
|
| `missing: true` | Field does NOT exist (alias for `exists: false`) | `{ email: { missing: true } }` |
|
||||||
|
| `missing: false` | Field exists (alias for `exists: true`) | `{ email: { missing: false } }` |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Find entities that have an email field
|
||||||
|
const withEmail = await brain.find({
|
||||||
|
where: { email: { exists: true } }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Pattern (In-Memory Only)
|
||||||
|
|
||||||
|
These operators work via the in-memory filter path. They are applied **after** the indexed query, so use them with other indexed operators for best performance.
|
||||||
|
|
||||||
|
| Operator | Description | Example |
|
||||||
|
|----------|-------------|---------|
|
||||||
|
| `matches` | Regex or string pattern match | `{ name: { matches: /^Dr\./ } }` |
|
||||||
|
| `startsWith` | String prefix | `{ name: { startsWith: 'John' } }` |
|
||||||
|
| `endsWith` | String suffix | `{ email: { endsWith: '@gmail.com' } }` |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const doctors = await brain.find({
|
||||||
|
where: {
|
||||||
|
type: NounType.Person, // Indexed — fast
|
||||||
|
name: { startsWith: 'Dr.' } // In-memory — applied after
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Logical
|
||||||
|
|
||||||
|
Combine multiple conditions:
|
||||||
|
|
||||||
|
| Operator | Description | Example |
|
||||||
|
|----------|-------------|---------|
|
||||||
|
| `allOf` | ALL sub-filters must match (AND) | `{ allOf: [{ status: 'active' }, { role: 'admin' }] }` |
|
||||||
|
| `anyOf` | ANY sub-filter must match (OR) | `{ anyOf: [{ role: 'admin' }, { role: 'owner' }] }` |
|
||||||
|
| `not` | Invert a filter | `{ not: { status: 'deleted' } }` |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Complex OR query
|
||||||
|
const adminsOrOwners = await brain.find({
|
||||||
|
where: {
|
||||||
|
anyOf: [
|
||||||
|
{ role: 'admin' },
|
||||||
|
{ role: 'owner' }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// NOT query
|
||||||
|
const notDeleted = await brain.find({
|
||||||
|
where: {
|
||||||
|
not: { status: 'deleted' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Combined AND + OR
|
||||||
|
const results = await brain.find({
|
||||||
|
where: {
|
||||||
|
allOf: [
|
||||||
|
{ department: 'engineering' },
|
||||||
|
{ anyOf: [
|
||||||
|
{ level: 'senior' },
|
||||||
|
{ yearsExperience: { greaterThan: 5 } }
|
||||||
|
]}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Indexed vs In-Memory Operators
|
||||||
|
|
||||||
|
Brainy's MetadataIndex supports a subset of operators natively for O(1) field lookups. Other operators fall back to in-memory filtering.
|
||||||
|
|
||||||
|
| Operator | MetadataIndex (Indexed) | In-Memory Fallback |
|
||||||
|
|----------|:-----------------------:|:------------------:|
|
||||||
|
| `equals` / `eq` | Yes | Yes |
|
||||||
|
| `notEquals` / `ne` | — | Yes |
|
||||||
|
| `greaterThan` / `gt` | Yes | Yes |
|
||||||
|
| `greaterThanOrEqual` / `gte` | Yes | Yes |
|
||||||
|
| `lessThan` / `lt` | Yes | Yes |
|
||||||
|
| `lessThanOrEqual` / `lte` | Yes | Yes |
|
||||||
|
| `between` | Yes | Yes |
|
||||||
|
| `oneOf` / `in` | Yes | Yes |
|
||||||
|
| `noneOf` | — | Yes |
|
||||||
|
| `contains` | Yes | Yes |
|
||||||
|
| `exists` / `missing` | Yes | Yes |
|
||||||
|
| `matches` | — | Yes |
|
||||||
|
| `startsWith` | — | Yes |
|
||||||
|
| `endsWith` | — | Yes |
|
||||||
|
| `allOf` | Partial | Yes |
|
||||||
|
| `anyOf` | Partial | Yes |
|
||||||
|
| `not` | — | Yes |
|
||||||
|
|
||||||
|
**Performance tip:** Combine indexed operators (equals, greaterThan, oneOf, between, contains, exists) with pattern operators for optimal speed — the index narrows results first, then patterns filter in memory.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Practical Examples
|
||||||
|
|
||||||
|
### Filter by entity type
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Using the type shorthand (recommended)
|
||||||
|
brain.find({ type: NounType.Person })
|
||||||
|
|
||||||
|
// Using where.noun directly
|
||||||
|
brain.find({ where: { noun: NounType.Person } })
|
||||||
|
|
||||||
|
// Multiple types
|
||||||
|
brain.find({ type: [NounType.Person, NounType.Agent] })
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by subtype
|
||||||
|
|
||||||
|
`subtype` is a top-level standard field — takes the column-store fast path, not the metadata fallback. Pair with `type` for the typical "Person who is an employee" query:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Equality on subtype:
|
||||||
|
brain.find({ type: NounType.Person, subtype: 'employee' })
|
||||||
|
|
||||||
|
// Set membership:
|
||||||
|
brain.find({ type: NounType.Person, subtype: ['employee', 'contractor'] })
|
||||||
|
|
||||||
|
// Operator-form predicates use `where`:
|
||||||
|
brain.find({
|
||||||
|
type: NounType.Person,
|
||||||
|
where: { subtype: { exists: true } }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
See the **[Subtypes & Facets guide](./guides/subtypes-and-facets.md)** for the full surface.
|
||||||
|
|
||||||
|
### Filter relationships by subtype (7.30+)
|
||||||
|
|
||||||
|
Verbs are first-class peers — `related()` and graph traversal both honor subtype filters on the fast path:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Filter relationships by VerbType subtype
|
||||||
|
const direct = await brain.related({
|
||||||
|
from: ceoId,
|
||||||
|
type: VerbType.ReportsTo,
|
||||||
|
subtype: 'direct'
|
||||||
|
})
|
||||||
|
|
||||||
|
// Set membership on verb subtype
|
||||||
|
const all = await brain.related({
|
||||||
|
from: ceoId,
|
||||||
|
type: VerbType.ReportsTo,
|
||||||
|
subtype: ['direct', 'dotted-line']
|
||||||
|
})
|
||||||
|
|
||||||
|
// Graph traversal — subtype filters traversal edges (depth-1 in 7.30 JS path;
|
||||||
|
// multi-hop subtype filtering lands on Cor native)
|
||||||
|
const reports = await brain.find({
|
||||||
|
connected: {
|
||||||
|
from: ceoId,
|
||||||
|
via: VerbType.ReportsTo,
|
||||||
|
subtype: 'direct',
|
||||||
|
depth: 1
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Combine semantic search with filters
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const results = await brain.find({
|
||||||
|
query: 'machine learning engineer', // Semantic search (on data)
|
||||||
|
type: NounType.Person, // Type filter (indexed)
|
||||||
|
where: {
|
||||||
|
department: 'engineering', // Exact match (indexed)
|
||||||
|
yearsExperience: { greaterThan: 3 } // Range filter (indexed)
|
||||||
|
},
|
||||||
|
limit: 10
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Temporal queries
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const lastWeek = Date.now() - 7 * 24 * 60 * 60 * 1000
|
||||||
|
const recentEntities = await brain.find({
|
||||||
|
where: {
|
||||||
|
createdAt: { greaterThan: lastWeek }
|
||||||
|
},
|
||||||
|
orderBy: 'createdAt',
|
||||||
|
order: 'desc',
|
||||||
|
limit: 50
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Graph + metadata combination
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const results = await brain.find({
|
||||||
|
connected: {
|
||||||
|
from: teamLeadId,
|
||||||
|
via: VerbType.WorksWith,
|
||||||
|
depth: 2
|
||||||
|
},
|
||||||
|
where: {
|
||||||
|
role: { oneOf: ['engineer', 'designer'] },
|
||||||
|
active: true
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## See Also
|
||||||
|
|
||||||
|
- [Data Model](./DATA_MODEL.md) — Entity structure, data vs metadata
|
||||||
|
- [API Reference](./api/README.md) — Complete API documentation
|
||||||
|
- [Find System](./FIND_SYSTEM.md) — Natural language find() details
|
||||||
211
docs/README.md
211
docs/README.md
|
|
@ -1,120 +1,129 @@
|
||||||
# Brainy Documentation
|
# Brainy Documentation
|
||||||
|
|
||||||
Welcome to the comprehensive documentation for Brainy, the multi-dimensional AI database with Triple Intelligence Engine.
|
> The multi-dimensional AI database with Triple Intelligence — vector search, graph traversal, and metadata filtering in one unified API.
|
||||||
|
|
||||||
## 📊 Implementation Status
|
## Quick Start
|
||||||
|
|
||||||
- ✅ **Production Ready**: Core features working today
|
|
||||||
- 🚧 **In Development**: Features coming soon
|
|
||||||
- 📅 **Roadmap**: See [ROADMAP.md](../ROADMAP.md)
|
|
||||||
|
|
||||||
## Quick Links
|
|
||||||
|
|
||||||
### Getting Started
|
|
||||||
- [Quick Start Guide](./guides/getting-started.md) - Get up and running in minutes
|
|
||||||
- [Enterprise for Everyone](./guides/enterprise-for-everyone.md) - **No limits, no tiers, everything free**
|
|
||||||
- [Natural Language Queries](./guides/natural-language.md) - Query with plain English
|
|
||||||
|
|
||||||
### Core Concepts
|
|
||||||
- [Zero Configuration](./architecture/zero-config.md) - **Auto-adapts to any environment**
|
|
||||||
- [Noun-Verb Taxonomy](./architecture/noun-verb-taxonomy.md) - **Revolutionary data model**
|
|
||||||
- [Triple Intelligence](./architecture/triple-intelligence.md) - Unified query system
|
|
||||||
- [Architecture Overview](./architecture/overview.md) - System design
|
|
||||||
|
|
||||||
### API Documentation
|
|
||||||
- [API Reference](./api/README.md) - Complete API documentation
|
|
||||||
- [TypeScript Types](./api/types.md) - Type definitions
|
|
||||||
|
|
||||||
### Advanced Topics
|
|
||||||
- [Augmentations System](./architecture/augmentations.md) - **Enterprise plugins & neural import**
|
|
||||||
- [Storage Architecture](./architecture/storage.md) - Storage adapter system
|
|
||||||
- [Performance Tuning](./guides/performance.md) - Optimization guide
|
|
||||||
- [Migration Guide](../MIGRATION.md) - Upgrading from 1.x
|
|
||||||
|
|
||||||
## What is Brainy?
|
|
||||||
|
|
||||||
Brainy is a next-generation AI database that combines:
|
|
||||||
- **Vector Search**: Semantic similarity using HNSW indexing
|
|
||||||
- **Graph Relationships**: Complex relationship mapping and traversal
|
|
||||||
- **Field Filtering**: Precise metadata filtering with O(1) lookups
|
|
||||||
- **Natural Language**: Query in plain English
|
|
||||||
|
|
||||||
## Key Features
|
|
||||||
|
|
||||||
### 🧠 Triple Intelligence Engine
|
|
||||||
All three intelligence types (vector, graph, field) work together in every query for optimal results.
|
|
||||||
|
|
||||||
### 📝 Noun-Verb Taxonomy
|
|
||||||
Model your data naturally as entities (nouns) and relationships (verbs) - no complex schemas needed.
|
|
||||||
|
|
||||||
### 🌍 Natural Language Queries
|
|
||||||
Ask questions in plain English and Brainy understands your intent:
|
|
||||||
```typescript
|
|
||||||
await brain.find("recent articles about AI with high ratings")
|
|
||||||
```
|
|
||||||
|
|
||||||
### ⚡ Production Ready
|
|
||||||
- Universal storage (FileSystem, S3, OPFS, Memory)
|
|
||||||
- Zero configuration with intelligent defaults
|
|
||||||
- Full TypeScript support
|
|
||||||
- Cross-platform compatibility
|
|
||||||
|
|
||||||
## Quick Example
|
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
import { BrainyData } from 'brainy'
|
import { Brainy, NounType, VerbType } from '@soulcraft/brainy'
|
||||||
|
|
||||||
// Initialize
|
const brain = new Brainy()
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
await brain.init()
|
||||||
|
|
||||||
// Add entities (nouns)
|
// Add entities — data is embedded for semantic search, metadata is indexed for filtering
|
||||||
const articleId = await brain.addNoun("Revolutionary AI Breakthrough", {
|
const id = await brain.add({
|
||||||
type: "article",
|
data: 'Revolutionary AI Breakthrough',
|
||||||
category: "technology",
|
type: NounType.Document,
|
||||||
rating: 4.8
|
metadata: { category: 'technology', rating: 4.8 }
|
||||||
})
|
})
|
||||||
|
|
||||||
const authorId = await brain.addNoun("Dr. Sarah Chen", {
|
// Search with Triple Intelligence
|
||||||
type: "person",
|
const results = await brain.find({
|
||||||
role: "researcher"
|
query: 'artificial intelligence', // Semantic search (on data)
|
||||||
|
where: { rating: { greaterThan: 4.0 } }, // Metadata filter
|
||||||
|
connected: { from: authorId, depth: 2 } // Graph traversal
|
||||||
})
|
})
|
||||||
|
|
||||||
// Create relationships (verbs)
|
|
||||||
await brain.addVerb(authorId, articleId, "authored", {
|
|
||||||
date: "2024-01-15",
|
|
||||||
contribution: "primary"
|
|
||||||
})
|
|
||||||
|
|
||||||
// Query naturally
|
|
||||||
const results = await brain.find("highly rated technology articles by researchers")
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Documentation Structure
|
---
|
||||||
|
|
||||||
```
|
## Core Documentation
|
||||||
docs/
|
|
||||||
├── README.md # This file
|
|
||||||
├── guides/ # User guides
|
|
||||||
│ ├── getting-started.md # Quick start guide
|
|
||||||
│ ├── natural-language.md # NLP query guide
|
|
||||||
│ └── performance.md # Performance tuning
|
|
||||||
├── architecture/ # Technical architecture
|
|
||||||
│ ├── overview.md # System overview
|
|
||||||
│ ├── noun-verb-taxonomy.md # Data model
|
|
||||||
│ ├── triple-intelligence.md # Query system
|
|
||||||
│ └── storage.md # Storage layer
|
|
||||||
└── api/ # API documentation
|
|
||||||
├── README.md # API overview
|
|
||||||
├── brainy-data.md # Main class
|
|
||||||
└── types.md # TypeScript types
|
|
||||||
```
|
|
||||||
|
|
||||||
## Community
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| **[API Reference](./api/README.md)** | Complete API documentation — **start here** |
|
||||||
|
| **[Data Model](./DATA_MODEL.md)** | Entity structure, data vs metadata, storage fields |
|
||||||
|
| **[Query Operators](./QUERY_OPERATORS.md)** | All BFO operators with examples and indexed/in-memory matrix |
|
||||||
|
| [Find System](./FIND_SYSTEM.md) | Natural language `find()` and hybrid search details |
|
||||||
|
| [Consistency Model](./concepts/consistency-model.md) | The Db API guarantees — snapshot isolation, atomic transactions, time travel |
|
||||||
|
|
||||||
- **GitHub**: [github.com/brainy-org/brainy](https://github.com/brainy-org/brainy)
|
---
|
||||||
- **Issues**: [Report bugs or request features](https://github.com/brainy-org/brainy/issues)
|
|
||||||
- **Discussions**: [Join the conversation](https://github.com/brainy-org/brainy/discussions)
|
## Architecture
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Architecture Overview](./architecture/overview.md) | High-level system design |
|
||||||
|
| [Triple Intelligence](./architecture/triple-intelligence.md) | Vector + Graph + Metadata unified query |
|
||||||
|
| [Noun-Verb Taxonomy](./architecture/noun-verb-taxonomy.md) | 42 nouns + 127 verbs type system |
|
||||||
|
| [Stage 3 Canonical Taxonomy](./STAGE3-CANONICAL-TAXONOMY.md) | Complete type reference |
|
||||||
|
| [Storage Architecture](./architecture/storage-architecture.md) | Storage adapters and optimization |
|
||||||
|
| [Index Architecture](./architecture/index-architecture.md) | Vector, Graph, and Metadata indexing |
|
||||||
|
| [Zero Configuration](./architecture/zero-config.md) | Auto-adapts to any environment |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Virtual Filesystem (VFS)
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [VFS Quick Start](./vfs/QUICK_START.md) | Get started in 30 seconds |
|
||||||
|
| [VFS Core](./vfs/VFS_CORE.md) | Core concepts and architecture |
|
||||||
|
| [VFS API Guide](./vfs/VFS_API_GUIDE.md) | Complete VFS API reference |
|
||||||
|
| [Common Patterns](./vfs/COMMON_PATTERNS.md) | VFS usage patterns |
|
||||||
|
|
||||||
|
See [vfs/](./vfs/) for the complete VFS documentation set.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Guides
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Import Anything](./guides/import-anything.md) | CSV, Excel, PDF, URL imports |
|
||||||
|
| [Snapshots & Time Travel](./guides/snapshots-and-time-travel.md) | Backups, restore, what-if analysis, audit trails |
|
||||||
|
| [Natural Language](./guides/natural-language.md) | Query in plain English |
|
||||||
|
| [Neural API](./guides/neural-api.md) | AI-powered features |
|
||||||
|
| [Enterprise for Everyone](./guides/enterprise-for-everyone.md) | No limits, no tiers |
|
||||||
|
| [Framework Integration](./guides/framework-integration.md) | React, Vue, Angular, Svelte |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Storage & Deployment
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Storage Architecture](./architecture/storage-architecture.md) | Filesystem and memory adapters, on-disk artifact layout, operator-layer backup |
|
||||||
|
| [Capacity Planning](./operations/capacity-planning.md) | Scale to millions of entities |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Plugins
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Plugins](./PLUGINS.md) | Plugin system overview — providers, `plugins` config, `brain.use()` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Performance & Scaling
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Performance](./PERFORMANCE.md) | Optimization techniques |
|
||||||
|
| [Scaling](./SCALING.md) | Scale to billions of entities |
|
||||||
|
| [Batching](./BATCHING.md) | Batch operations guide |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Migration & Reference
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [v3 to v4 Migration](./MIGRATION-V3-TO-V4.md) | Upgrade guide |
|
||||||
|
| [Release Guide](./RELEASE-GUIDE.md) | How to release new versions |
|
||||||
|
| [Production Architecture](./PRODUCTION_SERVICE_ARCHITECTURE.md) | Ops reference |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Internal
|
||||||
|
|
||||||
|
| Document | Description |
|
||||||
|
|----------|-------------|
|
||||||
|
| [Audit Report](./internal/AUDIT_REPORT.md) | Feature audit |
|
||||||
|
| [Honest Status](./internal/HONEST_STATUS.md) | Actual implementation status |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|
|
||||||
131
docs/RELEASE-GUIDE.md
Normal file
131
docs/RELEASE-GUIDE.md
Normal file
|
|
@ -0,0 +1,131 @@
|
||||||
|
# Brainy Release Guide
|
||||||
|
|
||||||
|
## Standard Semantic Versioning (Industry Guidelines)
|
||||||
|
|
||||||
|
### Official SemVer 2.0.0 says:
|
||||||
|
- **MAJOR**: Incompatible API changes (breaking changes)
|
||||||
|
- **MINOR**: Add functionality in backwards compatible manner
|
||||||
|
- **PATCH**: Backwards compatible bug fixes
|
||||||
|
|
||||||
|
## Our Approach for Brainy (More Conservative)
|
||||||
|
|
||||||
|
### We intentionally diverge from strict SemVer:
|
||||||
|
- **PATCH (2.3.0 → 2.3.1)**: Bug fixes, internal improvements, dependency updates
|
||||||
|
- **MINOR (2.3.0 → 2.4.0)**: New features, API changes, enhancements
|
||||||
|
- **MAJOR (3.0.0)**: Reserved for strategic platform shifts (manual decision)
|
||||||
|
|
||||||
|
### Why We Do This:
|
||||||
|
1. **User Trust**: Major versions signal huge changes and scare users
|
||||||
|
2. **Adoption**: People hesitate to upgrade major versions
|
||||||
|
3. **Flexibility**: We can evolve the API without version explosion
|
||||||
|
4. **Industry Practice**: Many successful projects (React, Vue) do this
|
||||||
|
|
||||||
|
## CRITICAL: Never Use "BREAKING CHANGE"
|
||||||
|
|
||||||
|
**"BREAKING CHANGE" in commits = Automatic major version = BAD!**
|
||||||
|
- Even if removing methods, just use `feat:` or `refactor:`
|
||||||
|
- Major versions are MANUAL decisions: `npm run release:major`
|
||||||
|
- Most API changes can be handled gracefully in minor versions
|
||||||
|
|
||||||
|
## Commit Message Guidelines
|
||||||
|
|
||||||
|
### ✅ CORRECT Examples:
|
||||||
|
```bash
|
||||||
|
# New features → MINOR bump
|
||||||
|
git commit -m "feat: add new model delivery system"
|
||||||
|
|
||||||
|
# Bug fixes → PATCH bump
|
||||||
|
git commit -m "fix: resolve model download timeout"
|
||||||
|
|
||||||
|
# Internal improvements → PATCH bump
|
||||||
|
git commit -m "refactor: simplify model manager logic"
|
||||||
|
git commit -m "perf: optimize model caching"
|
||||||
|
git commit -m "chore: remove unused dependency"
|
||||||
|
```
|
||||||
|
|
||||||
|
### ❌ AVOID These Mistakes:
|
||||||
|
```bash
|
||||||
|
# DON'T use BREAKING CHANGE for internal changes
|
||||||
|
git commit -m "feat: improve model delivery
|
||||||
|
|
||||||
|
BREAKING CHANGE: removed tar-stream dependency" # WRONG! This triggers
|
||||||
|
```
|
||||||
|
|
||||||
|
## Release Workflow Checklist
|
||||||
|
|
||||||
|
### Before Committing:
|
||||||
|
- [ ] Review commit message - no "BREAKING CHANGE" unless API changes
|
||||||
|
- [ ] Consider: Will users need to change their code? If NO → Not breaking
|
||||||
|
|
||||||
|
### Release Commands:
|
||||||
|
```bash
|
||||||
|
# Let standard-version figure it out from commits
|
||||||
|
npm run release # Recommended - auto-detects version
|
||||||
|
|
||||||
|
# Or be explicit:
|
||||||
|
npm run release:patch # 2.4.0 → 2.4.1 (fixes)
|
||||||
|
npm run release:minor # 2.4.0 → 2.5.0 (features)
|
||||||
|
npm run release:major # 2.4.0 → 3.0.0 (API changes only!)
|
||||||
|
```
|
||||||
|
|
||||||
|
### After Release:
|
||||||
|
```bash
|
||||||
|
git push --follow-tags origin main
|
||||||
|
npm publish
|
||||||
|
gh release create $(git describe --tags --abbrev=0) --generate-notes
|
||||||
|
```
|
||||||
|
|
||||||
|
## When to Use Major Version (3.0.0)
|
||||||
|
|
||||||
|
ONLY when we make changes like:
|
||||||
|
- Removing methods from the public API
|
||||||
|
- Changing method signatures (parameters, return types)
|
||||||
|
- Renaming public methods
|
||||||
|
- Changing default behaviors that break existing code
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
- ❌ `search(query, limit, options)` → `search(query, options)` (major)
|
||||||
|
- ✅ Adding `find()` method (minor - doesn't break existing code)
|
||||||
|
- ✅ Internal refactoring (patch - users don't see it)
|
||||||
|
|
||||||
|
## Quick Decision Tree
|
||||||
|
|
||||||
|
1. **Does this fix a bug?** → PATCH (fix:)
|
||||||
|
2. **Does this add new functionality?** → MINOR (feat:)
|
||||||
|
3. **Will users' existing code break?** → MAJOR (with BREAKING CHANGE)
|
||||||
|
4. **Is it internal/maintenance?** → PATCH (chore:/refactor:/perf:)
|
||||||
|
|
||||||
|
## Emergency: If Wrong Version is Released
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Deprecate wrong version on npm
|
||||||
|
npm deprecate @soulcraft/brainy@X.X.X "Incorrect version - use Y.Y.Y"
|
||||||
|
|
||||||
|
# 2. Fix version in package.json
|
||||||
|
# 3. Republish correct version
|
||||||
|
npm publish
|
||||||
|
|
||||||
|
# 4. Delete wrong GitHub tag/release
|
||||||
|
git push origin :vX.X.X
|
||||||
|
gh release delete vX.X.X --yes
|
||||||
|
|
||||||
|
# 5. Create correct tag/release
|
||||||
|
git tag vY.Y.Y
|
||||||
|
git push --tags
|
||||||
|
gh release create vY.Y.Y --generate-notes
|
||||||
|
```
|
||||||
|
|
||||||
|
## Remember:
|
||||||
|
- **Most releases should be MINOR or PATCH**
|
||||||
|
- **Major versions should be RARE**
|
||||||
|
- **When in doubt, it's probably MINOR**
|
||||||
|
- **NEVER use "BREAKING CHANGE" for internal changes**
|
||||||
|
## Hard Ordering Constraints (check before EVERY release)
|
||||||
|
|
||||||
|
- **Embedding model changes are SEQUENCED, not free.** No release may change the
|
||||||
|
embedding model (or its quantization/dimensions) before **vector model-version
|
||||||
|
stamping + hard-error-on-mismatch** ships. Stored vectors carry no model version
|
||||||
|
today; mixing vectors from two models silently corrupts every similarity
|
||||||
|
comparison. If a model bump is ever proposed, the stamping work moves ahead of
|
||||||
|
it in the schedule — coordinate with the native provider so both engines stamp
|
||||||
|
and enforce identically. (Registered with the native-provider team 2026-07-07.)
|
||||||
239
docs/SCALING.md
Normal file
239
docs/SCALING.md
Normal file
|
|
@ -0,0 +1,239 @@
|
||||||
|
# Brainy Scaling Guide
|
||||||
|
|
||||||
|
> **One Line Summary**: Single-node by design — Brainy scales by getting the most out of one machine plus operator-layer backup.
|
||||||
|
|
||||||
|
## Table of Contents
|
||||||
|
- [Quick Start](#quick-start)
|
||||||
|
- [How Brainy Scales](#how-brainy-scales)
|
||||||
|
- [Storage Configurations](#storage-configurations)
|
||||||
|
- [Scaling Patterns](#scaling-patterns)
|
||||||
|
- [Real World Examples](#real-world-examples)
|
||||||
|
|
||||||
|
## Quick Start
|
||||||
|
|
||||||
|
### In-Memory
|
||||||
|
```typescript
|
||||||
|
import Brainy from '@soulcraft/brainy'
|
||||||
|
const brain = new Brainy({ storage: { type: 'memory' } })
|
||||||
|
```
|
||||||
|
|
||||||
|
### On-Disk (Default for Node)
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: './brainy-data' }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
## How Brainy Scales
|
||||||
|
|
||||||
|
Brainy 8.0 is a **single-node library**. There is no cluster, no peer discovery, no S3 coordination. Scaling means:
|
||||||
|
|
||||||
|
- **Up**: give the process more RAM, CPU, and IOPS
|
||||||
|
- **Out**: stand up multiple independent Brainy instances behind your own service layer
|
||||||
|
- **Cold storage**: snapshot the on-disk artifact off-site so you can rehydrate elsewhere
|
||||||
|
|
||||||
|
The three knobs that matter most:
|
||||||
|
|
||||||
|
1. **`config.vector.recall`** — `'fast'`, `'balanced'`, or `'accurate'` (default `'balanced'`)
|
||||||
|
2. **`config.vector.persistMode`** — `'immediate'` for durability, `'deferred'` for throughput
|
||||||
|
|
||||||
|
The native vector provider (via the optional `@soulcraft/cor` package) extends this with a higher-performing index — and its own at-scale acceleration such as on-disk compressed indexing — when installed.
|
||||||
|
|
||||||
|
## Measured Performance
|
||||||
|
|
||||||
|
Numbers below are **measured** by `tests/benchmarks/find-composition-scale.js` (a single
|
||||||
|
Node 22 process, in-memory storage, 384-dim vectors, `balanced` recall). They are the
|
||||||
|
open-core (pure-TypeScript) path — what you get from `@soulcraft/brainy` with no native
|
||||||
|
provider installed. Run it yourself: `node --max-old-space-size=8192 tests/benchmarks/find-composition-scale.js 100000`.
|
||||||
|
|
||||||
|
`find()` query latency, p50 / p95 (200 queries each):
|
||||||
|
|
||||||
|
| Query | 5,000 entities | 100,000 entities |
|
||||||
|
|---|---|---|
|
||||||
|
| Vector similarity (`{ vector }`) | 0.8 / 1.3 ms | 1.4 / 4.7 ms |
|
||||||
|
| Graph 1-hop (`{ connected }`) | 0.5 / 0.7 ms | 0.7 / 0.8 ms |
|
||||||
|
| Metadata filter (`{ where }`, low-selectivity) | 0.7 / 1.2 ms | 23.5 / 30.1 ms |
|
||||||
|
| Vector + metadata | 7.7 / 8.3 ms | 78.8 / 93.8 ms |
|
||||||
|
|
||||||
|
What the shape tells you:
|
||||||
|
|
||||||
|
- **Vector and graph lookups scale ~logarithmically** — they barely move from 5k to 100k,
|
||||||
|
because HNSW search is ~O(ef·log n) and graph adjacency is O(degree).
|
||||||
|
- **Metadata-filtered paths scale with the size of the match set, not the database.** The
|
||||||
|
benchmark's `category` filter matches ~10% of rows (10,000 at 100k); the cost is
|
||||||
|
materializing that candidate set and running the vector search *inside* it (`find()` does
|
||||||
|
metadata-first hard filtering, then ranks within the candidates — see
|
||||||
|
[How find works](./FIND_SYSTEM.md)). A **high-selectivity** filter (few matches) is far
|
||||||
|
cheaper; a 10%-of-everything filter is the worst case. This candidate-restricted search is
|
||||||
|
precisely the path the native provider accelerates (Rust roaring-bitmap candidate
|
||||||
|
intersection).
|
||||||
|
- **Composition is correct, not lossy.** Combining vector + metadata + graph returns exactly
|
||||||
|
the entities satisfying all constraints — verified by
|
||||||
|
`tests/integration/find-triple-composition.test.ts`.
|
||||||
|
|
||||||
|
Memory: ~62 KB resident per entity at 100k (6.2 GB RSS for 100k × 384-dim including the HNSW
|
||||||
|
graph, metadata index, and 100k edges).
|
||||||
|
|
||||||
|
**Scale ceiling (open-core).** The pure-JS HNSW *build* cost (~100 inserts/s at 384-dim on
|
||||||
|
one core) makes the in-process open-core path most appropriate up to ~10⁵–10⁶ entities.
|
||||||
|
*Query* latency stays low well beyond that, but for the 10⁸–10¹⁰ regime install the native
|
||||||
|
provider (`@soulcraft/cor`, on-disk DiskANN) — same API, no code change. _Projected from
|
||||||
|
the two measured points, vector p50 at 1M is ~2 ms; metadata-heavy composition grows with
|
||||||
|
match-set size and is the path to move onto the native provider first._
|
||||||
|
|
||||||
|
## Storage Configurations
|
||||||
|
|
||||||
|
### Filesystem (Recommended for Production)
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: {
|
||||||
|
type: 'filesystem',
|
||||||
|
path: '/var/lib/brainy'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
- Stores everything in a sharded JSON tree under `path`
|
||||||
|
- Atomic writes via rename
|
||||||
|
- Survives process restarts
|
||||||
|
- Snapshot it off-site with `gsutil rsync`, `aws s3 sync`, `rclone`, or `tar` from your scheduler
|
||||||
|
|
||||||
|
### Memory
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({ storage: { type: 'memory' } })
|
||||||
|
```
|
||||||
|
- Zero I/O, fastest possible
|
||||||
|
- No persistence — process exit discards everything
|
||||||
|
- Use for tests and ephemeral caches
|
||||||
|
|
||||||
|
### Auto
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'auto', path: './data' }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
- Picks `filesystem` when running on Node with a writable `path`
|
||||||
|
- Falls back to `memory` otherwise
|
||||||
|
|
||||||
|
## Scaling Patterns
|
||||||
|
|
||||||
|
### Stage 1: Prototype (Memory)
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({ storage: { type: 'memory' } })
|
||||||
|
// Development, tests, <100K items
|
||||||
|
```
|
||||||
|
|
||||||
|
### Stage 2: Production (Filesystem)
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: '/var/lib/brainy' }
|
||||||
|
})
|
||||||
|
// Most production workloads up to ~10M entities on a single host
|
||||||
|
```
|
||||||
|
|
||||||
|
### Stage 3: Higher Throughput (Tune the Vector Index)
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: '/var/lib/brainy' },
|
||||||
|
vector: {
|
||||||
|
recall: 'fast', // Trade recall for latency
|
||||||
|
persistMode: 'deferred' // Batch persistence
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Stage 4: Multi-Instance (Operator-Layer)
|
||||||
|
Run multiple Brainy processes behind your own routing/service layer. Each process owns its own `path`. Sync each artifact off-site independently. Brainy itself does not coordinate between processes.
|
||||||
|
|
||||||
|
## Real World Examples
|
||||||
|
|
||||||
|
### Example 1: Single-Node App With Backup
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: '/var/lib/brainy' }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
Schedule (cron / systemd timer):
|
||||||
|
```bash
|
||||||
|
*/15 * * * * rclone sync /var/lib/brainy remote:brainy-backup
|
||||||
|
```
|
||||||
|
|
||||||
|
### Example 2: Tests
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({ storage: { type: 'memory' } })
|
||||||
|
// Fast, no cleanup needed between runs
|
||||||
|
```
|
||||||
|
|
||||||
|
### Example 3: Multi-Tenant Service
|
||||||
|
Spin up one Brainy instance per tenant, each in its own directory:
|
||||||
|
```typescript
|
||||||
|
function brainForTenant(tenantId: string) {
|
||||||
|
return new Brainy({
|
||||||
|
storage: {
|
||||||
|
type: 'filesystem',
|
||||||
|
path: `/var/lib/brainy/${tenantId}`
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
```
|
||||||
|
Your service layer handles routing and isolation; Brainy stays simple.
|
||||||
|
|
||||||
|
### Example 4: Higher Recall at Scale
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: '/var/lib/brainy' },
|
||||||
|
vector: {
|
||||||
|
recall: 'accurate'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
## Tuning Knobs Summary
|
||||||
|
|
||||||
|
| Setting | Values | When to change |
|
||||||
|
|---------|--------|----------------|
|
||||||
|
| `vector.recall` | `'fast'` / `'balanced'` / `'accurate'` | Trade recall for latency |
|
||||||
|
| `vector.persistMode` | `'immediate'` / `'deferred'` | Throughput vs. durability |
|
||||||
|
| `storage.cache.maxSize` | integer | Hot-path read cache size |
|
||||||
|
| `storage.cache.ttl` | ms | Cache freshness |
|
||||||
|
|
||||||
|
## Monitoring & Observability
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const stats = await brain.stats()
|
||||||
|
// {
|
||||||
|
// nounCount: 50000,
|
||||||
|
// verbCount: 80000,
|
||||||
|
// vectorIndex: { ... },
|
||||||
|
// storage: { used: '45GB' }
|
||||||
|
// }
|
||||||
|
```
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
### Issue: Slow queries
|
||||||
|
1. Switch to `vector.recall: 'fast'`
|
||||||
|
2. Increase the read cache (`storage.cache.maxSize`)
|
||||||
|
3. Consider the optional native vector provider via `@soulcraft/cor`
|
||||||
|
|
||||||
|
### Issue: Memory pressure
|
||||||
|
1. Reduce `storage.cache.maxSize`
|
||||||
|
2. Move to `vector.persistMode: 'deferred'` to batch writes
|
||||||
|
3. Consider the optional native vector provider via `@soulcraft/cor` for at-scale index acceleration
|
||||||
|
|
||||||
|
### Issue: Slow startup after a crash
|
||||||
|
1. Use `vector.persistMode: 'immediate'` so the index file stays in sync with storage
|
||||||
|
2. Verify backup integrity periodically
|
||||||
|
|
||||||
|
## Best Practices
|
||||||
|
|
||||||
|
1. **One process = one `path`** — never share a directory between processes
|
||||||
|
2. **Snapshot from your scheduler** — Brainy doesn't ship cloud SDKs; use `rclone` / `aws s3 sync` / `gsutil`
|
||||||
|
3. **Profile before tuning** — `recall: 'balanced'` is right for most workloads
|
||||||
|
4. **Install the native vector provider only when measured profiling shows it pays off**
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
- Brainy 8.0 is a **library**, not a cluster
|
||||||
|
- Storage adapters: `filesystem`, `memory`, `auto`
|
||||||
|
- Vector tuning: `recall`, `persistMode`
|
||||||
|
- Backup is an operator-layer concern — snapshot `path`
|
||||||
373
docs/STAGE3-CANONICAL-TAXONOMY.md
Normal file
373
docs/STAGE3-CANONICAL-TAXONOMY.md
Normal file
|
|
@ -0,0 +1,373 @@
|
||||||
|
# Brainy Stage 3: Canonical Taxonomy
|
||||||
|
|
||||||
|
**Status:** FINAL - This is the definitive, timeless taxonomy
|
||||||
|
**Total Types:** 169 (42 nouns + 127 verbs)
|
||||||
|
**Coverage:** 96-97% of all human knowledge
|
||||||
|
**Designed to last:** 20+ years without changes
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
- **Nouns:** 42 types
|
||||||
|
- **Verbs:** 127 types
|
||||||
|
- **Total:** 169 types
|
||||||
|
- **Previous (v5.x):** 71 types (31 nouns + 40 verbs)
|
||||||
|
- **Net Change:** +98 types (+11 nouns, +87 verbs)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Noun Types (42)
|
||||||
|
|
||||||
|
### Core Entity Types (7)
|
||||||
|
1. **person** - Individual human entities
|
||||||
|
2. **organization** - Collective entities, companies, institutions
|
||||||
|
3. **location** - Geographic and named spatial entities
|
||||||
|
4. **thing** - Discrete physical objects and artifacts
|
||||||
|
5. **concept** - Abstract ideas, principles, and intangibles
|
||||||
|
6. **event** - Temporal occurrences and happenings
|
||||||
|
7. **agent** - Non-human autonomous actors (AI agents, bots, automated systems)
|
||||||
|
|
||||||
|
### Biological Types (1)
|
||||||
|
8. **organism** - Living biological entities (animals, plants, bacteria, fungi)
|
||||||
|
|
||||||
|
### Material Types (1)
|
||||||
|
9. **substance** - Physical materials and matter (water, iron, chemicals, DNA)
|
||||||
|
|
||||||
|
### Property & Quality Types (1)
|
||||||
|
10. **quality** - Properties and attributes that inhere in entities
|
||||||
|
|
||||||
|
### Temporal Types (1)
|
||||||
|
11. **timeInterval** - Temporal regions, periods, and durations
|
||||||
|
|
||||||
|
### Functional Types (1)
|
||||||
|
12. **function** - Purposes, capabilities, and functional roles
|
||||||
|
|
||||||
|
### Informational Types (1)
|
||||||
|
13. **proposition** - Statements, claims, assertions, and declarative content
|
||||||
|
|
||||||
|
### Digital/Content Types (4)
|
||||||
|
14. **document** - Text-based files and written content
|
||||||
|
15. **media** - Non-text media files (audio, video, images)
|
||||||
|
16. **file** - Generic digital files and data blobs
|
||||||
|
17. **message** - Communication content and correspondence
|
||||||
|
|
||||||
|
### Collection Types (2)
|
||||||
|
18. **collection** - Groups and sets of items
|
||||||
|
19. **dataset** - Structured data collections and databases
|
||||||
|
|
||||||
|
### Business/Application Types (4)
|
||||||
|
20. **product** - Commercial products and offerings
|
||||||
|
21. **service** - Service offerings and intangible products
|
||||||
|
22. **task** - Actions, todos, and work items
|
||||||
|
23. **project** - Organized initiatives and programs
|
||||||
|
|
||||||
|
### Descriptive Types (6)
|
||||||
|
24. **process** - Workflows, procedures, and ongoing activities
|
||||||
|
25. **state** - Conditions, status, and situational contexts
|
||||||
|
26. **role** - Positions, responsibilities, and functional classifications
|
||||||
|
27. **language** - Natural and formal languages
|
||||||
|
28. **currency** - Monetary units and exchange mediums
|
||||||
|
29. **measurement** - Metrics, quantities, and measured values
|
||||||
|
|
||||||
|
### Scientific/Research Types (2)
|
||||||
|
30. **hypothesis** - Scientific theories, propositions, and conjectures
|
||||||
|
31. **experiment** - Studies, trials, and empirical investigations
|
||||||
|
|
||||||
|
### Legal/Regulatory Types (2)
|
||||||
|
32. **contract** - Legal agreements, terms, and binding documents
|
||||||
|
33. **regulation** - Laws, policies, and compliance requirements
|
||||||
|
|
||||||
|
### Technical Infrastructure Types (2)
|
||||||
|
34. **interface** - APIs, protocols, and connection points
|
||||||
|
35. **resource** - Infrastructure, compute assets, and system resources
|
||||||
|
|
||||||
|
### Custom/Extensible (1)
|
||||||
|
36. **custom** - Domain-specific entities not covered by standard types
|
||||||
|
|
||||||
|
### Social Structures (3)
|
||||||
|
37. **socialGroup** - Informal social groups and collectives
|
||||||
|
38. **institution** - Formal social structures and practices
|
||||||
|
39. **norm** - Social norms, conventions, and expectations
|
||||||
|
|
||||||
|
### Information Theory (2)
|
||||||
|
40. **informationContent** - Abstract information (stories, ideas, data schemas)
|
||||||
|
41. **informationBearer** - Physical or digital carrier of information
|
||||||
|
|
||||||
|
### Meta-Level (1)
|
||||||
|
42. **relationship** - Relationships as first-class entities for meta-level reasoning
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Verb Types (127)
|
||||||
|
|
||||||
|
### Foundational Ontological (3)
|
||||||
|
1. **instanceOf** - Individual to class relationship
|
||||||
|
2. **subclassOf** - Taxonomic hierarchy
|
||||||
|
3. **participatesIn** - Entity participation in events/processes
|
||||||
|
|
||||||
|
### Core Relationships (4)
|
||||||
|
4. **relatedTo** - Generic relationship (fallback)
|
||||||
|
5. **contains** - Containment relationship
|
||||||
|
6. **partOf** - Part-whole mereological relationship
|
||||||
|
7. **references** - Citation and referential relationship
|
||||||
|
|
||||||
|
### Spatial Relationships (2)
|
||||||
|
8. **locatedAt** - Spatial location relationship
|
||||||
|
9. **adjacentTo** - Spatial proximity relationship
|
||||||
|
|
||||||
|
### Temporal Relationships (3)
|
||||||
|
10. **precedes** - Temporal sequence (before)
|
||||||
|
11. **during** - Temporal containment
|
||||||
|
12. **occursAt** - Temporal location
|
||||||
|
|
||||||
|
### Causal & Dependency (5)
|
||||||
|
13. **causes** - Direct causal relationship
|
||||||
|
14. **enables** - Enablement without direct causation
|
||||||
|
15. **prevents** - Prevention relationship
|
||||||
|
16. **dependsOn** - Dependency relationship
|
||||||
|
17. **requires** - Necessity relationship
|
||||||
|
|
||||||
|
### Creation & Transformation (5)
|
||||||
|
18. **creates** - Creation relationship
|
||||||
|
19. **transforms** - Transformation relationship
|
||||||
|
20. **becomes** - State change relationship
|
||||||
|
21. **modifies** - Modification relationship
|
||||||
|
22. **consumes** - Consumption relationship
|
||||||
|
|
||||||
|
### Lifecycle Operations (1)
|
||||||
|
23. **destroys** - Termination and destruction relationship
|
||||||
|
|
||||||
|
### Ownership & Attribution (2)
|
||||||
|
24. **owns** - Ownership relationship
|
||||||
|
25. **attributedTo** - Attribution relationship
|
||||||
|
|
||||||
|
### Property & Quality (2)
|
||||||
|
26. **hasQuality** - Entity to quality attribution
|
||||||
|
27. **realizes** - Function realization relationship
|
||||||
|
|
||||||
|
### Effects & Experience (1)
|
||||||
|
28. **affects** - Patient/experiencer relationship
|
||||||
|
|
||||||
|
### Composition (2)
|
||||||
|
29. **composedOf** - Material composition
|
||||||
|
30. **inherits** - Inheritance relationship
|
||||||
|
|
||||||
|
### Social & Organizational (7)
|
||||||
|
31. **memberOf** - Membership relationship
|
||||||
|
32. **worksWith** - Professional collaboration
|
||||||
|
33. **friendOf** - Friendship relationship
|
||||||
|
34. **follows** - Following/subscription relationship
|
||||||
|
35. **likes** - Liking/favoriting relationship
|
||||||
|
36. **reportsTo** - Hierarchical reporting relationship
|
||||||
|
37. **mentors** - Mentorship relationship
|
||||||
|
38. **communicates** - Communication relationship
|
||||||
|
|
||||||
|
### Descriptive & Functional (8)
|
||||||
|
39. **describes** - Descriptive relationship
|
||||||
|
40. **defines** - Definition relationship
|
||||||
|
41. **categorizes** - Categorization relationship
|
||||||
|
42. **measures** - Measurement relationship
|
||||||
|
43. **evaluates** - Evaluation relationship
|
||||||
|
44. **uses** - Utilization relationship
|
||||||
|
45. **implements** - Implementation relationship
|
||||||
|
46. **extends** - Extension relationship
|
||||||
|
|
||||||
|
### Advanced Relationships (4)
|
||||||
|
47. **equivalentTo** - Equivalence/identity relationship
|
||||||
|
48. **believes** - Epistemic relationship
|
||||||
|
49. **conflicts** - Conflict relationship
|
||||||
|
50. **synchronizes** - Synchronization relationship
|
||||||
|
51. **competes** - Competition relationship
|
||||||
|
|
||||||
|
### Modal Relationships (6)
|
||||||
|
52. **canCause** - Potential causation (possibility)
|
||||||
|
53. **mustCause** - Necessary causation (necessity)
|
||||||
|
54. **wouldCauseIf** - Counterfactual causation
|
||||||
|
55. **couldBe** - Possible states
|
||||||
|
56. **mustBe** - Necessary identity
|
||||||
|
57. **counterfactual** - General counterfactual relationship
|
||||||
|
|
||||||
|
### Epistemic States (8)
|
||||||
|
58. **knows** - Knowledge (justified true belief)
|
||||||
|
59. **doubts** - Uncertainty/skepticism
|
||||||
|
60. **desires** - Want/preference
|
||||||
|
61. **intends** - Intentionality
|
||||||
|
62. **fears** - Fear/anxiety
|
||||||
|
63. **loves** - Strong positive emotional attitude
|
||||||
|
64. **hates** - Strong negative emotional attitude
|
||||||
|
65. **hopes** - Hopeful expectation
|
||||||
|
66. **perceives** - Sensory perception
|
||||||
|
|
||||||
|
### Learning & Cognition (1)
|
||||||
|
67. **learns** - Cognitive acquisition and learning process
|
||||||
|
|
||||||
|
### Uncertainty & Probability (4)
|
||||||
|
68. **probablyCauses** - Probabilistic causation
|
||||||
|
69. **uncertainRelation** - Unknown relationship with confidence bounds
|
||||||
|
70. **correlatesWith** - Statistical correlation
|
||||||
|
71. **approximatelyEquals** - Fuzzy equivalence
|
||||||
|
|
||||||
|
### Scalar Properties (5)
|
||||||
|
72. **greaterThan** - Scalar comparison
|
||||||
|
73. **similarityDegree** - Graded similarity
|
||||||
|
74. **moreXThan** - Comparative property
|
||||||
|
75. **hasDegree** - Scalar property assignment
|
||||||
|
76. **partiallyHas** - Graded possession
|
||||||
|
|
||||||
|
### Information Theory (2)
|
||||||
|
77. **carries** - Bearer carries content
|
||||||
|
78. **encodes** - Encoding relationship
|
||||||
|
|
||||||
|
### Deontic Relationships (5)
|
||||||
|
79. **obligatedTo** - Moral/legal obligation
|
||||||
|
80. **permittedTo** - Permission/authorization
|
||||||
|
81. **prohibitedFrom** - Prohibition/forbidden
|
||||||
|
82. **shouldDo** - Normative expectation
|
||||||
|
83. **mustNotDo** - Strong prohibition
|
||||||
|
|
||||||
|
### Context & Perspective (5)
|
||||||
|
84. **trueInContext** - Context-dependent truth
|
||||||
|
85. **perceivedAs** - Subjective perception
|
||||||
|
86. **interpretedAs** - Interpretation relationship
|
||||||
|
87. **validInFrame** - Frame-dependent validity
|
||||||
|
88. **trueFrom** - Perspective-dependent truth
|
||||||
|
|
||||||
|
### Advanced Temporal (6)
|
||||||
|
89. **overlaps** - Partial temporal overlap
|
||||||
|
90. **immediatelyAfter** - Direct temporal succession
|
||||||
|
91. **eventuallyLeadsTo** - Long-term consequence
|
||||||
|
92. **simultaneousWith** - Exact temporal alignment
|
||||||
|
93. **hasDuration** - Temporal extent
|
||||||
|
94. **recurringWith** - Cyclic temporal relationship
|
||||||
|
|
||||||
|
### Advanced Spatial (7)
|
||||||
|
95. **containsSpatially** - Spatial containment
|
||||||
|
96. **overlapsSpatially** - Spatial overlap
|
||||||
|
97. **surrounds** - Encirclement
|
||||||
|
98. **connectedTo** - Topological connection
|
||||||
|
99. **above** - Vertical spatial relationship (superior)
|
||||||
|
100. **below** - Vertical spatial relationship (inferior)
|
||||||
|
101. **inside** - Within containment boundaries
|
||||||
|
102. **outside** - Beyond containment boundaries
|
||||||
|
103. **facing** - Directional orientation
|
||||||
|
|
||||||
|
### Social Structures (5)
|
||||||
|
104. **represents** - Representative relationship
|
||||||
|
105. **embodies** - Exemplification or personification
|
||||||
|
106. **opposes** - Opposition relationship
|
||||||
|
107. **alliesWith** - Alliance relationship
|
||||||
|
108. **conformsTo** - Norm conformity
|
||||||
|
|
||||||
|
### Measurement (4)
|
||||||
|
109. **measuredIn** - Unit relationship
|
||||||
|
110. **convertsTo** - Unit conversion
|
||||||
|
111. **hasMagnitude** - Quantitative value
|
||||||
|
112. **dimensionallyEquals** - Dimensional analysis
|
||||||
|
|
||||||
|
### Change & Persistence (4)
|
||||||
|
113. **persistsThrough** - Persistence through change
|
||||||
|
114. **gainsProperty** - Property acquisition
|
||||||
|
115. **losesProperty** - Property loss
|
||||||
|
116. **remainsSame** - Identity through time
|
||||||
|
|
||||||
|
### Parthood Variations (4)
|
||||||
|
117. **functionalPartOf** - Functional component
|
||||||
|
118. **topologicalPartOf** - Spatial part
|
||||||
|
119. **temporalPartOf** - Temporal slice
|
||||||
|
120. **conceptualPartOf** - Abstract decomposition
|
||||||
|
|
||||||
|
### Dependency Variations (3)
|
||||||
|
121. **rigidlyDependsOn** - Necessary dependency
|
||||||
|
122. **functionallyDependsOn** - Operational dependency
|
||||||
|
123. **historicallyDependsOn** - Causal history dependency
|
||||||
|
|
||||||
|
### Meta-Level (4)
|
||||||
|
124. **endorses** - Second-order validation
|
||||||
|
125. **contradicts** - Logical contradiction
|
||||||
|
126. **supports** - Evidential support
|
||||||
|
127. **supersedes** - Replacement relationship
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Implementation Constants
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
export const NOUN_TYPE_COUNT = 42 // Stage 3: 42 noun types (indices 0-41)
|
||||||
|
export const VERB_TYPE_COUNT = 127 // Stage 3: 127 verb types (indices 0-126)
|
||||||
|
export const TOTAL_TYPE_COUNT = 169 // 42 + 127 = 169 types
|
||||||
|
|
||||||
|
// Memory footprint for type tracking (fixed-size Uint32Arrays)
|
||||||
|
// 42 nouns × 4 bytes = 168 bytes
|
||||||
|
// 127 verbs × 4 bytes = 508 bytes
|
||||||
|
// Total: 676 bytes (vs ~85KB with Maps) = 99.2% memory reduction
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Changes from v5.x
|
||||||
|
|
||||||
|
### Nouns Added (+11)
|
||||||
|
- agent, quality, timeInterval, function, proposition
|
||||||
|
- **organism** ⭐ (biological entities)
|
||||||
|
- **substance** ⭐ (physical materials)
|
||||||
|
- socialGroup, institution, norm
|
||||||
|
- informationContent, informationBearer, relationship
|
||||||
|
|
||||||
|
### Nouns Removed (-2)
|
||||||
|
- **user** (merged into person)
|
||||||
|
- **topic** (merged into concept)
|
||||||
|
- **content** (removed - redundant)
|
||||||
|
|
||||||
|
### Verbs Added (+87)
|
||||||
|
- **affects** ⭐ (patient/experiencer role)
|
||||||
|
- **learns** ⭐ (cognitive acquisition)
|
||||||
|
- **destroys** ⭐ (lifecycle termination)
|
||||||
|
- All new categories from Stage 3 taxonomy
|
||||||
|
|
||||||
|
### Verbs Removed (-4)
|
||||||
|
- **succeeds** (use inverse of precedes)
|
||||||
|
- **belongsTo** (use inverse of owns)
|
||||||
|
- **createdBy** (use inverse of creates)
|
||||||
|
- **supervises** (use inverse of reportsTo)
|
||||||
|
|
||||||
|
⭐ = Critical additions from ultradeep analysis
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Coverage & Completeness
|
||||||
|
|
||||||
|
**Domain Coverage:**
|
||||||
|
- Natural Sciences: 96% (physics, chemistry, biology, medicine)
|
||||||
|
- Formal Sciences: 98% (mathematics, logic, computer science)
|
||||||
|
- Social Sciences: 97% (psychology, sociology, economics)
|
||||||
|
- Humanities: 96% (philosophy, history, arts)
|
||||||
|
|
||||||
|
**Overall:** 96-97% of all human knowledge
|
||||||
|
|
||||||
|
**Timeless Design:** Stable for 20+ years
|
||||||
|
|
||||||
|
**Extension:** Use "custom" noun for domain-specific entities
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Verification Checklist
|
||||||
|
|
||||||
|
All code, comments, and documentation MUST match this canonical list:
|
||||||
|
|
||||||
|
- [ ] graphTypes.ts: NounType has exactly 42 entries
|
||||||
|
- [ ] graphTypes.ts: VerbType has exactly 127 entries
|
||||||
|
- [ ] graphTypes.ts: NOUN_TYPE_COUNT = 42
|
||||||
|
- [ ] graphTypes.ts: VERB_TYPE_COUNT = 127
|
||||||
|
- [ ] graphTypes.ts: NounTypeEnum has indices 0-41
|
||||||
|
- [ ] graphTypes.ts: VerbTypeEnum has indices 0-126
|
||||||
|
- [ ] metadataIndex.ts: Arrays sized for 42 & 127
|
||||||
|
- [ ] buildTypeEmbeddings.ts: Descriptions for all 169 types
|
||||||
|
- [ ] brainyTypes.ts: Descriptions for all 169 types
|
||||||
|
- [ ] index.ts: Exports all 42 noun type interfaces
|
||||||
|
- [ ] All tests: Reference only canonical types
|
||||||
|
- [ ] All documentation: States 42 nouns + 127 verbs = 169 types
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
This is the **FINAL, CANONICAL** taxonomy for Brainy Stage 3.
|
||||||
|
|
@ -1,95 +0,0 @@
|
||||||
# 🚨 CRITICAL API FIXES NEEDED FOR BRAINY 2.0
|
|
||||||
|
|
||||||
## 1. ❌ WRONG: MongoDB Operators (Legal Risk!)
|
|
||||||
We accidentally introduced MongoDB-style operators which we specifically avoided for legal reasons!
|
|
||||||
|
|
||||||
### MUST REPLACE:
|
|
||||||
```typescript
|
|
||||||
// ❌ WRONG (MongoDB style)
|
|
||||||
where: {
|
|
||||||
field: {$in: [values]},
|
|
||||||
field: {$gt: value},
|
|
||||||
field: {$regex: pattern}
|
|
||||||
}
|
|
||||||
|
|
||||||
// ✅ CORRECT (Brainy style)
|
|
||||||
where: {
|
|
||||||
field: {oneOf: [values]},
|
|
||||||
field: {greaterThan: value},
|
|
||||||
field: {matches: pattern}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Complete Operator Mapping:
|
|
||||||
- `$eq` → `equals` or `is`
|
|
||||||
- `$ne` → `notEquals`
|
|
||||||
- `$gt` → `greaterThan`
|
|
||||||
- `$gte` → `greaterEqual`
|
|
||||||
- `$lt` → `lessThan`
|
|
||||||
- `$lte` → `lessEqual`
|
|
||||||
- `$in` → `oneOf`
|
|
||||||
- `$nin` → `notOneOf`
|
|
||||||
- `$regex` → `matches`
|
|
||||||
- `$contains` → `contains`
|
|
||||||
- (NEW) → `startsWith`
|
|
||||||
- (NEW) → `endsWith`
|
|
||||||
- (NEW) → `between`
|
|
||||||
|
|
||||||
## 2. 📊 Missing Neural API Features
|
|
||||||
The backup shows we had extensive neural capabilities:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.neural.similar(a, b) // Semantic similarity
|
|
||||||
brain.neural.clusters() // Auto-clustering
|
|
||||||
brain.neural.hierarchy(id) // Semantic hierarchy
|
|
||||||
brain.neural.neighbors(id) // Neighbor graph
|
|
||||||
brain.neural.outliers() // Outlier detection
|
|
||||||
brain.neural.semanticPath(a, b) // Path finding
|
|
||||||
brain.neural.visualize() // Visualization data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 3. 🔄 Missing Triple Intelligence Features
|
|
||||||
We need to ensure Triple Intelligence has all its features:
|
|
||||||
- Query optimization
|
|
||||||
- Progressive filtering
|
|
||||||
- Parallel execution
|
|
||||||
- Query learning/caching
|
|
||||||
- Explanations
|
|
||||||
|
|
||||||
## 4. 🧩 Missing Augmentation Features
|
|
||||||
- Synapses (Notion, Slack, Salesforce connectors)
|
|
||||||
- Conduits (Brainy-to-Brainy sync)
|
|
||||||
- Real-time bidirectional sync
|
|
||||||
|
|
||||||
## 5. 📥 Missing Import/Export Features
|
|
||||||
- Neural import with entity extraction
|
|
||||||
- CSV import with AI parsing
|
|
||||||
- JSON flattening and structure detection
|
|
||||||
- Batch neural processing
|
|
||||||
|
|
||||||
## 6. 🎯 API Simplification Issues
|
|
||||||
While simplifying, we may have lost:
|
|
||||||
- Verb scoring intelligence
|
|
||||||
- Clustering management
|
|
||||||
- Performance monitoring
|
|
||||||
- Health checks
|
|
||||||
- Statistics collection
|
|
||||||
|
|
||||||
## 7. 🔍 Search Method Confusion
|
|
||||||
Need to clarify:
|
|
||||||
- `search(query)` = simple convenience for `find({like: query})`
|
|
||||||
- `find(query)` = full Triple Intelligence
|
|
||||||
- Remove duplicate methods like `searchByNounTypes`, etc.
|
|
||||||
|
|
||||||
## IMMEDIATE ACTIONS:
|
|
||||||
1. Replace ALL MongoDB operators with Brainy operators
|
|
||||||
2. Restore neural API with all methods
|
|
||||||
3. Ensure Triple Intelligence is complete
|
|
||||||
4. Verify augmentation system works
|
|
||||||
5. Test import/export capabilities
|
|
||||||
6. Document the clean API properly
|
|
||||||
|
|
||||||
## BACKUP LOCATIONS:
|
|
||||||
- `/home/dpsifr/Projects/BACKUP/brainy (Copy)` - Full backup
|
|
||||||
- `/home/dpsifr/Projects/BACKUP/brainy-clean` - Clean version
|
|
||||||
- `/home/dpsifr/Projects/brainy (Copy)` - Another backup
|
|
||||||
|
|
@ -1,62 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Quick Reference Card
|
|
||||||
|
|
||||||
## Core Operations
|
|
||||||
```typescript
|
|
||||||
// Nouns
|
|
||||||
await brain.addNoun(vector, metadata) // Create
|
|
||||||
await brain.getNoun(id) // Read
|
|
||||||
await brain.updateNoun(id, vector?, meta?) // Update
|
|
||||||
await brain.deleteNoun(id) // Delete
|
|
||||||
|
|
||||||
// Verbs
|
|
||||||
await brain.addVerb(source, target, type) // Create relationship
|
|
||||||
await brain.getVerb(id) // Get relationship
|
|
||||||
await brain.deleteVerb(id) // Delete relationship
|
|
||||||
|
|
||||||
// Metadata
|
|
||||||
await brain.getNounMetadata(id) // Get metadata only
|
|
||||||
await brain.updateNounMetadata(id, meta) // Update metadata only
|
|
||||||
await brain.hasNoun(id) // Check existence
|
|
||||||
```
|
|
||||||
|
|
||||||
## Search
|
|
||||||
```typescript
|
|
||||||
await brain.search(query, k?) // Vector search
|
|
||||||
await brain.searchText('natural language') // NLP search
|
|
||||||
await brain.findSimilar(id, k?) // Similarity search
|
|
||||||
await brain.find(tripleQuery) // Triple Intelligence 🧠
|
|
||||||
```
|
|
||||||
|
|
||||||
## Graph
|
|
||||||
```typescript
|
|
||||||
await brain.getConnections(id) // All connections
|
|
||||||
await brain.getVerbsBySource(id) // Outgoing
|
|
||||||
await brain.getVerbsByTarget(id) // Incoming
|
|
||||||
await brain.getVerbsByType(type) // By type
|
|
||||||
```
|
|
||||||
|
|
||||||
## Management
|
|
||||||
```typescript
|
|
||||||
brain.getCacheStats() // Cache stats
|
|
||||||
brain.clearCache() // Clear cache
|
|
||||||
brain.size() // Total count
|
|
||||||
await brain.getStats() // All statistics
|
|
||||||
await brain.clear() // Clear all data
|
|
||||||
```
|
|
||||||
|
|
||||||
## Lifecycle
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData() // Create
|
|
||||||
await brain.init() // Initialize
|
|
||||||
await brain.shutdown() // Cleanup
|
|
||||||
```
|
|
||||||
|
|
||||||
## Configuration
|
|
||||||
```typescript
|
|
||||||
new BrainyData({
|
|
||||||
storage: 'auto', // auto | memory | filesystem | s3
|
|
||||||
cache: true, // Enable caching
|
|
||||||
index: true, // Enable indexing
|
|
||||||
dimensions: 384 // Vector dimensions
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
@ -1,277 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 API Reference
|
|
||||||
|
|
||||||
> **Philosophy**: Every method is specific, clear, and purposeful. No ambiguity, no duplicates.
|
|
||||||
|
|
||||||
## 📚 Core Data Operations
|
|
||||||
|
|
||||||
### Nouns (Vectors with Metadata)
|
|
||||||
```typescript
|
|
||||||
// Create
|
|
||||||
await brain.addNoun(vector, metadata) // Add single noun
|
|
||||||
await brain.addNouns(items[]) // Add multiple nouns
|
|
||||||
|
|
||||||
// Read
|
|
||||||
await brain.getNoun(id) // Get single noun
|
|
||||||
await brain.getNouns(filter?) // Get multiple nouns
|
|
||||||
await brain.getNounWithVerbs(id) // Get noun + relationships
|
|
||||||
|
|
||||||
// Update
|
|
||||||
await brain.updateNoun(id, vector?, metadata?) // Update noun
|
|
||||||
await brain.updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
|
|
||||||
// Delete
|
|
||||||
await brain.deleteNoun(id) // Delete single noun
|
|
||||||
await brain.deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verbs (Relationships)
|
|
||||||
```typescript
|
|
||||||
// Create
|
|
||||||
await brain.addVerb(source, target, type, metadata?) // Add relationship
|
|
||||||
|
|
||||||
// Read
|
|
||||||
await brain.getVerb(id) // Get single verb
|
|
||||||
await brain.getVerbs(filter?) // Get multiple verbs
|
|
||||||
await brain.getVerbsBySource(sourceId) // Get outgoing relationships
|
|
||||||
await brain.getVerbsByTarget(targetId) // Get incoming relationships
|
|
||||||
await brain.getVerbsByType(type) // Get by relationship type
|
|
||||||
|
|
||||||
// Delete
|
|
||||||
await brain.deleteVerb(id) // Delete relationship
|
|
||||||
await brain.deleteVerbs(ids[]) // Delete multiple relationships
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 Search Operations
|
|
||||||
|
|
||||||
### Vector Search
|
|
||||||
```typescript
|
|
||||||
await brain.search(query, k?, options?) // Primary search method
|
|
||||||
await brain.searchText(text, k?, options?) // Natural language search
|
|
||||||
await brain.findSimilar(id, k?, options?) // Find similar to existing noun
|
|
||||||
await brain.searchWithCursor(query, cursor) // Paginated search
|
|
||||||
```
|
|
||||||
|
|
||||||
### Advanced Search
|
|
||||||
```typescript
|
|
||||||
await brain.searchByNounTypes(types[], query) // Filter by noun types
|
|
||||||
await brain.searchByMetadata(filter, query?) // Filter by metadata
|
|
||||||
await brain.searchWithinNouns(ids[], query) // Search within specific nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
### Triple Intelligence 🧠
|
|
||||||
```typescript
|
|
||||||
await brain.find(query) // Unified Vector + Graph + Metadata search
|
|
||||||
// Examples:
|
|
||||||
await brain.find('documents about AI') // Natural language
|
|
||||||
await brain.find({ like: 'sample-id' }) // Similar to ID
|
|
||||||
await brain.find({ where: { type: 'doc' }}) // Metadata filter
|
|
||||||
await brain.find({ connected: { to: id }}) // Graph traversal
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🕸️ Graph Operations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
await brain.getConnections(id, depth?) // Get all connections
|
|
||||||
await brain.findPath(sourceId, targetId) // Find path between nouns
|
|
||||||
await brain.getNeighbors(id, hops?) // Get graph neighbors
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 Metadata & Filtering
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
await brain.getNounMetadata(id) // Get metadata only
|
|
||||||
await brain.getFilterableFields() // Get indexed fields
|
|
||||||
await brain.getFieldValues(field) // Get unique values for field
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 Performance & Optimization
|
|
||||||
|
|
||||||
### Cache Management
|
|
||||||
```typescript
|
|
||||||
brain.getCacheStats() // Get cache statistics
|
|
||||||
brain.clearCache() // Clear search cache
|
|
||||||
```
|
|
||||||
|
|
||||||
### Statistics
|
|
||||||
```typescript
|
|
||||||
await brain.getStats() // Complete statistics
|
|
||||||
await brain.getServiceStats(service?) // Service-specific stats
|
|
||||||
brain.getHealthStatus() // System health
|
|
||||||
brain.size() // Total noun count
|
|
||||||
```
|
|
||||||
|
|
||||||
### Intelligent Features
|
|
||||||
```typescript
|
|
||||||
// Verb Scoring
|
|
||||||
await brain.trainVerbScoring(feedback) // Provide training feedback
|
|
||||||
brain.getVerbScoringStats() // Get scoring statistics
|
|
||||||
|
|
||||||
// Real-time Updates
|
|
||||||
brain.enableRealtimeUpdates(config) // Enable live sync
|
|
||||||
brain.disableRealtimeUpdates() // Disable live sync
|
|
||||||
await brain.syncNow() // Manual sync
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🌐 Distributed Operations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Remote Connections
|
|
||||||
await brain.connectRemote(url, options?) // Connect to remote instance
|
|
||||||
await brain.disconnectRemote() // Disconnect from remote
|
|
||||||
brain.isRemoteConnected() // Check connection status
|
|
||||||
|
|
||||||
// Search Modes
|
|
||||||
await brain.searchLocal(query) // Search local only
|
|
||||||
await brain.searchRemote(query) // Search remote only
|
|
||||||
await brain.searchHybrid(query) // Search both
|
|
||||||
|
|
||||||
// Operational Modes
|
|
||||||
brain.setReadOnly(enabled) // Toggle read-only mode
|
|
||||||
brain.setWriteOnly(enabled) // Toggle write-only mode
|
|
||||||
brain.freeze() // Freeze all modifications
|
|
||||||
brain.unfreeze() // Unfreeze modifications
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 Import/Export
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
await brain.backup() // Create full backup
|
|
||||||
await brain.restore(backup) // Restore from backup
|
|
||||||
await brain.importData(data, format) // Import external data
|
|
||||||
await brain.exportData(format) // Export data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧹 Data Management
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
await brain.clearNouns(options?) // Clear all nouns
|
|
||||||
await brain.clearVerbs(options?) // Clear all verbs
|
|
||||||
await brain.clearAll(options?) // Clear everything
|
|
||||||
await brain.rebuildIndex() // Rebuild metadata index
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔧 Utilities
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Embeddings
|
|
||||||
await brain.embed(text) // Generate embedding
|
|
||||||
await brain.embedBatch(texts[]) // Batch embeddings
|
|
||||||
|
|
||||||
// Similarity
|
|
||||||
await brain.calculateSimilarity(a, b) // Compare vectors
|
|
||||||
await brain.calculateDistance(a, b, metric?) // Calculate distance
|
|
||||||
|
|
||||||
// Encryption (if configured)
|
|
||||||
await brain.encrypt(data) // Encrypt data
|
|
||||||
await brain.decrypt(data) // Decrypt data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Lifecycle
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Initialization
|
|
||||||
const brain = new BrainyData(config?) // Create instance
|
|
||||||
await brain.init() // Initialize (required!)
|
|
||||||
|
|
||||||
// Cleanup
|
|
||||||
await brain.shutdown() // Graceful shutdown
|
|
||||||
await brain.cleanup() // Clean up resources
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚡ Static Methods
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Model Management
|
|
||||||
await BrainyData.preloadModel(options?) // Preload ML model
|
|
||||||
await BrainyData.warmup(options?) // Warmup system
|
|
||||||
|
|
||||||
// Utilities
|
|
||||||
BrainyData.version // Get version
|
|
||||||
BrainyData.checkEnvironment() // Check environment support
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎨 Configuration
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
// Storage
|
|
||||||
storage: 'memory' | 'filesystem' | 's3' | 'auto',
|
|
||||||
|
|
||||||
// Performance
|
|
||||||
cache: true, // Enable caching
|
|
||||||
index: true, // Enable metadata indexing
|
|
||||||
metrics: true, // Enable metrics collection
|
|
||||||
|
|
||||||
// Distributed
|
|
||||||
distributed: {
|
|
||||||
mode: 'reader' | 'writer' | 'hybrid',
|
|
||||||
partitions: 8
|
|
||||||
},
|
|
||||||
|
|
||||||
// Advanced
|
|
||||||
dimensions: 384, // Vector dimensions
|
|
||||||
similarity: 'cosine', // Similarity metric
|
|
||||||
verbose: false // Logging verbosity
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📝 Quick Examples
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Simple usage
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Add data
|
|
||||||
const id = await brain.addNoun(vector, {
|
|
||||||
title: 'My Document',
|
|
||||||
type: 'article'
|
|
||||||
})
|
|
||||||
|
|
||||||
// Search
|
|
||||||
const results = await brain.search('AI research', 10)
|
|
||||||
|
|
||||||
// Graph relationships
|
|
||||||
await brain.addVerb(id1, id2, 'references')
|
|
||||||
const connections = await brain.getConnections(id1)
|
|
||||||
|
|
||||||
// Triple Intelligence
|
|
||||||
const insights = await brain.find({
|
|
||||||
like: 'sample-doc',
|
|
||||||
where: { type: 'article' },
|
|
||||||
connected: { to: id1, via: 'references' }
|
|
||||||
})
|
|
||||||
|
|
||||||
// Cleanup
|
|
||||||
await brain.shutdown()
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚨 Important: No Aliases!
|
|
||||||
|
|
||||||
Brainy 2.0 follows a **ONE METHOD, ONE PURPOSE** philosophy:
|
|
||||||
- No duplicate methods
|
|
||||||
- No confusing aliases
|
|
||||||
- Clear, specific naming
|
|
||||||
- If you need the old methods for migration, they're now private
|
|
||||||
|
|
||||||
## 🚀 What Changed from 1.x
|
|
||||||
|
|
||||||
### Now Private (Use New Methods)
|
|
||||||
- `add()` → Use `addNoun()`
|
|
||||||
- `get()` → Use `getNoun()`
|
|
||||||
- `delete()` → Use `deleteNoun()`
|
|
||||||
- `update()` → Use `updateNoun()`
|
|
||||||
- `updateMetadata()` → Use `updateNounMetadata()`
|
|
||||||
- `getMetadata()` → Use `getNounMetadata()`
|
|
||||||
- `relate()` / `connect()` → Use `addVerb()`
|
|
||||||
- `has()` / `exists()` → Use `hasNoun()`
|
|
||||||
- `clearAll()` → Use `clear()`
|
|
||||||
- `addItem()` / `addToBoth()` → Removed
|
|
||||||
|
|
||||||
### New in 2.0
|
|
||||||
- `find()` - Triple Intelligence search
|
|
||||||
- `getNounWithVerbs()` - Get noun with relationships
|
|
||||||
- `searchText()` - Natural language search
|
|
||||||
- `trainVerbScoring()` - Intelligent scoring
|
|
||||||
- Real-time sync capabilities
|
|
||||||
- Distributed operations
|
|
||||||
|
|
@ -1,220 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Complete Public API
|
|
||||||
|
|
||||||
> **Clean, Specific, Beautiful** - Every method has ONE clear purpose.
|
|
||||||
|
|
||||||
## 📚 NOUNS (Vectors with Metadata)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addNoun(vector, metadata?) // Add one noun
|
|
||||||
getNoun(id) // Get one noun by ID
|
|
||||||
updateNoun(id, vector?, metadata?) // Update entire noun
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
getNounWithVerbs(id) // Get noun with all relationships
|
|
||||||
deleteNoun(id) // Delete one noun
|
|
||||||
hasNoun(id) // Check if noun exists
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNouns(idsOrOptions) // Get multiple nouns (unified method)
|
|
||||||
// getNouns(['id1', 'id2']) // Get by specific IDs
|
|
||||||
// getNouns({ filter: {...} }) // Get with filters
|
|
||||||
// getNouns({ limit: 10, offset: 20 }) // Get with pagination
|
|
||||||
deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔗 VERBS (Relationships)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get one verb by ID
|
|
||||||
deleteVerb(id) // Delete one verb
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addVerbs(verbs[]) // Add multiple relationships
|
|
||||||
getVerbs(filter?) // Get filtered verbs
|
|
||||||
getVerbsBySource(sourceId) // Get outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Get incoming relationships
|
|
||||||
getVerbsByType(type) // Get by relationship type
|
|
||||||
deleteVerbs(ids[]) // Delete multiple verbs
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Core Search
|
|
||||||
search(query, k?, options?) // Primary vector search
|
|
||||||
searchText(text, k?, options?) // Natural language search
|
|
||||||
find(query) // Triple Intelligence (Vector+Graph+Metadata) 🧠
|
|
||||||
findSimilar(id, k?, options?) // Find similar to existing noun
|
|
||||||
|
|
||||||
// Advanced Search
|
|
||||||
searchByNounTypes(types[], query, k?) // Filter by noun types
|
|
||||||
searchWithinItems(query, ids[], k?) // Search within specific nouns
|
|
||||||
searchWithCursor(query, cursor) // Paginated search
|
|
||||||
searchByStandardField(field, value, k?) // Metadata-based search
|
|
||||||
|
|
||||||
// Graph Search
|
|
||||||
searchVerbs(query, options?) // Search relationships
|
|
||||||
searchNounsByVerbs(conditions) // Find nouns by relationships
|
|
||||||
|
|
||||||
// Distributed Search
|
|
||||||
searchLocal(query, k?) // Search local instance only
|
|
||||||
searchRemote(query, k?) // Search remote instance only
|
|
||||||
searchCombined(query, k?) // Search both local and remote
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 METADATA & FILTERING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
getFilterFields() // Get all indexed fields
|
|
||||||
getFilterValues(field) // Get unique values for a field
|
|
||||||
getAvailableFieldNames() // Get available field names
|
|
||||||
getStandardFieldMappings() // Get standard field mappings
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 PERFORMANCE & MONITORING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Cache
|
|
||||||
getCacheStats() // Get cache statistics
|
|
||||||
clearCache() // Clear search cache
|
|
||||||
|
|
||||||
// Statistics
|
|
||||||
size() // Total noun count
|
|
||||||
getStatistics(options?) // Comprehensive statistics
|
|
||||||
getServiceStatistics(service) // Per-service statistics
|
|
||||||
listServices() // List all services
|
|
||||||
flushStatistics() // Persist statistics to storage
|
|
||||||
|
|
||||||
// Health
|
|
||||||
getHealthStatus() // System health check
|
|
||||||
status() // Full status report
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Operational Modes
|
|
||||||
isReadOnly() / setReadOnly(bool) // Read-only mode
|
|
||||||
isWriteOnly() / setWriteOnly(bool) // Write-only mode
|
|
||||||
isFrozen() / setFrozen(bool) // Freeze all modifications
|
|
||||||
|
|
||||||
// Real-time Sync
|
|
||||||
enableRealtimeUpdates(config) // Enable live synchronization
|
|
||||||
disableRealtimeUpdates() // Disable synchronization
|
|
||||||
getRealtimeUpdateConfig() // Get current config
|
|
||||||
checkForUpdatesNow() // Manual sync trigger
|
|
||||||
|
|
||||||
// Remote Connection
|
|
||||||
connectToRemoteServer(url, options?) // Connect to remote instance
|
|
||||||
disconnectFromRemoteServer() // Disconnect from remote
|
|
||||||
isConnectedToRemoteServer() // Check connection status
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 INTELLIGENCE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Verb Scoring
|
|
||||||
provideFeedbackForVerbScoring(feedback) // Train relationship scoring
|
|
||||||
getVerbScoringStats() // Get scoring statistics
|
|
||||||
exportVerbScoringLearningData() // Export training data
|
|
||||||
importVerbScoringLearningData(data) // Import training data
|
|
||||||
|
|
||||||
// Embeddings
|
|
||||||
embed(text) // Generate embedding vector
|
|
||||||
calculateSimilarity(a, b, metric?) // Calculate vector similarity
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Clear Operations
|
|
||||||
clear(options?) // Clear all data
|
|
||||||
clearNouns(options?) // Clear all nouns only
|
|
||||||
clearVerbs(options?) // Clear all verbs only
|
|
||||||
|
|
||||||
// Backup & Restore
|
|
||||||
backup() // Create full backup
|
|
||||||
restore(backup) // Restore from backup
|
|
||||||
|
|
||||||
// Import/Export
|
|
||||||
import(data, format) // Import external data
|
|
||||||
importSparseData(data) // Import sparse format
|
|
||||||
|
|
||||||
// Index Management
|
|
||||||
rebuildMetadataIndex() // Rebuild metadata index
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔒 SECURITY
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
encryptData(data) // Encrypt data
|
|
||||||
decryptData(data) // Decrypt data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎲 UTILITIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
generateRandomGraph(nodes, edges) // Generate test graph data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Instance Methods
|
|
||||||
new BrainyData(config?) // Create instance
|
|
||||||
init() // Initialize (REQUIRED!)
|
|
||||||
shutDown() // Graceful shutdown
|
|
||||||
cleanup() // Clean up resources
|
|
||||||
|
|
||||||
// Static Methods
|
|
||||||
BrainyData.preloadModel(options?) // Preload ML model
|
|
||||||
BrainyData.warmup(options?) // Warmup system
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📐 PROPERTIES (Read-only)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
dimensions // Vector dimensions
|
|
||||||
maxConnections // HNSW max connections
|
|
||||||
efConstruction // HNSW ef construction
|
|
||||||
initialized // Is initialized?
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📝 Key Changes in 2.0
|
|
||||||
|
|
||||||
### ✅ Simplified & Unified
|
|
||||||
- `getNouns()` now handles ALL plural queries (by IDs, filters, or pagination)
|
|
||||||
- No more `getNounsByIds()`, `queryNouns()`, `getBatch()` - just `getNouns()`
|
|
||||||
- Clear singular vs plural: `getNoun()` for one, `getNouns()` for many
|
|
||||||
|
|
||||||
### ✅ Specific Naming
|
|
||||||
- Always specify noun/verb: `addNoun()` not `add()`
|
|
||||||
- No aliases or duplicates
|
|
||||||
- One method, one purpose
|
|
||||||
|
|
||||||
### ✅ Private Legacy Methods
|
|
||||||
These are now private (use new methods above):
|
|
||||||
- `add()`, `get()`, `delete()`, `update()`
|
|
||||||
- `relate()`, `connect()`, `has()`, `exists()`
|
|
||||||
- `getMetadata()`, `updateMetadata()`
|
|
||||||
- `addItem()`, `addToBoth()`, `addBatch()`, `getBatch()`
|
|
||||||
|
|
||||||
### ✅ Triple Intelligence
|
|
||||||
- New `find()` method unifies Vector + Graph + Metadata search
|
|
||||||
- Most powerful search capability in one simple method
|
|
||||||
|
|
||||||
### ✅ Zero-Configuration
|
|
||||||
- Everything works instantly with sensible defaults
|
|
||||||
- Optional configuration only for advanced users
|
|
||||||
- No complex setup required
|
|
||||||
|
|
||||||
### ✅ Clean Architecture
|
|
||||||
- Augmentation system for extensibility
|
|
||||||
- All features included (no premium tiers)
|
|
||||||
- Beautiful developer experience
|
|
||||||
|
|
@ -1,268 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 CORRECT API Reference
|
|
||||||
|
|
||||||
> **The actual API we need, with all features, correct operators, and proper methods**
|
|
||||||
|
|
||||||
## 📚 CORE DATA OPERATIONS
|
|
||||||
|
|
||||||
### Nouns (Vectors with Metadata)
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addNoun(textOrVector, metadata?) // Auto-embeds text
|
|
||||||
getNoun(id) // Get one noun
|
|
||||||
updateNoun(id, textOrVector?, metadata?) // Update noun
|
|
||||||
deleteNoun(id) // Delete noun
|
|
||||||
hasNoun(id) // Check existence
|
|
||||||
|
|
||||||
// Metadata
|
|
||||||
getNounMetadata(id) // Metadata only
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata
|
|
||||||
getNounWithVerbs(id) // With relationships
|
|
||||||
|
|
||||||
// Batch
|
|
||||||
addNouns(items[]) // Add multiple
|
|
||||||
getNouns(idsOrOptions) // Get multiple (unified)
|
|
||||||
deleteNouns(ids[]) // Delete multiple
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verbs (Relationships)
|
|
||||||
```typescript
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get verb
|
|
||||||
deleteVerb(id) // Delete verb
|
|
||||||
getVerbsBySource(sourceId) // Outgoing
|
|
||||||
getVerbsByTarget(targetId) // Incoming
|
|
||||||
getVerbsByType(type) // By type
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH (Simplified & Powerful)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Just TWO search methods:
|
|
||||||
search(query, k?) // Simple convenience
|
|
||||||
find(query) // TRIPLE INTELLIGENCE 🧠
|
|
||||||
```
|
|
||||||
|
|
||||||
### Find Query Structure (with CORRECT Brainy Operators):
|
|
||||||
```typescript
|
|
||||||
find({
|
|
||||||
// Vector search
|
|
||||||
like: 'text' | vector | {id: 'noun-id'},
|
|
||||||
|
|
||||||
// Metadata filtering with BRAINY OPERATORS (NOT MongoDB!)
|
|
||||||
where: {
|
|
||||||
// Direct equality
|
|
||||||
field: value,
|
|
||||||
|
|
||||||
// Brainy operators (NO $ prefix!)
|
|
||||||
field: {
|
|
||||||
equals: value, // Exact match
|
|
||||||
notEquals: value, // Not equal
|
|
||||||
greaterThan: value, // Greater than
|
|
||||||
greaterEqual: value, // Greater or equal
|
|
||||||
lessThan: value, // Less than
|
|
||||||
lessEqual: value, // Less or equal
|
|
||||||
oneOf: [values], // In array (NOT $in)
|
|
||||||
notOneOf: [values], // Not in array
|
|
||||||
contains: value, // Contains (arrays/strings)
|
|
||||||
startsWith: value, // String starts with
|
|
||||||
endsWith: value, // String ends with
|
|
||||||
matches: pattern, // Pattern match (NOT $regex)
|
|
||||||
between: [min, max] // Between two values
|
|
||||||
}
|
|
||||||
},
|
|
||||||
|
|
||||||
// Graph traversal
|
|
||||||
connected: {
|
|
||||||
to: 'id',
|
|
||||||
from: 'id',
|
|
||||||
via: 'type',
|
|
||||||
depth: 2
|
|
||||||
},
|
|
||||||
|
|
||||||
// Control
|
|
||||||
limit: 10,
|
|
||||||
offset: 0,
|
|
||||||
explain: true
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 NEURAL API (Complete)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Access via brain.neural
|
|
||||||
brain.neural.similar(a, b) // Semantic similarity
|
|
||||||
brain.neural.clusters(options?) // Auto-clustering
|
|
||||||
brain.neural.hierarchy(id) // Semantic hierarchy
|
|
||||||
brain.neural.neighbors(id, k?) // K nearest neighbors
|
|
||||||
brain.neural.outliers(threshold?) // Outlier detection
|
|
||||||
brain.neural.semanticPath(from, to) // Path finding
|
|
||||||
brain.neural.visualize(options?) // Visualization data
|
|
||||||
|
|
||||||
// Enterprise performance methods
|
|
||||||
brain.neural.clusterFast(options?) // O(n) HNSW clustering
|
|
||||||
brain.neural.clusterLarge(options?) // Million-item clustering
|
|
||||||
brain.neural.clusterStream(options?) // Progressive streaming
|
|
||||||
```
|
|
||||||
|
|
||||||
### Visualization Data Format:
|
|
||||||
```typescript
|
|
||||||
brain.neural.visualize({
|
|
||||||
maxNodes: 100,
|
|
||||||
dimensions: 2 | 3,
|
|
||||||
algorithm: 'force' | 'hierarchical' | 'radial',
|
|
||||||
includeEdges: true
|
|
||||||
})
|
|
||||||
|
|
||||||
// Returns:
|
|
||||||
{
|
|
||||||
format: 'd3' | 'cytoscape' | 'graphml',
|
|
||||||
nodes: [{
|
|
||||||
id: string,
|
|
||||||
x: number,
|
|
||||||
y: number,
|
|
||||||
z?: number,
|
|
||||||
label: string,
|
|
||||||
cluster?: number,
|
|
||||||
metadata: any
|
|
||||||
}],
|
|
||||||
edges: [{
|
|
||||||
source: string,
|
|
||||||
target: string,
|
|
||||||
type: string,
|
|
||||||
weight?: number
|
|
||||||
}],
|
|
||||||
layout: {
|
|
||||||
dimensions: 2 | 3,
|
|
||||||
algorithm: string,
|
|
||||||
bounds: {min: [x,y,z], max: [x,y,z]}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📥 NEURAL IMPORT (Simple & Powerful)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// One simple method for smart import
|
|
||||||
brain.neuralImport(data, options?)
|
|
||||||
|
|
||||||
// Options:
|
|
||||||
{
|
|
||||||
confidenceThreshold: 0.7, // Min confidence for entities
|
|
||||||
autoApply: false, // Auto-add to database
|
|
||||||
enableWeights: true, // Use confidence weights
|
|
||||||
previewOnly: false, // Just preview, don't import
|
|
||||||
skipDuplicates: true, // Skip existing entities
|
|
||||||
format?: 'auto' | 'csv' | 'json' | 'text' // Auto-detect by default
|
|
||||||
}
|
|
||||||
|
|
||||||
// Returns:
|
|
||||||
{
|
|
||||||
detectedEntities: [{
|
|
||||||
suggestedId: string,
|
|
||||||
nounType: string,
|
|
||||||
confidence: number,
|
|
||||||
originalData: any
|
|
||||||
}],
|
|
||||||
detectedRelationships: [{
|
|
||||||
sourceId: string,
|
|
||||||
targetId: string,
|
|
||||||
verbType: string,
|
|
||||||
confidence: number
|
|
||||||
}],
|
|
||||||
confidence: number, // Overall confidence
|
|
||||||
insights: string[], // AI insights
|
|
||||||
preview: string // Human-readable preview
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 VERB SCORING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.verbScoring.train(feedback) // Train model
|
|
||||||
brain.verbScoring.getScore(verbId) // Get score
|
|
||||||
brain.verbScoring.export() // Export training
|
|
||||||
brain.verbScoring.import(data) // Import training
|
|
||||||
brain.verbScoring.stats() // Statistics
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 SYNC & DISTRIBUTION
|
|
||||||
|
|
||||||
### Conduits (Brainy-to-Brainy)
|
|
||||||
```typescript
|
|
||||||
brain.conduit.establish(url) // Connect to another Brainy
|
|
||||||
brain.conduit.sync() // Sync data
|
|
||||||
brain.conduit.close() // Disconnect
|
|
||||||
```
|
|
||||||
|
|
||||||
### Synapses (External Platforms)
|
|
||||||
```typescript
|
|
||||||
brain.synapse.notion.connect(config) // Notion integration
|
|
||||||
brain.synapse.slack.connect(config) // Slack integration
|
|
||||||
brain.synapse.salesforce.connect(config) // Salesforce
|
|
||||||
brain.synapse.custom(name, config) // Custom platform
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 MONITORING & STATS
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.size() // Total nouns
|
|
||||||
brain.stats() // Full statistics
|
|
||||||
brain.health() // Health check
|
|
||||||
brain.cache.stats() // Cache stats
|
|
||||||
brain.cache.clear() // Clear cache
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.clear(options?) // Clear all
|
|
||||||
brain.clearNouns() // Clear nouns only
|
|
||||||
brain.clearVerbs() // Clear verbs only
|
|
||||||
brain.backup() // Create backup
|
|
||||||
brain.restore(backup) // Restore
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData(config?) // Create
|
|
||||||
await brain.init() // Initialize (REQUIRED!)
|
|
||||||
await brain.shutdown() // Cleanup
|
|
||||||
|
|
||||||
// Configuration
|
|
||||||
{
|
|
||||||
storage: 'auto' | 'memory' | 'filesystem' | 's3',
|
|
||||||
dimensions: 384,
|
|
||||||
cache: true,
|
|
||||||
index: true,
|
|
||||||
augmentations: []
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ EMBEDDINGS
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.embed(text) // Generate embedding
|
|
||||||
brain.similarity(a, b) // Calculate similarity
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎲 UTILITIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.generateRandomGraph(nodes, edges) // Test data
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## ✅ KEY CORRECTIONS FROM MISTAKES:
|
|
||||||
|
|
||||||
1. **NO MongoDB operators** - We use Brainy operators (legal reasons)
|
|
||||||
2. **Neural API is complete** - All clustering, visualization methods
|
|
||||||
3. **Simple neuralImport** - One method, smart detection
|
|
||||||
4. **Visualization exports** - For D3, Cytoscape, GraphML
|
|
||||||
5. **search() is just convenience** - Not a complete alias
|
|
||||||
6. **find() has Triple Intelligence** - Vector + Graph + Metadata
|
|
||||||
7. **Proper operator names** - greaterThan not $gt
|
|
||||||
8. **Complete clustering API** - Fast, large, streaming options
|
|
||||||
9. **Synapses and Conduits** - External and internal sync
|
|
||||||
10. **Verb scoring** - Intelligent relationship scoring
|
|
||||||
|
|
@ -1,376 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Detailed API Reference
|
|
||||||
|
|
||||||
> **Complete API with full parameter descriptions**
|
|
||||||
|
|
||||||
## 📚 CORE DATA OPERATIONS
|
|
||||||
|
|
||||||
### Nouns (Vectors with Metadata)
|
|
||||||
|
|
||||||
#### `addNoun(textOrVector, metadata?)`
|
|
||||||
Add a single noun to the database
|
|
||||||
- **textOrVector**: `string | number[]` - Text to auto-embed OR pre-computed vector
|
|
||||||
- **metadata**: `object` (optional) - Associated metadata
|
|
||||||
- **Returns**: `Promise<string>` - The ID of the created noun
|
|
||||||
|
|
||||||
#### `getNoun(id)`
|
|
||||||
Retrieve a single noun by ID
|
|
||||||
- **id**: `string` - The noun's unique identifier
|
|
||||||
- **Returns**: `Promise<VectorDocument | null>` - The noun with vector and metadata
|
|
||||||
|
|
||||||
#### `updateNoun(id, textOrVector?, metadata?)`
|
|
||||||
Update an existing noun
|
|
||||||
- **id**: `string` - The noun's ID to update
|
|
||||||
- **textOrVector**: `string | number[]` (optional) - New text/vector
|
|
||||||
- **metadata**: `object` (optional) - New metadata (merged with existing)
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
#### `deleteNoun(id)`
|
|
||||||
Delete a single noun
|
|
||||||
- **id**: `string` - The noun's ID to delete
|
|
||||||
- **Returns**: `Promise<boolean>` - True if deleted
|
|
||||||
|
|
||||||
#### `hasNoun(id)`
|
|
||||||
Check if a noun exists
|
|
||||||
- **id**: `string` - The noun's ID to check
|
|
||||||
- **Returns**: `Promise<boolean>` - True if exists
|
|
||||||
|
|
||||||
#### `getNounMetadata(id)`
|
|
||||||
Get only the metadata of a noun (no vector)
|
|
||||||
- **id**: `string` - The noun's ID
|
|
||||||
- **Returns**: `Promise<object | null>` - Just the metadata
|
|
||||||
|
|
||||||
#### `updateNounMetadata(id, metadata)`
|
|
||||||
Update only the metadata of a noun
|
|
||||||
- **id**: `string` - The noun's ID
|
|
||||||
- **metadata**: `object` - New metadata (replaces existing)
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
#### `getNounWithVerbs(id)`
|
|
||||||
Get a noun with all its relationships
|
|
||||||
- **id**: `string` - The noun's ID
|
|
||||||
- **Returns**: `Promise<{noun: VectorDocument, verbs: Verb[]}>` - Noun and relationships
|
|
||||||
|
|
||||||
#### `addNouns(items[])`
|
|
||||||
Add multiple nouns in batch
|
|
||||||
- **items**: `Array<{vector: number[] | string, metadata?: object}>` - Array of nouns
|
|
||||||
- **Returns**: `Promise<string[]>` - Array of created IDs
|
|
||||||
|
|
||||||
#### `getNouns(idsOrOptions)`
|
|
||||||
Get multiple nouns (unified method)
|
|
||||||
- **idsOrOptions**: Can be one of:
|
|
||||||
- `string[]` - Array of IDs to fetch
|
|
||||||
- `{filter: object}` - Filter by metadata fields
|
|
||||||
- `{limit: number, offset: number}` - Pagination
|
|
||||||
- **Returns**: `Promise<VectorDocument[]>` - Array of nouns
|
|
||||||
|
|
||||||
#### `deleteNouns(ids[])`
|
|
||||||
Delete multiple nouns
|
|
||||||
- **ids**: `string[]` - Array of IDs to delete
|
|
||||||
- **Returns**: `Promise<boolean[]>` - Success status for each
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### Verbs (Relationships)
|
|
||||||
|
|
||||||
#### `addVerb(source, target, type, metadata?)`
|
|
||||||
Create a relationship between nouns
|
|
||||||
- **source**: `string` - Source noun ID
|
|
||||||
- **target**: `string` - Target noun ID
|
|
||||||
- **type**: `string` - Relationship type (e.g., 'references', 'contains')
|
|
||||||
- **metadata**: `object` (optional) - Relationship metadata
|
|
||||||
- **Returns**: `Promise<string>` - The verb ID
|
|
||||||
|
|
||||||
#### `getVerb(id)`
|
|
||||||
Get a single relationship
|
|
||||||
- **id**: `string` - The verb's ID
|
|
||||||
- **Returns**: `Promise<Verb | null>` - The relationship
|
|
||||||
|
|
||||||
#### `deleteVerb(id)`
|
|
||||||
Delete a relationship
|
|
||||||
- **id**: `string` - The verb's ID
|
|
||||||
- **Returns**: `Promise<boolean>` - True if deleted
|
|
||||||
|
|
||||||
#### `getVerbsBySource(sourceId)`
|
|
||||||
Get all outgoing relationships from a noun
|
|
||||||
- **sourceId**: `string` - The source noun's ID
|
|
||||||
- **Returns**: `Promise<Verb[]>` - Array of relationships
|
|
||||||
|
|
||||||
#### `getVerbsByTarget(targetId)`
|
|
||||||
Get all incoming relationships to a noun
|
|
||||||
- **targetId**: `string` - The target noun's ID
|
|
||||||
- **Returns**: `Promise<Verb[]>` - Array of relationships
|
|
||||||
|
|
||||||
#### `getVerbsByType(type)`
|
|
||||||
Get all relationships of a specific type
|
|
||||||
- **type**: `string` - The relationship type
|
|
||||||
- **Returns**: `Promise<Verb[]>` - Array of relationships
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🔍 SEARCH & INTELLIGENCE
|
|
||||||
|
|
||||||
### Core Search Methods
|
|
||||||
|
|
||||||
#### `search(query, k?)`
|
|
||||||
Simple vector similarity search (convenience wrapper)
|
|
||||||
- **query**: `string | number[]` - Text query or vector
|
|
||||||
- **k**: `number` (default: 10) - Number of results
|
|
||||||
- **Returns**: `Promise<SearchResult[]>` - Ranked results with scores
|
|
||||||
- **Note**: Equivalent to `find({like: query, limit: k})`
|
|
||||||
|
|
||||||
#### `find(query)`
|
|
||||||
**TRIPLE INTELLIGENCE** - The ultimate search method
|
|
||||||
- **query**: `object` - Complex query object supporting:
|
|
||||||
```typescript
|
|
||||||
{
|
|
||||||
// Vector similarity
|
|
||||||
like?: string | number[] | {id: string}, // Text, vector, or noun ID
|
|
||||||
|
|
||||||
// Metadata filtering
|
|
||||||
where?: {
|
|
||||||
field: value, // Exact match
|
|
||||||
field: {$in: [values]}, // In array
|
|
||||||
field: {$gt: value}, // Greater than
|
|
||||||
field: {$regex: pattern} // Pattern match
|
|
||||||
},
|
|
||||||
|
|
||||||
// Graph traversal
|
|
||||||
connected?: {
|
|
||||||
to?: string, // Target noun ID
|
|
||||||
from?: string, // Source noun ID
|
|
||||||
via?: string, // Relationship type
|
|
||||||
depth?: number // Traversal depth (default: 1)
|
|
||||||
},
|
|
||||||
|
|
||||||
// Control
|
|
||||||
limit?: number, // Max results (default: 10)
|
|
||||||
offset?: number, // Skip results
|
|
||||||
threshold?: number // Min similarity score
|
|
||||||
}
|
|
||||||
```
|
|
||||||
- **Returns**: `Promise<EnhancedSearchResult[]>` - Results with scores and explanations
|
|
||||||
|
|
||||||
#### `findSimilar(id, k?)`
|
|
||||||
Find nouns similar to an existing noun
|
|
||||||
- **id**: `string` - Reference noun ID
|
|
||||||
- **k**: `number` (default: 10) - Number of results
|
|
||||||
- **Returns**: `Promise<SearchResult[]>` - Similar nouns
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### Neural API
|
|
||||||
|
|
||||||
#### `neural.search(query, options?)`
|
|
||||||
Neural-enhanced semantic search
|
|
||||||
- **query**: `string` - Natural language query
|
|
||||||
- **options**: `{expand?: boolean, rerank?: boolean}` - Enhancement options
|
|
||||||
- **Returns**: `Promise<NeuralSearchResult[]>` - Enhanced results
|
|
||||||
|
|
||||||
#### `neural.cluster(options?)`
|
|
||||||
Automatic clustering of nouns
|
|
||||||
- **options**: `{k?: number, method?: 'kmeans'|'dbscan', minSize?: number}`
|
|
||||||
- **Returns**: `Promise<Cluster[]>` - Generated clusters
|
|
||||||
|
|
||||||
#### `neural.extract(text)`
|
|
||||||
Extract entities from text
|
|
||||||
- **text**: `string` - Text to analyze
|
|
||||||
- **Returns**: `Promise<{entities: Entity[], relationships: Relationship[]}>`
|
|
||||||
|
|
||||||
#### `neural.summarize(ids[])`
|
|
||||||
Generate summary from multiple nouns
|
|
||||||
- **ids**: `string[]` - Noun IDs to summarize
|
|
||||||
- **Returns**: `Promise<string>` - Generated summary
|
|
||||||
|
|
||||||
#### `neural.analyze(id)`
|
|
||||||
Deep analysis of a noun
|
|
||||||
- **id**: `string` - Noun ID to analyze
|
|
||||||
- **Returns**: `Promise<Analysis>` - Detailed analysis
|
|
||||||
|
|
||||||
#### `neural.compare(id1, id2)`
|
|
||||||
Semantic comparison of two nouns
|
|
||||||
- **id1**: `string` - First noun ID
|
|
||||||
- **id2**: `string` - Second noun ID
|
|
||||||
- **Returns**: `Promise<{similarity: number, differences: string[], commonalities: string[]}>`
|
|
||||||
|
|
||||||
#### `neural.topics(options?)`
|
|
||||||
Topic modeling across all nouns
|
|
||||||
- **options**: `{k?: number, method?: 'lda'|'nmf'}` - Topic extraction options
|
|
||||||
- **Returns**: `Promise<Topic[]>` - Discovered topics
|
|
||||||
|
|
||||||
#### `neural.patterns(options?)`
|
|
||||||
Pattern detection in data
|
|
||||||
- **options**: `{minSupport?: number, minConfidence?: number}`
|
|
||||||
- **Returns**: `Promise<Pattern[]>` - Detected patterns
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📥 IMPORT/EXPORT
|
|
||||||
|
|
||||||
### Neural Import
|
|
||||||
|
|
||||||
#### `neuralImport(data, options?)`
|
|
||||||
Smart AI-powered data import
|
|
||||||
- **data**: `any` - Data to import (auto-detects format)
|
|
||||||
- **options**: `{autoExtract?: boolean, autoRelate?: boolean, batchSize?: number}`
|
|
||||||
- **Returns**: `Promise<{nouns: string[], verbs: string[]}>`
|
|
||||||
|
|
||||||
#### `neuralImport.csv(file, options?)`
|
|
||||||
Import CSV with intelligent parsing
|
|
||||||
- **file**: `string | Buffer` - CSV file path or content
|
|
||||||
- **options**: `{headers?: boolean, delimiter?: string, embedColumns?: string[]}`
|
|
||||||
- **Returns**: `Promise<ImportResult>`
|
|
||||||
|
|
||||||
#### `neuralImport.json(data, options?)`
|
|
||||||
Import JSON with structure detection
|
|
||||||
- **data**: `object | string` - JSON data or string
|
|
||||||
- **options**: `{flatten?: boolean, keyPaths?: string[]}`
|
|
||||||
- **Returns**: `Promise<ImportResult>`
|
|
||||||
|
|
||||||
#### `neuralImport.text(text, options?)`
|
|
||||||
Import text with NLP processing
|
|
||||||
- **text**: `string` - Raw text
|
|
||||||
- **options**: `{chunk?: boolean, chunkSize?: number, extractEntities?: boolean}`
|
|
||||||
- **Returns**: `Promise<ImportResult>`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🔄 SYNC & DISTRIBUTION
|
|
||||||
|
|
||||||
### Real-time Sync
|
|
||||||
|
|
||||||
#### `sync.enable(config)`
|
|
||||||
Enable real-time synchronization
|
|
||||||
- **config**: `{url: string, interval?: number, bidirectional?: boolean}`
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
#### `sync.disable()`
|
|
||||||
Disable synchronization
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
#### `sync.now()`
|
|
||||||
Trigger manual sync
|
|
||||||
- **Returns**: `Promise<SyncResult>`
|
|
||||||
|
|
||||||
#### `sync.status()`
|
|
||||||
Get sync status
|
|
||||||
- **Returns**: `Promise<{enabled: boolean, lastSync: Date, pending: number}>`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### Remote Operations
|
|
||||||
|
|
||||||
#### `remote.connect(url, options?)`
|
|
||||||
Connect to remote Brainy instance
|
|
||||||
- **url**: `string` - Remote instance URL
|
|
||||||
- **options**: `{auth?: string, timeout?: number, retry?: boolean}`
|
|
||||||
- **Returns**: `Promise<Connection>`
|
|
||||||
|
|
||||||
#### `remote.search(query)`
|
|
||||||
Search remote instance
|
|
||||||
- **query**: `any` - Same as find() query
|
|
||||||
- **Returns**: `Promise<SearchResult[]>`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🧠 INTELLIGENCE FEATURES
|
|
||||||
|
|
||||||
### Verb Scoring
|
|
||||||
|
|
||||||
#### `verbScoring.train(feedback)`
|
|
||||||
Train the verb scoring model
|
|
||||||
- **feedback**: `{verbId: string, score: number, context?: object}`
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
#### `verbScoring.getScore(verbId)`
|
|
||||||
Get intelligent score for a verb
|
|
||||||
- **verbId**: `string` - The verb to score
|
|
||||||
- **Returns**: `Promise<number>` - Score between 0-1
|
|
||||||
|
|
||||||
#### `verbScoring.export()`
|
|
||||||
Export training data
|
|
||||||
- **Returns**: `Promise<TrainingData>`
|
|
||||||
|
|
||||||
#### `verbScoring.import(data)`
|
|
||||||
Import training data
|
|
||||||
- **data**: `TrainingData` - Previously exported data
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### Embeddings
|
|
||||||
|
|
||||||
#### `embed(text)`
|
|
||||||
Generate embedding vector for text
|
|
||||||
- **text**: `string` - Text to embed
|
|
||||||
- **Returns**: `Promise<number[]>` - Embedding vector
|
|
||||||
|
|
||||||
#### `embedBatch(texts[])`
|
|
||||||
Generate embeddings for multiple texts
|
|
||||||
- **texts**: `string[]` - Array of texts
|
|
||||||
- **Returns**: `Promise<number[][]>` - Array of vectors
|
|
||||||
|
|
||||||
#### `similarity(a, b, metric?)`
|
|
||||||
Calculate similarity between vectors or texts
|
|
||||||
- **a**: `string | number[]` - First item
|
|
||||||
- **b**: `string | number[]` - Second item
|
|
||||||
- **metric**: `'cosine' | 'euclidean' | 'manhattan'` (default: 'cosine')
|
|
||||||
- **Returns**: `Promise<number>` - Similarity score
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📊 MONITORING & PERFORMANCE
|
|
||||||
|
|
||||||
#### `size()`
|
|
||||||
Get total noun count
|
|
||||||
- **Returns**: `number` - Total nouns in database
|
|
||||||
|
|
||||||
#### `stats()`
|
|
||||||
Get comprehensive statistics
|
|
||||||
- **Returns**: `Promise<Statistics>` - Detailed stats
|
|
||||||
|
|
||||||
#### `health()`
|
|
||||||
System health check
|
|
||||||
- **Returns**: `Promise<{status: 'healthy'|'degraded'|'unhealthy', details: object}>`
|
|
||||||
|
|
||||||
#### `cache.stats()`
|
|
||||||
Get cache statistics
|
|
||||||
- **Returns**: `CacheStats` - Hit rates, size, etc.
|
|
||||||
|
|
||||||
#### `cache.clear()`
|
|
||||||
Clear all caches
|
|
||||||
- **Returns**: `void`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
#### `new BrainyData(config?)`
|
|
||||||
Create new Brainy instance
|
|
||||||
- **config**: `object` (optional)
|
|
||||||
```typescript
|
|
||||||
{
|
|
||||||
storage?: 'auto' | 'memory' | 'filesystem' | 's3',
|
|
||||||
dimensions?: number, // Vector dimensions (default: 384)
|
|
||||||
cache?: boolean, // Enable caching (default: true)
|
|
||||||
index?: boolean, // Enable indexing (default: true)
|
|
||||||
verbose?: boolean, // Verbose logging (default: false)
|
|
||||||
augmentations?: Augmentation[] // Custom augmentations
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
#### `init()`
|
|
||||||
Initialize the instance (REQUIRED!)
|
|
||||||
- **Returns**: `Promise<void>`
|
|
||||||
- **Note**: Must be called before any operations
|
|
||||||
|
|
||||||
#### `shutdown()`
|
|
||||||
Graceful shutdown
|
|
||||||
- **Returns**: `Promise<void>` - Saves state and closes connections
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📐 READ-ONLY PROPERTIES
|
|
||||||
|
|
||||||
- **dimensions**: `number` - Vector dimensions
|
|
||||||
- **initialized**: `boolean` - Whether init() was called
|
|
||||||
- **mode**: `string` - Current operational mode
|
|
||||||
|
|
@ -1,260 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Final Complete API
|
|
||||||
|
|
||||||
> **Clean, Powerful, Complete** - All features, beautifully organized.
|
|
||||||
|
|
||||||
## 📚 CORE DATA
|
|
||||||
|
|
||||||
### Nouns (Vectors with Metadata)
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addNoun(textOrVector, metadata?) // Add noun (auto-embeds text!)
|
|
||||||
getNoun(id) // Get one noun
|
|
||||||
updateNoun(id, textOrVector?, metadata?) // Update noun
|
|
||||||
deleteNoun(id) // Delete noun
|
|
||||||
hasNoun(id) // Check if exists
|
|
||||||
|
|
||||||
// Metadata Operations
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounWithVerbs(id) // Get noun with relationships
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNouns(idsOrOptions) // Get multiple nouns (unified)
|
|
||||||
deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verbs (Relationships)
|
|
||||||
```typescript
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get verb
|
|
||||||
deleteVerb(id) // Delete verb
|
|
||||||
getVerbsBySource(sourceId) // Outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Incoming relationships
|
|
||||||
getVerbsByType(type) // By relationship type
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH & INTELLIGENCE
|
|
||||||
|
|
||||||
### Core Search
|
|
||||||
```typescript
|
|
||||||
search(query, k?) // Simple vector search
|
|
||||||
find(query) // TRIPLE INTELLIGENCE 🧠
|
|
||||||
findSimilar(id, k?) // Find similar to noun
|
|
||||||
```
|
|
||||||
|
|
||||||
### Neural API
|
|
||||||
```typescript
|
|
||||||
neural.search(query) // Neural-enhanced search
|
|
||||||
neural.cluster(options?) // Automatic clustering
|
|
||||||
neural.extract(text) // Entity extraction
|
|
||||||
neural.summarize(ids[]) // Summarize nouns
|
|
||||||
neural.analyze(id) // Deep analysis
|
|
||||||
neural.compare(id1, id2) // Semantic comparison
|
|
||||||
neural.topics() // Topic modeling
|
|
||||||
neural.patterns() // Pattern detection
|
|
||||||
```
|
|
||||||
|
|
||||||
### Clustering
|
|
||||||
```typescript
|
|
||||||
clusters.create(options?) // Create clusters
|
|
||||||
clusters.get(id) // Get cluster
|
|
||||||
clusters.list() // List all clusters
|
|
||||||
clusters.addToCluster(clusterId, nounId) // Add to cluster
|
|
||||||
clusters.optimize() // Re-optimize clusters
|
|
||||||
clusters.suggest(nounId) // Suggest best cluster
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 INTELLIGENCE FEATURES
|
|
||||||
|
|
||||||
### Triple Intelligence
|
|
||||||
```typescript
|
|
||||||
tripleIntelligence.analyze(query) // Combined V+G+F analysis
|
|
||||||
tripleIntelligence.explain(results) // Explain search results
|
|
||||||
tripleIntelligence.optimize(query) // Query optimization
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verb Scoring
|
|
||||||
```typescript
|
|
||||||
verbScoring.train(feedback) // Train scoring model
|
|
||||||
verbScoring.getScore(verb) // Get verb score
|
|
||||||
verbScoring.export() // Export training data
|
|
||||||
verbScoring.import(data) // Import training data
|
|
||||||
verbScoring.stats() // Get statistics
|
|
||||||
```
|
|
||||||
|
|
||||||
### Embeddings
|
|
||||||
```typescript
|
|
||||||
embed(text) // Generate embedding
|
|
||||||
embedBatch(texts[]) // Batch embeddings
|
|
||||||
similarity(a, b) // Calculate similarity
|
|
||||||
distance(a, b, metric?) // Calculate distance
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📥 IMPORT/EXPORT
|
|
||||||
|
|
||||||
### Neural Import
|
|
||||||
```typescript
|
|
||||||
neuralImport(data, options?) // Smart data import
|
|
||||||
neuralImport.csv(file, options?) // Import CSV with AI
|
|
||||||
neuralImport.json(data, options?) // Import JSON with AI
|
|
||||||
neuralImport.text(text, options?) // Import text with NLP
|
|
||||||
neuralImport.url(url, options?) // Import from URL
|
|
||||||
neuralImport.batch(items[], options?) // Batch neural import
|
|
||||||
```
|
|
||||||
|
|
||||||
### Data Management
|
|
||||||
```typescript
|
|
||||||
import(data, format) // Standard import
|
|
||||||
export(format) // Export data
|
|
||||||
importSparse(data) // Import sparse format
|
|
||||||
backup() // Create backup
|
|
||||||
restore(backup) // Restore backup
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 SYNC & DISTRIBUTION
|
|
||||||
|
|
||||||
### Real-time Sync
|
|
||||||
```typescript
|
|
||||||
sync.enable(config) // Enable real-time sync
|
|
||||||
sync.disable() // Disable sync
|
|
||||||
sync.now() // Manual sync
|
|
||||||
sync.status() // Sync status
|
|
||||||
```
|
|
||||||
|
|
||||||
### Remote Operations
|
|
||||||
```typescript
|
|
||||||
remote.connect(url, options?) // Connect to remote
|
|
||||||
remote.disconnect() // Disconnect
|
|
||||||
remote.search(query) // Search remote
|
|
||||||
remote.sync() // Sync with remote
|
|
||||||
```
|
|
||||||
|
|
||||||
### Conduits (Brainy-to-Brainy)
|
|
||||||
```typescript
|
|
||||||
conduit.establish(url) // Create conduit
|
|
||||||
conduit.send(data) // Send via conduit
|
|
||||||
conduit.receive(callback) // Receive data
|
|
||||||
conduit.close() // Close conduit
|
|
||||||
```
|
|
||||||
|
|
||||||
### Synapses (External Platforms)
|
|
||||||
```typescript
|
|
||||||
synapse.notion.connect(config) // Connect to Notion
|
|
||||||
synapse.slack.connect(config) // Connect to Slack
|
|
||||||
synapse.salesforce.connect(config) // Connect to Salesforce
|
|
||||||
synapse.custom(platform, config) // Custom platform
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 ANALYTICS & MONITORING
|
|
||||||
|
|
||||||
### Statistics
|
|
||||||
```typescript
|
|
||||||
size() // Total count
|
|
||||||
stats() // Full statistics
|
|
||||||
stats.byService(service) // Per-service stats
|
|
||||||
stats.byType(type) // Per-type stats
|
|
||||||
health() // Health check
|
|
||||||
```
|
|
||||||
|
|
||||||
### Performance
|
|
||||||
```typescript
|
|
||||||
cache.stats() // Cache statistics
|
|
||||||
cache.clear() // Clear cache
|
|
||||||
cache.optimize() // Optimize cache
|
|
||||||
index.rebuild() // Rebuild index
|
|
||||||
index.optimize() // Optimize index
|
|
||||||
```
|
|
||||||
|
|
||||||
### Monitoring
|
|
||||||
```typescript
|
|
||||||
monitor.enable() // Enable monitoring
|
|
||||||
monitor.metrics() // Get metrics
|
|
||||||
monitor.alerts() // Get alerts
|
|
||||||
monitor.logs(options?) // Get logs
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
### Operational Modes
|
|
||||||
```typescript
|
|
||||||
setReadOnly(bool) // Read-only mode
|
|
||||||
setWriteOnly(bool) // Write-only mode
|
|
||||||
setFrozen(bool) // Freeze all changes
|
|
||||||
getMode() // Current mode
|
|
||||||
```
|
|
||||||
|
|
||||||
### Augmentations
|
|
||||||
```typescript
|
|
||||||
augmentations.register(augmentation) // Add augmentation
|
|
||||||
augmentations.list() // List all
|
|
||||||
augmentations.get(name) // Get by name
|
|
||||||
augmentations.enable(name) // Enable
|
|
||||||
augmentations.disable(name) // Disable
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔧 UTILITIES
|
|
||||||
|
|
||||||
### Data Operations
|
|
||||||
```typescript
|
|
||||||
clear(options?) // Clear all
|
|
||||||
clearNouns() // Clear nouns
|
|
||||||
clearVerbs() // Clear verbs
|
|
||||||
generateRandomGraph(nodes, edges) // Generate test data
|
|
||||||
```
|
|
||||||
|
|
||||||
### Field Management
|
|
||||||
```typescript
|
|
||||||
fields.list() // List indexed fields
|
|
||||||
fields.values(field) // Get unique values
|
|
||||||
fields.add(field) // Add field to index
|
|
||||||
fields.remove(field) // Remove from index
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Creation & Initialization
|
|
||||||
new BrainyData(config?) // Create instance
|
|
||||||
init() // Initialize (REQUIRED!)
|
|
||||||
warmup() // Warmup caches
|
|
||||||
|
|
||||||
// Cleanup
|
|
||||||
shutdown() // Graceful shutdown
|
|
||||||
cleanup() // Clean resources
|
|
||||||
|
|
||||||
// Static Methods
|
|
||||||
BrainyData.preloadModel() // Preload ML model
|
|
||||||
BrainyData.version // Get version
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📐 PROPERTIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
dimensions // Vector dimensions (readonly)
|
|
||||||
initialized // Is initialized (readonly)
|
|
||||||
mode // Current mode (readonly)
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🎯 Key Features Preserved
|
|
||||||
|
|
||||||
✅ **Neural Import** - Smart AI-powered data import
|
|
||||||
✅ **Clustering** - Automatic and manual clustering
|
|
||||||
✅ **Triple Intelligence** - Vector + Graph + Metadata combined
|
|
||||||
✅ **Verb Scoring** - Intelligent relationship scoring
|
|
||||||
✅ **Synapses** - External platform connectors
|
|
||||||
✅ **Conduits** - Brainy-to-Brainy sync
|
|
||||||
✅ **Neural API** - Advanced AI operations
|
|
||||||
✅ **Real-time Sync** - Live data synchronization
|
|
||||||
✅ **Monitoring** - Performance and health tracking
|
|
||||||
|
|
||||||
## 🚀 What's New in 2.0
|
|
||||||
|
|
||||||
1. **Auto-embedding** - `addNoun()` accepts text directly
|
|
||||||
2. **Unified `find()`** - One method for all complex queries
|
|
||||||
3. **Neural API** - Powerful AI operations built-in
|
|
||||||
4. **Augmentation System** - Extensible architecture
|
|
||||||
5. **Clean naming** - Specific noun/verb terminology
|
|
||||||
6. **No duplicates** - One method per operation
|
|
||||||
|
|
@ -1,243 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Final API Reference
|
|
||||||
|
|
||||||
> **The definitive API - Clean, Correct, Complete**
|
|
||||||
|
|
||||||
## ✅ KEY CORRECTIONS FROM REVIEW:
|
|
||||||
1. **Brainy Operators (NOT MongoDB)** - `greaterThan` not `$gt`
|
|
||||||
2. **Neural API is complete** - All methods available via `brain.neural`
|
|
||||||
3. **Code is correct** - Implementation uses right operators, just docs were wrong
|
|
||||||
4. **Nothing lost** - All features still present, just reorganized
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📚 CORE DATA OPERATIONS
|
|
||||||
|
|
||||||
### Nouns
|
|
||||||
```typescript
|
|
||||||
// Single
|
|
||||||
addNoun(textOrVector, metadata?) // Auto-embeds text!
|
|
||||||
getNoun(id)
|
|
||||||
updateNoun(id, textOrVector?, metadata?)
|
|
||||||
deleteNoun(id)
|
|
||||||
hasNoun(id)
|
|
||||||
|
|
||||||
// Metadata
|
|
||||||
getNounMetadata(id)
|
|
||||||
updateNounMetadata(id, metadata)
|
|
||||||
getNounWithVerbs(id)
|
|
||||||
|
|
||||||
// Batch
|
|
||||||
addNouns(items[])
|
|
||||||
getNouns(idsOrOptions) // Unified: IDs, filter, or pagination
|
|
||||||
deleteNouns(ids[])
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verbs
|
|
||||||
```typescript
|
|
||||||
addVerb(source, target, type, metadata?)
|
|
||||||
getVerb(id)
|
|
||||||
deleteVerb(id)
|
|
||||||
getVerbsBySource(sourceId)
|
|
||||||
getVerbsByTarget(targetId)
|
|
||||||
getVerbsByType(type)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH
|
|
||||||
|
|
||||||
Just TWO methods - simple and powerful:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
search(query, k?) // Convenience: same as find({like: query, limit: k})
|
|
||||||
find(query) // TRIPLE INTELLIGENCE: Vector + Graph + Metadata
|
|
||||||
```
|
|
||||||
|
|
||||||
### Find Query (with CORRECT Brainy Operators):
|
|
||||||
```typescript
|
|
||||||
find({
|
|
||||||
// Vector
|
|
||||||
like: 'text' | vector | {id: 'noun-id'},
|
|
||||||
|
|
||||||
// Fields (BRAINY operators, NOT MongoDB!)
|
|
||||||
where: {
|
|
||||||
field: value, // Direct equality
|
|
||||||
field: {
|
|
||||||
equals: value,
|
|
||||||
greaterThan: value, // NOT $gt
|
|
||||||
lessThan: value, // NOT $lt
|
|
||||||
greaterEqual: value,
|
|
||||||
lessEqual: value,
|
|
||||||
oneOf: [values], // NOT $in
|
|
||||||
notOneOf: [values], // NOT $nin
|
|
||||||
contains: value,
|
|
||||||
startsWith: value,
|
|
||||||
endsWith: value,
|
|
||||||
matches: pattern, // NOT $regex
|
|
||||||
between: [min, max]
|
|
||||||
}
|
|
||||||
},
|
|
||||||
|
|
||||||
// Graph
|
|
||||||
connected: {
|
|
||||||
to: 'id',
|
|
||||||
from: 'id',
|
|
||||||
via: 'type',
|
|
||||||
depth: 2
|
|
||||||
},
|
|
||||||
|
|
||||||
// Control
|
|
||||||
limit: 10,
|
|
||||||
offset: 0,
|
|
||||||
explain: true
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 NEURAL API
|
|
||||||
|
|
||||||
Complete and available via `brain.neural`:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.neural.similar(a, b) // Similarity 0-1
|
|
||||||
brain.neural.clusters() // Auto-clustering
|
|
||||||
brain.neural.hierarchy(id) // Semantic tree
|
|
||||||
brain.neural.neighbors(id, k?) // K-nearest
|
|
||||||
brain.neural.outliers(threshold?) // Outlier detection
|
|
||||||
brain.neural.semanticPath(from, to) // Path finding
|
|
||||||
brain.neural.visualize(options?) // For D3/Cytoscape/GraphML
|
|
||||||
|
|
||||||
// Performance
|
|
||||||
brain.neural.clusterFast() // O(n) HNSW
|
|
||||||
brain.neural.clusterLarge() // Million+ items
|
|
||||||
brain.neural.clusterStream() // Progressive
|
|
||||||
```
|
|
||||||
|
|
||||||
### Visualization Format:
|
|
||||||
```typescript
|
|
||||||
brain.neural.visualize({
|
|
||||||
maxNodes: 100,
|
|
||||||
dimensions: 2,
|
|
||||||
algorithm: 'force',
|
|
||||||
includeEdges: true
|
|
||||||
})
|
|
||||||
// Returns: {
|
|
||||||
// format: 'd3' | 'cytoscape' | 'graphml',
|
|
||||||
// nodes: [...], edges: [...], layout: {...}
|
|
||||||
// }
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📥 IMPORT
|
|
||||||
|
|
||||||
Simple, AI-powered:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.neuralImport(data, options?) // Auto-detects format!
|
|
||||||
// Options: {
|
|
||||||
// confidenceThreshold: 0.7,
|
|
||||||
// autoApply: false,
|
|
||||||
// skipDuplicates: true
|
|
||||||
// }
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 INTELLIGENCE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Verb Scoring
|
|
||||||
provideFeedbackForVerbScoring(feedback)
|
|
||||||
getVerbScoringStats()
|
|
||||||
exportVerbScoringLearningData()
|
|
||||||
importVerbScoringLearningData(data)
|
|
||||||
|
|
||||||
// Embeddings
|
|
||||||
embed(text) // Generate vector
|
|
||||||
calculateSimilarity(a, b, metric?) // Compare
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 SYNC
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Remote
|
|
||||||
connectToRemoteServer(url)
|
|
||||||
disconnectFromRemoteServer()
|
|
||||||
isConnectedToRemoteServer()
|
|
||||||
|
|
||||||
// Real-time
|
|
||||||
enableRealtimeUpdates(config)
|
|
||||||
disableRealtimeUpdates()
|
|
||||||
checkForUpdatesNow()
|
|
||||||
|
|
||||||
// Search modes
|
|
||||||
searchLocal(query, k?)
|
|
||||||
searchRemote(query, k?)
|
|
||||||
searchCombined(query, k?)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 MONITORING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
size() // Total nouns
|
|
||||||
getStatistics() // Full stats
|
|
||||||
getHealthStatus() // Health
|
|
||||||
getCacheStats() // Cache
|
|
||||||
clearCache() // Clear
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Modes
|
|
||||||
setReadOnly(bool)
|
|
||||||
setWriteOnly(bool)
|
|
||||||
setFrozen(bool)
|
|
||||||
|
|
||||||
// Augmentations
|
|
||||||
augmentations.register(aug)
|
|
||||||
augmentations.list()
|
|
||||||
augmentations.get(name)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
clear(options?) // Clear all
|
|
||||||
clearNouns() // Nouns only
|
|
||||||
clearVerbs() // Verbs only
|
|
||||||
backup() // Create backup
|
|
||||||
restore(backup) // Restore
|
|
||||||
rebuildMetadataIndex() // Rebuild index
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: 'auto', // auto | memory | filesystem | s3
|
|
||||||
dimensions: 384,
|
|
||||||
cache: true,
|
|
||||||
index: true
|
|
||||||
})
|
|
||||||
|
|
||||||
await brain.init() // REQUIRED!
|
|
||||||
await brain.shutdown() // Cleanup
|
|
||||||
|
|
||||||
// Static
|
|
||||||
BrainyData.preloadModel() // Preload
|
|
||||||
BrainyData.warmup() // Warmup
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## ✨ What Makes Brainy 2.0 Special:
|
|
||||||
|
|
||||||
1. **Zero-Config** - Works instantly, no setup
|
|
||||||
2. **Auto-Embedding** - Text automatically becomes vectors
|
|
||||||
3. **Triple Intelligence** - Vector + Graph + Metadata combined
|
|
||||||
4. **Brainy Operators** - Clean, legal, no MongoDB style
|
|
||||||
5. **Complete Neural API** - All clustering/viz features
|
|
||||||
6. **Simple Import** - One method, auto-detects everything
|
|
||||||
7. **Clean Architecture** - Augmentations for extensibility
|
|
||||||
|
|
||||||
## 🎯 Remember:
|
|
||||||
- **NO $operators** - We use readable names (legal requirement)
|
|
||||||
- **search() is simple** - Just wraps find({like: query})
|
|
||||||
- **find() is powerful** - Full Triple Intelligence
|
|
||||||
- **neural API complete** - All methods via brain.neural
|
|
||||||
- **Everything included** - No premium features, all MIT
|
|
||||||
|
|
@ -1,149 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Simplified Public API
|
|
||||||
|
|
||||||
> **Ultra-clean, Simple, Powerful** - Minimal methods, maximum capability.
|
|
||||||
|
|
||||||
## 📚 NOUNS (Data with Vectors)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addNoun(textOrVector, metadata?) // Add noun (auto-embeds text!)
|
|
||||||
getNoun(id) // Get one noun
|
|
||||||
updateNoun(id, textOrVector?, metadata?) // Update noun
|
|
||||||
deleteNoun(id) // Delete noun
|
|
||||||
hasNoun(id) // Check if exists
|
|
||||||
|
|
||||||
// Metadata Operations
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounWithVerbs(id) // Get noun with relationships
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNouns(idsOrOptions) // Get multiple nouns (unified)
|
|
||||||
deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔗 VERBS (Relationships)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Core Operations
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get verb
|
|
||||||
deleteVerb(id) // Delete verb
|
|
||||||
|
|
||||||
// Queries
|
|
||||||
getVerbsBySource(sourceId) // Outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Incoming relationships
|
|
||||||
getVerbsByType(type) // By relationship type
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH (One Method to Rule Them All)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// THE ONLY SEARCH METHODS YOU NEED:
|
|
||||||
search(query, k?) // Simple vector search (alias to find)
|
|
||||||
find(query) // TRIPLE INTELLIGENCE 🧠
|
|
||||||
```
|
|
||||||
|
|
||||||
### Find Query Examples:
|
|
||||||
```typescript
|
|
||||||
// Text search (auto-embeds)
|
|
||||||
find('documents about AI')
|
|
||||||
|
|
||||||
// Similar to existing noun
|
|
||||||
find({ like: 'noun-id-123' })
|
|
||||||
|
|
||||||
// Metadata filtering
|
|
||||||
find({ where: { type: 'article' }})
|
|
||||||
|
|
||||||
// Graph traversal
|
|
||||||
find({ connected: { to: 'id', via: 'references' }})
|
|
||||||
|
|
||||||
// Combined queries (Triple Intelligence!)
|
|
||||||
find({
|
|
||||||
like: 'sample-doc', // Vector similarity
|
|
||||||
where: { status: 'published' }, // Metadata filter
|
|
||||||
connected: { via: 'cites' }, // Graph relationships
|
|
||||||
limit: 10 // Pagination
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 METADATA
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
getFilterableFields() // Get indexed fields
|
|
||||||
getFieldValues(field) // Get unique values for field
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 PERFORMANCE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Cache
|
|
||||||
getCacheStats() // Cache statistics
|
|
||||||
clearCache() // Clear cache
|
|
||||||
|
|
||||||
// Stats
|
|
||||||
size() // Total count
|
|
||||||
getStatistics() // Full statistics
|
|
||||||
getHealthStatus() // Health check
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Modes
|
|
||||||
setReadOnly(bool) // Read-only mode
|
|
||||||
setWriteOnly(bool) // Write-only mode
|
|
||||||
setFrozen(bool) // Freeze all changes
|
|
||||||
|
|
||||||
// Remote Sync
|
|
||||||
connectRemote(url) // Connect to remote
|
|
||||||
disconnectRemote() // Disconnect
|
|
||||||
syncNow() // Manual sync
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
clear(options?) // Clear all
|
|
||||||
clearNouns() // Clear nouns
|
|
||||||
clearVerbs() // Clear verbs
|
|
||||||
backup() // Create backup
|
|
||||||
restore(backup) // Restore backup
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
new BrainyData(config?) // Create
|
|
||||||
init() // Initialize (REQUIRED!)
|
|
||||||
shutdown() // Cleanup
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🎯 Philosophy
|
|
||||||
|
|
||||||
### Why So Simple?
|
|
||||||
|
|
||||||
1. **`addNoun()` handles everything** - Text? Auto-embeds. Vector? Uses directly.
|
|
||||||
2. **`find()` is the ultimate search** - Combines vector, graph, and metadata search
|
|
||||||
3. **`search()` is just convenience** - Simple alias to `find()` for basic queries
|
|
||||||
4. **No duplicate methods** - One way to do each thing
|
|
||||||
|
|
||||||
### The Power of Find
|
|
||||||
|
|
||||||
The `find()` method is your Swiss Army knife:
|
|
||||||
- Text search → Auto-embeds and searches
|
|
||||||
- Vector search → `{ like: 'id' }` or `{ like: vector }`
|
|
||||||
- Metadata search → `{ where: { field: value }}`
|
|
||||||
- Graph search → `{ connected: { to/from: 'id' }}`
|
|
||||||
- Combine them all → Triple Intelligence!
|
|
||||||
|
|
||||||
### Zero Configuration
|
|
||||||
|
|
||||||
Everything just works:
|
|
||||||
- Text auto-embeds
|
|
||||||
- Vectors auto-index
|
|
||||||
- Metadata auto-indexes
|
|
||||||
- Relationships auto-optimize
|
|
||||||
|
|
@ -1,343 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Unified Public API
|
|
||||||
|
|
||||||
> **The complete, accurate API based on actual implementation**
|
|
||||||
|
|
||||||
## 📚 CORE DATA OPERATIONS
|
|
||||||
|
|
||||||
### Nouns (Vectors with Metadata)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === SINGLE OPERATIONS ===
|
|
||||||
addNoun(textOrVector, metadata?) // Add noun (auto-embeds text)
|
|
||||||
getNoun(id) // Get one noun
|
|
||||||
updateNoun(id, textOrVector?, metadata?) // Update noun
|
|
||||||
deleteNoun(id) // Delete noun
|
|
||||||
hasNoun(id) // Check if exists
|
|
||||||
|
|
||||||
// === METADATA OPERATIONS ===
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounWithVerbs(id) // Get noun with relationships
|
|
||||||
|
|
||||||
// === BATCH OPERATIONS ===
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNouns(idsOrOptions) // Get multiple (unified method)
|
|
||||||
// getNouns(['id1', 'id2']) // By IDs
|
|
||||||
// getNouns({filter: {...}}) // By filter
|
|
||||||
// getNouns({limit: 10, offset: 20}) // Paginated
|
|
||||||
deleteNouns(ids[]) // Delete multiple
|
|
||||||
```
|
|
||||||
|
|
||||||
### Verbs (Relationships)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === SINGLE OPERATIONS ===
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get verb
|
|
||||||
deleteVerb(id) // Delete verb
|
|
||||||
|
|
||||||
// === QUERY OPERATIONS ===
|
|
||||||
getVerbsBySource(sourceId) // Outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Incoming relationships
|
|
||||||
getVerbsByType(type) // By relationship type
|
|
||||||
getVerbs(filter?) // Get filtered verbs
|
|
||||||
deleteVerbs(ids[]) // Delete multiple
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH & INTELLIGENCE
|
|
||||||
|
|
||||||
### Primary Search Methods
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === TWO MAIN METHODS ===
|
|
||||||
search(query, k?) // Simple vector search
|
|
||||||
// Equivalent to: find({like: query, limit: k})
|
|
||||||
|
|
||||||
find(query) // TRIPLE INTELLIGENCE 🧠
|
|
||||||
// Combines Vector + Graph + Metadata search
|
|
||||||
```
|
|
||||||
|
|
||||||
### Find Query Structure
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
find({
|
|
||||||
// === VECTOR SEARCH ===
|
|
||||||
like: 'text query' | vector | {id: 'noun-id'},
|
|
||||||
similar: 'text' | vector, // Alternative to 'like'
|
|
||||||
|
|
||||||
// === FIELD FILTERING (Brainy Operators) ===
|
|
||||||
where: {
|
|
||||||
// Direct equality
|
|
||||||
field: value,
|
|
||||||
|
|
||||||
// Brainy operators (CORRECT - NO MongoDB $)
|
|
||||||
field: {
|
|
||||||
equals: value, // Exact match
|
|
||||||
is: value, // Same as equals
|
|
||||||
greaterThan: value, // Greater than
|
|
||||||
lessThan: value, // Less than
|
|
||||||
oneOf: [values], // In array (NOT $in)
|
|
||||||
contains: value // Array/string contains
|
|
||||||
// Note: Additional operators can be added
|
|
||||||
}
|
|
||||||
},
|
|
||||||
|
|
||||||
// === GRAPH TRAVERSAL ===
|
|
||||||
connected: {
|
|
||||||
to: 'id' | ['id1', 'id2'], // Target nodes
|
|
||||||
from: 'id' | ['id1', 'id2'], // Source nodes
|
|
||||||
type: 'type' | ['type1'], // Relationship types
|
|
||||||
depth: 2, // Traversal depth
|
|
||||||
direction: 'in' | 'out' | 'both'
|
|
||||||
},
|
|
||||||
|
|
||||||
// === CONTROL OPTIONS ===
|
|
||||||
limit: 10, // Max results
|
|
||||||
offset: 0, // Skip results
|
|
||||||
explain: false, // Add explanations
|
|
||||||
boost: 'recent' | 'popular' // Result boosting
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 NEURAL API
|
|
||||||
|
|
||||||
Access via `brain.neural`:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === SIMILARITY & CLUSTERING ===
|
|
||||||
brain.neural.similar(a, b, options?) // Semantic similarity (0-1)
|
|
||||||
brain.neural.clusters(input?) // Auto-clustering
|
|
||||||
brain.neural.hierarchy(id) // Semantic hierarchy tree
|
|
||||||
brain.neural.neighbors(id, options?) // K-nearest neighbors
|
|
||||||
|
|
||||||
// === ANALYSIS ===
|
|
||||||
brain.neural.outliers(threshold?) // Outlier detection
|
|
||||||
brain.neural.semanticPath(from, to) // Find semantic path
|
|
||||||
|
|
||||||
// === VISUALIZATION ===
|
|
||||||
brain.neural.visualize(options?) // Export for visualization
|
|
||||||
// options: {
|
|
||||||
// maxNodes: 100,
|
|
||||||
// dimensions: 2 | 3,
|
|
||||||
// algorithm: 'force' | 'hierarchical' | 'radial',
|
|
||||||
// includeEdges: true
|
|
||||||
// }
|
|
||||||
// Returns: {
|
|
||||||
// format: 'd3' | 'cytoscape' | 'graphml',
|
|
||||||
// nodes: [...], edges: [...], layout: {...}
|
|
||||||
// }
|
|
||||||
|
|
||||||
// === PERFORMANCE METHODS ===
|
|
||||||
brain.neural.clusterFast(options?) // O(n) HNSW clustering
|
|
||||||
brain.neural.clusterLarge(options?) // Million+ items
|
|
||||||
brain.neural.clusterStream(options?) // Progressive streaming
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📥 IMPORT & EXPORT
|
|
||||||
|
|
||||||
### Neural Import
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === SMART IMPORT (from cortex) ===
|
|
||||||
brain.neuralImport(filePath, options?) // AI-powered import
|
|
||||||
// options: {
|
|
||||||
// confidenceThreshold: 0.7,
|
|
||||||
// autoApply: false,
|
|
||||||
// enableWeights: true,
|
|
||||||
// previewOnly: false,
|
|
||||||
// skipDuplicates: true
|
|
||||||
// }
|
|
||||||
// Returns: {
|
|
||||||
// detectedEntities: [...],
|
|
||||||
// detectedRelationships: [...],
|
|
||||||
// confidence: 0.85,
|
|
||||||
// insights: [...],
|
|
||||||
// preview: "..."
|
|
||||||
// }
|
|
||||||
```
|
|
||||||
|
|
||||||
### Standard Import/Export
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import(data, format) // Standard import
|
|
||||||
importSparseData(data) // Sparse format import
|
|
||||||
backup() // Create full backup
|
|
||||||
restore(backup) // Restore from backup
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 VERB SCORING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === INTELLIGENT SCORING ===
|
|
||||||
provideFeedbackForVerbScoring(feedback) // Train model
|
|
||||||
getVerbScoringStats() // Get statistics
|
|
||||||
exportVerbScoringLearningData() // Export training
|
|
||||||
importVerbScoringLearningData(data) // Import training
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 SYNC & DISTRIBUTION
|
|
||||||
|
|
||||||
### Remote Operations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === REMOTE CONNECTION ===
|
|
||||||
connectToRemoteServer(url, options?) // Connect to remote
|
|
||||||
disconnectFromRemoteServer() // Disconnect
|
|
||||||
isConnectedToRemoteServer() // Check status
|
|
||||||
|
|
||||||
// === SEARCH MODES ===
|
|
||||||
searchLocal(query, k?) // Local only
|
|
||||||
searchRemote(query, k?) // Remote only
|
|
||||||
searchCombined(query, k?) // Both sources
|
|
||||||
```
|
|
||||||
|
|
||||||
### Real-time Sync
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
enableRealtimeUpdates(config) // Enable sync
|
|
||||||
disableRealtimeUpdates() // Disable sync
|
|
||||||
getRealtimeUpdateConfig() // Get config
|
|
||||||
checkForUpdatesNow() // Manual sync
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 MONITORING & STATS
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === STATISTICS ===
|
|
||||||
size() // Total noun count
|
|
||||||
getStatistics(options?) // Full statistics
|
|
||||||
getServiceStatistics(service) // Per-service stats
|
|
||||||
listServices() // List all services
|
|
||||||
flushStatistics() // Persist stats
|
|
||||||
|
|
||||||
// === HEALTH & CACHE ===
|
|
||||||
getHealthStatus() // System health
|
|
||||||
status() // Full status report
|
|
||||||
getCacheStats() // Cache statistics
|
|
||||||
clearCache() // Clear all caches
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
### Operational Modes
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === MODE CONTROL ===
|
|
||||||
isReadOnly() / setReadOnly(bool) // Read-only mode
|
|
||||||
isWriteOnly() / setWriteOnly(bool) // Write-only mode
|
|
||||||
isFrozen() / setFrozen(bool) // Freeze all changes
|
|
||||||
```
|
|
||||||
|
|
||||||
### Augmentations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === AUGMENTATION SYSTEM ===
|
|
||||||
augmentations.register(augmentation) // Add augmentation
|
|
||||||
augmentations.list() // List all
|
|
||||||
augmentations.get(name) // Get by name
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === CLEAR OPERATIONS ===
|
|
||||||
clear(options?) // Clear all data
|
|
||||||
clearNouns(options?) // Clear nouns only
|
|
||||||
clearVerbs(options?) // Clear verbs only
|
|
||||||
|
|
||||||
// === INDEX MANAGEMENT ===
|
|
||||||
rebuildMetadataIndex() // Rebuild index
|
|
||||||
getFilterFields() // Get indexed fields
|
|
||||||
getFilterValues(field) // Get unique values
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧬 EMBEDDINGS & SIMILARITY
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
embed(text) // Generate embedding
|
|
||||||
calculateSimilarity(a, b, metric?) // Calculate similarity
|
|
||||||
// metric: 'cosine' | 'euclidean' | 'manhattan'
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔒 SECURITY
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
encryptData(data) // Encrypt data
|
|
||||||
decryptData(data) // Decrypt data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎲 UTILITIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
generateRandomGraph(nodes, edges) // Generate test data
|
|
||||||
getAvailableFieldNames() // Get field names
|
|
||||||
getStandardFieldMappings() // Get field mappings
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// === INITIALIZATION ===
|
|
||||||
const brain = new BrainyData(config?) // Create instance
|
|
||||||
await brain.init() // Initialize (REQUIRED!)
|
|
||||||
await brain.shutDown() // Graceful shutdown
|
|
||||||
await brain.cleanup() // Clean resources
|
|
||||||
|
|
||||||
// === STATIC METHODS ===
|
|
||||||
BrainyData.preloadModel(options?) // Preload ML model
|
|
||||||
BrainyData.warmup(options?) // Warmup system
|
|
||||||
```
|
|
||||||
|
|
||||||
### Configuration Options
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
new BrainyData({
|
|
||||||
// Storage
|
|
||||||
storage: 'auto' | 'memory' | 'filesystem' | 's3' | {
|
|
||||||
adapter: 'custom',
|
|
||||||
// ... storage options
|
|
||||||
},
|
|
||||||
|
|
||||||
// Vector configuration
|
|
||||||
dimensions: 384, // Vector dimensions
|
|
||||||
similarity: 'cosine', // Similarity metric
|
|
||||||
|
|
||||||
// Performance
|
|
||||||
cache: true, // Enable caching
|
|
||||||
index: true, // Enable indexing
|
|
||||||
metrics: true, // Enable metrics
|
|
||||||
|
|
||||||
// Advanced
|
|
||||||
augmentations: [...], // Custom augmentations
|
|
||||||
verbose: false // Logging verbosity
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📐 READ-ONLY PROPERTIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.dimensions // Vector dimensions
|
|
||||||
brain.maxConnections // HNSW max connections
|
|
||||||
brain.efConstruction // HNSW ef construction
|
|
||||||
brain.initialized // Is initialized?
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## ✅ KEY POINTS:
|
|
||||||
|
|
||||||
1. **Brainy Operators** - NOT MongoDB style ($gt, $lt)
|
|
||||||
2. **Neural API** - Complete with visualization export
|
|
||||||
3. **Simple Search** - Just `search()` and `find()`
|
|
||||||
4. **Triple Intelligence** - Vector + Graph + Metadata in `find()`
|
|
||||||
5. **Auto-embedding** - `addNoun()` accepts text directly
|
|
||||||
6. **Unified Methods** - `getNouns()` handles all plural queries
|
|
||||||
7. **Clean Architecture** - Augmentation system for extensibility
|
|
||||||
|
|
||||||
## ⚠️ IMPORTANT NOTES:
|
|
||||||
|
|
||||||
- **NO MongoDB operators** - We use `greaterThan` not `$gt` (legal reasons)
|
|
||||||
- **Neural API is via brain.neural** - All clustering/viz methods available
|
|
||||||
- **search() is convenience** - Just wraps `find({like: query})`
|
|
||||||
- **find() is powerful** - Full Triple Intelligence capabilities
|
|
||||||
- **One import method** - `neuralImport()` auto-detects format
|
|
||||||
|
|
@ -1,227 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Complete Public API
|
|
||||||
|
|
||||||
> **ONE METHOD, ONE PURPOSE** - No duplicates, no aliases, just clean specific methods.
|
|
||||||
|
|
||||||
## 📚 NOUNS (Vectors with Metadata)
|
|
||||||
|
|
||||||
### Single Operations
|
|
||||||
```typescript
|
|
||||||
addNoun(vector, metadata?) // Add one noun
|
|
||||||
getNoun(id) // Get one noun
|
|
||||||
updateNoun(id, vector?, metadata?) // Update noun
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
deleteNoun(id) // Delete one noun
|
|
||||||
hasNoun(id) // Check if exists
|
|
||||||
getNounWithVerbs(id) // Get with relationships
|
|
||||||
```
|
|
||||||
|
|
||||||
### Batch Operations
|
|
||||||
```typescript
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNounsByIds(ids[]) // Get multiple by IDs
|
|
||||||
deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
queryNouns(options) // Query with filters/pagination
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔗 VERBS (Relationships)
|
|
||||||
|
|
||||||
### Single Operations
|
|
||||||
```typescript
|
|
||||||
addVerb(source, target, type, metadata?) // Add relationship
|
|
||||||
getVerb(id) // Get one verb
|
|
||||||
deleteVerb(id) // Delete one verb
|
|
||||||
```
|
|
||||||
|
|
||||||
### Batch Operations
|
|
||||||
```typescript
|
|
||||||
addVerbs(verbs[]) // Add multiple verbs
|
|
||||||
getVerbs(filter?) // Get filtered verbs
|
|
||||||
deleteVerbs(ids[]) // Delete multiple verbs
|
|
||||||
getVerbsBySource(sourceId) // Get outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Get incoming relationships
|
|
||||||
getVerbsByType(type) // Get by relationship type
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH
|
|
||||||
|
|
||||||
### Vector Search
|
|
||||||
```typescript
|
|
||||||
search(query, k?, options?) // Primary vector search
|
|
||||||
searchText(text, k?, options?) // Natural language search
|
|
||||||
findSimilar(id, k?, options?) // Find similar to noun
|
|
||||||
searchWithCursor(query, cursor) // Paginated search
|
|
||||||
```
|
|
||||||
|
|
||||||
### Advanced Search
|
|
||||||
```typescript
|
|
||||||
find(query) // Triple Intelligence 🧠
|
|
||||||
searchByNounTypes(types[], query) // Filter by noun types
|
|
||||||
searchWithinItems(query, ids[], k?) // Search within specific nouns
|
|
||||||
searchVerbs(query) // Search relationships
|
|
||||||
searchNounsByVerbs(conditions) // Graph-based noun search
|
|
||||||
searchByStandardField(field, value) // Metadata-based search
|
|
||||||
```
|
|
||||||
|
|
||||||
### Distributed Search
|
|
||||||
```typescript
|
|
||||||
searchLocal(query) // Local only
|
|
||||||
searchRemote(query) // Remote only
|
|
||||||
searchCombined(query) // Both sources
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 GRAPH OPERATIONS
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
getConnections(id, depth?) // Get all connections
|
|
||||||
getVerbsBySource(sourceId) // Outgoing edges
|
|
||||||
getVerbsByTarget(targetId) // Incoming edges
|
|
||||||
getVerbsByType(type) // Filter by type
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 PERFORMANCE & STATS
|
|
||||||
|
|
||||||
### Cache
|
|
||||||
```typescript
|
|
||||||
getCacheStats() // Cache statistics
|
|
||||||
clearCache() // Clear search cache
|
|
||||||
```
|
|
||||||
|
|
||||||
### Statistics
|
|
||||||
```typescript
|
|
||||||
size() // Total noun count
|
|
||||||
getStatistics(options?) // Complete statistics
|
|
||||||
getServiceStatistics(service) // Per-service stats
|
|
||||||
listServices() // List all services
|
|
||||||
flushStatistics() // Flush to storage
|
|
||||||
```
|
|
||||||
|
|
||||||
### Health & Status
|
|
||||||
```typescript
|
|
||||||
getHealthStatus() // System health
|
|
||||||
status() // Full status report
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔧 CONFIGURATION & MODES
|
|
||||||
|
|
||||||
### Operational Modes
|
|
||||||
```typescript
|
|
||||||
isReadOnly() / setReadOnly(bool) // Read-only mode
|
|
||||||
isWriteOnly() / setWriteOnly(bool) // Write-only mode
|
|
||||||
isFrozen() / setFrozen(bool) // Freeze modifications
|
|
||||||
```
|
|
||||||
|
|
||||||
### Real-time Updates
|
|
||||||
```typescript
|
|
||||||
enableRealtimeUpdates(config) // Enable live sync
|
|
||||||
disableRealtimeUpdates() // Disable sync
|
|
||||||
getRealtimeUpdateConfig() // Get config
|
|
||||||
checkForUpdatesNow() // Manual sync
|
|
||||||
```
|
|
||||||
|
|
||||||
### Remote Connection
|
|
||||||
```typescript
|
|
||||||
connectToRemoteServer(url, options?) // Connect remote
|
|
||||||
disconnectFromRemoteServer() // Disconnect
|
|
||||||
isConnectedToRemoteServer() // Check connection
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 INTELLIGENCE FEATURES
|
|
||||||
|
|
||||||
### Verb Scoring
|
|
||||||
```typescript
|
|
||||||
provideFeedbackForVerbScoring(feedback) // Train scoring
|
|
||||||
getVerbScoringStats() // Get statistics
|
|
||||||
exportVerbScoringLearningData() // Export training
|
|
||||||
importVerbScoringLearningData(data) // Import training
|
|
||||||
```
|
|
||||||
|
|
||||||
### Embeddings
|
|
||||||
```typescript
|
|
||||||
embed(text) // Generate embedding
|
|
||||||
calculateSimilarity(a, b) // Compare vectors
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
### Clear Operations
|
|
||||||
```typescript
|
|
||||||
clear(options?) // Clear all data
|
|
||||||
clearNouns(options?) // Clear all nouns
|
|
||||||
clearVerbs(options?) // Clear all verbs
|
|
||||||
```
|
|
||||||
|
|
||||||
### Import/Export
|
|
||||||
```typescript
|
|
||||||
backup() // Create backup
|
|
||||||
restore(backup) // Restore backup
|
|
||||||
import(data, format) // Import data
|
|
||||||
importSparseData(data) // Import sparse
|
|
||||||
```
|
|
||||||
|
|
||||||
### Index Management
|
|
||||||
```typescript
|
|
||||||
rebuildMetadataIndex() // Rebuild index
|
|
||||||
getFilterFields() // Get indexed fields
|
|
||||||
getFilterValues(field) // Get field values
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔒 SECURITY
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
encryptData(data) // Encrypt
|
|
||||||
decryptData(data) // Decrypt
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎲 UTILITIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
generateRandomGraph(nodes, edges) // Generate test data
|
|
||||||
getAvailableFieldNames() // Available fields
|
|
||||||
getStandardFieldMappings() // Field mappings
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
### Instance Methods
|
|
||||||
```typescript
|
|
||||||
init() // Initialize (required!)
|
|
||||||
shutDown() // Graceful shutdown
|
|
||||||
cleanup() // Clean resources
|
|
||||||
```
|
|
||||||
|
|
||||||
### Static Methods
|
|
||||||
```typescript
|
|
||||||
BrainyData.preloadModel(options?) // Preload ML model
|
|
||||||
BrainyData.warmup(options?) // Warmup system
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📐 PROPERTIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
dimensions // Vector dimensions (readonly)
|
|
||||||
maxConnections // HNSW max connections (readonly)
|
|
||||||
efConstruction // HNSW ef construction (readonly)
|
|
||||||
initialized // Is initialized (readonly)
|
|
||||||
```
|
|
||||||
|
|
||||||
## ❌ REMOVED/PRIVATE IN 2.0
|
|
||||||
|
|
||||||
These methods are now **private** - use the new specific methods above:
|
|
||||||
- ~~add()~~ → Use `addNoun()`
|
|
||||||
- ~~get()~~ → Use `getNoun()`
|
|
||||||
- ~~delete()~~ → Use `deleteNoun()`
|
|
||||||
- ~~update()~~ → Use `updateNoun()`
|
|
||||||
- ~~relate()~~ → Use `addVerb()`
|
|
||||||
- ~~connect()~~ → Use `addVerb()`
|
|
||||||
- ~~has()~~ → Use `hasNoun()`
|
|
||||||
- ~~exists()~~ → Use `hasNoun()`
|
|
||||||
- ~~getMetadata()~~ → Use `getNounMetadata()`
|
|
||||||
- ~~updateMetadata()~~ → Use `updateNounMetadata()`
|
|
||||||
- ~~clearAll()~~ → Use `clear()`
|
|
||||||
- ~~addItem()~~ → Removed
|
|
||||||
- ~~addToBoth()~~ → Removed
|
|
||||||
- ~~addBatch()~~ → Use `addNouns()`
|
|
||||||
- ~~getBatch()~~ → Use `getNounsByIds()`
|
|
||||||
- ~~getNouns()~~ with IDs → Use `getNounsByIds()`
|
|
||||||
- ~~getNouns()~~ with filters → Use `queryNouns()`
|
|
||||||
|
|
@ -1,185 +0,0 @@
|
||||||
# 🚨 CRITICAL API AUDIT - What We Changed & Lost
|
|
||||||
|
|
||||||
## 📅 Timeline of Changes (Friday-Saturday)
|
|
||||||
|
|
||||||
### Friday Changes:
|
|
||||||
1. Started unifying augmentation system to single BrainyAugmentation interface
|
|
||||||
2. Made old methods (add, get, delete) private
|
|
||||||
3. Created new specific methods (addNoun, getNoun, deleteNoun)
|
|
||||||
4. Started removing backward compatibility
|
|
||||||
|
|
||||||
### Saturday Changes:
|
|
||||||
1. Combined getNounsByIds and queryNouns into single getNouns method
|
|
||||||
2. Simplified search API (may have oversimplified!)
|
|
||||||
3. Accidentally introduced MongoDB operators ($gt, $in, etc.)
|
|
||||||
4. May have removed critical features while "simplifying"
|
|
||||||
|
|
||||||
## ❌ CRITICAL MISTAKES WE MADE:
|
|
||||||
|
|
||||||
### 1. **MongoDB Operators (LEGAL RISK!)**
|
|
||||||
```typescript
|
|
||||||
// ❌ WRONG - We accidentally added:
|
|
||||||
where: { field: {$gt: value} }
|
|
||||||
|
|
||||||
// ✅ CORRECT - Should be:
|
|
||||||
where: { field: {greaterThan: value} }
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Lost Neural API Methods**
|
|
||||||
```typescript
|
|
||||||
// ❌ MISSING - These were removed or not properly exposed:
|
|
||||||
brain.neural.similar(a, b)
|
|
||||||
brain.neural.clusters()
|
|
||||||
brain.neural.hierarchy(id)
|
|
||||||
brain.neural.neighbors(id)
|
|
||||||
brain.neural.outliers()
|
|
||||||
brain.neural.semanticPath(from, to)
|
|
||||||
brain.neural.visualize() // Critical for external tools!
|
|
||||||
brain.neural.clusterFast() // O(n) performance
|
|
||||||
brain.neural.clusterLarge() // Million-item support
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Lost Import Capabilities**
|
|
||||||
```typescript
|
|
||||||
// ❌ WRONG - We made it too complex:
|
|
||||||
neuralImport.csv()
|
|
||||||
neuralImport.json()
|
|
||||||
neuralImport.text()
|
|
||||||
|
|
||||||
// ✅ CORRECT - Should be ONE simple method:
|
|
||||||
brain.neuralImport(data, options?) // Auto-detects format!
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. **Lost Clustering for Visualization**
|
|
||||||
The visualization data format for external tools (D3, Cytoscape, GraphML) is missing!
|
|
||||||
```typescript
|
|
||||||
// ❌ MISSING - Critical for external visualization:
|
|
||||||
{
|
|
||||||
format: 'd3' | 'cytoscape' | 'graphml',
|
|
||||||
nodes: [...],
|
|
||||||
edges: [...],
|
|
||||||
layout: {...}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. **Oversimplified Search**
|
|
||||||
```typescript
|
|
||||||
// ❌ REMOVED too many methods:
|
|
||||||
searchByNounTypes()
|
|
||||||
searchWithinItems()
|
|
||||||
searchByStandardField()
|
|
||||||
searchVerbs()
|
|
||||||
searchNounsByVerbs()
|
|
||||||
|
|
||||||
// ✅ BUT this is actually OK if find() handles everything!
|
|
||||||
// Just need to ensure find() is complete
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 COMPARISON: Backup vs Current
|
|
||||||
|
|
||||||
### Methods in BACKUP but NOT in current:
|
|
||||||
```typescript
|
|
||||||
// From backup's brainyData.ts:
|
|
||||||
brain.neural // ❌ Not properly exposed
|
|
||||||
brain.visualize() // ❌ Missing
|
|
||||||
brain.clusters() // ❌ Missing
|
|
||||||
brain.similar() // ❌ Missing
|
|
||||||
brain.neuralImport() // ❌ Wrong implementation
|
|
||||||
|
|
||||||
// Operators in backup:
|
|
||||||
greaterThan, lessThan, equals // ❌ Replaced with $gt, $lt, $eq
|
|
||||||
oneOf, contains, matches // ❌ Replaced with $in, $contains, $regex
|
|
||||||
```
|
|
||||||
|
|
||||||
### Methods we ADDED (some good, some questionable):
|
|
||||||
```typescript
|
|
||||||
// New specific methods (GOOD ✅):
|
|
||||||
addNoun(), getNoun(), deleteNoun()
|
|
||||||
|
|
||||||
// Unified method (GOOD if complete ✅):
|
|
||||||
getNouns(idsOrOptions)
|
|
||||||
|
|
||||||
// But lost flexibility (BAD ❌):
|
|
||||||
- Can't do complex queries easily
|
|
||||||
- Lost specific search methods
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 Feature Comparison Table
|
|
||||||
|
|
||||||
| Feature | Backup | Current | Status |
|
|
||||||
|---------|---------|---------|---------|
|
|
||||||
| **Operators** | greaterThan, lessThan | $gt, $lt | ❌ WRONG |
|
|
||||||
| **Neural API** | Complete (10+ methods) | Missing/Hidden | ❌ BROKEN |
|
|
||||||
| **Clustering** | Full support | Missing | ❌ LOST |
|
|
||||||
| **Visualization** | D3/Cytoscape export | None | ❌ LOST |
|
|
||||||
| **Import** | Simple neuralImport() | Complex multi-method | ❌ WRONG |
|
|
||||||
| **Search** | Multiple specific | Simplified to 2 | ⚠️ OK if complete |
|
|
||||||
| **Verb Scoring** | Full intelligence | Partial | ⚠️ CHECK |
|
|
||||||
| **Synapses** | External connectors | Unknown | ⚠️ CHECK |
|
|
||||||
| **Conduits** | Brainy-to-Brainy | Partial | ⚠️ CHECK |
|
|
||||||
|
|
||||||
## 🔧 WHAT WE NEED TO FIX IMMEDIATELY:
|
|
||||||
|
|
||||||
### Priority 1 (CRITICAL):
|
|
||||||
1. **Replace ALL MongoDB operators with Brainy operators**
|
|
||||||
- This is a legal requirement!
|
|
||||||
- greaterThan not $gt
|
|
||||||
|
|
||||||
2. **Restore Neural API completely**
|
|
||||||
- brain.neural must have all methods
|
|
||||||
- Visualization MUST work for external tools
|
|
||||||
|
|
||||||
3. **Fix neuralImport to be simple**
|
|
||||||
- ONE method that auto-detects
|
|
||||||
- Not multiple complex methods
|
|
||||||
|
|
||||||
### Priority 2 (IMPORTANT):
|
|
||||||
4. **Restore clustering APIs**
|
|
||||||
- For visualization tools
|
|
||||||
- For analysis
|
|
||||||
|
|
||||||
5. **Verify Triple Intelligence is complete**
|
|
||||||
- find() must handle everything
|
|
||||||
- All operators must work
|
|
||||||
|
|
||||||
6. **Check augmentation system**
|
|
||||||
- Synapses (external)
|
|
||||||
- Conduits (internal)
|
|
||||||
|
|
||||||
## 🎯 RECOVERY PLAN:
|
|
||||||
|
|
||||||
### Step 1: Fix Operators (LEGAL REQUIREMENT)
|
|
||||||
- [ ] Find all $gt, $lt, $in, $regex references
|
|
||||||
- [ ] Replace with greaterThan, lessThan, oneOf, matches
|
|
||||||
- [ ] Update all documentation
|
|
||||||
|
|
||||||
### Step 2: Restore Neural API
|
|
||||||
- [ ] Ensure brain.neural is properly exposed
|
|
||||||
- [ ] All methods available: similar, clusters, hierarchy, etc.
|
|
||||||
- [ ] Visualization must return proper format
|
|
||||||
|
|
||||||
### Step 3: Fix Import
|
|
||||||
- [ ] Single neuralImport() method
|
|
||||||
- [ ] Auto-detection of format
|
|
||||||
- [ ] Simple options
|
|
||||||
|
|
||||||
### Step 4: Verify Nothing Lost
|
|
||||||
- [ ] Compare method-by-method with backup
|
|
||||||
- [ ] Test all features
|
|
||||||
- [ ] Update documentation
|
|
||||||
|
|
||||||
## 💡 LESSONS LEARNED:
|
|
||||||
|
|
||||||
1. **Don't oversimplify** - We lost important features
|
|
||||||
2. **Check legal requirements** - MongoDB operators were avoided for a reason
|
|
||||||
3. **Preserve all features** - Even if reorganizing
|
|
||||||
4. **Test against backup** - Always compare functionality
|
|
||||||
5. **Document changes** - Track what and why
|
|
||||||
|
|
||||||
## 🚀 NEXT ACTIONS:
|
|
||||||
|
|
||||||
1. STOP all other work
|
|
||||||
2. Fix operators IMMEDIATELY (legal risk)
|
|
||||||
3. Restore neural API completely
|
|
||||||
4. Test everything works
|
|
||||||
5. Document the final API properly
|
|
||||||
|
|
@ -1,210 +0,0 @@
|
||||||
# 🧠 Brainy 2.0 Final Public API
|
|
||||||
|
|
||||||
> **Clean, Specific, Beautiful** - Every method has ONE clear purpose.
|
|
||||||
|
|
||||||
## 📚 NOUNS (Vectors with Metadata)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addNoun(vector, metadata?) // Add one noun
|
|
||||||
getNoun(id) // Get one noun by ID
|
|
||||||
updateNoun(id, vector?, metadata?) // Update entire noun
|
|
||||||
updateNounMetadata(id, metadata) // Update metadata only
|
|
||||||
getNounMetadata(id) // Get metadata only
|
|
||||||
getNounWithVerbs(id) // Get noun with all relationships
|
|
||||||
deleteNoun(id) // Delete one noun
|
|
||||||
hasNoun(id) // Check if noun exists
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addNouns(items[]) // Add multiple nouns
|
|
||||||
getNouns(idsOrOptions) // Get multiple nouns (by IDs or query)
|
|
||||||
// getNouns(['id1', 'id2']) // Get by specific IDs
|
|
||||||
// getNouns({ filter: {...} }) // Get with filters
|
|
||||||
// getNouns({ limit: 10, offset: 20 }) // Get with pagination
|
|
||||||
deleteNouns(ids[]) // Delete multiple nouns
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔗 VERBS (Relationships)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Single Operations
|
|
||||||
addVerb(source, target, type, metadata?) // Create relationship
|
|
||||||
getVerb(id) // Get one verb by ID
|
|
||||||
deleteVerb(id) // Delete one verb
|
|
||||||
|
|
||||||
// Batch Operations
|
|
||||||
addVerbs(verbs[]) // Add multiple relationships
|
|
||||||
getVerbs(filter?) // Get filtered verbs
|
|
||||||
getVerbsBySource(sourceId) // Get outgoing relationships
|
|
||||||
getVerbsByTarget(targetId) // Get incoming relationships
|
|
||||||
getVerbsByType(type) // Get by relationship type
|
|
||||||
deleteVerbs(ids[]) // Delete multiple verbs
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 SEARCH
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Core Search
|
|
||||||
search(query, k?, options?) // Primary vector search
|
|
||||||
searchText(text, k?, options?) // Natural language search
|
|
||||||
find(query) // Triple Intelligence (Vector+Graph+Metadata) 🧠
|
|
||||||
findSimilar(id, k?, options?) // Find similar to existing noun
|
|
||||||
|
|
||||||
// Advanced Search
|
|
||||||
searchByNounTypes(types[], query, k?) // Filter by noun types
|
|
||||||
searchWithinItems(query, ids[], k?) // Search within specific nouns
|
|
||||||
searchWithCursor(query, cursor) // Paginated search
|
|
||||||
searchByStandardField(field, value, k?) // Metadata-based search
|
|
||||||
|
|
||||||
// Graph Search
|
|
||||||
searchVerbs(query, options?) // Search relationships
|
|
||||||
searchNounsByVerbs(conditions) // Find nouns by relationships
|
|
||||||
|
|
||||||
// Distributed Search
|
|
||||||
searchLocal(query, k?) // Search local instance only
|
|
||||||
searchRemote(query, k?) // Search remote instance only
|
|
||||||
searchCombined(query, k?) // Search both local and remote
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 METADATA & FILTERING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
getFilterFields() // Get all indexed fields
|
|
||||||
getFilterValues(field) // Get unique values for a field
|
|
||||||
getAvailableFieldNames() // Get available field names
|
|
||||||
getStandardFieldMappings() // Get standard field mappings
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 PERFORMANCE & MONITORING
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Cache
|
|
||||||
getCacheStats() // Get cache statistics
|
|
||||||
clearCache() // Clear search cache
|
|
||||||
|
|
||||||
// Statistics
|
|
||||||
size() // Total noun count
|
|
||||||
getStatistics(options?) // Comprehensive statistics
|
|
||||||
getServiceStatistics(service) // Per-service statistics
|
|
||||||
listServices() // List all services
|
|
||||||
flushStatistics() // Persist statistics to storage
|
|
||||||
|
|
||||||
// Health
|
|
||||||
getHealthStatus() // System health check
|
|
||||||
status() // Full status report
|
|
||||||
```
|
|
||||||
|
|
||||||
## ⚙️ CONFIGURATION
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Operational Modes
|
|
||||||
isReadOnly() / setReadOnly(bool) // Read-only mode
|
|
||||||
isWriteOnly() / setWriteOnly(bool) // Write-only mode
|
|
||||||
isFrozen() / setFrozen(bool) // Freeze all modifications
|
|
||||||
|
|
||||||
// Real-time Sync
|
|
||||||
enableRealtimeUpdates(config) // Enable live synchronization
|
|
||||||
disableRealtimeUpdates() // Disable synchronization
|
|
||||||
getRealtimeUpdateConfig() // Get current config
|
|
||||||
checkForUpdatesNow() // Manual sync trigger
|
|
||||||
|
|
||||||
// Remote Connection
|
|
||||||
connectToRemoteServer(url, options?) // Connect to remote instance
|
|
||||||
disconnectFromRemoteServer() // Disconnect from remote
|
|
||||||
isConnectedToRemoteServer() // Check connection status
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🧠 INTELLIGENCE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Verb Scoring
|
|
||||||
provideFeedbackForVerbScoring(feedback) // Train relationship scoring
|
|
||||||
getVerbScoringStats() // Get scoring statistics
|
|
||||||
exportVerbScoringLearningData() // Export training data
|
|
||||||
importVerbScoringLearningData(data) // Import training data
|
|
||||||
|
|
||||||
// Embeddings
|
|
||||||
embed(text) // Generate embedding vector
|
|
||||||
calculateSimilarity(a, b, metric?) // Calculate vector similarity
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 DATA MANAGEMENT
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Clear Operations
|
|
||||||
clear(options?) // Clear all data
|
|
||||||
clearNouns(options?) // Clear all nouns only
|
|
||||||
clearVerbs(options?) // Clear all verbs only
|
|
||||||
|
|
||||||
// Backup & Restore
|
|
||||||
backup() // Create full backup
|
|
||||||
restore(backup) // Restore from backup
|
|
||||||
|
|
||||||
// Import/Export
|
|
||||||
import(data, format) // Import external data
|
|
||||||
importSparseData(data) // Import sparse format
|
|
||||||
|
|
||||||
// Index Management
|
|
||||||
rebuildMetadataIndex() // Rebuild metadata index
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔒 SECURITY
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
encryptData(data) // Encrypt data
|
|
||||||
decryptData(data) // Decrypt data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎲 UTILITIES
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
generateRandomGraph(nodes, edges) // Generate test graph data
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 LIFECYCLE
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Instance Methods
|
|
||||||
new BrainyData(config?) // Create instance
|
|
||||||
init() // Initialize (REQUIRED!)
|
|
||||||
shutDown() // Graceful shutdown
|
|
||||||
cleanup() // Clean up resources
|
|
||||||
|
|
||||||
// Static Methods
|
|
||||||
BrainyData.preloadModel(options?) // Preload ML model
|
|
||||||
BrainyData.warmup(options?) // Warmup system
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📐 PROPERTIES (Read-only)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
dimensions // Vector dimensions
|
|
||||||
maxConnections // HNSW max connections
|
|
||||||
efConstruction // HNSW ef construction
|
|
||||||
initialized // Is initialized?
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📝 Key Changes in 2.0
|
|
||||||
|
|
||||||
### ✅ Simplified & Unified
|
|
||||||
- `getNouns()` now handles ALL plural queries (by IDs, filters, or pagination)
|
|
||||||
- No more `getNounsByIds()`, `queryNouns()`, `getBatch()` - just `getNouns()`
|
|
||||||
- Clear singular vs plural: `getNoun()` for one, `getNouns()` for many
|
|
||||||
|
|
||||||
### ✅ Specific Naming
|
|
||||||
- Always specify noun/verb: `addNoun()` not `add()`
|
|
||||||
- No aliases or duplicates
|
|
||||||
- One method, one purpose
|
|
||||||
|
|
||||||
### ✅ Private Legacy Methods
|
|
||||||
These are now private (use new methods above):
|
|
||||||
- `add()`, `get()`, `delete()`, `update()`
|
|
||||||
- `relate()`, `connect()`, `has()`, `exists()`
|
|
||||||
- `getMetadata()`, `updateMetadata()`
|
|
||||||
- `addItem()`, `addToBoth()`, `addBatch()`, `getBatch()`
|
|
||||||
|
|
||||||
### ✅ Triple Intelligence
|
|
||||||
- New `find()` method unifies Vector + Graph + Metadata search
|
|
||||||
- Most powerful search capability in one simple method
|
|
||||||
|
|
@ -1,28 +0,0 @@
|
||||||
# API Design Archive
|
|
||||||
|
|
||||||
This directory contains historical API design documents from the Brainy 2.0 development process.
|
|
||||||
|
|
||||||
## Purpose
|
|
||||||
These documents represent the iterative refinement of the Brainy 2.0 API during development. They are preserved here for historical reference and to document the design decisions made along the way.
|
|
||||||
|
|
||||||
## Current API Documentation
|
|
||||||
The definitive API documentation is now located at:
|
|
||||||
- **`/docs/api/README.md`** - The ONE official API reference
|
|
||||||
|
|
||||||
## Archived Documents
|
|
||||||
These documents show the evolution of the API design:
|
|
||||||
- Various iterations of API structure
|
|
||||||
- Exploration of different naming conventions
|
|
||||||
- Refinement of the Triple Intelligence concept
|
|
||||||
- Transition from MongoDB operators to Brainy operators
|
|
||||||
- Evolution of the Neural API
|
|
||||||
|
|
||||||
## Key Decisions Made
|
|
||||||
1. **Terminology**: Vector + Graph + Metadata (not Field)
|
|
||||||
2. **Operators**: Brainy operators (greaterThan, lessThan) not MongoDB ($gt, $lt)
|
|
||||||
3. **Methods**: Specific noun/verb naming (addNoun, getNoun)
|
|
||||||
4. **Unification**: Single getNouns() method instead of multiple variants
|
|
||||||
5. **Neural API**: Complete clustering and visualization features
|
|
||||||
|
|
||||||
## Note
|
|
||||||
These documents are archived, not deleted, to preserve the development history and rationale behind API decisions.
|
|
||||||
2387
docs/api/README.md
2387
docs/api/README.md
File diff suppressed because it is too large
Load diff
114
docs/architecture/PERFORMANCE_ANALYSIS.md
Normal file
114
docs/architecture/PERFORMANCE_ANALYSIS.md
Normal file
|
|
@ -0,0 +1,114 @@
|
||||||
|
# Brainy Performance Analysis & Optimization
|
||||||
|
|
||||||
|
## Current Issues Found
|
||||||
|
|
||||||
|
### 1. ❌ CRITICAL: notEquals Operator is O(n)
|
||||||
|
```javascript
|
||||||
|
// PROBLEM: Gets ALL items to filter
|
||||||
|
case 'notEquals':
|
||||||
|
const allItemIds = await this.getAllIds() // O(n) - TERRIBLE!
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. ❌ Soft Delete Performance
|
||||||
|
- Every query adds `deleted: { notEquals: true }`
|
||||||
|
- This makes EVERY query O(n) instead of O(log n)
|
||||||
|
|
||||||
|
### 3. ❌ exists Operator is Inefficient
|
||||||
|
```javascript
|
||||||
|
case 'exists':
|
||||||
|
// Scans all cache entries - O(n)
|
||||||
|
for (const [key, entry] of this.indexCache.entries()) {
|
||||||
|
if (entry.field === field) {
|
||||||
|
entry.ids.forEach(id => allIds.add(id))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. ⚠️ Query Optimizer Not Smart Enough
|
||||||
|
- `isSelectiveFilter()` needs to understand which filters are fast
|
||||||
|
- Should prioritize O(1) and O(log n) operations
|
||||||
|
|
||||||
|
## Performance Characteristics
|
||||||
|
|
||||||
|
### ✅ Fast Operations (Keep These)
|
||||||
|
| Operation | Complexity | Example |
|
||||||
|
|-----------|-----------|---------|
|
||||||
|
| Vector Search (HNSW) | O(log n) | `like: "query"` |
|
||||||
|
| Exact Match | O(1) | `where: { status: "active" }` |
|
||||||
|
| Deleted Filter (NEW) | O(1) | `where: { deleted: false }` |
|
||||||
|
| Range Query (sorted) | O(log n) | `where: { year: { gt: 2000 } }` |
|
||||||
|
| Graph Traversal | O(k) | `connected: { from: id }` |
|
||||||
|
|
||||||
|
### ❌ Slow Operations (Need Fixing)
|
||||||
|
| Operation | Current | Should Be | Fix |
|
||||||
|
|-----------|---------|-----------|-----|
|
||||||
|
| notEquals | O(n) | O(1) or O(log n) | Use complement index |
|
||||||
|
| exists | O(n) | O(1) | Maintain field existence bitmap |
|
||||||
|
| noneOf | O(n) | O(k) | Use set operations |
|
||||||
|
|
||||||
|
## Optimized Architecture
|
||||||
|
|
||||||
|
### Solution 1: Positive Indexing for Soft Delete ✅
|
||||||
|
```javascript
|
||||||
|
// Instead of: deleted !== true (O(n))
|
||||||
|
// Use: deleted === false (O(1))
|
||||||
|
where: { deleted: false }
|
||||||
|
|
||||||
|
// Ensure all items have deleted field
|
||||||
|
if (!metadata.deleted) metadata.deleted = false
|
||||||
|
```
|
||||||
|
|
||||||
|
### Solution 2: Complement Indices for notEquals
|
||||||
|
```javascript
|
||||||
|
class MetadataIndexManager {
|
||||||
|
// For common notEquals queries, maintain complement sets
|
||||||
|
private complementIndices: Map<string, Set<string>> = new Map()
|
||||||
|
|
||||||
|
// Example: Track non-deleted items separately
|
||||||
|
private activeItems: Set<string> = new Set()
|
||||||
|
private deletedItems: Set<string> = new Set()
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Solution 3: Field Existence Bitmap
|
||||||
|
```javascript
|
||||||
|
class FieldExistenceIndex {
|
||||||
|
private fieldBitmaps: Map<string, BitSet> = new Map()
|
||||||
|
|
||||||
|
hasField(id: string, field: string): boolean {
|
||||||
|
return this.fieldBitmaps.get(field)?.has(id) ?? false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Query Execution Strategy
|
||||||
|
|
||||||
|
### Progressive Search (When Metadata is Selective)
|
||||||
|
```
|
||||||
|
1. Field Filter (O(1) or O(log n)) → Small candidate set
|
||||||
|
2. Vector Search within candidates (O(k log k))
|
||||||
|
3. Fusion if needed
|
||||||
|
```
|
||||||
|
|
||||||
|
### Parallel Search (When Nothing is Selective)
|
||||||
|
```
|
||||||
|
1. Vector Search (O(log n)) → Top K results
|
||||||
|
2. Graph Traversal (O(m)) → Connected items
|
||||||
|
3. Field Filter (O(1)) → Metadata matches
|
||||||
|
4. Fusion: Intersection or Union
|
||||||
|
```
|
||||||
|
|
||||||
|
## Implementation Priority
|
||||||
|
|
||||||
|
1. **DONE** ✅ Fix soft delete to use `deleted: false`
|
||||||
|
2. **TODO** 🔧 Optimize notEquals for common fields
|
||||||
|
3. **TODO** 🔧 Add field existence index
|
||||||
|
4. **TODO** 🔧 Improve query optimizer intelligence
|
||||||
|
5. **TODO** 🔧 Add query explain mode for debugging
|
||||||
|
|
||||||
|
## Performance Targets
|
||||||
|
|
||||||
|
- Vector search: < 10ms for 1M items
|
||||||
|
- Metadata filter: < 1ms for exact match
|
||||||
|
- Combined query: < 20ms for complex queries
|
||||||
|
- Soft delete overhead: < 0.1ms (O(1))
|
||||||
242
docs/architecture/aggregation.md
Normal file
242
docs/architecture/aggregation.md
Normal file
|
|
@ -0,0 +1,242 @@
|
||||||
|
# Aggregation Architecture
|
||||||
|
|
||||||
|
> Write-time incremental aggregation with O(1) reads
|
||||||
|
|
||||||
|
## Design Principles
|
||||||
|
|
||||||
|
1. **Write-time computation** — aggregates update on every `add()`, `update()`, and `delete()`, not as batch jobs
|
||||||
|
2. **Incremental state** — running totals maintained per group, never rescanning the dataset
|
||||||
|
3. **Provider interface** — TypeScript engine is the default; plugins can replace it with native implementations
|
||||||
|
4. **Zero-allocation reads** — query results are computed from pre-aggregated state
|
||||||
|
|
||||||
|
## Component Overview
|
||||||
|
|
||||||
|
```
|
||||||
|
┌──────────────────────────────────────────────────────────┐
|
||||||
|
│ Brainy │
|
||||||
|
│ │
|
||||||
|
│ add() / update() / delete() │
|
||||||
|
│ │ │
|
||||||
|
│ ▼ │
|
||||||
|
│ ┌──────────────────────┐ ┌────────────────────────┐ │
|
||||||
|
│ │ AggregationIndex │ │ AggregateMaterializer │ │
|
||||||
|
│ │ │───▶│ (debounced writes) │ │
|
||||||
|
│ │ ├─ definitions Map │ └────────────────────────┘ │
|
||||||
|
│ │ ├─ states Map │ │
|
||||||
|
│ │ └─ staleMinMax Set │ ┌────────────────────────┐ │
|
||||||
|
│ │ │ │ timeWindows.ts │ │
|
||||||
|
│ │ Source filter ──────│───▶│ bucketTimestamp() │ │
|
||||||
|
│ │ Group key ──────────│───▶│ parseBucketRange() │ │
|
||||||
|
│ └──────────┬───────────┘ └────────────────────────┘ │
|
||||||
|
│ │ │
|
||||||
|
│ │ provider interface │
|
||||||
|
│ ▼ │
|
||||||
|
│ ┌──────────────────────┐ │
|
||||||
|
│ │ AggregationProvider │ (optional, registered by │
|
||||||
|
│ │ ├─ incrementalUpdate│ plugin like @soulcraft/cor) │
|
||||||
|
│ │ ├─ rebuildAggregate │ │
|
||||||
|
│ │ ├─ queryAggregate │ │
|
||||||
|
│ │ └─ serialize/restore│ │
|
||||||
|
│ └──────────────────────┘ │
|
||||||
|
└──────────────────────────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
## State Management
|
||||||
|
|
||||||
|
### Definitions
|
||||||
|
|
||||||
|
Registered via `brain.defineAggregate(def)`. Stored in a `Map<string, AggregateDefinition>` keyed by aggregate name. Persisted to storage under `__aggregation_definitions__` on flush.
|
||||||
|
|
||||||
|
### Group State
|
||||||
|
|
||||||
|
Each aggregate maintains a `Map<string, AggregateGroupState>` where keys are serialized group key values (e.g., `category=food|date=2024-01`). Each group holds per-metric `MetricState`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
interface MetricState {
|
||||||
|
sum: number // Running total
|
||||||
|
count: number // Entity count
|
||||||
|
min: number // Minimum (Infinity if empty)
|
||||||
|
max: number // Maximum (-Infinity if empty)
|
||||||
|
m2?: number // Welford's M2 for stddev/variance
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Change Detection
|
||||||
|
|
||||||
|
On restart, definition hashes (FNV-1a 32-bit) are compared with the persisted hash. If a definition changed (different groupBy, metrics, or source), the aggregate state is reset and must be rebuilt.
|
||||||
|
|
||||||
|
## Write-Time Update Flow
|
||||||
|
|
||||||
|
When `brain.add(entity)` is called:
|
||||||
|
|
||||||
|
```
|
||||||
|
1. For each registered aggregate:
|
||||||
|
├─ Source filter check (type, service, where)
|
||||||
|
│ └─ Skip if entity doesn't match
|
||||||
|
├─ Aggregate entity check
|
||||||
|
│ └─ Skip if entity.service === 'brainy:aggregation'
|
||||||
|
│ or entity.metadata.__aggregate is set
|
||||||
|
├─ Group key computation
|
||||||
|
│ └─ Extract groupBy fields from metadata
|
||||||
|
│ Apply time bucketing for windowed dimensions
|
||||||
|
└─ Metric update
|
||||||
|
└─ For each metric in the definition:
|
||||||
|
├─ count: increment count
|
||||||
|
├─ sum/avg: add value to sum, increment count
|
||||||
|
├─ min/max: compare and update
|
||||||
|
└─ stddev/variance: Welford's online update
|
||||||
|
```
|
||||||
|
|
||||||
|
### Update Handling
|
||||||
|
|
||||||
|
On `brain.update(entity)`, the engine reverses the old entity's contribution and applies the new entity's contribution. This correctly handles:
|
||||||
|
|
||||||
|
- **Value changes**: old amount=10, new amount=20 — sum adjusts by +10
|
||||||
|
- **Group key changes**: entity moves from category "food" to "drink" — both groups update
|
||||||
|
- **Source filter changes**: entity type changes from Event to Document — removed from matching aggregates
|
||||||
|
|
||||||
|
### Delete Handling
|
||||||
|
|
||||||
|
On `brain.remove(id)`, the engine reverses the entity's contribution:
|
||||||
|
|
||||||
|
- `count` and `sum` are decremented
|
||||||
|
- `min`/`max` may become stale (marked in `staleMinMax` for lazy recompute)
|
||||||
|
- Welford's M2 is updated with the inverse formula
|
||||||
|
- Empty groups (all metric counts at zero) are removed
|
||||||
|
|
||||||
|
## Algorithms
|
||||||
|
|
||||||
|
### Welford's Online Algorithm
|
||||||
|
|
||||||
|
Standard deviation and variance use Welford's numerically stable online algorithm with M2 tracking. This computes incrementally without storing individual values:
|
||||||
|
|
||||||
|
```
|
||||||
|
On add(x):
|
||||||
|
count += 1
|
||||||
|
oldMean = (sum - x) / (count - 1) // mean before this value
|
||||||
|
sum += x
|
||||||
|
mean = sum / count // mean after this value
|
||||||
|
M2 += (x - oldMean) * (x - mean)
|
||||||
|
|
||||||
|
On remove(x):
|
||||||
|
oldMean = sum / count
|
||||||
|
sum -= x
|
||||||
|
count -= 1
|
||||||
|
newMean = sum / count
|
||||||
|
M2 = max(0, M2 - (x - oldMean) * (x - newMean))
|
||||||
|
|
||||||
|
Sample variance = M2 / (count - 1)
|
||||||
|
Sample stddev = sqrt(variance)
|
||||||
|
```
|
||||||
|
|
||||||
|
M2 is clamped to zero on remove to prevent floating-point drift from producing negative values.
|
||||||
|
|
||||||
|
### MIN/MAX Handling
|
||||||
|
|
||||||
|
The TypeScript engine uses simple comparison for add operations and marks MIN/MAX as potentially stale on delete (since removing the current min/max value requires a rescan). Stale values are lazily recomputed on the next query.
|
||||||
|
|
||||||
|
The Cor native engine uses a `BTreeMap<OrderedFloat<f64>, u64>` that tracks the exact frequency of every value, providing precise MIN/MAX after any sequence of operations without rescanning.
|
||||||
|
|
||||||
|
### Time Window Bucketing
|
||||||
|
|
||||||
|
Timestamps (Unix milliseconds) are bucketed using UTC-based formatting:
|
||||||
|
|
||||||
|
| Granularity | Bucket Key | Algorithm |
|
||||||
|
|------------|-----------|-----------|
|
||||||
|
| `hour` | `2024-01-15T14` | UTC year-month-day-hour |
|
||||||
|
| `day` | `2024-01-15` | UTC year-month-day |
|
||||||
|
| `week` | `2024-W03` | ISO 8601 week (Monday start, week 1 contains first Thursday) |
|
||||||
|
| `month` | `2024-01` | UTC year-month |
|
||||||
|
| `quarter` | `2024-Q1` | `ceil((month) / 3)` |
|
||||||
|
| `year` | `2024` | UTC year |
|
||||||
|
| `{ seconds: N }` | ISO timestamp | `floor(timestamp / interval) * interval` |
|
||||||
|
|
||||||
|
Bucket keys can be parsed back into `{ start, end }` timestamp ranges via `parseBucketRange()`.
|
||||||
|
|
||||||
|
## Provider Interface
|
||||||
|
|
||||||
|
The `AggregationProvider` interface defines the contract between Brainy's `AggregationIndex` and plugin-provided native implementations:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
interface AggregationProvider {
|
||||||
|
defineAggregate?(def: AggregateDefinition): void
|
||||||
|
removeAggregate?(name: string): void
|
||||||
|
|
||||||
|
incrementalUpdate(
|
||||||
|
name: string,
|
||||||
|
def: AggregateDefinition,
|
||||||
|
entity: Record<string, unknown>,
|
||||||
|
op: 'add' | 'update' | 'delete',
|
||||||
|
prev?: Record<string, unknown>
|
||||||
|
): AggregateGroupState[]
|
||||||
|
|
||||||
|
computeGroupKey(
|
||||||
|
entity: Record<string, unknown>,
|
||||||
|
groupBy: GroupByDimension[]
|
||||||
|
): Record<string, string | number>
|
||||||
|
|
||||||
|
rebuildAggregate(
|
||||||
|
def: AggregateDefinition,
|
||||||
|
entities: Array<Record<string, unknown>>
|
||||||
|
): Map<string, AggregateGroupState>
|
||||||
|
|
||||||
|
queryAggregate(
|
||||||
|
state: Map<string, AggregateGroupState>,
|
||||||
|
params: AggregateQueryParams
|
||||||
|
): AggregateResult[]
|
||||||
|
|
||||||
|
restoreState?(data: string): void
|
||||||
|
serializeState?(): string
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
When a native provider is registered:
|
||||||
|
|
||||||
|
1. `AggregationIndex` delegates `incrementalUpdate()` to the provider instead of running TypeScript logic
|
||||||
|
2. Provider returns updated `AggregateGroupState[]` which are applied back into the state maps
|
||||||
|
3. Query execution is delegated via `queryAggregate()`
|
||||||
|
4. State serialization is delegated via `serializeState()`/`restoreState()`
|
||||||
|
|
||||||
|
Brainy retains ownership of the state maps and persistence. The provider handles computation.
|
||||||
|
|
||||||
|
## Materialization
|
||||||
|
|
||||||
|
The `AggregateMaterializer` converts aggregate group states into `NounType.Measurement` entities:
|
||||||
|
|
||||||
|
1. When an aggregate group is updated and `materialize` is enabled, `scheduleMaterialize()` is called
|
||||||
|
2. Materialization is debounced (default: 1000ms) to batch rapid updates during ingestion
|
||||||
|
3. On trigger, the materializer either creates or updates a `NounType.Measurement` entity
|
||||||
|
4. Materialized entities include `service: 'brainy:aggregation'` and `metadata.__aggregate` to prevent infinite loops
|
||||||
|
|
||||||
|
Materialized entities are automatically visible through:
|
||||||
|
- OData endpoints
|
||||||
|
- Google Sheets integration
|
||||||
|
- Server-Sent Events (SSE)
|
||||||
|
- Webhook notifications
|
||||||
|
|
||||||
|
## Persistence
|
||||||
|
|
||||||
|
### Storage Keys
|
||||||
|
|
||||||
|
| Key | Content |
|
||||||
|
|-----|---------|
|
||||||
|
| `__aggregation_definitions__` | Array of all definitions with FNV-1a hashes |
|
||||||
|
| `__aggregation_state_{name}__` | Per-aggregate group states (array of `AggregateGroupState`) |
|
||||||
|
| `__aggregation_native_state__` | Serialized native provider state (JSON string) |
|
||||||
|
|
||||||
|
### Lifecycle
|
||||||
|
|
||||||
|
1. **`init()`** — Load definitions, compare hashes, load matching state, restore native provider state
|
||||||
|
2. **Write operations** — Mark modified aggregates as dirty
|
||||||
|
3. **`flush()`** — Persist all dirty aggregate states and native provider state
|
||||||
|
4. **`close()`** — Flush and release resources
|
||||||
|
|
||||||
|
## Source Files
|
||||||
|
|
||||||
|
| File | Purpose |
|
||||||
|
|------|---------|
|
||||||
|
| `src/aggregation/AggregationIndex.ts` | Core engine: definitions, state, write hooks, query |
|
||||||
|
| `src/aggregation/materializer.ts` | Debounced materialization of results as entities |
|
||||||
|
| `src/aggregation/timeWindows.ts` | Time bucketing and bucket range parsing |
|
||||||
|
| `src/aggregation/index.ts` | Module exports |
|
||||||
|
| `src/types/brainy.types.ts` | Type definitions for all aggregation interfaces |
|
||||||
|
|
@ -1,359 +0,0 @@
|
||||||
# 🔍 Brainy 2.0 Augmentation System Architecture Audit (REVISED)
|
|
||||||
|
|
||||||
**Author**: Senior Architecture Review
|
|
||||||
**Date**: 2025-08-25
|
|
||||||
**Status**: 🟡 **WORKING BUT NEEDS MARKETPLACE FEATURES**
|
|
||||||
|
|
||||||
## Executive Summary
|
|
||||||
|
|
||||||
The augmentation system core execution is WORKING correctly through `AugmentationRegistry`. The system properly executes augmentations before/after operations. However, there's no discovery, installation, or marketplace integration for the brain-cloud registry vision.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🟢 What's Actually Working
|
|
||||||
|
|
||||||
### 1. Execution Mechanism ✅
|
|
||||||
The `AugmentationRegistry` class properly implements:
|
|
||||||
```typescript
|
|
||||||
async execute<T>(operation: string, params: any, mainOperation: () => Promise<T>): Promise<T>
|
|
||||||
```
|
|
||||||
- Chains augmentations correctly
|
|
||||||
- Respects timing (before/after/around)
|
|
||||||
- Handles operation filtering
|
|
||||||
- Works with all 27 augmentations
|
|
||||||
|
|
||||||
### 2. Registration System ✅
|
|
||||||
```typescript
|
|
||||||
brain.augmentations.register(augmentation)
|
|
||||||
```
|
|
||||||
- Two-phase initialization works (storage first)
|
|
||||||
- Context injection works
|
|
||||||
- Lifecycle management works
|
|
||||||
|
|
||||||
### 3. Clean Interface ✅
|
|
||||||
- 100% of augmentations use `BrainyAugmentation`
|
|
||||||
- `BaseAugmentation` provides solid foundation
|
|
||||||
- Proper TypeScript types
|
|
||||||
|
|
||||||
### 4. Auto-Configuration ✅
|
|
||||||
```typescript
|
|
||||||
new BrainyData({
|
|
||||||
cache: true, // Auto-registers CacheAugmentation
|
|
||||||
index: true, // Auto-registers IndexAugmentation
|
|
||||||
storage: 's3' // Auto-registers S3StorageAugmentation
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🟡 Missing for Marketplace Vision
|
|
||||||
|
|
||||||
### 1. No Package Discovery
|
|
||||||
**Current**: Manual registration only
|
|
||||||
```typescript
|
|
||||||
// Current (manual)
|
|
||||||
import { NotionSynapse } from './my-custom-synapse'
|
|
||||||
brain.augmentations.register(new NotionSynapse())
|
|
||||||
|
|
||||||
// Needed
|
|
||||||
await brain.discover('notion') // Search npm/brain-cloud
|
|
||||||
await brain.install('@soulcraft/notion-synapse')
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. No Installation Mechanism
|
|
||||||
**Current**: Must be bundled at build time
|
|
||||||
**Needed**: Dynamic installation
|
|
||||||
```typescript
|
|
||||||
interface AugmentationMarketplace {
|
|
||||||
search(query: string): Promise<Package[]>
|
|
||||||
install(packageId: string): Promise<void>
|
|
||||||
uninstall(packageId: string): Promise<void>
|
|
||||||
listInstalled(): Promise<Package[]>
|
|
||||||
checkUpdates(): Promise<Update[]>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. No Brain Cloud Registry Client
|
|
||||||
**Current**: No registry concept
|
|
||||||
**Needed**: Registry integration
|
|
||||||
```typescript
|
|
||||||
class BrainCloudRegistry {
|
|
||||||
private apiUrl = 'https://api.soulcraft.com/brain-cloud'
|
|
||||||
|
|
||||||
async search(query: string): Promise<AugmentationPackage[]> {
|
|
||||||
const response = await fetch(`${this.apiUrl}/augmentations/search?q=${query}`)
|
|
||||||
return response.json()
|
|
||||||
}
|
|
||||||
|
|
||||||
async getPackage(id: string): Promise<AugmentationPackage> {
|
|
||||||
const response = await fetch(`${this.apiUrl}/augmentations/${id}`)
|
|
||||||
return response.json()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. No License Management
|
|
||||||
**Current**: All augmentations free/bundled
|
|
||||||
**Needed**: License verification
|
|
||||||
```typescript
|
|
||||||
interface LicenseManager {
|
|
||||||
verify(packageId: string, licenseKey: string): Promise<boolean>
|
|
||||||
activate(packageId: string, licenseKey: string): Promise<void>
|
|
||||||
deactivate(packageId: string): Promise<void>
|
|
||||||
getStatus(packageId: string): Promise<LicenseStatus>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. No Version Management
|
|
||||||
**Current**: No versioning
|
|
||||||
**Needed**: Semver support
|
|
||||||
```typescript
|
|
||||||
interface VersionManager {
|
|
||||||
checkCompatibility(pkg: Package, brainyVersion: string): boolean
|
|
||||||
resolveConflicts(packages: Package[]): Package[]
|
|
||||||
upgrade(packageId: string, toVersion: string): Promise<void>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📋 Implementation Plan for Marketplace
|
|
||||||
|
|
||||||
### Phase 1: Local Package Discovery (1 week)
|
|
||||||
```typescript
|
|
||||||
class LocalPackageDiscovery {
|
|
||||||
async discover(): Promise<Package[]> {
|
|
||||||
// 1. Search node_modules for brainy augmentations
|
|
||||||
const packages = await glob('node_modules/@*/package.json')
|
|
||||||
|
|
||||||
// 2. Filter for brainy augmentations
|
|
||||||
return packages.filter(pkg => pkg.brainy?.type === 'augmentation')
|
|
||||||
}
|
|
||||||
|
|
||||||
async load(packageId: string): Promise<BrainyAugmentation> {
|
|
||||||
// Dynamic import
|
|
||||||
const module = await import(packageId)
|
|
||||||
return new module.default()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 2: NPM Integration (1 week)
|
|
||||||
```typescript
|
|
||||||
class NPMRegistry {
|
|
||||||
async search(query: string): Promise<Package[]> {
|
|
||||||
// Search npm for packages with brainy keyword
|
|
||||||
const response = await fetch(
|
|
||||||
`https://registry.npmjs.org/-/v1/search?text=${query}+keywords:brainy-augmentation`
|
|
||||||
)
|
|
||||||
return response.json()
|
|
||||||
}
|
|
||||||
|
|
||||||
async install(packageId: string): Promise<void> {
|
|
||||||
// Use npm programmatically
|
|
||||||
await exec(`npm install ${packageId}`)
|
|
||||||
|
|
||||||
// Auto-register after install
|
|
||||||
const aug = await this.load(packageId)
|
|
||||||
this.brain.augmentations.register(aug)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 3: Brain Cloud Registry (2 weeks)
|
|
||||||
```typescript
|
|
||||||
class BrainCloudMarketplace {
|
|
||||||
private registry = new BrainCloudRegistry()
|
|
||||||
private licenses = new LicenseManager()
|
|
||||||
private installer = new AugmentationInstaller()
|
|
||||||
|
|
||||||
async browse(category?: string): Promise<MarketplaceListing[]> {
|
|
||||||
const packages = await this.registry.list(category)
|
|
||||||
|
|
||||||
return packages.map(pkg => ({
|
|
||||||
...pkg,
|
|
||||||
installed: this.isInstalled(pkg.id),
|
|
||||||
licensed: this.isLicensed(pkg.id),
|
|
||||||
updates: this.hasUpdates(pkg.id)
|
|
||||||
}))
|
|
||||||
}
|
|
||||||
|
|
||||||
async purchase(packageId: string): Promise<void> {
|
|
||||||
// 1. Process payment
|
|
||||||
const license = await this.processPayment(packageId)
|
|
||||||
|
|
||||||
// 2. Activate license
|
|
||||||
await this.licenses.activate(packageId, license)
|
|
||||||
|
|
||||||
// 3. Install package
|
|
||||||
await this.install(packageId)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Phase 4: Developer Tools (1 week)
|
|
||||||
```typescript
|
|
||||||
// CLI for augmentation development
|
|
||||||
class AugmentationCLI {
|
|
||||||
async create(name: string): Promise<void> {
|
|
||||||
// Scaffold new augmentation project
|
|
||||||
await this.scaffold(name, 'augmentation-template')
|
|
||||||
}
|
|
||||||
|
|
||||||
async test(path: string): Promise<void> {
|
|
||||||
// Test augmentation locally
|
|
||||||
const aug = await this.load(path)
|
|
||||||
await this.runTests(aug)
|
|
||||||
}
|
|
||||||
|
|
||||||
async publish(path: string): Promise<void> {
|
|
||||||
// Publish to brain-cloud
|
|
||||||
const pkg = await this.package(path)
|
|
||||||
await this.registry.publish(pkg)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🏗️ Recommended Architecture
|
|
||||||
|
|
||||||
### 1. Augmentation Package Structure
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"name": "@soulcraft/notion-synapse",
|
|
||||||
"version": "1.0.0",
|
|
||||||
"main": "dist/index.js",
|
|
||||||
"types": "dist/index.d.ts",
|
|
||||||
"brainy": {
|
|
||||||
"type": "augmentation",
|
|
||||||
"class": "NotionSynapse",
|
|
||||||
"timing": "after",
|
|
||||||
"operations": ["addNoun", "updateNoun"],
|
|
||||||
"priority": 20,
|
|
||||||
"license": "premium",
|
|
||||||
"price": 9.99,
|
|
||||||
"compatibility": ">=2.0.0",
|
|
||||||
"dependencies": []
|
|
||||||
},
|
|
||||||
"keywords": ["brainy-augmentation", "notion", "sync"]
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Installation Flow
|
|
||||||
```typescript
|
|
||||||
// User flow
|
|
||||||
await brain.marketplace.search('notion')
|
|
||||||
// Returns: [@soulcraft/notion-synapse, @community/notion-sync, ...]
|
|
||||||
|
|
||||||
await brain.marketplace.install('@soulcraft/notion-synapse')
|
|
||||||
// 1. Check license (prompt for purchase if needed)
|
|
||||||
// 2. Check compatibility
|
|
||||||
// 3. Install dependencies
|
|
||||||
// 4. Download package
|
|
||||||
// 5. Load augmentation
|
|
||||||
// 6. Register with brain
|
|
||||||
// 7. Initialize
|
|
||||||
|
|
||||||
// Now it's working!
|
|
||||||
brain.augmentations.list()
|
|
||||||
// [..., { name: '@soulcraft/notion-synapse', enabled: true }]
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Discovery UI
|
|
||||||
```typescript
|
|
||||||
// Web UI component
|
|
||||||
<AugmentationMarketplace>
|
|
||||||
<SearchBar />
|
|
||||||
<Categories>
|
|
||||||
<Category name="Storage" count={12} />
|
|
||||||
<Category name="Sync" count={8} />
|
|
||||||
<Category name="AI" count={15} />
|
|
||||||
</Categories>
|
|
||||||
<Results>
|
|
||||||
<AugmentationCard
|
|
||||||
name="Notion Synapse"
|
|
||||||
author="Soulcraft"
|
|
||||||
rating={4.8}
|
|
||||||
installs={1200}
|
|
||||||
price={9.99}
|
|
||||||
onInstall={...}
|
|
||||||
/>
|
|
||||||
</Results>
|
|
||||||
</AugmentationMarketplace>
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🎯 Priority for 2.0 Release
|
|
||||||
|
|
||||||
### Must Have (Release Blockers)
|
|
||||||
- ✅ Working execution (DONE)
|
|
||||||
- ✅ Clean interface (DONE)
|
|
||||||
- ✅ Documentation (DONE)
|
|
||||||
- ⏳ Fix augmentationPipeline.ts removal
|
|
||||||
- ⏳ Test all 27 augmentations work
|
|
||||||
|
|
||||||
### Nice to Have (2.0.x)
|
|
||||||
- Local package discovery
|
|
||||||
- NPM integration
|
|
||||||
- Basic CLI tools
|
|
||||||
|
|
||||||
### Future (2.1+)
|
|
||||||
- Brain Cloud Registry
|
|
||||||
- License management
|
|
||||||
- Payment processing
|
|
||||||
- Marketplace UI
|
|
||||||
- Developer portal
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📊 Current State Assessment
|
|
||||||
|
|
||||||
| Component | Status | Notes |
|
|
||||||
|-----------|--------|-------|
|
|
||||||
| Core Execution | ✅ Working | AugmentationRegistry.execute() works |
|
|
||||||
| Registration | ✅ Working | Manual registration works |
|
|
||||||
| Auto-Config | ✅ Working | Cache, index, storage auto-register |
|
|
||||||
| Lifecycle | ✅ Working | Init, execute, shutdown work |
|
|
||||||
| Discovery | ❌ Missing | No package discovery |
|
|
||||||
| Installation | ❌ Missing | No dynamic installation |
|
|
||||||
| Marketplace | ❌ Missing | No registry client |
|
|
||||||
| Licensing | ❌ Missing | No license management |
|
|
||||||
| Versioning | ❌ Missing | No version checks |
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 💡 Recommendations
|
|
||||||
|
|
||||||
### For 2.0 Release
|
|
||||||
1. **Ship with manual registration** - It works!
|
|
||||||
2. **Document how to create augmentations** - Critical for adoption
|
|
||||||
3. **Create 2-3 example augmentations** - Show the patterns
|
|
||||||
4. **Add basic CLI for testing** - Help developers
|
|
||||||
|
|
||||||
### For 2.1 (Q1 2025)
|
|
||||||
1. **Add NPM discovery** - Find installed augmentations
|
|
||||||
2. **Dynamic loading** - Import augmentations at runtime
|
|
||||||
3. **Basic marketplace API** - List available augmentations
|
|
||||||
4. **Version checking** - Ensure compatibility
|
|
||||||
|
|
||||||
### For 3.0 (Q2 2025)
|
|
||||||
1. **Full marketplace** - Browse, search, install
|
|
||||||
2. **Payment integration** - Premium augmentations
|
|
||||||
3. **Developer portal** - Publish augmentations
|
|
||||||
4. **Enterprise features** - Private registries
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## ✅ Good News Summary
|
|
||||||
|
|
||||||
The augmentation system WORKS! The core architecture is solid:
|
|
||||||
- Execution mechanism is correct
|
|
||||||
- Registration works
|
|
||||||
- Lifecycle management works
|
|
||||||
- All 27 augmentations function properly
|
|
||||||
|
|
||||||
What's missing is the marketplace/discovery layer, which can be added incrementally without breaking the core system. The 2.0 release can ship with manual augmentation registration, and the marketplace features can be added in 2.1+.
|
|
||||||
|
|
||||||
**Recommendation: Ship 2.0 with current system, add marketplace in 2.1**
|
|
||||||
|
|
@ -4,10 +4,8 @@
|
||||||
|
|
||||||
## ✅ Actually Implemented Augmentations (12+)
|
## ✅ Actually Implemented Augmentations (12+)
|
||||||
|
|
||||||
### 1. WAL (Write-Ahead Logging) Augmentation ✅
|
|
||||||
Full implementation with crash recovery, checkpointing, and replay.
|
Full implementation with crash recovery, checkpointing, and replay.
|
||||||
```typescript
|
```typescript
|
||||||
import { WALAugmentation } from 'brainy'
|
|
||||||
// Fully working with all features documented
|
// Fully working with all features documented
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -75,10 +73,10 @@ import { MemoryStorageAugmentation } from 'brainy'
|
||||||
```
|
```
|
||||||
|
|
||||||
### 11. Server Search Augmentation ✅
|
### 11. Server Search Augmentation ✅
|
||||||
Distributed search capabilities.
|
Server-side search delegation over a conduit.
|
||||||
```typescript
|
```typescript
|
||||||
import { ServerSearchConduitAugmentation } from 'brainy'
|
import { ServerSearchConduitAugmentation } from 'brainy'
|
||||||
// Distributed query execution
|
// Forwards queries to a remote Brainy server
|
||||||
```
|
```
|
||||||
|
|
||||||
### 12. Neural Import Augmentation ✅
|
### 12. Neural Import Augmentation ✅
|
||||||
|
|
@ -102,7 +100,7 @@ await neuralImport.detectRelationships(entities)
|
||||||
await neuralImport.generateInsights(data)
|
await neuralImport.generateInsights(data)
|
||||||
```
|
```
|
||||||
|
|
||||||
### Distributed Operation Modes (Fully Implemented!)
|
### Operation Modes (Fully Implemented!)
|
||||||
```typescript
|
```typescript
|
||||||
// Read-only mode with optimized caching
|
// Read-only mode with optimized caching
|
||||||
const readerMode = new ReaderMode()
|
const readerMode = new ReaderMode()
|
||||||
|
|
@ -140,7 +138,7 @@ monitor.getThrottlingMetrics() // Rate limiting info
|
||||||
## 📊 Statistics System (Fully Working!)
|
## 📊 Statistics System (Fully Working!)
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const stats = await brain.getStatistics()
|
const stats = await brain.getStats()
|
||||||
// Returns comprehensive metrics:
|
// Returns comprehensive metrics:
|
||||||
{
|
{
|
||||||
nouns: {
|
nouns: {
|
||||||
|
|
@ -208,7 +206,7 @@ if (device === 'webgpu') {
|
||||||
|
|
||||||
// CUDA detection in Node:
|
// CUDA detection in Node:
|
||||||
if (device === 'cuda') {
|
if (device === 'cuda') {
|
||||||
// Requires ONNX Runtime GPU packages
|
// Future: GPU acceleration support
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -241,21 +239,16 @@ const cacheConfig = await getCacheAutoConfig()
|
||||||
|
|
||||||
## 🎨 How to Use Hidden Features
|
## 🎨 How to Use Hidden Features
|
||||||
|
|
||||||
### Enable Distributed Modes
|
### Enable Reader / Writer Modes
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
mode: 'reader', // or 'writer' or 'hybrid'
|
mode: 'reader' // or 'writer' or 'hybrid'
|
||||||
distributed: {
|
|
||||||
role: 'reader',
|
|
||||||
cacheStrategy: 'aggressive',
|
|
||||||
prefetch: true
|
|
||||||
}
|
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
||||||
### Use Neural Import
|
### Use Neural Import
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new NeuralImportAugmentation({
|
new NeuralImportAugmentation({
|
||||||
confidenceThreshold: 0.7,
|
confidenceThreshold: 0.7,
|
||||||
|
|
@ -271,7 +264,7 @@ await brain.neuralImport('data.csv')
|
||||||
### Access Statistics
|
### Access Statistics
|
||||||
```typescript
|
```typescript
|
||||||
// Get comprehensive stats
|
// Get comprehensive stats
|
||||||
const stats = await brain.getStatistics()
|
const stats = await brain.getStats()
|
||||||
|
|
||||||
// Get specific service stats
|
// Get specific service stats
|
||||||
const nounStats = await brain.getStatistics({
|
const nounStats = await brain.getStatistics({
|
||||||
|
|
@ -287,7 +280,7 @@ const freshStats = await brain.getStatistics({
|
||||||
## 📝 What Needs Documentation
|
## 📝 What Needs Documentation
|
||||||
|
|
||||||
These features EXIST but need better docs:
|
These features EXIST but need better docs:
|
||||||
1. Distributed operation modes
|
1. Reader / writer operation modes
|
||||||
2. Neural import full API
|
2. Neural import full API
|
||||||
3. 3-level cache configuration
|
3. 3-level cache configuration
|
||||||
4. Performance monitoring API
|
4. Performance monitoring API
|
||||||
|
|
@ -299,7 +292,7 @@ These features EXIST but need better docs:
|
||||||
## 💡 The Truth
|
## 💡 The Truth
|
||||||
|
|
||||||
Brainy is MORE powerful than its own documentation suggests! Most "missing" features are actually implemented but hidden or not properly exposed. The codebase contains sophisticated systems for:
|
Brainy is MORE powerful than its own documentation suggests! Most "missing" features are actually implemented but hidden or not properly exposed. The codebase contains sophisticated systems for:
|
||||||
- Distributed operations
|
- Reader / writer operation modes
|
||||||
- AI-powered import
|
- AI-powered import
|
||||||
- Advanced caching
|
- Advanced caching
|
||||||
- Performance monitoring
|
- Performance monitoring
|
||||||
|
|
|
||||||
|
|
@ -15,7 +15,7 @@ High-performance deduplication for streaming data ingestion.
|
||||||
```typescript
|
```typescript
|
||||||
import { EntityRegistryAugmentation } from 'brainy'
|
import { EntityRegistryAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new EntityRegistryAugmentation({
|
new EntityRegistryAugmentation({
|
||||||
maxCacheSize: 100000, // Track up to 100k unique entities
|
maxCacheSize: 100000, // Track up to 100k unique entities
|
||||||
|
|
@ -26,8 +26,8 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
|
|
||||||
// Automatically prevents duplicate entities
|
// Automatically prevents duplicate entities
|
||||||
await brain.addNoun("Same content", { id: "123" }) // Added
|
await brain.add("Same content", { id: "123" }) // Added
|
||||||
await brain.addNoun("Same content", { id: "123" }) // Skipped (duplicate)
|
await brain.add("Same content", { id: "123" }) // Skipped (duplicate)
|
||||||
```
|
```
|
||||||
|
|
||||||
**Benefits:**
|
**Benefits:**
|
||||||
|
|
@ -36,17 +36,13 @@ await brain.addNoun("Same content", { id: "123" }) // Skipped (duplicate)
|
||||||
- Custom hash field selection
|
- Custom hash field selection
|
||||||
- Perfect for real-time data streams
|
- Perfect for real-time data streams
|
||||||
|
|
||||||
### 2. WAL (Write-Ahead Logging) Augmentation ✅ Available
|
|
||||||
|
|
||||||
Enterprise-grade durability and crash recovery.
|
Enterprise-grade durability and crash recovery.
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
import { WALAugmentation } from 'brainy'
|
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new WALAugmentation({
|
|
||||||
walPath: './wal', // WAL directory
|
|
||||||
checkpointInterval: 1000, // Checkpoint every 1000 operations
|
checkpointInterval: 1000, // Checkpoint every 1000 operations
|
||||||
compression: true, // Enable log compression
|
compression: true, // Enable log compression
|
||||||
maxLogSize: 100 * 1024 * 1024 // 100MB max log size
|
maxLogSize: 100 * 1024 * 1024 // 100MB max log size
|
||||||
|
|
@ -55,13 +51,10 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
|
|
||||||
// All operations are now durably logged
|
// All operations are now durably logged
|
||||||
await brain.addNoun("Critical data") // Written to WAL before storage
|
|
||||||
|
|
||||||
// Recover from crash
|
// Recover from crash
|
||||||
const recovered = new BrainyData({
|
const recovered = new Brainy({
|
||||||
augmentations: [new WALAugmentation({ recover: true })]
|
|
||||||
})
|
})
|
||||||
await recovered.init() // Automatically replays WAL
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Features:**
|
**Features:**
|
||||||
|
|
@ -78,7 +71,7 @@ AI-powered relationship strength calculation.
|
||||||
```typescript
|
```typescript
|
||||||
import { IntelligentVerbScoringAugmentation } from 'brainy'
|
import { IntelligentVerbScoringAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new IntelligentVerbScoringAugmentation({
|
new IntelligentVerbScoringAugmentation({
|
||||||
factors: {
|
factors: {
|
||||||
|
|
@ -92,8 +85,8 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
|
|
||||||
// Relationships automatically get intelligent scores
|
// Relationships automatically get intelligent scores
|
||||||
await brain.addVerb(user1, product1, "viewed", { timestamp: Date.now() })
|
await brain.relate(user1, product1, "viewed", { timestamp: Date.now() })
|
||||||
await brain.addVerb(user1, product1, "purchased", { timestamp: Date.now() })
|
await brain.relate(user1, product1, "purchased", { timestamp: Date.now() })
|
||||||
// Automatically calculates relationship strength based on multiple factors
|
// Automatically calculates relationship strength based on multiple factors
|
||||||
|
|
||||||
// Query using intelligent scores
|
// Query using intelligent scores
|
||||||
|
|
@ -118,7 +111,7 @@ Automatically extracts and registers entities from text.
|
||||||
```typescript
|
```typescript
|
||||||
import { AutoRegisterEntitiesAugmentation } from 'brainy'
|
import { AutoRegisterEntitiesAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new AutoRegisterEntitiesAugmentation({
|
new AutoRegisterEntitiesAugmentation({
|
||||||
types: ['person', 'organization', 'location', 'product'],
|
types: ['person', 'organization', 'location', 'product'],
|
||||||
|
|
@ -129,7 +122,7 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
|
|
||||||
// Automatically extracts and registers entities
|
// Automatically extracts and registers entities
|
||||||
await brain.addNoun(
|
await brain.add(
|
||||||
"Apple CEO Tim Cook announced the new iPhone 15 in Cupertino",
|
"Apple CEO Tim Cook announced the new iPhone 15 in Cupertino",
|
||||||
{ type: "news" }
|
{ type: "news" }
|
||||||
)
|
)
|
||||||
|
|
@ -154,7 +147,7 @@ Optimizes bulk operations for maximum throughput.
|
||||||
```typescript
|
```typescript
|
||||||
import { BatchProcessingAugmentation } from 'brainy'
|
import { BatchProcessingAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new BatchProcessingAugmentation({
|
new BatchProcessingAugmentation({
|
||||||
batchSize: 100,
|
batchSize: 100,
|
||||||
|
|
@ -167,7 +160,7 @@ const brain = new BrainyData({
|
||||||
|
|
||||||
// Operations are automatically batched
|
// Operations are automatically batched
|
||||||
for (let i = 0; i < 10000; i++) {
|
for (let i = 0; i < 10000; i++) {
|
||||||
await brain.addNoun(`Item ${i}`) // Internally batched
|
await brain.add(`Item ${i}`) // Internally batched
|
||||||
}
|
}
|
||||||
// Processes in optimized batches of 100
|
// Processes in optimized batches of 100
|
||||||
```
|
```
|
||||||
|
|
@ -185,7 +178,7 @@ Intelligent multi-level caching system.
|
||||||
```typescript
|
```typescript
|
||||||
import { CachingAugmentation } from 'brainy'
|
import { CachingAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new CachingAugmentation({
|
new CachingAugmentation({
|
||||||
levels: {
|
levels: {
|
||||||
|
|
@ -218,7 +211,7 @@ Reduces storage size while maintaining query performance.
|
||||||
```typescript
|
```typescript
|
||||||
import { CompressionAugmentation } from 'brainy'
|
import { CompressionAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new CompressionAugmentation({
|
new CompressionAugmentation({
|
||||||
algorithm: 'brotli',
|
algorithm: 'brotli',
|
||||||
|
|
@ -230,7 +223,7 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
|
|
||||||
// Data automatically compressed/decompressed
|
// Data automatically compressed/decompressed
|
||||||
await brain.addNoun(largeDocument) // Compressed before storage
|
await brain.add(largeDocument) // Compressed before storage
|
||||||
const doc = await brain.getNoun(id) // Decompressed on retrieval
|
const doc = await brain.getNoun(id) // Decompressed on retrieval
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -247,7 +240,7 @@ Real-time performance monitoring and metrics.
|
||||||
```typescript
|
```typescript
|
||||||
import { MonitoringAugmentation } from 'brainy'
|
import { MonitoringAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new MonitoringAugmentation({
|
new MonitoringAugmentation({
|
||||||
metrics: ['operations', 'latency', 'cache', 'memory'],
|
metrics: ['operations', 'latency', 'cache', 'memory'],
|
||||||
|
|
@ -286,7 +279,7 @@ brain.on('metrics', (metrics) => {
|
||||||
```typescript
|
```typescript
|
||||||
import { NeuralImportAugmentation } from 'brainy'
|
import { NeuralImportAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new NeuralImportAugmentation({
|
new NeuralImportAugmentation({
|
||||||
autoStructure: true,
|
autoStructure: true,
|
||||||
|
|
@ -382,7 +375,7 @@ import { Augmentation } from 'brainy'
|
||||||
class CustomAugmentation extends Augmentation {
|
class CustomAugmentation extends Augmentation {
|
||||||
name = 'CustomAugmentation'
|
name = 'CustomAugmentation'
|
||||||
|
|
||||||
async onInit(brain: BrainyData): Promise<void> {
|
async onInit(brain: Brainy): Promise<void> {
|
||||||
// Initialize augmentation
|
// Initialize augmentation
|
||||||
console.log('Custom augmentation initialized')
|
console.log('Custom augmentation initialized')
|
||||||
}
|
}
|
||||||
|
|
@ -415,7 +408,7 @@ class CustomAugmentation extends Augmentation {
|
||||||
}
|
}
|
||||||
|
|
||||||
// Use custom augmentation
|
// Use custom augmentation
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [new CustomAugmentation()]
|
augmentations: [new CustomAugmentation()]
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
@ -427,7 +420,7 @@ const brain = new BrainyData({
|
||||||
```typescript
|
```typescript
|
||||||
interface AugmentationHooks {
|
interface AugmentationHooks {
|
||||||
// Initialization
|
// Initialization
|
||||||
onInit(brain: BrainyData): Promise<void>
|
onInit(brain: Brainy): Promise<void>
|
||||||
onShutdown(): Promise<void>
|
onShutdown(): Promise<void>
|
||||||
|
|
||||||
// Noun operations
|
// Noun operations
|
||||||
|
|
@ -466,7 +459,7 @@ interface AugmentationHooks {
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Combine multiple augmentations
|
// Combine multiple augmentations
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
// Order matters - executed in sequence
|
// Order matters - executed in sequence
|
||||||
new EntityRegistryAugmentation(), // Deduplication first
|
new EntityRegistryAugmentation(), // Deduplication first
|
||||||
|
|
@ -474,7 +467,6 @@ const brain = new BrainyData({
|
||||||
new IntelligentVerbScoringAugmentation(), // Scoring
|
new IntelligentVerbScoringAugmentation(), // Scoring
|
||||||
new CompressionAugmentation(), // Compression
|
new CompressionAugmentation(), // Compression
|
||||||
new CachingAugmentation(), // Caching
|
new CachingAugmentation(), // Caching
|
||||||
new WALAugmentation(), // Durability
|
|
||||||
new MonitoringAugmentation() // Monitoring last
|
new MonitoringAugmentation() // Monitoring last
|
||||||
]
|
]
|
||||||
})
|
})
|
||||||
|
|
|
||||||
388
docs/architecture/data-storage-architecture.md
Normal file
388
docs/architecture/data-storage-architecture.md
Normal file
|
|
@ -0,0 +1,388 @@
|
||||||
|
# Brainy Data Storage Architecture (8.0)
|
||||||
|
|
||||||
|
**Complete on-disk reference for the 8.0 layout.**
|
||||||
|
|
||||||
|
This document describes what a Brainy 8.0 data directory actually contains: the
|
||||||
|
canonical entity records, the system area, the generational-MVCC bookkeeping,
|
||||||
|
the column store, the blob area, and the lock files — plus how the in-memory
|
||||||
|
indexes rebuild from them. The authoritative design records are
|
||||||
|
[ADR-001 (generational MVCC)](../ADR-001-generational-mvcc.md) and
|
||||||
|
[index-architecture.md](./index-architecture.md); this document is the on-disk
|
||||||
|
map that ties them together.
|
||||||
|
|
||||||
|
8.0 removed the 7.x copy-on-write subsystem (`_cow/`, `branches/{branch}/…`
|
||||||
|
paths) and the cloud/OPFS storage adapters. The two storage backends are
|
||||||
|
**filesystem** and **memory**; both speak the same path vocabulary (memory
|
||||||
|
storage keys its internal map by the identical path strings).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Directory Tree
|
||||||
|
|
||||||
|
A real 8.0 filesystem store (`path`, default `./brainy-data`):
|
||||||
|
|
||||||
|
```
|
||||||
|
brainy-data/
|
||||||
|
│
|
||||||
|
├── entities/ # Canonical records (current state)
|
||||||
|
│ ├── nouns/
|
||||||
|
│ │ └── {shard}/ # 256 shards: first 2 hex chars of the UUID
|
||||||
|
│ │ └── {id}/ # One directory per entity UUID
|
||||||
|
│ │ ├── vectors.json.gz # Embedding + HNSW node state
|
||||||
|
│ │ └── metadata.json.gz # Everything else (type, subtype, data, fields, _rev)
|
||||||
|
│ └── verbs/
|
||||||
|
│ └── {shard}/
|
||||||
|
│ └── {id}/
|
||||||
|
│ ├── vectors.json.gz # Relationship embedding (when present)
|
||||||
|
│ └── metadata.json.gz # sourceId, targetId, verb, subtype, weight, data…
|
||||||
|
│
|
||||||
|
├── _system/ # System singletons + bucketed system keys
|
||||||
|
│ ├── generation.json(.gz) # { generation, updatedAt } — the write watermark
|
||||||
|
│ ├── manifest.json # { version, generation, … } — MVCC commit point
|
||||||
|
│ ├── tx-log.jsonl # One line per committed transact() batch (append-only)
|
||||||
|
│ ├── counts.json # Entity/verb totals
|
||||||
|
│ ├── type-statistics.json.gz # Per-NounType counts
|
||||||
|
│ ├── subtype-statistics.json.gz # Per-(NounType, subtype) counts
|
||||||
|
│ ├── verb-subtype-statistics.json.gz # Per-(VerbType, subtype) counts
|
||||||
|
│ ├── statistics.json # Aggregate statistics blob (counts, index sizes)
|
||||||
|
│ ├── hnsw-system.json # Vector-index entry point + max level
|
||||||
|
│ ├── __metadata_field_registry__.json.gz # Which metadata fields are indexed
|
||||||
|
│ ├── brainy:entityIdMapper.json.gz # UUID ↔ u64 mapping for native index providers
|
||||||
|
│ └── idx/
|
||||||
|
│ └── {bucket}/ # 256 buckets: FNV-1a hash of the key
|
||||||
|
│ ├── __metadata_field_index__field_{name}.json.gz # Sparse field indexes
|
||||||
|
│ ├── __chunk__*.json.gz # Metadata-index roaring-bitmap chunks
|
||||||
|
│ ├── __sparse_index__*.json.gz# Zone maps + bloom filters
|
||||||
|
│ └── graph-lsm-verbs-{source|target}-*.json.gz # Graph LSM SSTables + manifest
|
||||||
|
│
|
||||||
|
├── _generations/ # MVCC history (written ONLY by transact())
|
||||||
|
│ └── {N}/ # One directory per committed generation N
|
||||||
|
│ ├── tx.json # The generation-N delta (immutable)
|
||||||
|
│ └── prev/
|
||||||
|
│ └── {id}.json # Before-image of each touched record (immutable)
|
||||||
|
│
|
||||||
|
├── _column_index/ # Column store manifests (one dir per field)
|
||||||
|
│ └── {field}/
|
||||||
|
│ └── MANIFEST.json.gz # Run list + zone metadata for that column
|
||||||
|
│
|
||||||
|
├── _blobs/ # Binary blob area (`<key>.bin` convention)
|
||||||
|
│ ├── _column_index/
|
||||||
|
│ │ └── {field}/
|
||||||
|
│ │ └── L0-000001.bin # Column-store runs (level-0 segments)
|
||||||
|
│ └── … # VFS file content and other binary blobs
|
||||||
|
│
|
||||||
|
└── locks/ # Process coordination (NEVER snapshotted)
|
||||||
|
├── _writer.lock # Single-writer lock: pid, hostname, heartbeat
|
||||||
|
├── _flush_requests/ # Reader→writer flush RPC (.req files)
|
||||||
|
└── _flush_responses/ # Writer acks (.ack files)
|
||||||
|
```
|
||||||
|
|
||||||
|
Most JSON objects are gzip-compressed (`.json.gz`) — compression is on by
|
||||||
|
default for filesystem storage (`storage.options.compression`, zlib level 6).
|
||||||
|
A few hot singletons (`manifest.json`, `counts.json`, `hnsw-system.json`,
|
||||||
|
`tx-log.jsonl`) are written uncompressed for cheap partial reads and appends.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Canonical Entity Records
|
||||||
|
|
||||||
|
Each entity (noun) and relationship (verb) is **two files** under one
|
||||||
|
ID-first directory. The split keeps vector I/O (large, append-mostly) separate
|
||||||
|
from metadata I/O (small, read-heavy).
|
||||||
|
|
||||||
|
### Noun vector file — `entities/nouns/{shard}/{id}/vectors.json.gz`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "421d92e7-4241-470a-80f4-4b39414e7a83",
|
||||||
|
"vector": [-0.139, -0.056, 0.028, "…384 dims…"],
|
||||||
|
"connections": { "0": ["neighbor-uuid…"] },
|
||||||
|
"level": 0
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The HNSW node state (`connections`, `level`) is persisted with the vector so
|
||||||
|
the vector index can rebuild without recomputing the graph.
|
||||||
|
|
||||||
|
### Noun metadata file — `entities/nouns/{shard}/{id}/metadata.json.gz`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"data": "React is a JavaScript library for building user interfaces",
|
||||||
|
"noun": "concept",
|
||||||
|
"subtype": "cli-add",
|
||||||
|
"createdAt": 1781198053726,
|
||||||
|
"updatedAt": 1781198053726,
|
||||||
|
"_rev": 1
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `noun` is the NounType. **Type lives in metadata, not in the path** — lookup
|
||||||
|
by ID is a single path construction, no type needed.
|
||||||
|
- `subtype` is the per-product sub-classification (required on write by
|
||||||
|
default in 8.0).
|
||||||
|
- `_rev` increments on every write and backs `update({ ifRev })` CAS.
|
||||||
|
- Consumer metadata fields sit alongside the standard ones.
|
||||||
|
|
||||||
|
### Verb files — `entities/verbs/{shard}/{id}/…`
|
||||||
|
|
||||||
|
Same two-file split. The verb metadata record carries the graph edge:
|
||||||
|
`sourceId`, `targetId`, `verb` (VerbType), `subtype`, `weight`, `data`,
|
||||||
|
`metadata`, timestamps, `_rev`. Verb IDs are Brainy-generated UUIDs by
|
||||||
|
contract (8.0 rejects caller-supplied verb ids) because native graph providers
|
||||||
|
intern the raw UUID bytes as u64 handles.
|
||||||
|
|
||||||
|
### Path construction
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const shard = id.substring(0, 2) // '42'
|
||||||
|
const metadataPath = `entities/nouns/${shard}/${id}/metadata.json`
|
||||||
|
const vectorsPath = `entities/nouns/${shard}/${id}/vectors.json`
|
||||||
|
// Verbs: same shape under entities/verbs/
|
||||||
|
```
|
||||||
|
|
||||||
|
Implemented in `src/storage/baseStorage.ts` (path generators) and
|
||||||
|
`src/storage/sharding.ts` (`getShardIdFromUuid`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. The `_system/` Area
|
||||||
|
|
||||||
|
Two kinds of keys live here, resolved by `BaseStorage.parsePath()`:
|
||||||
|
|
||||||
|
1. **Singletons** — well-known keys written at `_system/<key>.json`:
|
||||||
|
`counts`, `statistics`, `type-statistics`, `hnsw-system`,
|
||||||
|
`__metadata_field_registry__`, `brainy:entityIdMapper`, plus the MVCC
|
||||||
|
trio (`generation.json`, `manifest.json`, `tx-log.jsonl`).
|
||||||
|
2. **Bucketed system keys** — everything else (field indexes, bitmap chunks,
|
||||||
|
sparse-index segments, graph LSM SSTables) hashes into one of 256
|
||||||
|
`_system/idx/{bucket}/` directories via FNV-1a, so no single directory
|
||||||
|
accumulates unbounded entries.
|
||||||
|
|
||||||
|
Notable singletons:
|
||||||
|
|
||||||
|
| File | Contents |
|
||||||
|
|------|----------|
|
||||||
|
| `generation.json` | `{ generation, updatedAt }` — monotonic watermark, bumped by **every** write batch |
|
||||||
|
| `manifest.json` | MVCC commit point: highest *committed* generation (see §4) |
|
||||||
|
| `tx-log.jsonl` | One JSON line per committed `transact()` batch: generation, timestamp, `meta` |
|
||||||
|
| `type-statistics.json.gz` | Per-NounType counts (backs `brain.counts.byType`) |
|
||||||
|
| `subtype-statistics.json.gz` | `{ counts: { [type]: { [subtype]: n } }, updatedAt }` (contract-bound shape) |
|
||||||
|
| `verb-subtype-statistics.json.gz` | Same shape for VerbTypes |
|
||||||
|
| `hnsw-system.json` | `{ entryPointId, maxLevel }` — per-node state lives in each entity's `vectors.json` |
|
||||||
|
| `brainy:entityIdMapper.json.gz` | UUID ↔ u64 interning table for the BigInt provider contract |
|
||||||
|
| `__metadata_field_registry__.json.gz` | Registry of indexed metadata field names |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Generational MVCC (`_generations/` + the `_system` trio)
|
||||||
|
|
||||||
|
Full design in [ADR-001](../ADR-001-generational-mvcc.md). The on-disk shape:
|
||||||
|
|
||||||
|
```
|
||||||
|
_system/generation.json { generation, updatedAt } atomic tmp+rename
|
||||||
|
_system/manifest.json { version, generation, … } atomic tmp+rename — THE commit point
|
||||||
|
_system/tx-log.jsonl one line per committed transact() append-only
|
||||||
|
_generations/{N}/tx.json the generation-N delta immutable once written
|
||||||
|
_generations/{N}/prev/{id}.json before-image of {id} immutable once written
|
||||||
|
```
|
||||||
|
|
||||||
|
Commit protocol (writer side): stage before-images and the delta under
|
||||||
|
`_generations/N/`, fsync, apply the delta to the canonical `entities/…`
|
||||||
|
records, then atomically rename `manifest.json` to publish generation N. The
|
||||||
|
tx-log line is appended last (advisory). Crash recovery on open discards any
|
||||||
|
`_generations/{N}` newer than the manifest.
|
||||||
|
|
||||||
|
Two write classes share the generation clock:
|
||||||
|
|
||||||
|
- **Single-operation writes** (`add`/`update`/`remove`/`relate` outside
|
||||||
|
`transact()`) bump `generation.json` so watermarks and `_rev` CAS stay sound,
|
||||||
|
but write **no** history — they are not visible to `db.since()` and remain
|
||||||
|
visible through earlier pins.
|
||||||
|
- **`transact()` batches** write the full `_generations/{N}` record and a
|
||||||
|
tx-log line, and are the unit of time travel (`brain.asOf()`).
|
||||||
|
|
||||||
|
Snapshots (`db.persist(path)`) hard-link the entire store **except `locks/`**
|
||||||
|
into a self-contained directory openable via `Brainy.load(path)`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Column Store (`_column_index/` + `_blobs/_column_index/`)
|
||||||
|
|
||||||
|
The metadata index persists per-field columnar runs for O(log n) range and
|
||||||
|
membership queries at scale:
|
||||||
|
|
||||||
|
- `_column_index/{field}/MANIFEST.json.gz` — the run list and zone metadata
|
||||||
|
for one field (`createdAt`, `subtype`, `noun`, `_rev`, consumer fields,
|
||||||
|
`__words__` for tokenized text…).
|
||||||
|
- `_blobs/_column_index/{field}/L0-NNNNNN.bin` — the actual level-0 run
|
||||||
|
segments, stored through the shared `_blobs/<key>.bin` binary convention.
|
||||||
|
|
||||||
|
Sparse per-field indexes, roaring-bitmap chunks, and zone-map/bloom segments
|
||||||
|
additionally live as bucketed keys under `_system/idx/` (see §3). Which path
|
||||||
|
serves a given `where` clause is the query planner's decision — inspect it
|
||||||
|
with `brainy inspect explain <dir> --where '…'`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Blob Area (`_blobs/`)
|
||||||
|
|
||||||
|
`_blobs/<key>.bin` is the flat binary-blob convention shared by every storage
|
||||||
|
adapter (`saveBinaryBlob`/`getBinaryBlob` in the storage contract):
|
||||||
|
|
||||||
|
- **VFS file content** — VFS entities are regular nouns (path, ownership, and
|
||||||
|
timestamps in entity metadata); the file *bytes* are blobs.
|
||||||
|
- **Column-store runs** (under the `_column_index/` key prefix, §5).
|
||||||
|
- Any other binary payload an index provider persists.
|
||||||
|
|
||||||
|
Writes use unique temp names + rename, so concurrent writers of the same key
|
||||||
|
cannot tear each other's blobs.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Locks (`locks/`)
|
||||||
|
|
||||||
|
```
|
||||||
|
locks/_writer.lock # single-writer lock: { pid, hostname, startedAt, heartbeat, version }
|
||||||
|
locks/_flush_requests/ # readers drop <uuid>.req to ask the writer to flush
|
||||||
|
locks/_flush_responses/ # writer answers with <uuid>.ack
|
||||||
|
```
|
||||||
|
|
||||||
|
- One **writer** per data directory, enforced at `init()`; stale locks (dead
|
||||||
|
PID / stale heartbeat) are reclaimed automatically.
|
||||||
|
- Read-only processes (`Brainy.openReadOnly()`, the `brainy inspect` CLI
|
||||||
|
family) can ask the live writer to flush via the request/response files, so
|
||||||
|
out-of-process diagnostics see fresh state.
|
||||||
|
- `locks/` is excluded from snapshots (`SNAPSHOT_EXCLUDED_TOP_DIRS` in
|
||||||
|
`src/storage/adapters/fileSystemStorage.ts`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. In-Memory Indexes and What Rebuilds From What
|
||||||
|
|
||||||
|
| Index | In memory | Persisted state | Rebuild source |
|
||||||
|
|-------|-----------|-----------------|----------------|
|
||||||
|
| **Vector (HNSW)** | Graph of vector connections | `_system/hnsw-system.json` + per-entity `vectors.json` | Walk entity vector files; lazy mode loads structure only and pages vectors on demand |
|
||||||
|
| **Metadata index** | Field → value bitmaps + column-store readers | `_system/idx/` chunks + `_column_index/` manifests + `_blobs/_column_index/` runs | Loaded directly; full rebuild re-scans entity metadata |
|
||||||
|
| **Graph adjacency** | sourceId/targetId → verb-id LSM trees | `graph-lsm-verbs-{source,target}-*` SSTables under `_system/idx/` | Loaded from SSTables; full rebuild re-scans verb metadata |
|
||||||
|
| **Counts/statistics** | Per-type and per-subtype maps | `_system/{type,subtype,verb-subtype}-statistics.json.gz`, `counts.json` | Recomputable by scanning entities (`brainy inspect repair`) |
|
||||||
|
|
||||||
|
A pluggable index provider (the 8.0 plugin contract in
|
||||||
|
`@soulcraft/brainy/plugin`) may replace any of the JS implementations; the
|
||||||
|
persisted formats above are contract-bound so JS and native implementations
|
||||||
|
can interleave on the same directory.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Sharding Strategy
|
||||||
|
|
||||||
|
**Entities:** first 2 hex characters of the UUID → 256 uniform shards.
|
||||||
|
Deterministic, configuration-free, and keeps per-directory entry counts low
|
||||||
|
(at 1M entities: ~3,900 directories per shard). Paginated whole-store walks
|
||||||
|
(`getNouns`/`getVerbs`) iterate shards `00`–`ff` in order.
|
||||||
|
|
||||||
|
**System keys:** FNV-1a hash of the key → 256 `_system/idx/` buckets. Same
|
||||||
|
motivation, different keyspace (system keys are not UUIDs).
|
||||||
|
|
||||||
|
**What is never sharded:** the `_system/` singletons, `_generations/{N}`
|
||||||
|
directories (keyed by generation number), `_column_index/{field}` manifests
|
||||||
|
(keyed by field name), and `locks/`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Durability and Atomicity
|
||||||
|
|
||||||
|
- **Per-object atomicity:** every JSON object and blob is written to a unique
|
||||||
|
temp file then `rename()`d — readers never observe torn objects.
|
||||||
|
- **Transaction atomicity:** the `manifest.json` rename is the single commit
|
||||||
|
point for `transact()` batches (§4); everything staged before it is
|
||||||
|
discarded by crash recovery if the rename never lands.
|
||||||
|
- **Compression:** gzip per object (`.json.gz`), transparent to all readers.
|
||||||
|
Native index providers that mmap binary formats use the uncompressed
|
||||||
|
`_blobs/` area instead.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. `clear()` Semantics
|
||||||
|
|
||||||
|
`brain.clear()` removes all entities, relationships, indexes, statistics, and
|
||||||
|
MVCC history, then re-resolves every index exactly as `init()` does —
|
||||||
|
including plugin-provided vector/metadata/id-mapper factories and VFS root
|
||||||
|
re-creation. The data directory afterwards contains a fresh, empty store (the
|
||||||
|
writer lock remains held by the running process).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Common Scenarios
|
||||||
|
|
||||||
|
### Adding an entity
|
||||||
|
|
||||||
|
```
|
||||||
|
brain.add({ data, type, subtype })
|
||||||
|
1. Generate UUID → shard = first 2 hex chars
|
||||||
|
2. Embed data → 384-dim vector
|
||||||
|
3. Write entities/nouns/{shard}/{id}/vectors.json.gz (vector + HNSW node state)
|
||||||
|
4. Write entities/nouns/{shard}/{id}/metadata.json.gz (type/subtype/data/fields, _rev: 1)
|
||||||
|
5. Update in-memory indexes (HNSW insert, metadata index, statistics)
|
||||||
|
6. Bump _system/generation.json (no _generations/ entry — single-op write)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Committing a transaction
|
||||||
|
|
||||||
|
```
|
||||||
|
await brain.transact(tx => { tx.add(…); tx.update(…) })
|
||||||
|
1. Stage _generations/{N}/prev/{id}.json before-images + tx.json delta; fsync
|
||||||
|
2. Apply the delta to canonical entities/… records
|
||||||
|
3. Atomic-rename _system/manifest.json → generation N is committed
|
||||||
|
4. Append one line to _system/tx-log.jsonl
|
||||||
|
```
|
||||||
|
|
||||||
|
### Cold start
|
||||||
|
|
||||||
|
```
|
||||||
|
await brain.init()
|
||||||
|
1. Acquire locks/_writer.lock (or open read-only)
|
||||||
|
2. Crash recovery: drop _generations/{N} newer than manifest.json
|
||||||
|
3. Load _system singletons (counts, statistics, field registry, id mapper)
|
||||||
|
4. Vector index: hnsw-system.json + entity vectors.json (lazy mode if large)
|
||||||
|
5. Graph adjacency: load LSM SSTables from _system/idx/
|
||||||
|
6. Metadata index: column-store manifests + bitmap chunks on demand
|
||||||
|
```
|
||||||
|
|
||||||
|
### Snapshot and restore
|
||||||
|
|
||||||
|
```
|
||||||
|
const db = brain.now(); await db.persist('/backups/today'); await db.release()
|
||||||
|
→ hard-links everything except locks/ into a self-contained directory
|
||||||
|
|
||||||
|
await Brainy.load('/backups/today') // open snapshot read-only as a Db
|
||||||
|
await brain.restore('/backups/today', { confirm: true }) // replace store state
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Summary
|
||||||
|
|
||||||
|
- **Two backends** (filesystem, memory), one path vocabulary.
|
||||||
|
- **Two files per entity** under ID-first `entities/{kind}/{shard}/{id}/`.
|
||||||
|
- **Type and subtype are metadata**, not directory structure; type queries go
|
||||||
|
through the metadata index, not the filesystem.
|
||||||
|
- **`_system/`** holds singletons plus 256 hash buckets of index state.
|
||||||
|
- **`_generations/` + `manifest.json` + `tx-log.jsonl`** implement
|
||||||
|
generational MVCC. History is per-write: every `add()`/`update()`/`remove()`/
|
||||||
|
`relate()` gets its own generation, and `transact()` groups several ops into
|
||||||
|
one atomic generation. (Single-op retention has been the model since 8.0;
|
||||||
|
you never need to route a write through `transact()` just to keep its history.)
|
||||||
|
- **`_column_index/` + `_blobs/`** hold the columnar metadata runs and binary
|
||||||
|
blobs (VFS content included).
|
||||||
|
- **`locks/`** coordinates the single writer and reader flush requests, and
|
||||||
|
never travels with snapshots.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
- [ADR-001 — Generational MVCC](../ADR-001-generational-mvcc.md)
|
||||||
|
- [Index Architecture](./index-architecture.md)
|
||||||
|
- [Consistency Model](../concepts/consistency-model.md)
|
||||||
|
- [VFS Guide](../vfs/README.md)
|
||||||
538
docs/architecture/finite-type-system.md
Normal file
538
docs/architecture/finite-type-system.md
Normal file
|
|
@ -0,0 +1,538 @@
|
||||||
|
# 🎯 Brainy's Finite Noun/Verb Type System
|
||||||
|
|
||||||
|
> **Why Brainy's Finite Type System is Revolutionary for Knowledge Graphs at Billion Scale**
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Brainy introduces a **finite type system** that sits between traditional schemaless NoSQL and rigid relational databases. This approach unlocks unprecedented optimization opportunities while maintaining semantic flexibility.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Three-Way Comparison
|
||||||
|
|
||||||
|
### 1. Traditional NoSQL (Schemaless)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Complete freedom, zero optimization
|
||||||
|
{
|
||||||
|
id: '123',
|
||||||
|
randomField1: 'value',
|
||||||
|
anotherWeirdKey: 42,
|
||||||
|
whoKnowsWhatElse: { nested: 'chaos' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Problems:**
|
||||||
|
- ❌ No index optimization possible
|
||||||
|
- ❌ Tools can't understand data structure
|
||||||
|
- ❌ Incompatible augmentations/extensions
|
||||||
|
- ❌ Memory explosion with billions of unique keys
|
||||||
|
- ❌ No semantic understanding
|
||||||
|
- ❌ Query planning impossible
|
||||||
|
|
||||||
|
### 2. Traditional Relational (Rigid Schema)
|
||||||
|
|
||||||
|
```sql
|
||||||
|
CREATE TABLE entities (
|
||||||
|
id UUID PRIMARY KEY,
|
||||||
|
field1 VARCHAR(255),
|
||||||
|
field2 INTEGER,
|
||||||
|
...
|
||||||
|
field50 TEXT
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
**Problems:**
|
||||||
|
- ❌ Must define schema upfront
|
||||||
|
- ❌ Schema migrations are painful
|
||||||
|
- ❌ Can't handle heterogeneous data
|
||||||
|
- ❌ Requires restart for schema changes
|
||||||
|
- ❌ Fixed columns waste space
|
||||||
|
|
||||||
|
### 3. Brainy's Finite Type System (Semantic Structure)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Finite noun types (extensible but constrained)
|
||||||
|
type NounType =
|
||||||
|
| 'person' | 'place' | 'organization' | 'document'
|
||||||
|
| 'event' | 'concept' | 'thing' | ...
|
||||||
|
|
||||||
|
// Finite verb types (semantic relationships)
|
||||||
|
type VerbType =
|
||||||
|
| 'relatedTo' | 'contains' | 'isA' | 'causedBy'
|
||||||
|
| 'precedes' | 'influences' | ...
|
||||||
|
|
||||||
|
// Example usage
|
||||||
|
const entity = {
|
||||||
|
id: '123',
|
||||||
|
nounType: 'person', // Finite! Known type
|
||||||
|
vector: [...], // Semantic embedding
|
||||||
|
metadata: {
|
||||||
|
noun: 'person', // Required type field
|
||||||
|
name: 'Alice', // Custom fields allowed
|
||||||
|
occupation: 'Engineer' // Flexible metadata
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- ✅ **Index Optimization**: Fixed-size Uint32Arrays for type tracking (99.76% memory reduction)
|
||||||
|
- ✅ **Semantic Understanding**: Types have meaning, not just structure
|
||||||
|
- ✅ **Tool Compatibility**: All augmentations understand core types
|
||||||
|
- ✅ **Concept Extraction**: NLP can map text to known types
|
||||||
|
- ✅ **Explicit Types**: Clear type specification in API
|
||||||
|
- ✅ **Query Optimization**: Type-aware query planning
|
||||||
|
- ✅ **Flexible Metadata**: Any fields within typed structure
|
||||||
|
- ✅ **Billion-Scale Ready**: Type tracking scales linearly
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Revolutionary Benefits in Detail
|
||||||
|
|
||||||
|
### 1. Index Optimization at Billion Scale
|
||||||
|
|
||||||
|
**The Problem**: Traditional NoSQL stores arbitrary field names in indexes:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Memory explosion with unique keys
|
||||||
|
Map<string, Set<string>> {
|
||||||
|
"user_preference_notification_email_enabled": Set(['id1', 'id2', ...]),
|
||||||
|
"customer_shipping_address_line_1": Set(['id3', 'id4', ...]),
|
||||||
|
// Billions of unique, unpredictable keys!
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Brainy's Solution**: Fixed noun/verb types enable fixed-size tracking:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// 99.76% memory reduction with Uint32Arrays
|
||||||
|
class TypeAwareMetadataIndex {
|
||||||
|
// Fixed size: nounTypes × verbTypes × fieldCount
|
||||||
|
private nounTypeBitmaps: RoaringBitmap32[] // One per noun type
|
||||||
|
private verbTypeBitmaps: RoaringBitmap32[] // One per verb type
|
||||||
|
|
||||||
|
// Example: 100 noun types × 50 verb types = 5KB overhead
|
||||||
|
// vs 500MB+ for arbitrary keys!
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Real-World Impact (PROJECTED - not yet benchmarked)**:
|
||||||
|
- **Before**: 500MB memory for 1M entities with diverse keys
|
||||||
|
- **After**: PROJECTED 1.2MB memory for same dataset (385x reduction - calculated from Uint32Array size, not measured)
|
||||||
|
- **Scales to billions**: Memory grows with entity count, not key diversity
|
||||||
|
|
||||||
|
### 2. Explicit Type System
|
||||||
|
|
||||||
|
**The Design**: Specify types clearly in your API calls:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { Brainy, NounType, VerbType } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
// Add entity with explicit type
|
||||||
|
await brain.add({
|
||||||
|
data: { name: 'Alice', role: 'CEO of Acme Corp' },
|
||||||
|
type: NounType.Person // Explicit type specification
|
||||||
|
})
|
||||||
|
|
||||||
|
// Query with type filtering
|
||||||
|
await brain.find({
|
||||||
|
query: 'Alice',
|
||||||
|
type: NounType.Person // Type-optimized search
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why Explicit Types?**:
|
||||||
|
1. **Deterministic**: You control exactly how entities are classified
|
||||||
|
2. **Predictable**: No inference surprises or edge cases
|
||||||
|
3. **Fast**: No neural processing overhead on every add/query
|
||||||
|
4. **Smaller**: No embedded keyword models needed
|
||||||
|
|
||||||
|
**Real-World Use Case**:
|
||||||
|
```typescript
|
||||||
|
// Import data with known types
|
||||||
|
await brain.add({
|
||||||
|
data: { name: 'Apple Inc.', industry: 'Technology' },
|
||||||
|
type: NounType.Organization
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.add({
|
||||||
|
data: { name: 'Cupertino', country: 'USA' },
|
||||||
|
type: NounType.Location
|
||||||
|
})
|
||||||
|
|
||||||
|
// Create relationship
|
||||||
|
await brain.relate({
|
||||||
|
from: appleId,
|
||||||
|
to: cupertinoId,
|
||||||
|
type: VerbType.LocatedIn
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Tool & Augmentation Compatibility
|
||||||
|
|
||||||
|
**The Problem with Schemaless**: Every tool must handle infinite variations:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Incompatible tools
|
||||||
|
const tool1Data = { type: 'person', name: 'Alice' }
|
||||||
|
const tool2Data = { kind: 'human', fullName: 'Alice' }
|
||||||
|
const tool3Data = { entity_type: 'individual', person_name: 'Alice' }
|
||||||
|
|
||||||
|
// Tools can't understand each other!
|
||||||
|
```
|
||||||
|
|
||||||
|
**Brainy's Solution**: Finite types create a common language:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// All tools/augmentations understand core types
|
||||||
|
interface NounMetadata {
|
||||||
|
noun: NounType // Agreed-upon type system
|
||||||
|
// ... custom fields
|
||||||
|
}
|
||||||
|
|
||||||
|
// Augmentation 1: Adds caching for 'person' entities
|
||||||
|
class PersonCacheAugmentation {
|
||||||
|
execute(op, params) {
|
||||||
|
if (params.noun?.metadata?.noun === 'person') {
|
||||||
|
// All person entities are understood!
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Augmentation 2: Enriches 'organization' entities
|
||||||
|
class OrgEnrichmentAugmentation {
|
||||||
|
execute(op, params) {
|
||||||
|
if (params.noun?.metadata?.noun === 'organization') {
|
||||||
|
// Fetch industry data, employees, etc.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Augmentations compose seamlessly!
|
||||||
|
```
|
||||||
|
|
||||||
|
**Ecosystem Benefits**:
|
||||||
|
- Third-party augmentations are **interoperable**
|
||||||
|
- Type-specific optimizations are **portable**
|
||||||
|
- Query builders understand **semantic structure**
|
||||||
|
- Visualization tools render **type-appropriate** displays
|
||||||
|
- Import/export tools map to **universal types**
|
||||||
|
|
||||||
|
### 4. Concept Extraction & NLP Integration
|
||||||
|
|
||||||
|
**Traditional Approach**: Extract entities, ignore types:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Generic NER (Named Entity Recognition)
|
||||||
|
"Alice works at Google"
|
||||||
|
// → ['Alice', 'Google'] // What are these?
|
||||||
|
```
|
||||||
|
|
||||||
|
**Brainy's Approach**: Extract **typed** concepts:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { NaturalLanguageProcessor } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
const nlp = new NaturalLanguageProcessor()
|
||||||
|
const concepts = await nlp.extractConcepts("Alice works at Google in San Francisco")
|
||||||
|
|
||||||
|
// Returns typed entities:
|
||||||
|
[
|
||||||
|
{ text: 'Alice', nounType: 'person', confidence: 0.95 },
|
||||||
|
{ text: 'Google', nounType: 'organization', confidence: 0.98 },
|
||||||
|
{ text: 'San Francisco', nounType: 'place', confidence: 0.92 }
|
||||||
|
]
|
||||||
|
|
||||||
|
// And typed relationships:
|
||||||
|
[
|
||||||
|
{
|
||||||
|
from: 'Alice',
|
||||||
|
to: 'Google',
|
||||||
|
verbType: 'worksAt',
|
||||||
|
confidence: 0.88
|
||||||
|
},
|
||||||
|
{
|
||||||
|
from: 'Google',
|
||||||
|
to: 'San Francisco',
|
||||||
|
verbType: 'locatedIn',
|
||||||
|
confidence: 0.85
|
||||||
|
}
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Downstream Benefits**:
|
||||||
|
- **Smart Clustering**: Group by semantic type, not arbitrary keys
|
||||||
|
- **Type-Aware Queries**: "Find all organizations in California"
|
||||||
|
- **Relationship Reasoning**: "Who works at companies in SF?"
|
||||||
|
- **Automatic Ontology**: Types form natural hierarchy
|
||||||
|
|
||||||
|
### 5. Query Optimization & Planning
|
||||||
|
|
||||||
|
**The Problem**: Schemaless queries are guesswork:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- MongoDB: No idea what fields exist
|
||||||
|
db.collection.find({ someField: 'value' })
|
||||||
|
// Full collection scan!
|
||||||
|
```
|
||||||
|
|
||||||
|
**Brainy's Solution**: Type-aware query planning:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Query planner knows types exist!
|
||||||
|
brain.find({
|
||||||
|
where: { noun: 'person' } // Type index lookup: O(1)!
|
||||||
|
})
|
||||||
|
|
||||||
|
// Multi-type queries are optimized
|
||||||
|
brain.find({
|
||||||
|
where: {
|
||||||
|
noun: ['person', 'organization'], // Bitmap union
|
||||||
|
location: 'California' // Then filter
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Relationship traversal is type-aware
|
||||||
|
brain.find({
|
||||||
|
verb: 'worksAt', // Verb type index
|
||||||
|
sourceType: 'person', // Source noun type index
|
||||||
|
targetType: 'organization' // Target noun type index
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Query Performance**:
|
||||||
|
- **Type Filtering**: O(1) bitmap intersection
|
||||||
|
- **Join Planning**: Type-aware join order optimization
|
||||||
|
- **Index Selection**: Automatic best index for type
|
||||||
|
- **Cardinality Estimation**: Type statistics guide planning
|
||||||
|
|
||||||
|
### 6. Architecture & Development Benefits
|
||||||
|
|
||||||
|
#### Memory-Efficient Type Tracking
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Traditional approach: Map per field
|
||||||
|
class TraditionalIndex {
|
||||||
|
private fieldIndexes: Map<string, Map<any, Set<string>>>
|
||||||
|
// Memory: O(unique_fields × unique_values × entities)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Brainy approach: Fixed Uint32Array per type
|
||||||
|
class TypeAwareIndex {
|
||||||
|
private nounTypeTracking: Uint32Array // Fixed size!
|
||||||
|
private typeIndexes: RoaringBitmap32[] // One per type
|
||||||
|
// Memory: O(noun_types) + O(entities_per_type)
|
||||||
|
// PROJECTED: 385x smaller at billion scale (calculated from architecture, not benchmarked)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Type-Driven Code Organization
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Natural code structure follows types
|
||||||
|
/src
|
||||||
|
/nouns
|
||||||
|
/person
|
||||||
|
personStorage.ts // Type-specific storage
|
||||||
|
personQueries.ts // Type-specific queries
|
||||||
|
personAugmentation.ts // Type-specific logic
|
||||||
|
/organization
|
||||||
|
orgStorage.ts
|
||||||
|
orgQueries.ts
|
||||||
|
orgAugmentation.ts
|
||||||
|
/verbs
|
||||||
|
/worksAt
|
||||||
|
worksAtValidation.ts // Relationship rules
|
||||||
|
worksAtInference.ts // Type inference
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Type Safety in TypeScript
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Compiler-enforced type correctness
|
||||||
|
function processPerson(noun: Noun) {
|
||||||
|
if (noun.metadata.noun === 'person') {
|
||||||
|
// TypeScript narrows type!
|
||||||
|
const name: string = noun.metadata.name // Safe access
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Exhaustive type checking
|
||||||
|
function processNoun(noun: Noun) {
|
||||||
|
switch (noun.metadata.noun) {
|
||||||
|
case 'person': return handlePerson(noun)
|
||||||
|
case 'place': return handlePlace(noun)
|
||||||
|
case 'organization': return handleOrg(noun)
|
||||||
|
// Compiler error if missing cases!
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Public API: Type System
|
||||||
|
|
||||||
|
The type system is **fully public** for developers and augmentation authors:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import {
|
||||||
|
NounType,
|
||||||
|
VerbType,
|
||||||
|
getNounTypes,
|
||||||
|
getVerbTypes,
|
||||||
|
BrainyTypes,
|
||||||
|
suggestType
|
||||||
|
} from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
// Get all available noun types
|
||||||
|
const nounTypes = getNounTypes()
|
||||||
|
// → ['Person', 'Organization', 'Location', 'Thing', 'Concept', ...]
|
||||||
|
|
||||||
|
// Get all available verb types
|
||||||
|
const verbTypes = getVerbTypes()
|
||||||
|
// → ['RelatedTo', 'Contains', 'CreatedBy', 'LocatedIn', ...]
|
||||||
|
|
||||||
|
// Use types directly
|
||||||
|
await brain.add({
|
||||||
|
data: { name: 'Alice' },
|
||||||
|
type: NounType.Person
|
||||||
|
})
|
||||||
|
|
||||||
|
// Query by type
|
||||||
|
await brain.find({
|
||||||
|
type: NounType.Person,
|
||||||
|
where: { name: 'Alice' }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use Cases**:
|
||||||
|
- **Type-Safe Code**: Use TypeScript enums for compile-time checking
|
||||||
|
- **Import Tools**: Specify entity types during data import
|
||||||
|
- **Query Builders**: Filter by known types
|
||||||
|
- **Augmentations**: Type-specific processing pipelines
|
||||||
|
- **Visualization**: Type-appropriate rendering
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Real-World Performance Comparison
|
||||||
|
|
||||||
|
### Scenario: 1 Billion Entities with Rich Metadata
|
||||||
|
|
||||||
|
| Aspect | NoSQL (Schemaless) | Relational (Fixed) | Brainy (Finite Types) |
|
||||||
|
|--------|-------------------|-------------------|----------------------|
|
||||||
|
| **Memory (Indexes)** | 500GB+ | 250GB | 1.3GB |
|
||||||
|
| **Type Lookup** | Full scan | O(log n) | O(1) bitmap |
|
||||||
|
| **Add New Type** | Zero cost | Schema migration! | Register type |
|
||||||
|
| **Query Planning** | Impossible | Table statistics | Type statistics |
|
||||||
|
| **Tool Compatibility** | None | SQL only | Full ecosystem |
|
||||||
|
| **Semantic Understanding** | None | None | Built-in |
|
||||||
|
| **Concept Extraction** | Manual | Manual | Via SmartExtractor |
|
||||||
|
| **Flexibility** | Infinite | Zero | Optimal balance |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Design Principles
|
||||||
|
|
||||||
|
### 1. Finite but Extensible
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Core types are finite
|
||||||
|
const coreNounTypes = [
|
||||||
|
'person', 'place', 'organization', 'thing', ...
|
||||||
|
]
|
||||||
|
|
||||||
|
// But easily extended
|
||||||
|
brain.registerNounType('chemical_compound', {
|
||||||
|
keywords: ['molecule', 'compound', 'element'],
|
||||||
|
synonyms: ['substance', 'material'],
|
||||||
|
parentType: 'thing'
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### 1a. Subtypes — sub-classification without hierarchy
|
||||||
|
|
||||||
|
The 42-type taxonomy is intentionally coarse. Per-product vocabulary fits on the **`subtype`** axis — a top-level standard string field on every entity. Flat by design — no hierarchy, no parent chain, no recursive resolution. That preserves the Uint32Array-backed O(1) type stats while giving consumers a place to put `'employee'` / `'customer'` / `'invoice'` / `'milestone'` without burning a slot in the global enum.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Same NounType, different subtypes:
|
||||||
|
await brain.add({ type: NounType.Person, subtype: 'employee' })
|
||||||
|
await brain.add({ type: NounType.Person, subtype: 'customer' })
|
||||||
|
await brain.add({ type: NounType.Document, subtype: 'invoice' })
|
||||||
|
|
||||||
|
// Fast path — column-store hit, not metadata fallback:
|
||||||
|
await brain.find({ type: NounType.Person, subtype: 'employee' })
|
||||||
|
|
||||||
|
// Per-NounType-per-subtype counts maintained incrementally:
|
||||||
|
brain.counts.bySubtype(NounType.Person)
|
||||||
|
// → { employee: 12, customer: 847 }
|
||||||
|
```
|
||||||
|
|
||||||
|
Subtype has its own statistics rollup (`_system/subtype-statistics.json`) maintained alongside `nounCountsByType`, so per-subtype counts stay O(1) at billion scale.
|
||||||
|
|
||||||
|
The same principle applies to **VerbTypes** (7.30+): the 127-verb taxonomy is intentionally coarse, and `subtype` is the per-product axis for relationships too. A `ReportsTo` relationship might carry `subtype: 'direct'` vs `'dotted-line'`; a `RelatedTo` edge might carry `'spouse'` / `'colleague'`. Verb-side rollup lives at `_system/verb-subtype-statistics.json` with identical shape to the noun-side rollup. Per-VerbType-per-subtype counts are O(1) via `brain.counts.byRelationshipSubtype()`. Brainy's design is fully symmetric — nouns and verbs are first-class peers with identical capability surfaces.
|
||||||
|
|
||||||
|
Full guide: **[Subtypes & Facets](../guides/subtypes-and-facets.md)**.
|
||||||
|
|
||||||
|
### 2. Semantic not Structural
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// NOT structural types
|
||||||
|
type Person = {
|
||||||
|
name: string
|
||||||
|
age: number
|
||||||
|
// Fixed structure
|
||||||
|
}
|
||||||
|
|
||||||
|
// Semantic types
|
||||||
|
type Noun = {
|
||||||
|
nounType: 'person', // Semantic meaning!
|
||||||
|
metadata: {
|
||||||
|
noun: 'person', // Required type
|
||||||
|
// Any custom fields!
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Optimizable yet Flexible
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Optimized type tracking
|
||||||
|
const typeIndex = new RoaringBitmap32() // 99.76% smaller!
|
||||||
|
|
||||||
|
// Flexible metadata
|
||||||
|
const metadata = {
|
||||||
|
noun: 'person', // Required type
|
||||||
|
customField1: 'value', // Your fields
|
||||||
|
customField2: 123, // Any structure
|
||||||
|
nested: { ... } // Full flexibility
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Conclusion
|
||||||
|
|
||||||
|
Brainy's **Finite Noun/Verb Type System** is revolutionary because it achieves the impossible:
|
||||||
|
|
||||||
|
1. ✅ **Billion-scale performance** (99.76% memory reduction)
|
||||||
|
2. ✅ **Semantic understanding** (NLP integration)
|
||||||
|
3. ✅ **Tool compatibility** (ecosystem interoperability)
|
||||||
|
4. ✅ **Query optimization** (type-aware planning)
|
||||||
|
5. ✅ **Concept extraction** (via SmartExtractor for imports)
|
||||||
|
6. ✅ **Developer experience** (clean architecture)
|
||||||
|
7. ✅ **Flexibility** (metadata freedom within types)
|
||||||
|
|
||||||
|
It's not schemaless chaos. It's not rigid relational constraints. It's **semantic structure** - the perfect balance for knowledge graphs at scale.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Further Reading
|
||||||
|
|
||||||
|
- [Storage Architecture](./storage-architecture.md) - How types enable billion-scale storage
|
||||||
|
- [Augmentation System](./augmentations.md) - Building type-aware augmentations
|
||||||
|
- [Query Optimization](../api/query-optimization.md) - Type-aware query planning
|
||||||
|
- [Import Flow](../guides/import-flow.md) - How types work in the import pipeline
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*Brainy's finite type system: The foundation of billion-scale, semantically-aware knowledge graphs.*
|
||||||
941
docs/architecture/index-architecture.md
Normal file
941
docs/architecture/index-architecture.md
Normal file
|
|
@ -0,0 +1,941 @@
|
||||||
|
# Index Architecture
|
||||||
|
|
||||||
|
Brainy uses a sophisticated **3-tier index architecture** that enables "Triple Intelligence" - the unified combination of vector similarity, graph relationships, and metadata filtering. This document provides a comprehensive architectural overview of how these indexes work internally and coordinate with each other.
|
||||||
|
|
||||||
|
## Overview: The Three Main Indexes + Sub-Indexes
|
||||||
|
|
||||||
|
Brainy has **3 main indexes** at the top level, each with multiple sub-indexes managed automatically:
|
||||||
|
|
||||||
|
### Main Indexes (Level 1)
|
||||||
|
|
||||||
|
| Index | Purpose | Data Structure | Complexity | File Location | rebuild() Method |
|
||||||
|
|-------|---------|----------------|------------|---------------|------------------|
|
||||||
|
| **TypeAwareVectorIndex** | Type-aware vector similarity search | 42 type-specific hierarchical graphs | O(log n) search | `src/hnsw/typeAwareHNSWIndex.ts` | ✅ Line 403 |
|
||||||
|
| **MetadataIndexManager** | Fast metadata filtering | Chunked sparse indices with bloom filters + zone maps + roaring bitmaps | O(1) exact, O(log n) ranges | `src/utils/metadataIndex.ts` | ✅ Line 2318 |
|
||||||
|
| **GraphAdjacencyIndex** | Relationship traversal | 2 verb-id LSM-trees + tombstone-filtered adjacency derivation | O(degree) per hop | `src/graph/graphAdjacencyIndex.ts` | ✅ Line 389 |
|
||||||
|
|
||||||
|
### Sub-Indexes (Level 2)
|
||||||
|
|
||||||
|
**TypeAwareVectorIndex contains:**
|
||||||
|
- **42 type-specific vector indexes** - One per NounType (automatically rebuilt via parent)
|
||||||
|
|
||||||
|
**MetadataIndexManager contains:**
|
||||||
|
- **ChunkManager** - Adaptive chunked sparse indexing
|
||||||
|
- **EntityIdMapper** - UUID ↔ integer mapping for roaring bitmaps
|
||||||
|
- **FieldTypeInference** - DuckDB-inspired value-based field type detection
|
||||||
|
- **Field Sparse Indexes** - Per-field sparse indexes with roaring bitmaps (dynamic count)
|
||||||
|
- **Sorted Indexes** - Support orderBy queries (automatically maintained)
|
||||||
|
- **Word Index (`__words__`)** - Text search via FNV-1a word hashes
|
||||||
|
|
||||||
|
**GraphAdjacencyIndex contains:**
|
||||||
|
- **lsmTreeSource** - Source → Targets (outgoing edges)
|
||||||
|
- **lsmTreeTarget** - Target → Sources (incoming edges)
|
||||||
|
- **lsmTreeVerbsBySource** - Source → Verb IDs
|
||||||
|
- **lsmTreeVerbsByTarget** - Target → Verb IDs
|
||||||
|
|
||||||
|
All indexes share a **UnifiedCache** for coordinated memory management, ensuring fair resource allocation and preventing any single index from monopolizing memory.
|
||||||
|
|
||||||
|
## 1. MetadataIndex - Fast Field Filtering
|
||||||
|
|
||||||
|
**Purpose**: Enable O(1) field-value lookups and O(log n) range queries on metadata fields using adaptive chunked sparse indexing.
|
||||||
|
|
||||||
|
### Internal Architecture
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class MetadataIndexManager {
|
||||||
|
// Chunked sparse indices: field → SparseIndex (replaces flat files)
|
||||||
|
private sparseIndices = new Map<string, SparseIndex>()
|
||||||
|
|
||||||
|
// Chunk management
|
||||||
|
private chunkManager: ChunkManager
|
||||||
|
private chunkingStrategy: AdaptiveChunkingStrategy
|
||||||
|
|
||||||
|
// Lightweight field statistics
|
||||||
|
private fieldIndexes = new Map<string, FieldIndexData>() // value → count
|
||||||
|
private fieldStats = new Map<string, FieldStats>() // cardinality tracking
|
||||||
|
|
||||||
|
// Type-field affinity for NLP understanding
|
||||||
|
private typeFieldAffinity = new Map<string, Map<string, number>>()
|
||||||
|
|
||||||
|
// Shared memory management
|
||||||
|
private unifiedCache: UnifiedCache
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key Data Structures
|
||||||
|
|
||||||
|
#### Chunked Sparse Index
|
||||||
|
```typescript
|
||||||
|
// SparseIndex: Directory of chunks for a field
|
||||||
|
// Example: field="status"
|
||||||
|
class SparseIndex {
|
||||||
|
field: string
|
||||||
|
chunks: ChunkDescriptor[] // Metadata about each chunk
|
||||||
|
bloomFilters: BloomFilter[] // Fast membership testing
|
||||||
|
}
|
||||||
|
|
||||||
|
// ChunkDescriptor: Metadata about a chunk
|
||||||
|
interface ChunkDescriptor {
|
||||||
|
chunkId: number
|
||||||
|
valueCount: number // How many unique values in this chunk
|
||||||
|
idCount: number // Total entity IDs
|
||||||
|
zoneMap: ZoneMap // Min/max for range queries
|
||||||
|
lastUpdated: number
|
||||||
|
}
|
||||||
|
|
||||||
|
// Actual chunk data stored separately
|
||||||
|
class ChunkData {
|
||||||
|
chunkId: number
|
||||||
|
field: string
|
||||||
|
entries: Map<value, RoaringBitmap32> // ~50 values per chunk (roaring bitmaps!)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance**:
|
||||||
|
- O(1) exact lookup with bloom filters (1% false positive rate)
|
||||||
|
- O(log n) range queries with zone maps
|
||||||
|
- 630x file reduction (560k flat files → 89 chunk files)
|
||||||
|
|
||||||
|
#### Roaring Bitmap Optimization
|
||||||
|
|
||||||
|
**Problem Solved**: JavaScript `Set<string>` for storing entity IDs was inefficient:
|
||||||
|
- Memory overhead: ~40 bytes per UUID string (36 chars + overhead)
|
||||||
|
- Slow intersection: JavaScript array filtering for multi-field queries
|
||||||
|
- No hardware acceleration
|
||||||
|
|
||||||
|
**Solution**: Replace `Set<string>` with `RoaringBitmap32` (WebAssembly implementation) for 90% memory savings and hardware-accelerated operations. Uses `roaring-wasm` package for universal compatibility (Node.js, browsers, serverless) without requiring native compilation.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// EntityIdMapper: UUID ↔ Integer mapping
|
||||||
|
class EntityIdMapper {
|
||||||
|
private uuidToInt = new Map<string, number>()
|
||||||
|
private intToUuid = new Map<number, string>()
|
||||||
|
private nextId = 1
|
||||||
|
|
||||||
|
getOrAssign(uuid: string): number {
|
||||||
|
// O(1) mapping: UUIDs → integers for bitmap storage
|
||||||
|
let intId = this.uuidToInt.get(uuid)
|
||||||
|
if (!intId) {
|
||||||
|
intId = this.nextId++
|
||||||
|
this.uuidToInt.set(uuid, intId)
|
||||||
|
this.intToUuid.set(intId, uuid)
|
||||||
|
}
|
||||||
|
return intId
|
||||||
|
}
|
||||||
|
|
||||||
|
intsIterableToUuids(ints: Iterable<number>): string[] {
|
||||||
|
// Convert bitmap results back to UUIDs
|
||||||
|
const result: string[] = []
|
||||||
|
for (const intId of ints) {
|
||||||
|
const uuid = this.intToUuid.get(intId)
|
||||||
|
if (uuid) result.push(uuid)
|
||||||
|
}
|
||||||
|
return result
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ChunkData now uses RoaringBitmap32 instead of Set<string>
|
||||||
|
class ChunkData {
|
||||||
|
chunkId: number
|
||||||
|
field: string
|
||||||
|
entries: Map<string, RoaringBitmap32> // value → bitmap of integer IDs
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key Benefits**:
|
||||||
|
- **90% memory savings**: Roaring bitmaps compress much better than UUID strings
|
||||||
|
- **Hardware-accelerated operations**: SIMD instructions (AVX2/SSE4.2) for ultra-fast bitmap AND/OR
|
||||||
|
- **Portable serialization**: Cross-platform compatible format (Java/Go/Node.js)
|
||||||
|
- **Lazy conversion**: UUIDs converted to integers only once, not per query
|
||||||
|
|
||||||
|
**Multi-Field Intersection (THE BIG WIN!)**:
|
||||||
|
```typescript
|
||||||
|
// Before: JavaScript array filtering
|
||||||
|
async getIdsForFilter(filter: {status: 'active', role: 'admin'}): Promise<string[]> {
|
||||||
|
// 1. Fetch UUID arrays for each field
|
||||||
|
const statusIds = await this.getIds('status', 'active') // ["uuid1", "uuid2", ...]
|
||||||
|
const roleIds = await this.getIds('role', 'admin') // ["uuid2", "uuid3", ...]
|
||||||
|
|
||||||
|
// 2. JavaScript intersection (SLOW!)
|
||||||
|
return statusIds.filter(id => roleIds.includes(id)) // O(n*m) array filtering
|
||||||
|
}
|
||||||
|
|
||||||
|
// After: Roaring bitmap intersection
|
||||||
|
async getIdsForMultipleFields(pairs: [{field, value}, ...]): Promise<string[]> {
|
||||||
|
// 1. Fetch roaring bitmaps (integers, not UUIDs)
|
||||||
|
const bitmaps: RoaringBitmap32[] = []
|
||||||
|
for (const {field, value} of pairs) {
|
||||||
|
const bitmap = await this.getBitmapFromChunks(field, value)
|
||||||
|
if (!bitmap) return [] // Short-circuit if any field has no matches
|
||||||
|
bitmaps.push(bitmap)
|
||||||
|
}
|
||||||
|
|
||||||
|
// 2. Hardware-accelerated intersection (FAST! AVX2/SSE4.2 SIMD)
|
||||||
|
const result = RoaringBitmap32.and(...bitmaps) // O(1) hardware operation!
|
||||||
|
|
||||||
|
// 3. Convert final bitmap to UUIDs (once, not per-field)
|
||||||
|
return this.idMapper.intsIterableToUuids(result)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance Impact**:
|
||||||
|
- Multi-field intersection: **1.4x average speedup**, up to 3.3x on 10K entities
|
||||||
|
- Memory usage: **90% reduction** (17.17 MB → 2.01 MB for 100K entities)
|
||||||
|
- Hardware acceleration: SIMD instructions make bitmap operations nearly free
|
||||||
|
|
||||||
|
**Benchmark Results** — example output from a single run of `tests/performance/roaring-bitmap-benchmark.ts` (1,000 queries per size, one machine; absolute times vary by hardware, the relative speedup and memory savings are the durable signal):
|
||||||
|
| Dataset Size | Operation | Set Time | Roaring Time | Speedup | Memory Savings |
|
||||||
|
|--------------|-----------|----------|--------------|---------|----------------|
|
||||||
|
| 10,000 entities | 3-field intersection | 3.74ms | 1.14ms | **3.3x faster** | 90% |
|
||||||
|
| 100,000 entities | 3-field intersection | 2.60ms | 1.78ms | **1.5x faster** | 88% |
|
||||||
|
|
||||||
|
**Implementation**: See `src/utils/entityIdMapper.ts` and benchmark at `tests/performance/roaring-bitmap-benchmark.ts`
|
||||||
|
|
||||||
|
#### Bloom Filter (Probabilistic Membership Testing)
|
||||||
|
```typescript
|
||||||
|
class BloomFilter {
|
||||||
|
bits: Uint8Array // Bit array
|
||||||
|
size: number // Total bits
|
||||||
|
hashCount: number // Number of hash functions (FNV-1a, DJB2)
|
||||||
|
|
||||||
|
mightContain(value): boolean // ~1% false positive, 0% false negative
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use case**: Quickly skip chunks that definitely don't contain a value
|
||||||
|
|
||||||
|
#### Zone Map (Range Query Optimization)
|
||||||
|
```typescript
|
||||||
|
interface ZoneMap {
|
||||||
|
min: any | null // Minimum value in chunk
|
||||||
|
max: any | null // Maximum value in chunk
|
||||||
|
count: number // Number of entries
|
||||||
|
hasNulls: boolean // Whether chunk contains null values
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use case**: Skip entire chunks during range queries (ClickHouse-inspired)
|
||||||
|
|
||||||
|
#### Type-Field Affinity
|
||||||
|
```typescript
|
||||||
|
// Tracks which fields are commonly used with which types
|
||||||
|
// Example:
|
||||||
|
// typeFieldAffinity.get('character') → {
|
||||||
|
// 'name': 127, // 127 characters have a 'name' field
|
||||||
|
// 'age': 89, // 89 characters have an 'age' field
|
||||||
|
// 'alignment': 45 // 45 characters have an 'alignment' field
|
||||||
|
// }
|
||||||
|
```
|
||||||
|
|
||||||
|
**Use case**: Enables NLP to understand "find characters named John" → knows 'name' is a character field
|
||||||
|
|
||||||
|
#### Word Index (`__words__`) -
|
||||||
|
```typescript
|
||||||
|
// Special field for text/keyword search
|
||||||
|
// Entity text content is tokenized and indexed as word hashes
|
||||||
|
|
||||||
|
// Tokenization:
|
||||||
|
// "David Smith is a software engineer" → ["david", "smith", "is", "software", "engineer"]
|
||||||
|
|
||||||
|
// Word Hashing (FNV-1a):
|
||||||
|
// "david" → hashWord("david") → 1234567 (int32)
|
||||||
|
// "smith" → hashWord("smith") → 9876543 (int32)
|
||||||
|
|
||||||
|
// Index structure (same as other fields):
|
||||||
|
// __words__ → 1234567 → RoaringBitmap{entity1, entity5, ...}
|
||||||
|
// __words__ → 9876543 → RoaringBitmap{entity1, entity3, ...}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Design Decisions**:
|
||||||
|
- **Max 50 words per entity**: Prevents index bloat for large documents
|
||||||
|
- **FNV-1a hashing**: Fast, low collision rate, int32 output
|
||||||
|
- **Min word length 2 chars**: Filters out noise words
|
||||||
|
- **Lowercase normalization**: Case-insensitive matching
|
||||||
|
- **Automatic integration**: Words extracted via `extractIndexableFields()`
|
||||||
|
|
||||||
|
**Hybrid Search**: Text results combined with vector results using Reciprocal Rank Fusion (RRF):
|
||||||
|
```typescript
|
||||||
|
// RRF formula: score(d) = sum(1 / (k + rank(d)))
|
||||||
|
// where k = 60 (standard constant)
|
||||||
|
// alpha = weight for semantic (0 = text only, 1 = semantic only)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Query Algorithm
|
||||||
|
|
||||||
|
**Exact Match Query**:
|
||||||
|
```typescript
|
||||||
|
async getIds(field: string, value: any): Promise<string[]> {
|
||||||
|
// 1. Load sparse index for field
|
||||||
|
const sparseIndex = await this.loadSparseIndex(field)
|
||||||
|
|
||||||
|
// 2. Find candidate chunks using bloom filters
|
||||||
|
const candidateChunks = sparseIndex.findChunksForValue(value)
|
||||||
|
// → Bloom filter checks all chunks (~1ms)
|
||||||
|
// → Returns only chunks that *might* contain value
|
||||||
|
|
||||||
|
// 3. Load candidate chunks and collect IDs
|
||||||
|
const results = []
|
||||||
|
for (const chunkId of candidateChunks) {
|
||||||
|
const chunk = await this.chunkManager.loadChunk(field, chunkId)
|
||||||
|
const ids = chunk.entries.get(value)
|
||||||
|
if (ids) results.push(...ids)
|
||||||
|
}
|
||||||
|
|
||||||
|
return results
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Range Query**:
|
||||||
|
```typescript
|
||||||
|
async getIdsForRange(field: string, min: any, max: any): Promise<string[]> {
|
||||||
|
// 1. Load sparse index for field
|
||||||
|
const sparseIndex = await this.loadSparseIndex(field)
|
||||||
|
|
||||||
|
// 2. Find candidate chunks using zone maps
|
||||||
|
const candidateChunks = sparseIndex.findChunksForRange(min, max)
|
||||||
|
// → Check zoneMap.min and zoneMap.max for each chunk
|
||||||
|
// → Skip chunks where max < min or min > max
|
||||||
|
|
||||||
|
// 3. Load chunks and filter values
|
||||||
|
const results = []
|
||||||
|
for (const chunkId of candidateChunks) {
|
||||||
|
const chunk = await this.chunkManager.loadChunk(field, chunkId)
|
||||||
|
for (const [value, ids] of chunk.entries) {
|
||||||
|
if (value >= min && value <= max) {
|
||||||
|
results.push(...ids)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return results
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits**:
|
||||||
|
- Bloom filters: Skip 99% of irrelevant chunks (exact match)
|
||||||
|
- Zone maps: Skip entire chunks that fall outside range
|
||||||
|
- Adaptive chunking: ~50 values per chunk optimizes I/O
|
||||||
|
- Immediate flushing: No need for dirty tracking or batch writes
|
||||||
|
|
||||||
|
### Temporal Bucketing
|
||||||
|
|
||||||
|
**Problem Solved**: High-cardinality timestamp fields created massive file pollution.
|
||||||
|
- Example: 575 entities with unique timestamps → 358,407 index files (98.7% pollution!)
|
||||||
|
|
||||||
|
**Solution**: Automatic bucketing of temporal fields to 1-minute intervals.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// In normalizeValue(value, field):
|
||||||
|
if (field && typeof value === 'number') {
|
||||||
|
const fieldLower = field.toLowerCase()
|
||||||
|
const isTemporal = fieldLower.includes('time') ||
|
||||||
|
fieldLower.includes('date') ||
|
||||||
|
fieldLower.includes('accessed') ||
|
||||||
|
fieldLower.includes('modified') ||
|
||||||
|
fieldLower.includes('created') ||
|
||||||
|
fieldLower.includes('updated')
|
||||||
|
|
||||||
|
if (isTemporal) {
|
||||||
|
// Bucket to 1-minute intervals
|
||||||
|
const bucketSize = 60000 // milliseconds
|
||||||
|
const bucketed = Math.floor(value / bucketSize) * bucketSize
|
||||||
|
return bucketed.toString()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits**:
|
||||||
|
- ✅ Reduces 575 unique timestamps → ~10 buckets
|
||||||
|
- ✅ File count: 358,407 → ~4,600 (98.7% reduction)
|
||||||
|
- ✅ Zero configuration - automatic field name detection
|
||||||
|
- ✅ Still enables range queries (not excluded like before)
|
||||||
|
- ✅ 1-minute precision sufficient for most use cases
|
||||||
|
|
||||||
|
**Field Name Detection**: Automatically buckets fields with these keywords:
|
||||||
|
- `time`, `date`, `accessed`, `modified`, `created`, `updated`
|
||||||
|
- Examples: `timestamp`, `createdAt`, `lastModified`, `birthdate`, `eventTime`
|
||||||
|
|
||||||
|
### Operations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Add to index (src/brainy.ts:387)
|
||||||
|
await this.metadataIndex.addToIndex(id, metadata)
|
||||||
|
|
||||||
|
// Query exact match
|
||||||
|
const ids = await this.metadataIndex.getIds('status', 'active')
|
||||||
|
|
||||||
|
// Query range
|
||||||
|
const ids = await this.metadataIndex.getIdsForFilter({
|
||||||
|
publishDate: { greaterThan: 1640995200000 }
|
||||||
|
})
|
||||||
|
|
||||||
|
// Filter discovery (what values exist for a field)
|
||||||
|
const values = await this.metadataIndex.getFilterValues('status')
|
||||||
|
// → ['active', 'archived', 'draft']
|
||||||
|
|
||||||
|
// Statistics (O(1))
|
||||||
|
const totalEntities = this.metadataIndex.getTotalEntityCount()
|
||||||
|
const typeBreakdown = this.metadataIndex.getAllEntityCounts()
|
||||||
|
// → Map { 'character': 127, 'item': 89, 'location': 45 }
|
||||||
|
```
|
||||||
|
|
||||||
|
### Excluded Fields
|
||||||
|
|
||||||
|
Some fields are excluded from indexing to prevent pollution:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const DEFAULT_EXCLUDE_FIELDS = [
|
||||||
|
'id', // Primary key (redundant to index)
|
||||||
|
'uuid', // Alternative primary key
|
||||||
|
'vector', // High-dimensional data
|
||||||
|
'embedding', // Same as vector
|
||||||
|
'content', // Large text content
|
||||||
|
'description', // Large text content
|
||||||
|
'metadata', // Nested object (too large)
|
||||||
|
'data' // Generic nested object
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Note**: Timestamp fields like `modified`, `accessed`, `created` are NO LONGER excluded as of they are indexed with automatic bucketing.
|
||||||
|
|
||||||
|
## 2. Vector Index - Vector Similarity Search
|
||||||
|
|
||||||
|
**Purpose**: O(log n) semantic similarity search using vector embeddings.
|
||||||
|
|
||||||
|
The default JS implementation is `JsHnswVectorIndex`; an optional native acceleration package (`@soulcraft/cor`) can register a higher-performing `VectorIndexProvider` through the plugin system. The public API stays the same either way.
|
||||||
|
|
||||||
|
### Internal Architecture
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class JsHnswVectorIndex {
|
||||||
|
// Per-noun indexes for efficiency
|
||||||
|
private nouns: Map<string, HNSWNoun> = new Map()
|
||||||
|
|
||||||
|
// Global entry point for search
|
||||||
|
private entryPointId: string | null = null
|
||||||
|
private maxLevel = 0
|
||||||
|
|
||||||
|
// Shared memory management
|
||||||
|
private unifiedCache: UnifiedCache
|
||||||
|
private storage: BaseStorage | null = null
|
||||||
|
}
|
||||||
|
|
||||||
|
// Each noun has its own HNSW graph
|
||||||
|
class HNSWNoun {
|
||||||
|
noun: string
|
||||||
|
nodes: Map<string, HNSWNode>
|
||||||
|
entryPointId: string | null
|
||||||
|
maxLevel: number
|
||||||
|
}
|
||||||
|
|
||||||
|
// Each node in the graph
|
||||||
|
class HNSWNode {
|
||||||
|
id: string
|
||||||
|
vector: Vector | null // Lazy-loaded from storage
|
||||||
|
level: number
|
||||||
|
connections: Map<number, string[]> // level → neighbor IDs
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Hierarchical Graph Structure
|
||||||
|
|
||||||
|
The default vector index builds a multi-layered graph:
|
||||||
|
|
||||||
|
```
|
||||||
|
Layer 2: [entry] ←→ [node1] (sparse, long-range connections)
|
||||||
|
↓ ↓
|
||||||
|
Layer 1: [entry] ←→ [node1] ←→ [node2] ←→ [node3] (medium density)
|
||||||
|
↓ ↓ ↓ ↓
|
||||||
|
Layer 0: [entry] ←→ [node1] ←→ [node2] ←→ [node3] ←→ [node4] ←→ [node5] (dense, all nodes)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Search Algorithm**:
|
||||||
|
1. Start at entry point in top layer
|
||||||
|
2. Greedy search for nearest neighbor in current layer
|
||||||
|
3. Move down to next layer with found neighbor
|
||||||
|
4. Repeat until reaching layer 0
|
||||||
|
5. Return k nearest neighbors
|
||||||
|
|
||||||
|
**Complexity**: O(log n) due to hierarchical structure
|
||||||
|
|
||||||
|
### Adaptive Vector Loading
|
||||||
|
|
||||||
|
Vectors are lazy-loaded on demand based on memory availability:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
private async getVectorSafe(noun: HNSWNoun): Promise<Vector> {
|
||||||
|
// Check UnifiedCache first
|
||||||
|
const cached = this.unifiedCache.get(noun.id)
|
||||||
|
if (cached) return cached
|
||||||
|
|
||||||
|
// Load from storage if memory available
|
||||||
|
if (this.unifiedCache.canCache()) {
|
||||||
|
const vector = await this.storage.loadVector(noun.id)
|
||||||
|
this.unifiedCache.set(noun.id, vector)
|
||||||
|
return vector
|
||||||
|
}
|
||||||
|
|
||||||
|
// Load transiently if memory pressure
|
||||||
|
return await this.storage.loadVector(noun.id)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Operations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Add entity (src/brainy.ts:add)
|
||||||
|
await this.index.addEntity(id, vector, noun)
|
||||||
|
|
||||||
|
// Search for similar vectors
|
||||||
|
const results = await this.index.search(queryVector, k, threshold)
|
||||||
|
// Returns: Array<{id: string, similarity: number}>
|
||||||
|
|
||||||
|
// Rebuild from storage
|
||||||
|
await this.index.rebuild()
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. GraphAdjacencyIndex - O(1) Relationship Traversal
|
||||||
|
|
||||||
|
**Purpose**: Constant-time neighbor lookups regardless of graph size.
|
||||||
|
|
||||||
|
### Internal Architecture
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class GraphAdjacencyIndex {
|
||||||
|
// O(1) bidirectional lookups
|
||||||
|
private sourceIndex = new Map<string, Set<string>>() // sourceId → targetIds
|
||||||
|
private targetIndex = new Map<string, Set<string>>() // targetId → sourceIds
|
||||||
|
|
||||||
|
// Full relationship data
|
||||||
|
private verbIndex = new Map<string, GraphVerb>() // verbId → metadata
|
||||||
|
|
||||||
|
// Statistics
|
||||||
|
private relationshipCountsByType = new Map<string, number>()
|
||||||
|
|
||||||
|
// Shared memory
|
||||||
|
private unifiedCache: UnifiedCache
|
||||||
|
private storage: BaseStorage
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key Innovation: Bidirectional Adjacency
|
||||||
|
|
||||||
|
**Core Insight**: Store BOTH directions of each relationship for O(1) lookups.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Example: Alice KNOWS Bob
|
||||||
|
// verbId = "verb-123"
|
||||||
|
|
||||||
|
// Source index: Alice → Bob
|
||||||
|
sourceIndex.set('alice', Set(['bob']))
|
||||||
|
|
||||||
|
// Target index: Bob ← Alice
|
||||||
|
targetIndex.set('bob', Set(['alice']))
|
||||||
|
|
||||||
|
// Full metadata
|
||||||
|
verbIndex.set('verb-123', {
|
||||||
|
id: 'verb-123',
|
||||||
|
verb: 'knows',
|
||||||
|
source: 'alice',
|
||||||
|
target: 'bob',
|
||||||
|
metadata: { since: 2020 }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**Result**: Finding Alice's friends OR Bob's friends is O(1) - just one Map lookup!
|
||||||
|
|
||||||
|
### Operations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Add relationship (src/brainy.ts:relate)
|
||||||
|
await this.graphIndex.addRelationship(verbId, sourceId, targetId, verb)
|
||||||
|
|
||||||
|
// Get neighbors (O(1) per hop)
|
||||||
|
const outgoing = await this.graphIndex.getNeighbors(id, 'out') // Who does id point to?
|
||||||
|
const incoming = await this.graphIndex.getNeighbors(id, 'in') // Who points to id?
|
||||||
|
const both = await this.graphIndex.getNeighbors(id, 'both') // All neighbors
|
||||||
|
|
||||||
|
// Get relationships
|
||||||
|
const verbs = await this.graphIndex.getRelationships(sourceId, targetId)
|
||||||
|
|
||||||
|
// Statistics (O(1))
|
||||||
|
const totalRelationships = this.graphIndex.getTotalRelationshipCount()
|
||||||
|
const byType = this.graphIndex.getRelationshipCountsByType()
|
||||||
|
// → Map { 'knows': 45, 'created': 23, 'located_at': 12 }
|
||||||
|
```
|
||||||
|
|
||||||
|
### Graph Traversal
|
||||||
|
|
||||||
|
The index supports multi-hop traversal:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Find all entities within 2 hops
|
||||||
|
const reachable = await this.graphIndex.traverse({
|
||||||
|
startId: 'alice',
|
||||||
|
depth: 2,
|
||||||
|
direction: 'out'
|
||||||
|
})
|
||||||
|
// Complexity: O(V + E) breadth-first search, but each neighbor lookup is O(1)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Shared Memory Management: UnifiedCache
|
||||||
|
|
||||||
|
All three main indexes share a single **UnifiedCache** instance for coordinated memory management.
|
||||||
|
|
||||||
|
### Architecture
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class UnifiedCache {
|
||||||
|
private cache: Map<string, CachedItem> = new Map()
|
||||||
|
private maxSize: number
|
||||||
|
private currentSize: number = 0
|
||||||
|
private evictionPolicy: 'LRU' | 'LFU' = 'LRU'
|
||||||
|
}
|
||||||
|
|
||||||
|
// Each index gets the same cache instance
|
||||||
|
const unifiedCache = new UnifiedCache({ maxSize: 1000 })
|
||||||
|
this.metadataIndex = new MetadataIndexManager(storage, { unifiedCache })
|
||||||
|
this.vectorIndex = new JsHnswVectorIndex(storage, { unifiedCache })
|
||||||
|
this.graphIndex = new GraphAdjacencyIndex(storage, { unifiedCache })
|
||||||
|
```
|
||||||
|
|
||||||
|
### Benefits
|
||||||
|
|
||||||
|
1. **Fair Resource Allocation**: All indexes compete for the same memory pool
|
||||||
|
2. **Prevents Monopolization**: No single index can starve others of memory
|
||||||
|
3. **Coordinated Eviction**: LRU eviction across all cached items system-wide
|
||||||
|
4. **Memory Pressure Handling**: Automatic cache shrinking when memory is tight
|
||||||
|
5. **Adaptive Loading**: Indexes load data transiently under memory pressure
|
||||||
|
|
||||||
|
### Cache Key Patterns
|
||||||
|
|
||||||
|
Each index uses different key prefixes:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Metadata index
|
||||||
|
cache.set(`meta:${field}:${value}`, indexEntry)
|
||||||
|
|
||||||
|
// Vector index
|
||||||
|
cache.set(`vector:${id}`, vectorData)
|
||||||
|
|
||||||
|
// Graph index
|
||||||
|
cache.set(`graph:${sourceId}`, neighbors)
|
||||||
|
|
||||||
|
// Deleted items (no caching needed - uses Set)
|
||||||
|
```
|
||||||
|
|
||||||
|
## How Indexes Work Together
|
||||||
|
|
||||||
|
### 1. Entity Creation (`brainy.add()`)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts:add()
|
||||||
|
async add(params: AddParams): Promise<string> {
|
||||||
|
const id = generateId()
|
||||||
|
const vector = await this.embedder(params.content)
|
||||||
|
|
||||||
|
// Add to metadata index (field filtering)
|
||||||
|
await this.metadataIndex.addToIndex(id, params.metadata)
|
||||||
|
|
||||||
|
// Add to vector index (vector search)
|
||||||
|
await this.index.addEntity(id, vector, params.noun)
|
||||||
|
|
||||||
|
// Relationships added via separate relate() calls
|
||||||
|
|
||||||
|
return id
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Entity Search (`brainy.find()`)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts:find()
|
||||||
|
async find(query: FindQuery): Promise<Result[]> {
|
||||||
|
let results: Result[] = []
|
||||||
|
|
||||||
|
// Step 1: Metadata filtering (fast pre-filter)
|
||||||
|
if (query.where) {
|
||||||
|
const filteredIds = await this.metadataIndex.getIdsForFilter(query.where)
|
||||||
|
results = await this.getEntitiesByIds(filteredIds)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Step 2: Vector similarity search (semantic ranking)
|
||||||
|
if (query.like) {
|
||||||
|
const queryVector = await this.embedder(query.like)
|
||||||
|
const vectorResults = await this.index.search(queryVector, query.limit)
|
||||||
|
|
||||||
|
// Intersect or union with metadata results
|
||||||
|
results = this.combineResults(results, vectorResults)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Step 3: Graph traversal (relationship filtering)
|
||||||
|
if (query.connected) {
|
||||||
|
const connectedIds = await this.graphIndex.traverse(query.connected)
|
||||||
|
results = results.filter(r => connectedIds.includes(r.id))
|
||||||
|
}
|
||||||
|
|
||||||
|
return results
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Entity Update (`brainy.update()`)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts:update()
|
||||||
|
async update(params: UpdateParams): Promise<void> {
|
||||||
|
const existing = await this.get(params.id)
|
||||||
|
|
||||||
|
// Update metadata index (remove old, add new)
|
||||||
|
await this.metadataIndex.removeFromIndex(params.id, existing.metadata)
|
||||||
|
await this.metadataIndex.addToIndex(params.id, params.metadata)
|
||||||
|
|
||||||
|
// Update vector index (re-embed if content changed)
|
||||||
|
if (params.content) {
|
||||||
|
const newVector = await this.embedder(params.content)
|
||||||
|
await this.index.updateEntity(params.id, newVector)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Graph relationships unchanged (managed separately)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Statistics (`brainy.stats()`)
|
||||||
|
|
||||||
|
All indexes provide O(1) statistics:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts:stats()
|
||||||
|
async stats(): Promise<Statistics> {
|
||||||
|
return {
|
||||||
|
// From metadata index
|
||||||
|
entities: this.metadataIndex.getTotalEntityCount(),
|
||||||
|
entityTypes: this.metadataIndex.getAllEntityCounts(),
|
||||||
|
|
||||||
|
// From graph index
|
||||||
|
relationships: this.graphIndex.getTotalRelationshipCount(),
|
||||||
|
relationshipTypes: this.graphIndex.getRelationshipCountsByType(),
|
||||||
|
|
||||||
|
// From vector index
|
||||||
|
vectorIndexSize: this.index.getSize()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5. Index Rebuilding (Lazy Loading Support)
|
||||||
|
|
||||||
|
**Two modes of index loading:**
|
||||||
|
|
||||||
|
#### Mode 1: Auto-Rebuild on init() (default)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts:init()
|
||||||
|
async init(): Promise<void> {
|
||||||
|
// When disableAutoRebuild: false (default)
|
||||||
|
const metadataStats = await this.metadataIndex.getStats()
|
||||||
|
const vectorIndexSize = this.index.size()
|
||||||
|
const graphIndexSize = await this.graphIndex.size()
|
||||||
|
|
||||||
|
if (metadataStats.totalEntries === 0 ||
|
||||||
|
vectorIndexSize === 0 ||
|
||||||
|
graphIndexSize === 0) {
|
||||||
|
// Rebuild all indexes in parallel
|
||||||
|
await Promise.all([
|
||||||
|
metadataStats.totalEntries === 0 ? this.metadataIndex.rebuild() : Promise.resolve(),
|
||||||
|
vectorIndexSize === 0 ? this.index.rebuild() : Promise.resolve(),
|
||||||
|
graphIndexSize === 0 ? this.graphIndex.rebuild() : Promise.resolve()
|
||||||
|
])
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Mode 2: Lazy Loading on First Query
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// When disableAutoRebuild: true
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem' },
|
||||||
|
disableAutoRebuild: true // Enable lazy loading
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.init() // Returns instantly, indexes empty
|
||||||
|
|
||||||
|
// First query triggers lazy rebuild
|
||||||
|
const results = await brain.find({ limit: 10 })
|
||||||
|
// → Calls ensureIndexesLoaded() (line 4617)
|
||||||
|
// → Rebuilds all 3 main indexes with concurrency control
|
||||||
|
// → Subsequent queries are instant (0ms check)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance:**
|
||||||
|
- First query with lazy loading: ~50-200ms rebuild (1K-10K entities)
|
||||||
|
- Concurrent queries: Wait for same rebuild (mutex prevents duplicates)
|
||||||
|
- Subsequent queries: 0ms check (instant)
|
||||||
|
|
||||||
|
See [initialization-and-rebuild.md](./initialization-and-rebuild.md) for detailed lazy loading implementation.
|
||||||
|
|
||||||
|
## Triple Intelligence Integration
|
||||||
|
|
||||||
|
The **TripleIntelligenceSystem** (`src/triple/TripleIntelligenceSystem.ts`) combines all three core indexes:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
class TripleIntelligenceSystem {
|
||||||
|
constructor(
|
||||||
|
private metadataIndex: MetadataIndexManager,
|
||||||
|
private vectorIndex: VectorIndexProvider,
|
||||||
|
private graphIndex: GraphAdjacencyIndex,
|
||||||
|
private embedder: EmbedderFunction,
|
||||||
|
private storage: BaseStorage
|
||||||
|
) {}
|
||||||
|
|
||||||
|
async query(nlpQuery: string): Promise<Result[]> {
|
||||||
|
// Parse natural language
|
||||||
|
const parsed = await this.parseQuery(nlpQuery)
|
||||||
|
|
||||||
|
// Execute across all three indexes
|
||||||
|
const [metadataResults, vectorResults, graphResults] = await Promise.all([
|
||||||
|
this.metadataIndex.getIdsForFilter(parsed.filters),
|
||||||
|
this.vectorIndex.search(parsed.vector, parsed.limit),
|
||||||
|
this.graphIndex.traverse(parsed.graphConstraints)
|
||||||
|
])
|
||||||
|
|
||||||
|
// Fuse results with weighted scoring
|
||||||
|
return this.fuseResults(metadataResults, vectorResults, graphResults)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Performance Characteristics
|
||||||
|
|
||||||
|
### Operation Complexity by Index
|
||||||
|
|
||||||
|
| Operation | MetadataIndexManager | TypeAwareVectorIndex | GraphAdjacencyIndex |
|
||||||
|
|-----------|---------------------|-------------------|---------------------|
|
||||||
|
| **Add** | O(1) per field | O(log n) | O(1) |
|
||||||
|
| **Remove** | O(1) per field | O(log n) | O(1) |
|
||||||
|
| **Exact lookup** | O(1) | N/A | O(1) |
|
||||||
|
| **Range query** | O(log n) + O(k) | N/A | N/A |
|
||||||
|
| **Similarity search** | N/A | O(log n) | N/A |
|
||||||
|
| **Neighbor lookup** | N/A | N/A | O(1) |
|
||||||
|
| **Statistics** | O(1) | O(1) | O(1) |
|
||||||
|
| **Rebuild** | O(n) | O(n) | O(n) |
|
||||||
|
|
||||||
|
Where:
|
||||||
|
- n = total number of entities
|
||||||
|
- k = number of matching results
|
||||||
|
|
||||||
|
**Note**: All 3 main indexes have rebuild() methods that load persisted data (O(n)) rather than recomputing (which would be O(n log n) for the vector index).
|
||||||
|
|
||||||
|
### Memory Footprint
|
||||||
|
|
||||||
|
| Index | Per-Entity Memory | Notes |
|
||||||
|
|-------|-------------------|-------|
|
||||||
|
| **MetadataIndexManager** | ~100 bytes | Depends on field count and cardinality (RoaringBitmap32 compression) |
|
||||||
|
| **TypeAwareVectorIndex** | ~1.5 KB | Vector (384 dims × 4 bytes) + graph connections across 42 type-specific indexes |
|
||||||
|
| **GraphAdjacencyIndex** | ~50 bytes per relationship | Bidirectional verb-id references in 2 LSM-trees |
|
||||||
|
|
||||||
|
**Total overhead**: ~1.6 KB per entity + ~50 bytes per relationship
|
||||||
|
|
||||||
|
**Sub-index memory:**
|
||||||
|
- ChunkManager: ~20 bytes per chunk descriptor
|
||||||
|
- EntityIdMapper: ~32 bytes per UUID mapping (50-90% savings vs Set\<string\>)
|
||||||
|
- LSM-trees: ~200 bytes per relationship (SSTable storage)
|
||||||
|
|
||||||
|
### Scalability
|
||||||
|
|
||||||
|
All indexes scale gracefully. The cost of each stage is governed by its algorithmic complexity, not a fixed millisecond figure — absolute latency depends on hardware, embedding model, and storage backend. Only the graph adjacency index carries a committed scale assertion:
|
||||||
|
|
||||||
|
| Query stage | Complexity | Scaling behavior |
|
||||||
|
|-------------|------------|------------------|
|
||||||
|
| Metadata filter (exact) | O(1) | Constant — independent of dataset size |
|
||||||
|
| Metadata filter (range) | O(log n) + O(k) | Sub-linear; k = matching results |
|
||||||
|
| Vector search (HNSW) | O(log n) | Degrades gracefully via hierarchical layers |
|
||||||
|
| Graph hop | O(1) | Measured <1 ms per neighbor lookup, validated up to 1M relationships (`tests/performance/graph-scale-performance.test.ts:238`) |
|
||||||
|
| Combined query | O(log n) | Bounded by the vector stage; metadata and graph stages stay O(1)/O(log n) |
|
||||||
|
|
||||||
|
**Key observations**:
|
||||||
|
- Graph queries stay O(1) regardless of scale
|
||||||
|
- Metadata filtering scales sub-linearly
|
||||||
|
- Vector search degrades gracefully due to the hierarchical index
|
||||||
|
- Combined queries remain fast even at scale
|
||||||
|
|
||||||
|
## Best Practices
|
||||||
|
|
||||||
|
### When to Use Each Index
|
||||||
|
|
||||||
|
**MetadataIndex**:
|
||||||
|
- Filtering by exact field values (status, type, category)
|
||||||
|
- Range queries on numeric/temporal fields (dates, prices, counts)
|
||||||
|
- Field discovery (what filters are available)
|
||||||
|
- Type-based querying (find all characters, all items)
|
||||||
|
|
||||||
|
**Vector Index**:
|
||||||
|
- Semantic similarity search ("find similar documents")
|
||||||
|
- Content-based retrieval ("find posts about AI")
|
||||||
|
- Fuzzy matching (when exact matches aren't required)
|
||||||
|
- Recommendation systems (find related items)
|
||||||
|
|
||||||
|
**GraphAdjacencyIndex**:
|
||||||
|
- Relationship queries ("who knows whom")
|
||||||
|
- Path finding ("how are these entities connected")
|
||||||
|
- Network analysis ("find communities")
|
||||||
|
- Multi-hop traversal ("friends of friends")
|
||||||
|
|
||||||
|
**Note**: Soft-delete functionality is not currently integrated. Brainy uses hard deletes via storage layer.
|
||||||
|
|
||||||
|
### Query Optimization
|
||||||
|
|
||||||
|
1. **Start with metadata filters** - They're fastest and most selective
|
||||||
|
2. **Use graph constraints** - O(1) lookups significantly reduce search space
|
||||||
|
3. **Vector search last** - Most expensive, best used on pre-filtered set
|
||||||
|
4. **Leverage temporal bucketing** - Timestamp range queries work efficiently
|
||||||
|
5. **Monitor statistics** - Use O(1) stats methods for cardinality estimation
|
||||||
|
|
||||||
|
### Memory Management
|
||||||
|
|
||||||
|
1. **Configure UnifiedCache appropriately** - Balance between speed and memory
|
||||||
|
2. **Use lazy loading** - Vector index loads vectors on-demand
|
||||||
|
3. **Monitor cache hit rates** - Adjust cache size if hit rate is low
|
||||||
|
4. **Consider storage adapter** - Memory = fastest, filesystem = persistent
|
||||||
|
|
||||||
|
## Related Documentation
|
||||||
|
|
||||||
|
- [Find System](../FIND_SYSTEM.md) - Query-centric view of index usage
|
||||||
|
- [Triple Intelligence](./triple-intelligence.md) - Advanced query system
|
||||||
|
- [Storage Architecture](./storage-architecture.md) - Storage layer details
|
||||||
|
- [Performance Guide](../PERFORMANCE.md) - Performance tuning
|
||||||
|
- [Overview](./overview.md) - High-level architecture
|
||||||
|
|
||||||
|
## Summary: Index Hierarchy
|
||||||
|
|
||||||
|
### Level 1: Main Indexes (3)
|
||||||
|
All have rebuild() methods and are covered by lazy loading:
|
||||||
|
1. **TypeAwareVectorIndex** - `src/hnsw/typeAwareHNSWIndex.ts:403`
|
||||||
|
2. **MetadataIndexManager** - `src/utils/metadataIndex.ts:2318`
|
||||||
|
3. **GraphAdjacencyIndex** - `src/graph/graphAdjacencyIndex.ts:389`
|
||||||
|
|
||||||
|
### Level 2: Sub-Indexes (~50+)
|
||||||
|
Automatically managed by parent rebuild():
|
||||||
|
- **42 type-specific vector indexes** (one per NounType)
|
||||||
|
- **6 metadata components** (ChunkManager, EntityIdMapper, FieldTypeInference, Field Sparse Indexes, Sorted Indexes)
|
||||||
|
- **2 LSM-trees** (lsmTreeVerbsBySource, lsmTreeVerbsByTarget — the verb set is the single adjacency source of truth; neighbor reads derive from live verbs so removals are honored)
|
||||||
|
- **In-memory graph structures** (sourceIndex, targetIndex, verbIndex)
|
||||||
|
|
||||||
|
### Lazy Loading
|
||||||
|
- **Mode 1**: Auto-rebuild on init() (default)
|
||||||
|
- **Mode 2**: Lazy rebuild on first query (when `disableAutoRebuild: true`)
|
||||||
|
- **Concurrency-safe**: Mutex prevents duplicate rebuilds
|
||||||
|
- **Performance**: First query ~50-200ms, subsequent queries instant
|
||||||
|
|
||||||
|
### Total Functional Index Count
|
||||||
|
- **3 main indexes** with independent rebuild() methods
|
||||||
|
- **~50+ sub-components** managed automatically
|
||||||
|
- **All covered** by rebuildIndexesIfNeeded() or built-in lazy initialization
|
||||||
|
|
||||||
|
## Version History
|
||||||
|
|
||||||
|
- **v5.7.7** (November 2025): Added production-scale lazy loading with concurrency control. Fixed critical bug where `disableAutoRebuild: true` left indexes empty forever. Added `ensureIndexesLoaded()` helper and `getIndexStatus()` diagnostic.
|
||||||
|
- **v3.43.0** (October 2025): Migrated from `roaring` (native C++) to `roaring-wasm` (WebAssembly) for universal compatibility. No API changes - maintains identical RoaringBitmap32 interface. Benefits: works in all environments (Node.js, browsers, serverless) without build tools, zero compilation errors, simpler developer experience. 90% memory savings and hardware-accelerated operations unchanged.
|
||||||
|
- **v3.42.0** (October 2025): Replaced flat file indexing with adaptive chunked sparse indexing. Bloom filters + zone maps for O(1) exact match and O(log n) range queries. 630x file reduction (560k → 89 files). Removed dual code paths.
|
||||||
|
- **v3.41.0** (October 2025): Added automatic temporal bucketing to MetadataIndex
|
||||||
|
- **v3.40.0** (October 2025): Enhanced batch processing for imports
|
||||||
|
- **v3.0.0** (September 2025): Introduced 3-tier index architecture with UnifiedCache
|
||||||
713
docs/architecture/initialization-and-rebuild.md
Normal file
713
docs/architecture/initialization-and-rebuild.md
Normal file
|
|
@ -0,0 +1,713 @@
|
||||||
|
# Initialization and Rebuild Processes
|
||||||
|
|
||||||
|
This document explains how Brainy's four indexes (MetadataIndex, vector index, GraphAdjacencyIndex, DeletedItemsIndex) initialize and rebuild from persisted storage.
|
||||||
|
|
||||||
|
## Core Principle: All Indexes Are Disk-Based
|
||||||
|
|
||||||
|
**KEY INSIGHT**: All indexes in Brainy are already disk-based. There is no need for snapshots or separate backup mechanisms. Initialization simply loads the right amount of data from storage into memory based on available resources.
|
||||||
|
|
||||||
|
### What Gets Persisted
|
||||||
|
|
||||||
|
| Index | Persisted Data | Storage Method | Since Version |
|
||||||
|
|-------|---------------|----------------|---------------|
|
||||||
|
| **MetadataIndex** | Field registry + chunked sparse indices with bloom filters + zone maps | `storage.saveMetadata()` | v3.42.0 (chunks), v4.2.1 (registry) |
|
||||||
|
| **Vector Index** | Vector embeddings + graph connections | `storage.saveHNSWData()` + `storage.saveHNSWSystem()` | v3.35.0 |
|
||||||
|
| **GraphAdjacencyIndex** | Relationships via LSM-tree SSTables | LSM-tree auto-persistence | v3.44.0 |
|
||||||
|
| **DeletedItemsIndex** | Set of deleted IDs | `storage.saveDeletedItems()` | v3.0.0 |
|
||||||
|
|
||||||
|
#### MetadataIndex Persistence Details
|
||||||
|
|
||||||
|
The MetadataIndex now persists two components:
|
||||||
|
|
||||||
|
1. **Field Registry** (`__metadata_field_registry__`): Directory of indexed fields for O(1) discovery
|
||||||
|
- Size: ~4-8KB (50-200 fields typical)
|
||||||
|
- Enables instant cold starts by discovering persisted indices
|
||||||
|
- Auto-saved during every flush operation
|
||||||
|
|
||||||
|
2. **Sparse Indices** (`__sparse_index__<field>`): Per-field index directories
|
||||||
|
- Contains chunk metadata, zone maps, and bloom filters
|
||||||
|
- Lazy-loaded via UnifiedCache on first query
|
||||||
|
|
||||||
|
3. **Chunks** (`__metadata_chunk__<field>_<chunkId>`): Actual inverted index data
|
||||||
|
- Roaring bitmaps for compressed entity ID storage
|
||||||
|
- Loaded on-demand based on query patterns
|
||||||
|
|
||||||
|
All storage operations use the **StorageAdapter** interface, which works with FileSystem and Memory backends.
|
||||||
|
|
||||||
|
## Initialization Process
|
||||||
|
|
||||||
|
### 1. Lazy Initialization Pattern
|
||||||
|
|
||||||
|
All indexes use lazy initialization - they don't load data until first use:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Example: GraphAdjacencyIndex
|
||||||
|
class GraphAdjacencyIndex {
|
||||||
|
private initialized = false
|
||||||
|
|
||||||
|
private async ensureInitialized(): Promise<void> {
|
||||||
|
if (this.initialized) return
|
||||||
|
|
||||||
|
// Initialize LSM-trees from storage
|
||||||
|
await this.lsmTreeSource.init()
|
||||||
|
await this.lsmTreeTarget.init()
|
||||||
|
|
||||||
|
this.initialized = true
|
||||||
|
}
|
||||||
|
|
||||||
|
// Every public method calls ensureInitialized() first
|
||||||
|
async getNeighbors(id: string): Promise<string[]> {
|
||||||
|
await this.ensureInitialized() // Lazy init!
|
||||||
|
// ... actual logic
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits**:
|
||||||
|
- Zero-cost abstraction: No initialization overhead if index not used
|
||||||
|
- Faster startup: Indexes initialize in parallel on first use
|
||||||
|
- Lower memory: Only used indexes consume memory
|
||||||
|
|
||||||
|
### 2. Brain Initialization Flow
|
||||||
|
|
||||||
|
When you create a `Brain` instance and call `init()`, behavior depends on the `disableAutoRebuild` configuration:
|
||||||
|
|
||||||
|
#### Mode 1: Auto-Rebuild on init() (Default)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/brainy.ts (lines 192-237)
|
||||||
|
async init(): Promise<void> {
|
||||||
|
const initStartTime = Date.now()
|
||||||
|
|
||||||
|
// STEP 1: Initialize storage and unified cache
|
||||||
|
await this.storage.init()
|
||||||
|
|
||||||
|
// STEP 2: Check index sizes (lazy initialization triggers here)
|
||||||
|
const metadataStats = await this.metadataIndex.getStats()
|
||||||
|
const vectorIndexSize = this.index.size()
|
||||||
|
const graphIndexSize = await this.graphIndex.size()
|
||||||
|
|
||||||
|
// STEP 3: Rebuild empty indexes from storage in parallel
|
||||||
|
if (metadataStats.totalEntries === 0 ||
|
||||||
|
vectorIndexSize === 0 ||
|
||||||
|
graphIndexSize === 0) {
|
||||||
|
|
||||||
|
const rebuildStartTime = Date.now()
|
||||||
|
await Promise.all([
|
||||||
|
metadataStats.totalEntries === 0
|
||||||
|
? this.metadataIndex.rebuild()
|
||||||
|
: Promise.resolve(),
|
||||||
|
vectorIndexSize === 0
|
||||||
|
? this.index.rebuild()
|
||||||
|
: Promise.resolve(),
|
||||||
|
graphIndexSize === 0
|
||||||
|
? this.graphIndex.rebuild()
|
||||||
|
: Promise.resolve()
|
||||||
|
])
|
||||||
|
|
||||||
|
const rebuildDuration = Date.now() - rebuildStartTime
|
||||||
|
console.log(`✅ All indexes rebuilt in ${rebuildDuration}ms`)
|
||||||
|
}
|
||||||
|
|
||||||
|
// STEP 4: Log statistics
|
||||||
|
const stats = await this.stats()
|
||||||
|
console.log(`📊 Brain initialized with ${stats.entities} entities`)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Timeline** (typical cold start with 10K entities):
|
||||||
|
- 0-50ms: Storage adapter initialization
|
||||||
|
- 50-100ms: Field registry loading (O(1) discovery of persisted indices)
|
||||||
|
- 100-200ms: Index lazy initialization (LSM-tree loading)
|
||||||
|
- 200-500ms: Cache warming (preload common fields)
|
||||||
|
- **No rebuild needed!** Registry discovers existing indices
|
||||||
|
- Total: ~0.5-1 second (instant cold starts)
|
||||||
|
|
||||||
|
**Timeline** (cold start WITHOUT field registry - first run only):
|
||||||
|
- 0-50ms: Storage adapter initialization
|
||||||
|
- 50-100ms: Index lazy initialization
|
||||||
|
- 100-2000ms: One-time rebuild to create indices
|
||||||
|
- Total: ~1-3 seconds (one time only)
|
||||||
|
|
||||||
|
#### Mode 2: Lazy Loading on First Query
|
||||||
|
|
||||||
|
When `disableAutoRebuild: true`, indexes remain empty after init() and rebuild on first query:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// User code
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem' },
|
||||||
|
disableAutoRebuild: true // Enable lazy loading
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.init() // Returns instantly (0-10ms)
|
||||||
|
|
||||||
|
// First query triggers lazy rebuild
|
||||||
|
const results = await brain.find({ limit: 10 })
|
||||||
|
// → Calls ensureIndexesLoaded() internally (brainy.ts:4617)
|
||||||
|
// → Rebuilds all 3 main indexes with concurrency control
|
||||||
|
// → Returns results (~50-200ms total for 1K-10K entities)
|
||||||
|
|
||||||
|
// Subsequent queries are instant
|
||||||
|
const more = await brain.find({ limit: 100 }) // 0ms check, instant
|
||||||
|
```
|
||||||
|
|
||||||
|
**ensureIndexesLoaded() Implementation** (brainy.ts:4617-4664):
|
||||||
|
```typescript
|
||||||
|
private async ensureIndexesLoaded(): Promise<void> {
|
||||||
|
// Fast path: Already loaded
|
||||||
|
if (this.lazyRebuildCompleted) {
|
||||||
|
return // 0ms
|
||||||
|
}
|
||||||
|
|
||||||
|
// Concurrency control: Wait for in-progress rebuild
|
||||||
|
if (this.lazyRebuildInProgress && this.lazyRebuildPromise) {
|
||||||
|
await this.lazyRebuildPromise // Wait for same rebuild
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// Check if storage has data
|
||||||
|
const entities = await this.storage.getNouns({ pagination: { limit: 1 } })
|
||||||
|
const hasData = (entities.totalCount && entities.totalCount > 0) || entities.items.length > 0
|
||||||
|
|
||||||
|
if (!hasData) {
|
||||||
|
this.lazyRebuildCompleted = true
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// Start lazy rebuild with mutex
|
||||||
|
this.lazyRebuildInProgress = true
|
||||||
|
this.lazyRebuildPromise = this.rebuildIndexesIfNeeded(true)
|
||||||
|
.then(() => {
|
||||||
|
this.lazyRebuildCompleted = true
|
||||||
|
})
|
||||||
|
.finally(() => {
|
||||||
|
this.lazyRebuildInProgress = false
|
||||||
|
this.lazyRebuildPromise = null
|
||||||
|
})
|
||||||
|
|
||||||
|
await this.lazyRebuildPromise
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Lazy Loading Performance:**
|
||||||
|
- First query: ~50-200ms (1K-10K entities) - triggers rebuild
|
||||||
|
- Concurrent queries: Wait for same rebuild (mutex prevents duplicates)
|
||||||
|
- Subsequent queries: 0ms check (instant)
|
||||||
|
- Zero-config: Works automatically, no code changes needed
|
||||||
|
|
||||||
|
**Use Cases for Lazy Loading:**
|
||||||
|
- **Serverless/Edge**: Minimize cold start time, indexes load on demand
|
||||||
|
- **Development**: Faster restarts during development
|
||||||
|
- **Large datasets**: Defer index loading until actually needed
|
||||||
|
- **Read-heavy workloads**: Write operations don't wait for index rebuild
|
||||||
|
|
||||||
|
## Rebuild Process
|
||||||
|
|
||||||
|
### What "Rebuild" Actually Means
|
||||||
|
|
||||||
|
**IMPORTANT**: "Rebuild" does NOT mean recomputing data. It means:
|
||||||
|
1. **Load persisted data** from storage (vector index connections, metadata chunks, LSM-tree SSTables)
|
||||||
|
2. **Populate in-memory structures** (Maps, Sets, graphs)
|
||||||
|
3. **Apply adaptive caching** (preload vectors if small dataset, lazy load if large)
|
||||||
|
|
||||||
|
**Complexity**: O(N) - linear scan through storage, NOT O(N log N) recomputation!
|
||||||
|
|
||||||
|
### 1. Vector Index Rebuild (Correct Pattern)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/hnsw/hnswIndex.ts (lines 809-947)
|
||||||
|
public async rebuild(options: {
|
||||||
|
lazy?: boolean
|
||||||
|
batchSize?: number
|
||||||
|
onProgress?: (loaded: number, total: number) => void
|
||||||
|
} = {}): Promise<void> {
|
||||||
|
// STEP 1: Clear in-memory structures
|
||||||
|
this.clear()
|
||||||
|
|
||||||
|
// STEP 2: Load system data (entry point, max level)
|
||||||
|
const systemData = await this.storage.getHNSWSystem()
|
||||||
|
this.entryPointId = systemData.entryPointId
|
||||||
|
this.maxLevel = systemData.maxLevel
|
||||||
|
|
||||||
|
// STEP 3: Determine preloading strategy (adaptive caching)
|
||||||
|
const totalNouns = await this.storage.getNounCount()
|
||||||
|
const vectorMemory = totalNouns * 384 * 4 // 384 dims × 4 bytes
|
||||||
|
const availableCache = this.unifiedCache.getRemainingCapacity()
|
||||||
|
const shouldPreload = vectorMemory < availableCache * 0.3
|
||||||
|
|
||||||
|
// STEP 4: Load entities with persisted vector index connections
|
||||||
|
let hasMore = true
|
||||||
|
let cursor: string | undefined = undefined
|
||||||
|
|
||||||
|
while (hasMore) {
|
||||||
|
const result = await this.storage.getNouns({
|
||||||
|
pagination: { limit: 1000, cursor }
|
||||||
|
})
|
||||||
|
|
||||||
|
for (const nounData of result.items) {
|
||||||
|
// Load vector graph data from storage (NOT recomputed!)
|
||||||
|
const hnswData = await this.storage.getHNSWData(nounData.id)
|
||||||
|
|
||||||
|
// Create noun with restored connections
|
||||||
|
const noun: HNSWNoun = {
|
||||||
|
id: nounData.id,
|
||||||
|
vector: shouldPreload ? nounData.vector : [], // Adaptive!
|
||||||
|
connections: new Map(),
|
||||||
|
level: hnswData.level
|
||||||
|
}
|
||||||
|
|
||||||
|
// Restore connections from persisted data
|
||||||
|
for (const [levelStr, nounIds] of Object.entries(hnswData.connections)) {
|
||||||
|
const level = parseInt(levelStr, 10)
|
||||||
|
noun.connections.set(level, new Set<string>(nounIds))
|
||||||
|
}
|
||||||
|
|
||||||
|
// Just add to memory (no recomputation!)
|
||||||
|
this.nouns.set(nounData.id, noun)
|
||||||
|
}
|
||||||
|
|
||||||
|
hasMore = result.hasMore
|
||||||
|
cursor = result.nextCursor
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key Points**:
|
||||||
|
- ✅ Loads vector index connections from storage via `getHNSWData()`
|
||||||
|
- ✅ Uses adaptive caching (preload vectors if < 30% of available cache)
|
||||||
|
- ✅ O(N) complexity - just loads existing data
|
||||||
|
- ❌ Does NOT call `addItem()` which would recompute connections (O(N log N))
|
||||||
|
|
||||||
|
### 2. TypeAwareVectorIndex Rebuild (Fixed in v3.45.0)
|
||||||
|
|
||||||
|
**Critical Architectural Fix**: The type-aware vector index previously had TWO major bugs:
|
||||||
|
|
||||||
|
1. **Bug #1**: Called `addItem()` during rebuild → O(N log N) recomputation instead of O(N) loading
|
||||||
|
2. **Bug #2**: Loaded ALL nouns 31 times in parallel (once per type) → O(31*N) complexity causing timeouts
|
||||||
|
|
||||||
|
Both were fixed in v3.45.0 by loading ALL nouns ONCE and routing to correct type indexes:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/hnsw/typeAwareHNSWIndex.ts (lines 379-571)
|
||||||
|
public async rebuild(options?: {
|
||||||
|
lazy?: boolean
|
||||||
|
batchSize?: number
|
||||||
|
onProgress?: (loaded: number, total: number) => void
|
||||||
|
}): Promise<void> {
|
||||||
|
// STEP 1: Clear all type-specific indexes
|
||||||
|
for (const index of this.typeIndexes.values()) {
|
||||||
|
index.clear()
|
||||||
|
}
|
||||||
|
|
||||||
|
// STEP 2: Determine preloading strategy (same as vector index)
|
||||||
|
const totalNouns = await this.storage.getNounCount()
|
||||||
|
const vectorMemory = totalNouns * 384 * 4
|
||||||
|
const availableCache = this.unifiedCache.getRemainingCapacity()
|
||||||
|
const shouldPreload = vectorMemory < availableCache * 0.3
|
||||||
|
|
||||||
|
// STEP 3: Load entities grouped by type
|
||||||
|
for (const nounType of ALL_NOUN_TYPES) {
|
||||||
|
const index = this.getOrCreateIndex(nounType)
|
||||||
|
let hasMore = true
|
||||||
|
let cursor: string | undefined = undefined
|
||||||
|
|
||||||
|
while (hasMore) {
|
||||||
|
const result = await this.storage.getNouns({
|
||||||
|
type: nounType,
|
||||||
|
pagination: { limit: 1000, cursor }
|
||||||
|
})
|
||||||
|
|
||||||
|
for (const nounData of result.items) {
|
||||||
|
// CORRECT: Load persisted vector index data (not recomputed!)
|
||||||
|
const hnswData = await this.storage.getHNSWData(nounData.id)
|
||||||
|
|
||||||
|
const noun = {
|
||||||
|
id: nounData.id,
|
||||||
|
vector: shouldPreload ? nounData.vector : [],
|
||||||
|
connections: new Map(),
|
||||||
|
level: hnswData.level
|
||||||
|
}
|
||||||
|
|
||||||
|
// Restore connections from storage
|
||||||
|
for (const [levelStr, nounIds] of Object.entries(hnswData.connections)) {
|
||||||
|
const level = parseInt(levelStr, 10)
|
||||||
|
noun.connections.set(level, new Set<string>(nounIds))
|
||||||
|
}
|
||||||
|
|
||||||
|
// Add to in-memory index (no recomputation!)
|
||||||
|
index.nouns.set(nounData.id, noun)
|
||||||
|
}
|
||||||
|
|
||||||
|
hasMore = result.hasMore
|
||||||
|
cursor = result.nextCursor
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Bug Fix**: Changed from `index.addItem()` (recomputation) to direct `nouns.set()` (restoration).
|
||||||
|
|
||||||
|
**Performance Impact**: 200-600x speedup (5 minutes → 500ms for 10K entities)
|
||||||
|
|
||||||
|
**Correct Pattern**:
|
||||||
|
```typescript
|
||||||
|
// Load ALL nouns ONCE (not 31 times!)
|
||||||
|
while (hasMore) {
|
||||||
|
const result = await storage.getNounsWithPagination({ limit: 1000, cursor })
|
||||||
|
|
||||||
|
for (const noun of result.items) {
|
||||||
|
const type = noun.nounType || noun.metadata?.noun
|
||||||
|
const index = this.getIndexForType(type)
|
||||||
|
|
||||||
|
// Load persisted HNSW data
|
||||||
|
const hnswData = await storage.getHNSWData(noun.id)
|
||||||
|
|
||||||
|
// Restore connections (not recompute!)
|
||||||
|
const restoredNoun = {
|
||||||
|
id: noun.id,
|
||||||
|
vector: shouldPreload ? noun.vector : [],
|
||||||
|
connections: restoreConnections(hnswData),
|
||||||
|
level: hnswData.level
|
||||||
|
}
|
||||||
|
|
||||||
|
// Add to correct type index
|
||||||
|
index.nouns.set(noun.id, restoredNoun)
|
||||||
|
}
|
||||||
|
|
||||||
|
cursor = result.nextCursor
|
||||||
|
hasMore = result.hasMore
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance Improvements**:
|
||||||
|
- 31x speedup: Load nouns ONCE instead of 31 times (O(N) vs O(31*N))
|
||||||
|
- 200-600x speedup: Load from storage instead of recomputing (O(N) vs O(N log N))
|
||||||
|
- **Combined**: ~6000x speedup! (150 minutes → 1.5 seconds for 10K entities)
|
||||||
|
|
||||||
|
### 3. MetadataIndex Rebuild (v4.2.1+ with Field Registry)
|
||||||
|
|
||||||
|
**v4.2.1 Critical Fix**: Field registry persistence eliminates unnecessary rebuilds!
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/utils/metadataIndex.ts (lines 202-216)
|
||||||
|
async init(): Promise<void> {
|
||||||
|
// STEP 1: Load field registry to discover persisted indices
|
||||||
|
// This is THE KEY FIX - O(1) discovery of existing indices
|
||||||
|
await this.loadFieldRegistry()
|
||||||
|
|
||||||
|
// If registry found, fieldIndexes Map is now populated
|
||||||
|
// getStats() will return totalEntries > 0 → skips rebuild!
|
||||||
|
|
||||||
|
// STEP 2: Initialize EntityIdMapper
|
||||||
|
await this.idMapper.init()
|
||||||
|
|
||||||
|
// STEP 3: Warm cache with discovered fields
|
||||||
|
await this.warmCache()
|
||||||
|
}
|
||||||
|
|
||||||
|
async loadFieldRegistry(): Promise<void> {
|
||||||
|
const registry = await this.storage.getMetadata('__metadata_field_registry__')
|
||||||
|
|
||||||
|
if (registry?.fields) {
|
||||||
|
// Populate fieldIndexes Map from discovered fields
|
||||||
|
// Sparse indices are lazy-loaded when first accessed
|
||||||
|
for (const field of registry.fields) {
|
||||||
|
this.fieldIndexes.set(field, {
|
||||||
|
values: {},
|
||||||
|
lastUpdated: registry.lastUpdated
|
||||||
|
})
|
||||||
|
}
|
||||||
|
// Result: getStats() now returns totalEntries > 0
|
||||||
|
// → Brain skips rebuild, cold start in 2-3 seconds!
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Rebuild Only Happens If**:
|
||||||
|
1. **First run** (no field registry exists yet)
|
||||||
|
2. **Registry corruption** (rare)
|
||||||
|
3. **Explicit rebuild request** (manual operation)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Only runs if field registry not found
|
||||||
|
async rebuild(): Promise<void> {
|
||||||
|
// STEP 1: Clear in-memory structures
|
||||||
|
this.fieldIndexes.clear()
|
||||||
|
|
||||||
|
// STEP 2: Load all entity metadata and rebuild indices
|
||||||
|
// Sequential batching (25/batch) to prevent socket exhaustion
|
||||||
|
// After rebuild: Field registry saved during next flush()
|
||||||
|
|
||||||
|
// One-time cost: ~2-3 seconds for 1K entities
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Performance Comparison**:
|
||||||
|
|
||||||
|
| Version | Cold Start | Discovery Method | Rebuild Needed? |
|
||||||
|
|---------|------------|------------------|-----------------|
|
||||||
|
| v4.2.0 | 8-9 min | None (always rebuild) | Always |
|
||||||
|
| v4.2.1 | 2-3 sec | Field registry O(1) | First run only |
|
||||||
|
|
||||||
|
**Key Points**:
|
||||||
|
- ✅ Field registry enables O(1) discovery (4-8KB file)
|
||||||
|
- ✅ Sparse indices lazy-loaded on first query
|
||||||
|
- ✅ Bloom filters + zone maps loaded for fast filtering
|
||||||
|
- ✅ One-time rebuild on first run, then instant restarts forever
|
||||||
|
- ✅ Automatic: No configuration needed
|
||||||
|
|
||||||
|
### 4. GraphAdjacencyIndex Rebuild
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// src/graph/graphAdjacencyIndex.ts (lines 279-336)
|
||||||
|
async rebuild(): Promise<void> {
|
||||||
|
// STEP 1: Clear in-memory caches
|
||||||
|
this.verbIndex.clear()
|
||||||
|
this.relationshipCountsByType.clear()
|
||||||
|
|
||||||
|
// STEP 2: Load all verbs from storage
|
||||||
|
let hasMore = true
|
||||||
|
let cursor: string | undefined = undefined
|
||||||
|
|
||||||
|
while (hasMore) {
|
||||||
|
const result = await this.storage.getVerbs({
|
||||||
|
pagination: { limit: 1000, cursor }
|
||||||
|
})
|
||||||
|
|
||||||
|
for (const verb of result.items) {
|
||||||
|
// Add to index (which updates LSM-trees)
|
||||||
|
await this.addVerb(verb)
|
||||||
|
}
|
||||||
|
|
||||||
|
hasMore = result.hasMore
|
||||||
|
cursor = result.nextCursor
|
||||||
|
}
|
||||||
|
|
||||||
|
// Note: LSM-trees (lsmTreeSource, lsmTreeTarget) are already
|
||||||
|
// initialized from persisted SSTables during ensureInitialized()
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key Points**:
|
||||||
|
- ✅ LSM-tree SSTables already loaded during `init()`
|
||||||
|
- ✅ Rebuild just repopulates verb cache
|
||||||
|
- ✅ O(E) complexity where E = number of edges
|
||||||
|
|
||||||
|
## Adaptive Memory Management
|
||||||
|
|
||||||
|
### Strategy: Preload vs Lazy Load
|
||||||
|
|
||||||
|
All indexes use the **UnifiedCache** to determine memory allocation:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Decision logic (in all indexes)
|
||||||
|
const totalDataSize = estimateDataSize()
|
||||||
|
const availableCache = unifiedCache.getRemainingCapacity()
|
||||||
|
|
||||||
|
if (totalDataSize < availableCache * 0.3) {
|
||||||
|
// PRELOAD: Dataset is small relative to available memory
|
||||||
|
// Load everything into memory for maximum performance
|
||||||
|
shouldPreload = true
|
||||||
|
} else {
|
||||||
|
// LAZY LOAD: Dataset is large
|
||||||
|
// Load on-demand with LRU eviction
|
||||||
|
shouldPreload = false
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Thresholds**:
|
||||||
|
- **< 30% of available cache**: Preload all vectors
|
||||||
|
- **> 30% of available cache**: Lazy load on demand
|
||||||
|
|
||||||
|
**Example** (default 100MB cache):
|
||||||
|
- 10K entities × 1.5KB = 15MB → **Preload** (15MB < 30MB)
|
||||||
|
- 100K entities × 1.5KB = 150MB → **Lazy load** (150MB > 30MB)
|
||||||
|
|
||||||
|
### UnifiedCache Integration
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// All indexes share the same cache
|
||||||
|
const unifiedCache = getGlobalCache() // Singleton, 100MB default
|
||||||
|
|
||||||
|
// MetadataIndex
|
||||||
|
this.unifiedCache = unifiedCache
|
||||||
|
|
||||||
|
// Vector index
|
||||||
|
this.unifiedCache = unifiedCache
|
||||||
|
|
||||||
|
// GraphAdjacencyIndex
|
||||||
|
this.unifiedCache = unifiedCache
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits**:
|
||||||
|
- Fair resource allocation across indexes
|
||||||
|
- Prevents any single index from monopolizing memory
|
||||||
|
- Coordinated LRU eviction system-wide
|
||||||
|
|
||||||
|
## Performance Characteristics
|
||||||
|
|
||||||
|
### Rebuild Times (Typical Hardware)
|
||||||
|
|
||||||
|
| Dataset Size | Metadata | Vector | Graph | Total (Parallel) |
|
||||||
|
|--------------|----------|------|-------|------------------|
|
||||||
|
| 1K entities | 50ms | 100ms | 30ms | **150ms** |
|
||||||
|
| 10K entities | 200ms | 500ms | 150ms | **600ms** |
|
||||||
|
| 100K entities | 1s | 3s | 1s | **3.5s** |
|
||||||
|
| 1M entities | 8s | 25s | 10s | **28s** |
|
||||||
|
|
||||||
|
**Note**: Parallel rebuild means total time ≈ max(individual times), not sum.
|
||||||
|
|
||||||
|
### Memory Overhead
|
||||||
|
|
||||||
|
| Index | In-Memory Overhead | Disk Storage |
|
||||||
|
|-------|-------------------|--------------|
|
||||||
|
| **MetadataIndex** | ~100 bytes/entity | ~500 bytes/entity (chunks) |
|
||||||
|
| **Vector Index** | ~200 bytes/entity (no vectors) | ~1.5 KB/entity (vectors + connections) |
|
||||||
|
| **GraphAdjacencyIndex** | ~128 bytes/relationship | ~200 bytes/relationship (LSM-tree) |
|
||||||
|
| **DeletedItemsIndex** | ~40 bytes/deleted ID | ~50 bytes/deleted ID |
|
||||||
|
|
||||||
|
**Total overhead** (lazy loading):
|
||||||
|
- **In-memory**: ~300 bytes per entity + ~128 bytes per relationship
|
||||||
|
- **On-disk**: ~2 KB per entity + ~200 bytes per relationship
|
||||||
|
|
||||||
|
### O(N) vs O(N log N) Comparison
|
||||||
|
|
||||||
|
**Before fix** (TypeAwareVectorIndex bug):
|
||||||
|
```typescript
|
||||||
|
// BAD: Recomputes vector index connections during rebuild
|
||||||
|
for (const noun of nouns) {
|
||||||
|
await index.addItem(noun) // O(log N) per item → O(N log N) total
|
||||||
|
}
|
||||||
|
// 10K entities: ~5 minutes
|
||||||
|
```
|
||||||
|
|
||||||
|
**After fix** (correct pattern):
|
||||||
|
```typescript
|
||||||
|
// GOOD: Loads connections from storage
|
||||||
|
for (const noun of nouns) {
|
||||||
|
const hnswData = await storage.getHNSWData(noun.id) // O(1) per item
|
||||||
|
noun.connections = restoreConnections(hnswData) // O(1) per item
|
||||||
|
index.nouns.set(noun.id, noun) // O(1) per item
|
||||||
|
}
|
||||||
|
// 10K entities: ~500ms (600x faster!)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Common Patterns
|
||||||
|
|
||||||
|
### Cold Start (Empty Storage)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brain({ storage })
|
||||||
|
|
||||||
|
// First init: All indexes are empty
|
||||||
|
await brain.init()
|
||||||
|
// → No rebuild needed, indexes start empty
|
||||||
|
|
||||||
|
// Add data
|
||||||
|
await brain.add({ content: 'Hello', noun: 'message' })
|
||||||
|
|
||||||
|
// Second init: Indexes populated
|
||||||
|
const brain2 = new Brain({ storage })
|
||||||
|
await brain2.init()
|
||||||
|
// → Rebuilds all indexes from storage (~1-3s for 10K entities)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Warm Start (Storage Already Populated)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brain({ storage })
|
||||||
|
|
||||||
|
// Init with existing data
|
||||||
|
await brain.init()
|
||||||
|
// → Detects non-empty storage
|
||||||
|
// → Rebuilds indexes in parallel
|
||||||
|
// → Uses adaptive caching (preload if small, lazy if large)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Manual Rebuild
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brain({ storage })
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Force rebuild (e.g., after data corruption)
|
||||||
|
await brain.metadataIndex.rebuild()
|
||||||
|
await brain.index.rebuild()
|
||||||
|
await brain.graphIndex.rebuild()
|
||||||
|
```
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
### Slow Rebuild Times
|
||||||
|
|
||||||
|
**Symptom**: Rebuild takes minutes instead of seconds
|
||||||
|
|
||||||
|
**Diagnosis**:
|
||||||
|
```typescript
|
||||||
|
// Check if rebuild is recomputing instead of loading
|
||||||
|
console.time('rebuild')
|
||||||
|
await brain.index.rebuild()
|
||||||
|
console.timeEnd('rebuild')
|
||||||
|
|
||||||
|
// For 10K entities:
|
||||||
|
// - Expected: 500-800ms (loading from storage)
|
||||||
|
// - Bug: 5-10 minutes (recomputing vector index connections)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Solution**: Ensure index is loading from storage, not calling `addItem()` during rebuild.
|
||||||
|
|
||||||
|
### High Memory Usage
|
||||||
|
|
||||||
|
**Symptom**: Memory usage exceeds expectations
|
||||||
|
|
||||||
|
**Diagnosis**:
|
||||||
|
```typescript
|
||||||
|
// Check if vectors are being preloaded
|
||||||
|
const stats = brain.index.getStats()
|
||||||
|
console.log('Preloaded vectors:', stats.preloadedVectors)
|
||||||
|
|
||||||
|
// Expected:
|
||||||
|
// - Small dataset (< 30% cache): Most vectors preloaded
|
||||||
|
// - Large dataset (> 30% cache): Few vectors preloaded
|
||||||
|
```
|
||||||
|
|
||||||
|
**Solution**: Adjust `UnifiedCache` size or force lazy loading:
|
||||||
|
```typescript
|
||||||
|
const brain = new Brain({
|
||||||
|
storage,
|
||||||
|
cache: { maxSize: 50 * 1024 * 1024 } // 50MB cache
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Missing Data After Rebuild
|
||||||
|
|
||||||
|
**Symptom**: Entities disappear after restart
|
||||||
|
|
||||||
|
**Diagnosis**:
|
||||||
|
```typescript
|
||||||
|
// Check storage persistence
|
||||||
|
const nouns = await storage.getNouns({ pagination: { limit: 10 } })
|
||||||
|
console.log('Nouns in storage:', nouns.items.length)
|
||||||
|
|
||||||
|
// If empty: Storage not persisting
|
||||||
|
// If populated: Rebuild not loading correctly
|
||||||
|
```
|
||||||
|
|
||||||
|
**Solution**: Verify storage adapter is configured correctly (e.g., FileSystem path exists).
|
||||||
|
|
||||||
|
## Related Documentation
|
||||||
|
|
||||||
|
- [Index Architecture](./index-architecture.md) - Data structures and operations
|
||||||
|
- [Storage Architecture](./storage-architecture.md) - Storage layer details
|
||||||
|
- [Performance Guide](../PERFORMANCE.md) - Performance tuning
|
||||||
|
- [Scaling Guide](../SCALING.md) - Large dataset optimization
|
||||||
|
|
||||||
|
## Version History
|
||||||
|
|
||||||
|
- **v5.7.7** (November 2025): Added production-scale lazy loading with `ensureIndexesLoaded()` helper. Fixed critical bug where `disableAutoRebuild: true` left indexes empty forever. Added concurrency control (mutex) to prevent duplicate rebuilds from concurrent queries. Added `getIndexStatus()` diagnostic method. Zero-config operation - works automatically.
|
||||||
|
- **v3.45.0** (October 2025): Fixed type-aware vector index `rebuild()` to load from storage instead of recomputing. Removed all snapshot code (unnecessary with correct rebuild pattern). 200-600x speedup.
|
||||||
|
- **v3.44.0** (October 2025): GraphAdjacencyIndex migrated to LSM-tree storage for billion-scale relationships
|
||||||
|
- **v3.42.0** (October 2025): MetadataIndex migrated to chunked sparse indexing
|
||||||
|
- **v3.35.0** (August 2025): Vector index connections first persisted to storage
|
||||||
|
- **v3.0.0** (September 2025): Initial 3-tier index architecture
|
||||||
134
docs/architecture/multiprocess-storage-mixin.md
Normal file
134
docs/architecture/multiprocess-storage-mixin.md
Normal file
|
|
@ -0,0 +1,134 @@
|
||||||
|
# Design note: multi-process storage mixin
|
||||||
|
|
||||||
|
**Status:** Proposed (future minor)
|
||||||
|
**Owner:** Brainy core
|
||||||
|
**Filed:** 2026-05-15
|
||||||
|
**Companion:** [`concepts/storage-adapters`](../concepts/storage-adapters.md)
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Brainy 7.21 added seven storage-adapter methods to support multi-process
|
||||||
|
safety:
|
||||||
|
|
||||||
|
```
|
||||||
|
supportsMultiProcessLocking()
|
||||||
|
acquireWriterLock(opts)
|
||||||
|
releaseWriterLock()
|
||||||
|
readWriterLock()
|
||||||
|
startFlushRequestWatcher(cb)
|
||||||
|
stopFlushRequestWatcher()
|
||||||
|
requestFlushOverFilesystem(timeoutMs)
|
||||||
|
```
|
||||||
|
|
||||||
|
They live on `BaseStorage` as no-op defaults and are overridden on
|
||||||
|
`FileSystemStorage` with real implementations. Any adapter extending
|
||||||
|
`FileSystemStorage` (e.g. Cor's `MmapFileSystemStorage`) inherits the
|
||||||
|
real ones for free.
|
||||||
|
|
||||||
|
This works correctly today. The question is whether the methods *belong*
|
||||||
|
on `BaseStorage`.
|
||||||
|
|
||||||
|
## The case for moving them out
|
||||||
|
|
||||||
|
`BaseStorage` already mixes several concerns:
|
||||||
|
- entity / verb CRUD primitives
|
||||||
|
- generational record hooks (8.0 MVCC)
|
||||||
|
- type-statistics tracking
|
||||||
|
- count persistence
|
||||||
|
- multi-process safety (new)
|
||||||
|
|
||||||
|
Adapters that have no notion of multi-process semantics — `MemoryStorage`,
|
||||||
|
cloud adapters (S3, GCS, R2, Azure, OPFS) — still carry seven inherited
|
||||||
|
no-ops on their prototype chain. A reader can't tell from the class
|
||||||
|
declaration whether a given adapter participates in the locking protocol;
|
||||||
|
it has to call `supportsMultiProcessLocking()` and trust the answer.
|
||||||
|
|
||||||
|
A cleaner separation:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
interface MultiProcessSafeStorage {
|
||||||
|
supportsMultiProcessLocking(): boolean
|
||||||
|
acquireWriterLock(opts?: { force?: boolean }): Promise<WriterLockInfo | null>
|
||||||
|
releaseWriterLock(): Promise<void>
|
||||||
|
readWriterLock(): Promise<WriterLockInfo | null>
|
||||||
|
startFlushRequestWatcher(cb: () => Promise<void>): void
|
||||||
|
stopFlushRequestWatcher(): void
|
||||||
|
requestFlushOverFilesystem(timeoutMs: number): Promise<boolean>
|
||||||
|
}
|
||||||
|
|
||||||
|
function isMultiProcessSafe(s: BaseStorage): s is BaseStorage & MultiProcessSafeStorage {
|
||||||
|
return typeof (s as any).supportsMultiProcessLocking === 'function'
|
||||||
|
&& (s as any).supportsMultiProcessLocking()
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Brainy's call sites become:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
if (this.config.mode !== 'reader' && isMultiProcessSafe(this.storage)) {
|
||||||
|
await this.storage.acquireWriterLock({ force: this.config.force })
|
||||||
|
// ... TypeScript narrows the rest correctly ...
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Benefits:
|
||||||
|
- Type system enforces the capability — no more `(this.storage as any).X()`.
|
||||||
|
- Adapters that opt out (memory, cloud) are visibly clean.
|
||||||
|
- `hasStorageMethod()` defensive helper can stay (still guards
|
||||||
|
build/install artifacts) but doesn't carry the conceptual weight of
|
||||||
|
"did the plugin implement the interface."
|
||||||
|
- ADR-style trail for future capability additions: each new capability
|
||||||
|
gets its own interface, opted into explicitly.
|
||||||
|
|
||||||
|
## The case against doing it now
|
||||||
|
|
||||||
|
- Breaking change for any adapter that already overrides these methods.
|
||||||
|
`FileSystemStorage` is the only one in-tree, but Cor's
|
||||||
|
`MmapFileSystemStorage` inherits from it — interface relocation would
|
||||||
|
ripple through the plugin ecosystem.
|
||||||
|
- The current state works. The real failure modes seen in the field
|
||||||
|
were build/install artifacts, not type-system failures.
|
||||||
|
- 7.22.0 just shipped a clean fix. Stacking another refactor before
|
||||||
|
consumers absorb it adds churn without urgency.
|
||||||
|
- The `hasStorageMethod()` guard accomplishes the same runtime safety the
|
||||||
|
interface narrowing would in TypeScript-aware code.
|
||||||
|
|
||||||
|
## Recommendation
|
||||||
|
|
||||||
|
**Defer.** Keep the current architecture through the 7.x line. Revisit
|
||||||
|
when:
|
||||||
|
- A second multi-process capability lands (e.g. distributed-readers
|
||||||
|
coordination) and the natural surface area is more than seven
|
||||||
|
methods. Five+ becomes the moment a separate interface earns its
|
||||||
|
keep.
|
||||||
|
- A v8 major is on the table for unrelated reasons. Bundle the
|
||||||
|
interface extraction with that release so consumers absorb both
|
||||||
|
changes in one upgrade.
|
||||||
|
|
||||||
|
Until then:
|
||||||
|
- Document the inheritance contract (done — see
|
||||||
|
[`concepts/storage-adapters`](../concepts/storage-adapters.md)).
|
||||||
|
- Keep `hasStorageMethod()` as the runtime guard.
|
||||||
|
- Don't add new methods to `BaseStorage` defaults without re-evaluating
|
||||||
|
the surface-area boundary.
|
||||||
|
|
||||||
|
## Migration sketch (when we do it)
|
||||||
|
|
||||||
|
For reference, a clean migration path:
|
||||||
|
|
||||||
|
1. Add `MultiProcessSafeStorage` interface to `src/storage/coreTypes.ts`.
|
||||||
|
2. Move the seven method signatures from `BaseStorage` to the new
|
||||||
|
interface. Default implementations stay on `BaseStorage` but only as
|
||||||
|
private helpers consumed by `FileSystemStorage`'s explicit
|
||||||
|
implementations.
|
||||||
|
3. `FileSystemStorage implements MultiProcessSafeStorage` becomes
|
||||||
|
explicit; methods get the `public` modifier with full JSDoc.
|
||||||
|
4. Brainy call sites switch from `hasStorageMethod` to
|
||||||
|
`isMultiProcessSafe` type-guard. Keep `hasStorageMethod` for
|
||||||
|
build/install artifact protection.
|
||||||
|
5. Document the new contract in `concepts/storage-adapters.md`.
|
||||||
|
6. Major-version-bump the `@soulcraft/brainy` peerDep range expected by
|
||||||
|
plugins.
|
||||||
|
|
||||||
|
Estimated work: ~half a day of code, ~2 hours of doc/example updates,
|
||||||
|
ecosystem coordination via the platform handoff.
|
||||||
File diff suppressed because it is too large
Load diff
|
|
@ -4,17 +4,16 @@ Brainy is a multi-dimensional AI database that combines vector similarity, graph
|
||||||
|
|
||||||
## Core Components
|
## Core Components
|
||||||
|
|
||||||
### BrainyData (Main Entry Point)
|
### Brainy (Main Entry Point)
|
||||||
The central orchestrator that manages all subsystems:
|
The central orchestrator that manages all subsystems:
|
||||||
- **HNSW Index**: O(log n) vector similarity search
|
- **4-Index Architecture**: MetadataIndex, vector index, GraphAdjacencyIndex, DeletedItemsIndex (see [Index Architecture](./index-architecture.md))
|
||||||
- **Storage System**: Universal storage adapters (FileSystem, S3, OPFS, Memory)
|
- **Storage System**: FileSystem and Memory adapters
|
||||||
- **Metadata Index**: O(1) field lookups with inverted indexing
|
|
||||||
- **Augmentation System**: Extensible plugin architecture
|
- **Augmentation System**: Extensible plugin architecture
|
||||||
- **Triple Intelligence**: Unified query engine
|
- **Triple Intelligence**: Unified query engine
|
||||||
|
|
||||||
### Triple Intelligence Engine
|
### Triple Intelligence Engine
|
||||||
Brainy's revolutionary feature that unifies three types of search:
|
Brainy's revolutionary feature that unifies three types of search:
|
||||||
- **Vector Search**: Semantic similarity using HNSW indexing
|
- **Vector Search**: Semantic similarity via the pluggable vector index
|
||||||
- **Graph Traversal**: Relationship-based queries
|
- **Graph Traversal**: Relationship-based queries
|
||||||
- **Field Filtering**: Precise metadata filtering with O(1) performance
|
- **Field Filtering**: Precise metadata filtering with O(1) performance
|
||||||
|
|
||||||
|
|
@ -40,16 +39,16 @@ brainy-data/
|
||||||
│ ├── __entity_registry__.json
|
│ ├── __entity_registry__.json
|
||||||
│ └── __metadata_index__*.json
|
│ └── __metadata_index__*.json
|
||||||
├── verbs/ # Relationship storage
|
├── verbs/ # Relationship storage
|
||||||
├── wal/ # Write-Ahead Logging
|
|
||||||
└── locks/ # Concurrent access control
|
└── locks/ # Concurrent access control
|
||||||
```
|
```
|
||||||
|
|
||||||
### HNSW Index
|
### Vector Index
|
||||||
Hierarchical Navigable Small World index for efficient vector search:
|
Pluggable vector index (`VectorIndexProvider`) for efficient nearest-neighbor search. The default JS implementation, `JsHnswVectorIndex`, uses a hierarchical graph:
|
||||||
- **Performance**: O(log n) search complexity
|
- **Performance**: O(log n) search complexity
|
||||||
- **Memory Efficient**: Product quantization support
|
- **Configurable recall**: `fast` / `balanced` / `accurate` presets trade recall for latency
|
||||||
- **Scalable**: Handles millions of vectors
|
- **Scalable**: Handles millions of vectors per process
|
||||||
- **Persistent**: Serializable to storage
|
- **Persistent**: Serializable to storage
|
||||||
|
- **Swappable**: Replace with a native implementation (such as `@soulcraft/cor`) via the plugin system without changing application code
|
||||||
|
|
||||||
### Metadata Index Manager
|
### Metadata Index Manager
|
||||||
High-performance field indexing system:
|
High-performance field indexing system:
|
||||||
|
|
@ -61,7 +60,7 @@ High-performance field indexing system:
|
||||||
## Performance Characteristics
|
## Performance Characteristics
|
||||||
|
|
||||||
### Operation Complexity
|
### Operation Complexity
|
||||||
- **Vector Search**: O(log n) via HNSW
|
- **Vector Search**: O(log n) via the vector index
|
||||||
- **Field Filtering**: O(1) via inverted indexes
|
- **Field Filtering**: O(1) via inverted indexes
|
||||||
- **Graph Traversal**: O(V + E) for breadth-first search
|
- **Graph Traversal**: O(V + E) for breadth-first search
|
||||||
- **Add Operation**: O(log n) for index insertion
|
- **Add Operation**: O(log n) for index insertion
|
||||||
|
|
@ -83,20 +82,18 @@ High-performance field indexing system:
|
||||||
Brainy's extensible plugin architecture allows for powerful enhancements:
|
Brainy's extensible plugin architecture allows for powerful enhancements:
|
||||||
|
|
||||||
### Core Augmentations
|
### Core Augmentations
|
||||||
- **WAL (Write-Ahead Logging)**: Durability and crash recovery
|
|
||||||
- **Entity Registry**: High-speed deduplication for streaming data
|
- **Entity Registry**: High-speed deduplication for streaming data
|
||||||
- **Batch Processing**: Optimized bulk operations
|
- **Batch Processing**: Optimized bulk operations
|
||||||
- **Connection Pool**: Efficient resource management
|
|
||||||
- **Request Deduplicator**: Prevents duplicate processing
|
- **Request Deduplicator**: Prevents duplicate processing
|
||||||
|
|
||||||
### Creating Custom Augmentations
|
### Creating Custom Augmentations
|
||||||
```typescript
|
```typescript
|
||||||
class CustomAugmentation extends BrainyAugmentation {
|
class CustomAugmentation extends BrainyAugmentation {
|
||||||
async onInit(brain: BrainyData): Promise<void> {
|
async onInit(brain: Brainy): Promise<void> {
|
||||||
// Initialize augmentation
|
// Initialize augmentation
|
||||||
}
|
}
|
||||||
|
|
||||||
async onAdd(item: any, brain: BrainyData): Promise<any> {
|
async onAdd(item: any, brain: Brainy): Promise<any> {
|
||||||
// Process item before adding
|
// Process item before adding
|
||||||
return item
|
return item
|
||||||
}
|
}
|
||||||
|
|
@ -114,11 +111,14 @@ Multi-layered caching for optimal performance:
|
||||||
## Integration Points
|
## Integration Points
|
||||||
|
|
||||||
### Key Objects for Extensions
|
### Key Objects for Extensions
|
||||||
- `brain.index`: Access HNSW vector index
|
- `brain.index`: Access the vector index
|
||||||
- `brain.metadataIndex`: Access field indexing
|
- `brain.metadataIndex`: Access field indexing
|
||||||
|
- `brain.graphIndex`: Access graph adjacency index
|
||||||
- `brain.storage`: Access storage layer
|
- `brain.storage`: Access storage layer
|
||||||
- `brain.augmentations`: Access augmentation manager
|
- `brain.augmentations`: Access augmentation manager
|
||||||
|
|
||||||
|
For detailed information about each index, see [Index Architecture](./index-architecture.md).
|
||||||
|
|
||||||
### Event System
|
### Event System
|
||||||
```typescript
|
```typescript
|
||||||
brain.on('add', (item) => console.log('Item added:', item))
|
brain.on('add', (item) => console.log('Item added:', item))
|
||||||
|
|
@ -144,6 +144,7 @@ brain.on('error', (error) => console.error('Error:', error))
|
||||||
|
|
||||||
## Next Steps
|
## Next Steps
|
||||||
|
|
||||||
|
- [Index Architecture](./index-architecture.md) - Deep dive into the 4-index system
|
||||||
- [Storage Architecture](./storage-architecture.md) - Deep dive into storage system
|
- [Storage Architecture](./storage-architecture.md) - Deep dive into storage system
|
||||||
- [Triple Intelligence](./triple-intelligence.md) - Advanced query system
|
- [Triple Intelligence](./triple-intelligence.md) - Advanced query system
|
||||||
- [API Reference](../api/README.md) - Complete API documentation
|
- [API Reference](../api/README.md) - Complete API documentation
|
||||||
|
|
@ -1,86 +1,120 @@
|
||||||
# Storage Architecture
|
# Storage Architecture
|
||||||
|
|
||||||
Brainy implements a sophisticated, unified storage system that works across all environments (Node.js, Browser, Edge Workers) with enterprise-grade features like metadata indexing, entity registry, and write-ahead logging.
|
> **Updated**: Metadata/vector separation, UUID-based sharding, on-disk artifact for operator-layer backup
|
||||||
|
|
||||||
## Storage Structure
|
## Storage Structure
|
||||||
|
|
||||||
|
### Architecture: Metadata/Vector Separation
|
||||||
|
|
||||||
|
Entities and relationships are split into **2 separate files** for optimal performance at billion-entity scale:
|
||||||
|
|
||||||
```
|
```
|
||||||
brainy-data/
|
brainy-data/
|
||||||
├── _system/ # System management
|
├── _system/ # System metadata (not sharded)
|
||||||
│ └── statistics.json # Performance metrics and statistics
|
│ ├── statistics.json # Performance metrics
|
||||||
├── nouns/ # Primary entity storage
|
│ ├── __metadata_field_index__*.json # Field indexes
|
||||||
│ └── {uuid}.json # Individual entity documents
|
│ └── __metadata_sorted_index__*.json # Sorted indexes
|
||||||
├── metadata/ # Metadata and indexing system
|
│
|
||||||
│ ├── {uuid}.json # Entity metadata
|
├── entities/
|
||||||
│ ├── __entity_registry__.json # Entity deduplication registry
|
│ ├── nouns/
|
||||||
│ ├── __metadata_field_index__field_{field}.json # Field discovery
|
│ │ ├── vectors/ # Vector graph data (sharded by UUID)
|
||||||
│ └── __metadata_index__{field}_{value}_chunk{n}.json # Value indexes
|
│ │ │ ├── 00/ # Shard 00 (first 2 hex digits)
|
||||||
├── verbs/ # Relationship/action storage
|
│ │ │ │ ├── 00123456-....json # Vector + graph connections
|
||||||
│ └── {uuid}.json # Relationship documents
|
│ │ │ │ └── 00abcdef-....json
|
||||||
├── wal/ # Write-Ahead Logging
|
│ │ │ ├── 01/ ... ff/ # 256 shards total
|
||||||
│ └── wal_{timestamp}_{id}.wal # Transaction logs
|
│ │ │
|
||||||
└── locks/ # Concurrent access control
|
│ │ └── metadata/ # Business data (sharded by UUID)
|
||||||
└── {resource}.lock # Resource locks
|
│ │ ├── 00/
|
||||||
|
│ │ │ ├── 00123456-....json # Entity metadata only
|
||||||
|
│ │ │ └── 00abcdef-....json
|
||||||
|
│ │ ├── 01/ ... ff/
|
||||||
|
│ │
|
||||||
|
│ └── verbs/
|
||||||
|
│ ├── vectors/ # Relationship vectors (sharded)
|
||||||
|
│ │ ├── 00/ ... ff/
|
||||||
|
│ │
|
||||||
|
│ └── metadata/ # Relationship data (sharded)
|
||||||
|
│ ├── 00/ ... ff/
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Why Split Metadata and Vectors?
|
||||||
|
|
||||||
|
**Performance at scale:**
|
||||||
|
- **Vector search operations**: Only load vectors (4KB) during search, not metadata (2-10KB)
|
||||||
|
- **Filtering**: Only load metadata during filtering, not vectors
|
||||||
|
- **Pagination**: Load metadata IDs first, fetch vectors/metadata on-demand
|
||||||
|
- **Result**: 60-70% reduction in I/O for typical queries at million-entity scale
|
||||||
|
|
||||||
|
### UUID-Based Sharding (256 Shards)
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
```typescript
|
||||||
|
const uuid = "3fa85f64-5717-4562-b3fc-2c963f66afa6"
|
||||||
|
const shard = uuid.substring(0, 2) // "3f"
|
||||||
|
|
||||||
|
// Vector path: entities/nouns/vectors/3f/3fa85f64-....json
|
||||||
|
// Metadata path: entities/nouns/metadata/3f/3fa85f64-....json
|
||||||
|
```
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- **Uniform distribution**: ~3,900 entities per shard (at 1M scale)
|
||||||
|
- **Filesystem optimization**: avoids huge flat directories that bog down `readdir`
|
||||||
|
- **Parallel operations**: walk 256 shards in parallel
|
||||||
|
- **Predictable**: Deterministic shard assignment
|
||||||
|
|
||||||
## Storage Adapters
|
## Storage Adapters
|
||||||
|
|
||||||
Brainy provides multiple storage adapters with identical APIs:
|
Brainy 8.0 ships two adapters, both implementing the same `StorageAdapter` interface:
|
||||||
|
|
||||||
### FileSystem Storage (Node.js)
|
### FileSystem Storage (Node.js, default)
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
storage: {
|
storage: {
|
||||||
type: 'filesystem',
|
type: 'filesystem',
|
||||||
path: './data'
|
path: './data'
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
- **Use case**: Server applications, CLI tools
|
- **Use case**: Server applications, CLI tools, single-node deployments
|
||||||
- **Performance**: Direct file I/O
|
- **Performance**: Direct file I/O
|
||||||
- **Persistence**: Permanent on disk
|
- **Persistence**: Permanent on disk
|
||||||
|
- **Features**:
|
||||||
### S3 Compatible Storage
|
- **Batch Delete**: Efficient bulk deletion with retries
|
||||||
```typescript
|
- **UUID Sharding**: Automatic 256-shard distribution
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 's3',
|
|
||||||
bucket: 'my-brainy-data',
|
|
||||||
region: 'us-east-1',
|
|
||||||
credentials: {
|
|
||||||
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
|
|
||||||
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY
|
|
||||||
}
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
- **Use case**: Distributed applications, cloud deployments
|
|
||||||
- **Performance**: Network dependent, with intelligent caching
|
|
||||||
- **Persistence**: Cloud storage durability
|
|
||||||
|
|
||||||
### Origin Private File System (Browser)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'opfs'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
- **Use case**: Browser applications, PWAs
|
|
||||||
- **Performance**: Near-native file system speed
|
|
||||||
- **Persistence**: Permanent in browser (with quota limits)
|
|
||||||
|
|
||||||
### Memory Storage
|
### Memory Storage
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
storage: {
|
storage: {
|
||||||
type: 'memory'
|
type: 'memory'
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
- **Use case**: Testing, temporary processing
|
- **Use case**: Tests, ephemeral workloads, single-process caches
|
||||||
- **Performance**: Fastest possible
|
- **Performance**: No I/O — all data lives in process memory
|
||||||
- **Persistence**: Volatile (lost on restart)
|
- **Persistence**: None — data is lost when the process exits
|
||||||
|
|
||||||
|
### Auto
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: {
|
||||||
|
type: 'auto',
|
||||||
|
path: './data'
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
`'auto'` picks `'filesystem'` when running on Node.js with a writable `path`, and falls back to `'memory'` otherwise.
|
||||||
|
|
||||||
|
## Backup and Off-Site Replication
|
||||||
|
|
||||||
|
Brainy 8.0 does not embed cloud SDKs. The on-disk artifact at `path` is a plain directory tree of JSON files, so backup is an operator-layer concern. Typical patterns:
|
||||||
|
|
||||||
|
- `gsutil rsync -r ./data gs://my-bucket/brainy-data`
|
||||||
|
- `aws s3 sync ./data s3://my-bucket/brainy-data`
|
||||||
|
- `rclone sync ./data remote:brainy-data`
|
||||||
|
- Periodic `tar` snapshots to any object store
|
||||||
|
|
||||||
|
Run these from your scheduler (cron, systemd timer, k8s CronJob) — Brainy itself only reads and writes the local directory.
|
||||||
|
|
||||||
## Metadata Indexing System
|
## Metadata Indexing System
|
||||||
|
|
||||||
|
|
@ -144,43 +178,39 @@ High-performance deduplication system for streaming data:
|
||||||
- **Cache**: LRU with configurable TTL
|
- **Cache**: LRU with configurable TTL
|
||||||
- **Sync**: Periodic or on-demand
|
- **Sync**: Periodic or on-demand
|
||||||
|
|
||||||
## Write-Ahead Logging (WAL)
|
## Durability
|
||||||
|
|
||||||
Ensures durability and enables recovery:
|
Brainy persists writes to disk through the filesystem adapter. Each save is a rename-based atomic write of a JSON file under the appropriate shard. Operators that need point-in-time recovery should snapshot `path` (see [Backup and Off-Site Replication](#backup-and-off-site-replication)).
|
||||||
|
|
||||||
### WAL Entry Format
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"timestamp": 1699564234567,
|
|
||||||
"operation": "add",
|
|
||||||
"data": {
|
|
||||||
"id": "550e8400-e29b-41d4-a716-446655440000",
|
|
||||||
"content": "...",
|
|
||||||
"metadata": {}
|
|
||||||
},
|
|
||||||
"checksum": "sha256:..."
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Recovery Process
|
|
||||||
1. On startup, check for WAL files
|
|
||||||
2. Replay operations from last checkpoint
|
|
||||||
3. Verify checksums for integrity
|
|
||||||
4. Clean up processed WAL files
|
|
||||||
|
|
||||||
## Storage Optimization
|
## Storage Optimization
|
||||||
|
|
||||||
### Compression
|
### 1. Batch Operations
|
||||||
- **JSON**: Automatic minification
|
|
||||||
- **Vectors**: Float32 to Uint8 quantization option
|
|
||||||
- **Indexes**: Binary format for large datasets
|
|
||||||
|
|
||||||
### Caching Strategy
|
|
||||||
```typescript
|
```typescript
|
||||||
// Configure caching per storage type
|
// Efficient batch delete
|
||||||
const brain = new BrainyData({
|
await storage.batchDelete([
|
||||||
|
'entities/nouns/vectors/00/00123456-....json',
|
||||||
|
'entities/nouns/metadata/00/00123456-....json'
|
||||||
|
// ...
|
||||||
|
])
|
||||||
|
|
||||||
|
// Batch writes for performance
|
||||||
|
await brain.addBatch([
|
||||||
|
{ content: "item1", metadata: {} },
|
||||||
|
{ content: "item2", metadata: {} },
|
||||||
|
{ content: "item3", metadata: {} }
|
||||||
|
])
|
||||||
|
// Single transaction, optimized I/O
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Caching Strategy
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Configure caching
|
||||||
|
const brain = new Brainy({
|
||||||
storage: {
|
storage: {
|
||||||
type: 'filesystem',
|
type: 'filesystem',
|
||||||
|
path: './data',
|
||||||
cache: {
|
cache: {
|
||||||
enabled: true,
|
enabled: true,
|
||||||
maxSize: 1000, // Maximum cached items
|
maxSize: 1000, // Maximum cached items
|
||||||
|
|
@ -191,17 +221,6 @@ const brain = new BrainyData({
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
||||||
### Batch Operations
|
|
||||||
```typescript
|
|
||||||
// Batch writes for performance
|
|
||||||
await brain.addBatch([
|
|
||||||
{ content: "item1", metadata: {} },
|
|
||||||
{ content: "item2", metadata: {} },
|
|
||||||
{ content: "item3", metadata: {} }
|
|
||||||
])
|
|
||||||
// Single transaction, optimized I/O
|
|
||||||
```
|
|
||||||
|
|
||||||
## Concurrent Access
|
## Concurrent Access
|
||||||
|
|
||||||
### Locking Mechanism
|
### Locking Mechanism
|
||||||
|
|
@ -220,58 +239,39 @@ await brain.storage.withLock('resource-id', async () => {
|
||||||
|
|
||||||
## Migration and Backup
|
## Migration and Backup
|
||||||
|
|
||||||
### Export Data
|
Backup and restore go through the Db API — see
|
||||||
|
[Snapshots & Time Travel](../guides/snapshots-and-time-travel.md) for the full
|
||||||
|
recipe book.
|
||||||
|
|
||||||
|
### Snapshot (backup)
|
||||||
```typescript
|
```typescript
|
||||||
// Export entire database
|
// Instant, self-contained snapshot (hard links on filesystem storage)
|
||||||
const backup = await brain.export({
|
const db = brain.now()
|
||||||
format: 'json',
|
await db.persist('/backups/2026-06-11')
|
||||||
includeVectors: true,
|
await db.release()
|
||||||
includeIndexes: false
|
|
||||||
})
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Import Data
|
### Restore
|
||||||
```typescript
|
```typescript
|
||||||
// Import from backup
|
// Replace the store's entire state from a snapshot (destructive — confirm required)
|
||||||
await brain.import(backup, {
|
await brain.restore('/backups/2026-06-11', { confirm: true })
|
||||||
mode: 'merge', // or 'replace'
|
|
||||||
validateSchema: true
|
|
||||||
})
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Storage Migration
|
### Move to a new directory
|
||||||
```typescript
|
```typescript
|
||||||
// Migrate between storage types
|
// A snapshot directory is a complete store: restore it into a fresh brain
|
||||||
const oldBrain = new BrainyData({ storage: { type: 'filesystem' } })
|
const brain = new Brainy({ storage: { type: 'filesystem', path: './new' } })
|
||||||
const newBrain = new BrainyData({ storage: { type: 's3' } })
|
await brain.init()
|
||||||
|
await brain.restore('/backups/2026-06-11', { confirm: true })
|
||||||
await oldBrain.init()
|
|
||||||
await newBrain.init()
|
|
||||||
|
|
||||||
// Transfer all data
|
|
||||||
const data = await oldBrain.export()
|
|
||||||
await newBrain.import(data)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Performance Tuning
|
## Performance Tuning
|
||||||
|
|
||||||
### Storage-Specific Optimizations
|
### FileSystem Optimizations
|
||||||
|
- **Directory sharding**: 256 shards spread files across subdirectories
|
||||||
#### FileSystem
|
|
||||||
- **Directory sharding**: Split files across subdirectories
|
|
||||||
- **Async I/O**: Non-blocking file operations
|
- **Async I/O**: Non-blocking file operations
|
||||||
- **Buffer pooling**: Reuse buffers for efficiency
|
- **Buffer pooling**: Reuse buffers for efficiency
|
||||||
|
|
||||||
#### S3
|
|
||||||
- **Multipart uploads**: For large objects
|
|
||||||
- **Request batching**: Combine small operations
|
|
||||||
- **CDN integration**: Edge caching for reads
|
|
||||||
|
|
||||||
#### OPFS
|
|
||||||
- **Quota management**: Monitor and request increases
|
|
||||||
- **Worker offloading**: Heavy operations in workers
|
|
||||||
- **Transaction batching**: Group operations
|
|
||||||
|
|
||||||
### Monitoring
|
### Monitoring
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
|
|
@ -290,23 +290,29 @@ console.log(stats)
|
||||||
## Best Practices
|
## Best Practices
|
||||||
|
|
||||||
### Choose the Right Adapter
|
### Choose the Right Adapter
|
||||||
1. **Development**: Memory or FileSystem
|
1. **Development & tests**: `memory` for speed, `filesystem` when you need persistence
|
||||||
2. **Production Server**: FileSystem or S3
|
2. **Single-process production**: `filesystem` with off-site backup via `gsutil` / `aws s3 sync` / `rclone`
|
||||||
3. **Browser Apps**: OPFS or Memory
|
3. **Horizontal scaling**: Brainy runs in one process — there is no built-in cluster. Run independent instances behind a service layer and replicate the on-disk artifact with your operator tooling; or run many reader processes against one shared store with a single writer
|
||||||
4. **Distributed**: S3 with caching
|
|
||||||
|
|
||||||
### Optimize for Your Use Case
|
### Optimize for Your Use Case
|
||||||
1. **Read-heavy**: Enable aggressive caching
|
1. **Read-heavy**: Enable caching and let the OS page cache do its job
|
||||||
2. **Write-heavy**: Use WAL and batching
|
2. **Write-heavy**: Batch operations and tune the cache `maxSize`
|
||||||
3. **Real-time**: Memory with periodic persistence
|
3. **Real-time**: FileSystem with periodic snapshots
|
||||||
4. **Archival**: S3 with compression
|
4. **Archival**: Snapshot `path` to cold object storage on a schedule
|
||||||
|
5. **Large-scale**: Rely on metadata/vector separation + UUID sharding
|
||||||
|
|
||||||
### Monitor and Maintain
|
### Monitor and Maintain
|
||||||
1. Regular statistics collection
|
1. Regular statistics collection
|
||||||
2. WAL cleanup scheduling
|
2. Watch disk usage and shard balance
|
||||||
3. Index optimization
|
3. Index optimization
|
||||||
4. Cache tuning based on hit rates
|
4. Cache tuning based on hit rates
|
||||||
|
5. Verify backup runs (test restore quarterly)
|
||||||
|
|
||||||
## API Reference
|
## API Reference
|
||||||
|
|
||||||
See the [Storage API](../api/storage.md) for complete method documentation.
|
See the [Storage API](../api/storage.md) for complete method documentation.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Last Updated**: 2026
|
||||||
|
**Key Features**: Metadata/vector separation, UUID sharding, filesystem-and-memory adapters, operator-layer backup
|
||||||
|
|
|
||||||
|
|
@ -1,3 +1,16 @@
|
||||||
|
---
|
||||||
|
title: Triple Intelligence
|
||||||
|
slug: concepts/triple-intelligence
|
||||||
|
public: true
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 1
|
||||||
|
description: Unified vector similarity, graph traversal, and metadata filtering in one query. Auto-optimizes between parallel execution and progressive filtering.
|
||||||
|
next:
|
||||||
|
- concepts/noun-types
|
||||||
|
- api/reference
|
||||||
|
---
|
||||||
|
|
||||||
# Triple Intelligence System
|
# Triple Intelligence System
|
||||||
|
|
||||||
The Triple Intelligence System is Brainy's revolutionary query engine that unifies vector similarity, graph relationships, and metadata filtering into a single, optimized query interface.
|
The Triple Intelligence System is Brainy's revolutionary query engine that unifies vector similarity, graph relationships, and metadata filtering into a single, optimized query interface.
|
||||||
|
|
@ -10,29 +23,37 @@ Traditional databases force you to choose between vector search, graph traversal
|
||||||
|
|
||||||
### Unified Query Structure
|
### Unified Query Structure
|
||||||
|
|
||||||
```typescript
|
`find()` accepts a single `FindParams` object (or a natural-language string). One
|
||||||
interface TripleQuery {
|
object combines all three intelligences:
|
||||||
// Vector/Semantic search
|
|
||||||
like?: string | Vector | any
|
|
||||||
similar?: string | Vector | any
|
|
||||||
|
|
||||||
// Graph/Relationship search
|
```typescript
|
||||||
|
interface FindParams {
|
||||||
|
// Vector intelligence — semantic similarity
|
||||||
|
query?: string // Natural-language / semantic query (embedded, matched via HNSW + text index)
|
||||||
|
vector?: number[] // Pre-computed embedding for direct vector search
|
||||||
|
|
||||||
|
// Metadata intelligence — structured field filters
|
||||||
|
type?: NounType | NounType[] // Filter by entity type
|
||||||
|
subtype?: string | string[] // Filter by per-product subtype
|
||||||
|
where?: Record<string, any> // Field predicates with bare operators (gte, lt, in, contains, exists…)
|
||||||
|
|
||||||
|
// Graph intelligence — relationship traversal
|
||||||
connected?: {
|
connected?: {
|
||||||
to?: string | string[]
|
to?: string // Reachable to this entity
|
||||||
from?: string | string[]
|
from?: string // Reachable from this entity
|
||||||
type?: string | string[]
|
via?: VerbType | VerbType[] // Relationship type(s) to traverse (alias: type)
|
||||||
depth?: number
|
depth?: number // Max traversal depth (default: 1)
|
||||||
direction?: 'in' | 'out' | 'both'
|
direction?: 'in' | 'out' | 'both'
|
||||||
}
|
}
|
||||||
|
|
||||||
// Field/Attribute search
|
// Proximity — nearest neighbours of a known entity
|
||||||
where?: Record<string, any>
|
near?: { id: string; threshold?: number }
|
||||||
|
|
||||||
// Advanced options
|
// Control
|
||||||
limit?: number
|
limit?: number // Max results (default: 10)
|
||||||
boost?: 'recent' | 'popular' | 'verified' | string
|
offset?: number // Skip N results
|
||||||
explain?: boolean
|
orderBy?: string // Field to sort by (e.g. 'createdAt')
|
||||||
threshold?: number
|
order?: 'asc' | 'desc' // Sort direction
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -55,16 +76,16 @@ const articles = await brain.find("verified articles by John Smith about machine
|
||||||
|
|
||||||
#### Simple Vector Search
|
#### Simple Vector Search
|
||||||
```typescript
|
```typescript
|
||||||
const results = await brain.search("machine learning concepts")
|
const results = await brain.find("machine learning concepts")
|
||||||
```
|
```
|
||||||
|
|
||||||
#### Combined Intelligence Query
|
#### Combined Intelligence Query
|
||||||
```typescript
|
```typescript
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
like: "neural networks",
|
query: "neural networks",
|
||||||
where: {
|
where: {
|
||||||
category: "research",
|
category: "research",
|
||||||
year: { $gte: 2023 }
|
year: { gte: 2023 }
|
||||||
},
|
},
|
||||||
connected: {
|
connected: {
|
||||||
to: "deep-learning-team",
|
to: "deep-learning-team",
|
||||||
|
|
@ -96,8 +117,8 @@ All three search types execute simultaneously:
|
||||||
```typescript
|
```typescript
|
||||||
// Parallel execution for balanced query
|
// Parallel execution for balanced query
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
like: "AI research", // ~1000 potential matches
|
query: "AI research", // ~1000 potential matches
|
||||||
where: { type: "paper" }, // ~500 potential matches
|
where: { kind: "paper" }, // ~500 potential matches
|
||||||
connected: { to: "stanford" } // ~200 potential matches
|
connected: { to: "stanford" } // ~200 potential matches
|
||||||
})
|
})
|
||||||
// All three execute in parallel, results fused
|
// All three execute in parallel, results fused
|
||||||
|
|
@ -113,7 +134,7 @@ Operations chain for maximum efficiency:
|
||||||
// Progressive execution for selective query
|
// Progressive execution for selective query
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
where: { userId: "user123" }, // Very selective (1-10 matches)
|
where: { userId: "user123" }, // Very selective (1-10 matches)
|
||||||
like: "recent posts", // Applied to filtered set
|
query: "recent posts", // Applied to filtered set
|
||||||
limit: 5
|
limit: 5
|
||||||
})
|
})
|
||||||
// Metadata filter first, then vector search on results
|
// Metadata filter first, then vector search on results
|
||||||
|
|
@ -148,15 +169,15 @@ Brainy includes 220+ embedded patterns for natural language understanding:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Natural language automatically parsed
|
// Natural language automatically parsed
|
||||||
const results = await brain.search(
|
const results = await brain.find(
|
||||||
"show me recent AI papers from Stanford published this year"
|
"show me recent AI papers from Stanford published this year"
|
||||||
)
|
)
|
||||||
// Automatically converts to:
|
// Automatically converts to:
|
||||||
// {
|
// {
|
||||||
// like: "AI papers",
|
// query: "AI papers",
|
||||||
// where: {
|
// where: {
|
||||||
// institution: "Stanford",
|
// institution: "Stanford",
|
||||||
// published: { $gte: "2024-01-01" }
|
// published: { gte: "2024-01-01" }
|
||||||
// }
|
// }
|
||||||
// }
|
// }
|
||||||
```
|
```
|
||||||
|
|
@ -175,11 +196,11 @@ The NLP processor identifies query intent:
|
||||||
|
|
||||||
Successful execution plans are cached:
|
Successful execution plans are cached:
|
||||||
```typescript
|
```typescript
|
||||||
// First query: 50ms (plan generation + execution)
|
// First call parses the natural-language query and builds an execution plan
|
||||||
await brain.search("machine learning papers")
|
await brain.find("machine learning papers")
|
||||||
|
|
||||||
// Subsequent similar queries: 10ms (cached plan)
|
// A structurally similar query reuses that plan, skipping plan generation
|
||||||
await brain.search("deep learning papers")
|
await brain.find("deep learning papers")
|
||||||
```
|
```
|
||||||
|
|
||||||
### Self-Optimization
|
### Self-Optimization
|
||||||
|
|
@ -200,50 +221,45 @@ Triple Intelligence leverages all available indexes:
|
||||||
|
|
||||||
### Explain Mode
|
### Explain Mode
|
||||||
|
|
||||||
Understand how your query was executed:
|
Diagnose how a query's `where` fields map to the index. Run `brain.explain()`
|
||||||
|
first whenever `find()` returns surprising or empty results:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const results = await brain.find({
|
const plan = await brain.explain({
|
||||||
like: "quantum computing",
|
query: "quantum computing",
|
||||||
where: { category: "research" },
|
where: { category: "research" }
|
||||||
explain: true
|
|
||||||
})
|
})
|
||||||
|
|
||||||
console.log(results[0].explanation)
|
console.log(plan.fieldPlan)
|
||||||
// {
|
// [
|
||||||
// plan: "field-first-progressive",
|
// { field: 'category', path: 'column-store', notes: '...' }
|
||||||
// timing: {
|
// ]
|
||||||
// fieldFilter: 2,
|
|
||||||
// vectorSearch: 8,
|
console.log(plan.warnings)
|
||||||
// fusion: 1
|
// e.g. ['Field "category" has no index entries. find() will return [] silently...']
|
||||||
// },
|
|
||||||
// selectivity: {
|
|
||||||
// field: 0.1,
|
|
||||||
// vector: 0.3
|
|
||||||
// }
|
|
||||||
// }
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Boosting
|
### Result Ordering
|
||||||
|
|
||||||
Apply custom ranking boosts:
|
Sort results by any stored field with `orderBy` / `order`:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
like: "news articles",
|
query: "news articles",
|
||||||
boost: 'recent', // Boost recent items
|
where: { verified: true },
|
||||||
where: { verified: true }
|
orderBy: 'createdAt', // Newest first
|
||||||
|
order: 'desc'
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
||||||
### Threshold Control
|
### Similarity Threshold
|
||||||
|
|
||||||
Set minimum similarity thresholds:
|
Find the nearest neighbours of a known entity and keep only close matches with
|
||||||
|
`near`:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
like: "exact match needed",
|
near: { id: anchorId, threshold: 0.9 }, // Only results >= 0.9 similarity
|
||||||
threshold: 0.9, // Only very similar results
|
|
||||||
limit: 10
|
limit: 10
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
@ -270,7 +286,7 @@ const results = await brain.find({
|
||||||
```typescript
|
```typescript
|
||||||
// Find similar content with constraints
|
// Find similar content with constraints
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
like: query,
|
query: searchText,
|
||||||
where: {
|
where: {
|
||||||
status: 'published',
|
status: 'published',
|
||||||
language: 'en'
|
language: 'en'
|
||||||
|
|
@ -285,7 +301,7 @@ const results = await brain.find({
|
||||||
connected: {
|
connected: {
|
||||||
to: itemId,
|
to: itemId,
|
||||||
depth: 2,
|
depth: 2,
|
||||||
type: 'similar'
|
via: VerbType.RelatedTo
|
||||||
},
|
},
|
||||||
limit: 20
|
limit: 20
|
||||||
})
|
})
|
||||||
|
|
@ -296,10 +312,11 @@ const results = await brain.find({
|
||||||
// Recent items matching criteria
|
// Recent items matching criteria
|
||||||
const results = await brain.find({
|
const results = await brain.find({
|
||||||
where: {
|
where: {
|
||||||
timestamp: { $gte: Date.now() - 86400000 }
|
timestamp: { gte: Date.now() - 86400000 }
|
||||||
},
|
},
|
||||||
like: "trending topics",
|
query: "trending topics",
|
||||||
boost: 'recent'
|
orderBy: 'timestamp',
|
||||||
|
order: 'desc'
|
||||||
})
|
})
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,769 +1,157 @@
|
||||||
|
---
|
||||||
|
title: Zero Configuration
|
||||||
|
slug: concepts/zero-config
|
||||||
|
public: false
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 3
|
||||||
|
description: Brainy auto-detects storage, initializes embeddings, and builds indexes — no configuration required. Works in Node.js and Bun (server-only since 8.0).
|
||||||
|
next:
|
||||||
|
- getting-started/installation
|
||||||
|
- guides/storage-adapters
|
||||||
|
---
|
||||||
|
|
||||||
# Zero Configuration & Auto-Adaptation
|
# Zero Configuration & Auto-Adaptation
|
||||||
|
|
||||||
> **Current Status**: Basic zero-config is fully functional. Advanced auto-adaptation features are in development.
|
> **"Zero config by default, fully tunable when you need it."** Construct a
|
||||||
|
> `Brainy()` with no options and it picks sensible, environment-aware defaults.
|
||||||
|
> Every default below is overridable through the constructor — see the
|
||||||
|
> [API Reference](../api/README.md#configuration).
|
||||||
|
|
||||||
## Overview
|
## Overview
|
||||||
|
|
||||||
Brainy is designed with **"Zero Config by Default, Infinite Tunability"** philosophy. It automatically detects your environment, adapts to available resources, learns from usage patterns, and optimizes itself for your specific workload—all without any configuration.
|
Brainy 8.0 is server-only (Node.js 22+ / Bun). With no configuration it:
|
||||||
|
|
||||||
## Zero Configuration Magic
|
- selects a storage adapter from the runtime,
|
||||||
|
- initializes the embedding model (all-MiniLM-L6-v2, 384 dimensions),
|
||||||
|
- builds and maintains the metadata, graph, and vector indexes,
|
||||||
|
- sizes its caches and write buffers to the detected memory budget,
|
||||||
|
- chooses a persistence mode that matches the storage backend, and
|
||||||
|
- quiets its own logging when it detects a production environment.
|
||||||
|
|
||||||
### Instant Start
|
There is no public config-generation function — adaptation happens inside the
|
||||||
|
constructor and `init()`.
|
||||||
|
|
||||||
|
## Instant Start
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
import { BrainyData } from 'brainy'
|
import { Brainy } from '@soulcraft/brainy'
|
||||||
|
|
||||||
// That's it. No config needed.
|
// That's it. No config needed.
|
||||||
const brain = new BrainyData()
|
const brain = new Brainy()
|
||||||
await brain.init()
|
await brain.init()
|
||||||
|
|
||||||
// Brainy automatically:
|
await brain.add({ data: 'First entity', type: 'concept' })
|
||||||
// ✓ Detects environment (Node.js, Browser, Edge, Deno)
|
const results = await brain.find('first')
|
||||||
// ✓ Chooses optimal storage (FileSystem, OPFS, Memory)
|
|
||||||
// ✓ Downloads required models (if needed)
|
|
||||||
// ✓ Configures vector dimensions (384 optimal)
|
|
||||||
// ✓ Sets up indexing strategies
|
|
||||||
// ✓ Enables appropriate augmentations
|
|
||||||
// ✓ Configures caching layers
|
|
||||||
// ✓ Optimizes for your hardware
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Environment Detection ✅ Available
|
## What Auto-Adaptation Covers
|
||||||
|
|
||||||
Brainy automatically detects and adapts to your runtime:
|
### 1. Storage auto-detection
|
||||||
|
|
||||||
|
With no `storage` option, Brainy uses `type: 'auto'`:
|
||||||
|
|
||||||
|
- **Filesystem** when running on a runtime with a writable Node filesystem and a
|
||||||
|
resolvable root directory. This is the default for typical Node/Bun servers and
|
||||||
|
persists across restarts.
|
||||||
|
- **In-memory** otherwise (no filesystem access, or an explicit memory request).
|
||||||
|
Fast, zero I/O, discarded on process exit — ideal for tests and ephemeral
|
||||||
|
caches.
|
||||||
|
|
||||||
|
8.0 ships exactly two storage adapters — `memory` and `filesystem` — plus the
|
||||||
|
`auto` selector that resolves to one of them. See
|
||||||
|
[Storage Adapters](../concepts/storage-adapters.md) for the full contract.
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Brainy's environment detection
|
// Explicit override when you want a specific root
|
||||||
const environment = {
|
const brain = new Brainy({
|
||||||
// Runtime detection
|
storage: { type: 'filesystem', path: './brainy-data' }
|
||||||
isNode: typeof process !== 'undefined',
|
|
||||||
isBrowser: typeof window !== 'undefined',
|
|
||||||
isDeno: typeof Deno !== 'undefined',
|
|
||||||
isEdge: typeof EdgeRuntime !== 'undefined',
|
|
||||||
isWebWorker: typeof WorkerGlobalScope !== 'undefined',
|
|
||||||
|
|
||||||
// Capability detection
|
|
||||||
hasFileSystem: /* auto-detected */,
|
|
||||||
hasIndexedDB: /* auto-detected */,
|
|
||||||
hasOPFS: /* auto-detected */,
|
|
||||||
hasWebGPU: /* auto-detected */,
|
|
||||||
hasWASM: /* auto-detected */,
|
|
||||||
|
|
||||||
// Resource detection
|
|
||||||
cpuCores: /* auto-detected */,
|
|
||||||
memory: /* auto-detected */,
|
|
||||||
storage: /* auto-detected */
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Auto-Adaptive Storage ✅ Available
|
|
||||||
|
|
||||||
> **Current**: Brainy automatically selects the best storage adapter for your environment.
|
|
||||||
|
|
||||||
### Storage Selection Logic
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Brainy's intelligent storage selection
|
|
||||||
async function autoSelectStorage() {
|
|
||||||
// Server environments
|
|
||||||
if (environment.isNode) {
|
|
||||||
if (await hasWritePermission('./data')) {
|
|
||||||
return 'filesystem' // Best for servers
|
|
||||||
} else if (process.env.S3_BUCKET) {
|
|
||||||
return 's3' // Cloud deployment
|
|
||||||
} else {
|
|
||||||
return 'memory' // Fallback for restricted environments
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Browser environments
|
|
||||||
if (environment.isBrowser) {
|
|
||||||
if (await navigator.storage.estimate() > 1GB) {
|
|
||||||
return 'opfs' // Best for modern browsers
|
|
||||||
} else if (indexedDB) {
|
|
||||||
return 'indexeddb' // Fallback for older browsers
|
|
||||||
} else {
|
|
||||||
return 'memory' // In-memory for restricted contexts
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Edge environments
|
|
||||||
if (environment.isEdge) {
|
|
||||||
return 'kv' // Use edge KV stores (Cloudflare, Vercel)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Storage Migration
|
|
||||||
|
|
||||||
Brainy seamlessly migrates between storage types:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Start with memory storage (development)
|
|
||||||
const brain = new BrainyData() // Auto-selects memory
|
|
||||||
|
|
||||||
// Later, migrate to production storage
|
|
||||||
await brain.migrate({
|
|
||||||
to: 'filesystem',
|
|
||||||
path: './production-data'
|
|
||||||
})
|
})
|
||||||
// All data seamlessly transferred
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Learning & Optimization 🚧 Coming Soon
|
### 2. HNSW quality from the `recall` preset
|
||||||
|
|
||||||
> **Note**: These features are planned for Q2 2025. Currently, Brainy uses static optimizations.
|
Vector-index quality comes from a single preset rather than hand-tuned graph
|
||||||
|
parameters. `config.vector.recall` accepts `'fast'`, `'balanced'`, or
|
||||||
### Query Pattern Learning 🚧 Planned
|
`'accurate'` and defaults to `'balanced'`. The preset maps internally to the
|
||||||
|
HNSW construction and search parameters (`M` / `efConstruction` / `efSearch`),
|
||||||
Brainy learns from your query patterns and optimizes accordingly:
|
so you trade recall against latency with one knob instead of three.
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Brainy observes query patterns
|
const brain = new Brainy({
|
||||||
class QueryPatternLearner {
|
vector: { recall: 'fast' } // favor latency over recall
|
||||||
analyze(queries: Query[]) {
|
})
|
||||||
return {
|
|
||||||
// Frequency analysis
|
|
||||||
mostCommonFields: this.getTopFields(queries),
|
|
||||||
avgResultSize: this.getAvgSize(queries),
|
|
||||||
temporalPatterns: this.getTimePatterns(queries),
|
|
||||||
|
|
||||||
// Relationship analysis
|
|
||||||
commonTraversals: this.getGraphPatterns(queries),
|
|
||||||
typicalDepth: this.getAvgDepth(queries),
|
|
||||||
|
|
||||||
// Performance analysis
|
|
||||||
slowQueries: this.getSlowQueries(queries),
|
|
||||||
cacheability: this.getCacheability(queries)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Automatic optimizations based on learning:
|
|
||||||
// - Creates indexes for frequently queried fields
|
|
||||||
// - Pre-computes common graph traversals
|
|
||||||
// - Adjusts cache sizes based on working set
|
|
||||||
// - Optimizes vector search parameters
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Auto-Indexing 🚧 Planned
|
The default JS index is `JsHnswVectorIndex`. An optional native acceleration
|
||||||
|
provider (the `@soulcraft/cor` package) can replace it with a
|
||||||
|
higher-performing implementation; the public knobs stay the same. Quantization
|
||||||
|
and other index-internal acceleration are the native provider's concern, not a
|
||||||
|
Brainy configuration option.
|
||||||
|
|
||||||
Brainy automatically creates indexes based on usage:
|
### 3. Persistence mode follows the backend
|
||||||
|
|
||||||
|
`config.vector.persistMode` accepts `'immediate'` or `'deferred'`. Left unset,
|
||||||
|
Brainy chooses for you:
|
||||||
|
|
||||||
|
- **Immediate** on filesystem storage, so the index file stays in lock-step with
|
||||||
|
the data and survives a crash.
|
||||||
|
- **Deferred** on in-memory storage, where there is nothing durable to sync to,
|
||||||
|
so writes are batched for throughput.
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// No manual index configuration needed
|
const brain = new Brainy({
|
||||||
await brain.find({ where: { category: "tech" } }) // First query
|
vector: { persistMode: 'deferred' } // batch persistence for write-heavy loads
|
||||||
// Brainy notices 'category' field usage
|
})
|
||||||
|
|
||||||
await brain.find({ where: { category: "science" } }) // Second query
|
|
||||||
// Pattern detected - auto-creates category index
|
|
||||||
|
|
||||||
await brain.find({ where: { category: "tech" } }) // Third query
|
|
||||||
// Now using index - 100x faster!
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Adaptive Caching 🚧 Planned
|
### 4. Memory-aware cache and buffer sizing
|
||||||
|
|
||||||
Cache strategies adapt to your access patterns:
|
Brainy reads the container's memory budget — `CLOUD_RUN_MEMORY`, `MEMORY_LIMIT`,
|
||||||
|
or the cgroup memory limit when running in a container — and sizes its read
|
||||||
|
caches and write buffers to fit. On a small instance it stays conservative; on a
|
||||||
|
large one it uses more of the available headroom. Query-result limits are capped
|
||||||
|
against the same budget (roughly 25 KB per result) to keep a single oversized
|
||||||
|
query from exhausting memory.
|
||||||
|
|
||||||
|
You can pin the cache explicitly:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
class AdaptiveCache {
|
const brain = new Brainy({
|
||||||
async adapt(metrics: AccessMetrics) {
|
cache: { maxSize: 10000, ttl: 3_600_000 }
|
||||||
if (metrics.hitRate < 0.3) {
|
})
|
||||||
// Low hit rate - switch strategy
|
|
||||||
this.strategy = 'lfu' // Least Frequently Used
|
|
||||||
} else if (metrics.workingSet > this.size) {
|
|
||||||
// Working set too large - increase size
|
|
||||||
this.size = Math.min(metrics.workingSet * 1.5, maxMemory)
|
|
||||||
} else if (metrics.temporalLocality > 0.8) {
|
|
||||||
// High temporal locality - use time-based eviction
|
|
||||||
this.strategy = 'ttl'
|
|
||||||
this.ttl = metrics.avgAccessInterval * 2
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Performance Auto-Scaling 🚧 Coming Soon
|
### 5. Logging quiets in production
|
||||||
|
|
||||||
### Dynamic Batch Sizing
|
Brainy detects production-style environments (for example `NODE_ENV` set to a
|
||||||
|
non-development value) and reduces its own log verbosity automatically. This is
|
||||||
Brainy adjusts batch sizes based on system load:
|
logging-only behavior — it does not change indexing, storage, or query results.
|
||||||
|
|
||||||
```typescript
|
|
||||||
class DynamicBatcher {
|
|
||||||
calculateOptimalBatch() {
|
|
||||||
const cpuUsage = process.cpuUsage()
|
|
||||||
const memoryUsage = process.memoryUsage()
|
|
||||||
|
|
||||||
if (cpuUsage < 30 && memoryUsage < 50) {
|
|
||||||
return 1000 // System idle - large batches
|
|
||||||
} else if (cpuUsage < 60 && memoryUsage < 70) {
|
|
||||||
return 100 // Moderate load - medium batches
|
|
||||||
} else {
|
|
||||||
return 10 // High load - small batches
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Automatically applied during bulk operations
|
|
||||||
for (const item of millionItems) {
|
|
||||||
await brain.addNoun(item) // Internally batched optimally
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Memory Management
|
|
||||||
|
|
||||||
Automatic memory pressure handling:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class MemoryManager {
|
|
||||||
async handlePressure() {
|
|
||||||
const usage = process.memoryUsage()
|
|
||||||
const available = os.freemem()
|
|
||||||
|
|
||||||
if (available < 100 * 1024 * 1024) { // Less than 100MB free
|
|
||||||
// Emergency mode
|
|
||||||
await this.flushCaches()
|
|
||||||
await this.compactIndexes()
|
|
||||||
await this.offloadToDisk()
|
|
||||||
} else if (usage.heapUsed / usage.heapTotal > 0.9) {
|
|
||||||
// Preventive mode
|
|
||||||
await this.reduceCacheSizes()
|
|
||||||
await this.pauseBackgroundTasks()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Connection Pooling
|
|
||||||
|
|
||||||
Automatic connection management for storage backends:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class ConnectionPool {
|
|
||||||
async getOptimalPoolSize() {
|
|
||||||
// Adapts based on workload
|
|
||||||
const metrics = await this.getMetrics()
|
|
||||||
|
|
||||||
if (metrics.waitTime > 100) {
|
|
||||||
// Queries waiting - increase pool
|
|
||||||
this.size = Math.min(this.size * 1.5, this.maxSize)
|
|
||||||
} else if (metrics.idleConnections > this.size * 0.5) {
|
|
||||||
// Too many idle - decrease pool
|
|
||||||
this.size = Math.max(this.size * 0.7, this.minSize)
|
|
||||||
}
|
|
||||||
|
|
||||||
return this.size
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Model Auto-Selection
|
|
||||||
|
|
||||||
### Embedding Model Selection
|
|
||||||
|
|
||||||
Brainy chooses the best embedding model for your use case:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
async function autoSelectModel(data: Sample[]) {
|
|
||||||
const analysis = {
|
|
||||||
languages: detectLanguages(data),
|
|
||||||
domainSpecific: detectDomain(data),
|
|
||||||
averageLength: getAvgLength(data),
|
|
||||||
requiresMultilingual: languages.length > 1
|
|
||||||
}
|
|
||||||
|
|
||||||
if (analysis.requiresMultilingual) {
|
|
||||||
return 'multilingual-e5-base' // Handles 100+ languages
|
|
||||||
} else if (analysis.domainSpecific === 'code') {
|
|
||||||
return 'codebert-base' // Optimized for code
|
|
||||||
} else if (analysis.averageLength > 512) {
|
|
||||||
return 'all-mpnet-base-v2' // Better for long text
|
|
||||||
} else {
|
|
||||||
return 'all-MiniLM-L6-v2' // Fast and efficient default
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Model Downloading
|
|
||||||
|
|
||||||
Models are automatically downloaded when needed:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// First use - model auto-downloads
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init() // Downloads model if not cached
|
|
||||||
|
|
||||||
// Intelligent model caching
|
|
||||||
const modelCache = {
|
|
||||||
location: process.env.MODEL_CACHE || '~/.brainy/models',
|
|
||||||
maxSize: 5 * 1024 * 1024 * 1024, // 5GB max
|
|
||||||
strategy: 'lru', // Least recently used eviction
|
|
||||||
|
|
||||||
// CDN selection based on location
|
|
||||||
cdn: await selectFastestCDN([
|
|
||||||
'https://cdn.brainy.io',
|
|
||||||
'https://brainy.b-cdn.net',
|
|
||||||
'https://models.huggingface.co'
|
|
||||||
])
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Workload Detection
|
|
||||||
|
|
||||||
### Pattern Recognition
|
|
||||||
|
|
||||||
Brainy identifies your workload type and optimizes:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
enum WorkloadType {
|
|
||||||
OLTP = 'oltp', // Many small transactions
|
|
||||||
OLAP = 'olap', // Analytical queries
|
|
||||||
STREAMING = 'streaming', // Real-time ingestion
|
|
||||||
BATCH = 'batch', // Bulk processing
|
|
||||||
HYBRID = 'hybrid' // Mixed workload
|
|
||||||
}
|
|
||||||
|
|
||||||
class WorkloadDetector {
|
|
||||||
detect(metrics: OperationMetrics): WorkloadType {
|
|
||||||
if (metrics.writesPerSecond > 1000 && metrics.avgWriteSize < 1024) {
|
|
||||||
return WorkloadType.STREAMING
|
|
||||||
} else if (metrics.avgQueryComplexity > 0.8 && metrics.avgResultSize > 10000) {
|
|
||||||
return WorkloadType.OLAP
|
|
||||||
} else if (metrics.batchOperations > metrics.singleOperations) {
|
|
||||||
return WorkloadType.BATCH
|
|
||||||
} else if (metrics.writeReadRatio > 0.3 && metrics.writeReadRatio < 0.7) {
|
|
||||||
return WorkloadType.HYBRID
|
|
||||||
} else {
|
|
||||||
return WorkloadType.OLTP
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Optimization Strategies
|
|
||||||
|
|
||||||
Different optimizations for different workloads:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class WorkloadOptimizer {
|
|
||||||
optimize(workload: WorkloadType) {
|
|
||||||
switch (workload) {
|
|
||||||
case WorkloadType.STREAMING:
|
|
||||||
return {
|
|
||||||
entityRegistry: true, // Deduplication
|
|
||||||
batchSize: 1000,
|
|
||||||
walEnabled: true,
|
|
||||||
cacheSize: 'small',
|
|
||||||
indexStrategy: 'lazy'
|
|
||||||
}
|
|
||||||
|
|
||||||
case WorkloadType.OLAP:
|
|
||||||
return {
|
|
||||||
entityRegistry: false,
|
|
||||||
batchSize: 10000,
|
|
||||||
walEnabled: false,
|
|
||||||
cacheSize: 'large',
|
|
||||||
indexStrategy: 'eager',
|
|
||||||
parallelQueries: true
|
|
||||||
}
|
|
||||||
|
|
||||||
case WorkloadType.BATCH:
|
|
||||||
return {
|
|
||||||
entityRegistry: false,
|
|
||||||
batchSize: 50000,
|
|
||||||
walEnabled: false,
|
|
||||||
cacheSize: 'minimal',
|
|
||||||
indexStrategy: 'deferred'
|
|
||||||
}
|
|
||||||
|
|
||||||
default:
|
|
||||||
return this.defaultConfig
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Hardware Adaptation 🚧 Coming Soon
|
|
||||||
|
|
||||||
> **Note**: GPU acceleration and hardware optimization planned for Q3 2025.
|
|
||||||
|
|
||||||
### CPU Optimization
|
|
||||||
|
|
||||||
Adapts to available CPU resources:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class CPUAdapter {
|
|
||||||
async optimize() {
|
|
||||||
const cores = os.cpus().length
|
|
||||||
const type = os.cpus()[0].model
|
|
||||||
|
|
||||||
// Parallel processing based on cores
|
|
||||||
this.parallelism = Math.max(1, cores - 1) // Leave one core free
|
|
||||||
|
|
||||||
// SIMD detection for vector operations
|
|
||||||
if (type.includes('Intel') || type.includes('AMD')) {
|
|
||||||
this.enableSIMD = await checkSIMDSupport()
|
|
||||||
}
|
|
||||||
|
|
||||||
// Thread pool sizing
|
|
||||||
this.threadPoolSize = cores * 2 // Optimal for I/O bound
|
|
||||||
|
|
||||||
// Vector search optimization
|
|
||||||
if (cores >= 8) {
|
|
||||||
this.hnswConstruction = 200 // Higher quality index
|
|
||||||
this.hnswSearch = 100 // More accurate search
|
|
||||||
} else {
|
|
||||||
this.hnswConstruction = 100 // Balanced
|
|
||||||
this.hnswSearch = 50 // Faster search
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Memory Adaptation
|
|
||||||
|
|
||||||
Intelligent memory allocation:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class MemoryAdapter {
|
|
||||||
async configure() {
|
|
||||||
const totalMemory = os.totalmem()
|
|
||||||
const availableMemory = os.freemem()
|
|
||||||
|
|
||||||
// Allocate based on available memory
|
|
||||||
const allocation = {
|
|
||||||
cache: Math.min(availableMemory * 0.25, 2 * GB),
|
|
||||||
vectors: Math.min(availableMemory * 0.30, 4 * GB),
|
|
||||||
indexes: Math.min(availableMemory * 0.20, 2 * GB),
|
|
||||||
working: Math.min(availableMemory * 0.25, 2 * GB)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Adjust for low memory systems
|
|
||||||
if (totalMemory < 4 * GB) {
|
|
||||||
allocation.cache *= 0.5
|
|
||||||
allocation.vectors *= 0.7
|
|
||||||
this.enableSwapping = true
|
|
||||||
}
|
|
||||||
|
|
||||||
return allocation
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### GPU Acceleration
|
|
||||||
|
|
||||||
Automatic GPU detection and utilization:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class GPUAdapter {
|
|
||||||
async detect() {
|
|
||||||
// WebGPU in browsers
|
|
||||||
if (navigator?.gpu) {
|
|
||||||
const adapter = await navigator.gpu.requestAdapter()
|
|
||||||
return {
|
|
||||||
available: true,
|
|
||||||
type: 'webgpu',
|
|
||||||
memory: adapter.limits.maxBufferSize,
|
|
||||||
compute: adapter.limits.maxComputeWorkgroupsPerDimension
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// CUDA in Node.js
|
|
||||||
if (process.platform === 'linux' || process.platform === 'win32') {
|
|
||||||
const hasCuda = await checkCudaSupport()
|
|
||||||
if (hasCuda) {
|
|
||||||
return {
|
|
||||||
available: true,
|
|
||||||
type: 'cuda',
|
|
||||||
memory: await getCudaMemory(),
|
|
||||||
compute: await getCudaCores()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
return { available: false }
|
|
||||||
}
|
|
||||||
|
|
||||||
async optimize(gpu: GPUInfo) {
|
|
||||||
if (gpu.available) {
|
|
||||||
// Offload vector operations to GPU
|
|
||||||
this.vectorOps = 'gpu'
|
|
||||||
this.embeddingGeneration = 'gpu'
|
|
||||||
this.matrixMultiplication = 'gpu'
|
|
||||||
|
|
||||||
// Larger batch sizes for GPU
|
|
||||||
this.batchSize = gpu.memory > 8 * GB ? 10000 : 1000
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Network Adaptation
|
|
||||||
|
|
||||||
### Bandwidth Detection
|
|
||||||
|
|
||||||
Optimizes for available network bandwidth:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class NetworkAdapter {
|
|
||||||
async measureBandwidth() {
|
|
||||||
const testSize = 1 * MB
|
|
||||||
const start = Date.now()
|
|
||||||
await this.transfer(testSize)
|
|
||||||
const duration = Date.now() - start
|
|
||||||
|
|
||||||
const bandwidth = (testSize / duration) * 1000 // bytes/sec
|
|
||||||
|
|
||||||
if (bandwidth < 1 * MB) {
|
|
||||||
// Low bandwidth - optimize
|
|
||||||
this.compression = 'aggressive'
|
|
||||||
this.batchTransfers = true
|
|
||||||
this.cacheRemote = true
|
|
||||||
} else if (bandwidth > 100 * MB) {
|
|
||||||
// High bandwidth
|
|
||||||
this.compression = 'minimal'
|
|
||||||
this.parallelTransfers = true
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Latency Optimization
|
|
||||||
|
|
||||||
Adapts to network latency:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class LatencyOptimizer {
|
|
||||||
async optimize() {
|
|
||||||
const latency = await this.measureLatency()
|
|
||||||
|
|
||||||
if (latency > 100) { // High latency
|
|
||||||
// Batch operations
|
|
||||||
this.minBatchSize = 100
|
|
||||||
|
|
||||||
// Aggressive prefetching
|
|
||||||
this.prefetchDepth = 3
|
|
||||||
|
|
||||||
// Local caching
|
|
||||||
this.cacheStrategy = 'aggressive'
|
|
||||||
|
|
||||||
// Connection pooling
|
|
||||||
this.connectionPool = Math.min(latency / 10, 50)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Cloud Provider Detection 🚧 Coming Soon
|
|
||||||
|
|
||||||
> **Note**: Cloud provider auto-detection planned for Q3 2025.
|
|
||||||
|
|
||||||
### Automatic Cloud Optimization
|
|
||||||
|
|
||||||
Detects and optimizes for cloud providers:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class CloudDetector {
|
|
||||||
async detect() {
|
|
||||||
// AWS Detection
|
|
||||||
if (process.env.AWS_REGION || await canReachMetadata('169.254.169.254')) {
|
|
||||||
return {
|
|
||||||
provider: 'aws',
|
|
||||||
instance: await getEC2InstanceType(),
|
|
||||||
region: process.env.AWS_REGION,
|
|
||||||
services: {
|
|
||||||
storage: 's3',
|
|
||||||
cache: 'elasticache',
|
|
||||||
compute: 'lambda'
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Google Cloud Detection
|
|
||||||
if (process.env.GOOGLE_CLOUD_PROJECT || await canReachMetadata('metadata.google.internal')) {
|
|
||||||
return {
|
|
||||||
provider: 'gcp',
|
|
||||||
instance: await getGCEInstanceType(),
|
|
||||||
region: process.env.GOOGLE_CLOUD_REGION,
|
|
||||||
services: {
|
|
||||||
storage: 'gcs',
|
|
||||||
cache: 'memorystore',
|
|
||||||
compute: 'cloud-run'
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Vercel Edge Detection
|
|
||||||
if (process.env.VERCEL) {
|
|
||||||
return {
|
|
||||||
provider: 'vercel',
|
|
||||||
region: process.env.VERCEL_REGION,
|
|
||||||
services: {
|
|
||||||
storage: 'vercel-kv',
|
|
||||||
cache: 'edge-config',
|
|
||||||
compute: 'edge-runtime'
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Development vs Production
|
|
||||||
|
|
||||||
### Automatic Environment Detection
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class EnvironmentDetector {
|
|
||||||
detect() {
|
|
||||||
const indicators = {
|
|
||||||
// Development indicators
|
|
||||||
isDevelopment:
|
|
||||||
process.env.NODE_ENV === 'development' ||
|
|
||||||
process.env.DEBUG ||
|
|
||||||
process.argv.includes('--dev') ||
|
|
||||||
isLocalhost() ||
|
|
||||||
hasDevTools(),
|
|
||||||
|
|
||||||
// Test indicators
|
|
||||||
isTest:
|
|
||||||
process.env.NODE_ENV === 'test' ||
|
|
||||||
process.env.CI ||
|
|
||||||
isTestRunner(),
|
|
||||||
|
|
||||||
// Production indicators
|
|
||||||
isProduction:
|
|
||||||
process.env.NODE_ENV === 'production' ||
|
|
||||||
process.env.VERCEL ||
|
|
||||||
process.env.NETLIFY ||
|
|
||||||
!isLocalhost()
|
|
||||||
}
|
|
||||||
|
|
||||||
return indicators
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Different defaults for different environments
|
|
||||||
const config = environment.isProduction ? {
|
|
||||||
storage: 'filesystem',
|
|
||||||
wal: true,
|
|
||||||
monitoring: true,
|
|
||||||
compression: true,
|
|
||||||
caching: 'aggressive'
|
|
||||||
} : {
|
|
||||||
storage: 'memory',
|
|
||||||
wal: false,
|
|
||||||
monitoring: false,
|
|
||||||
compression: false,
|
|
||||||
caching: 'minimal'
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Error Recovery
|
|
||||||
|
|
||||||
### Automatic Fallbacks
|
|
||||||
|
|
||||||
Brainy automatically recovers from errors:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class AutoRecovery {
|
|
||||||
async handleStorageFailure() {
|
|
||||||
try {
|
|
||||||
await this.primaryStorage.write(data)
|
|
||||||
} catch (error) {
|
|
||||||
console.warn('Primary storage failed, trying fallback')
|
|
||||||
|
|
||||||
// Try secondary storage
|
|
||||||
if (this.secondaryStorage) {
|
|
||||||
await this.secondaryStorage.write(data)
|
|
||||||
} else {
|
|
||||||
// Fall back to memory
|
|
||||||
await this.memoryStorage.write(data)
|
|
||||||
|
|
||||||
// Schedule retry
|
|
||||||
this.scheduleRetry(data)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
async handleModelFailure() {
|
|
||||||
try {
|
|
||||||
return await this.primaryModel.embed(text)
|
|
||||||
} catch (error) {
|
|
||||||
// Fall back to simpler model
|
|
||||||
return await this.fallbackModel.embed(text)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Configuration Override
|
## Configuration Override
|
||||||
|
|
||||||
While zero-config is default, you can override when needed:
|
Zero-config is the default, not a ceiling. Every adaptive decision above has an
|
||||||
|
explicit constructor option:
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Explicit configuration when needed
|
const brain = new Brainy({
|
||||||
const brain = new BrainyData({
|
storage: { type: 'filesystem', path: '/var/lib/brainy' },
|
||||||
// Override auto-detection
|
vector: {
|
||||||
storage: {
|
recall: 'accurate',
|
||||||
type: 'filesystem',
|
persistMode: 'immediate'
|
||||||
path: '/custom/path'
|
|
||||||
},
|
},
|
||||||
|
cache: { maxSize: 50000, ttl: 600_000 }
|
||||||
// Override auto-optimization
|
|
||||||
optimization: {
|
|
||||||
autoIndex: false,
|
|
||||||
autoCache: false,
|
|
||||||
autoBatch: false
|
|
||||||
},
|
|
||||||
|
|
||||||
// Override auto-scaling
|
|
||||||
scaling: {
|
|
||||||
maxMemory: 2 * GB,
|
|
||||||
maxConnections: 100,
|
|
||||||
maxBatchSize: 1000
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## Monitoring Auto-Adaptation
|
|
||||||
|
|
||||||
Brainy provides visibility into its auto-adaptation:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
brain.on('adaptation', (event) => {
|
|
||||||
console.log(`Brainy adapted: ${event.type}`)
|
|
||||||
console.log(`Reason: ${event.reason}`)
|
|
||||||
console.log(`Before: ${JSON.stringify(event.before)}`)
|
|
||||||
console.log(`After: ${JSON.stringify(event.after)}`)
|
|
||||||
})
|
})
|
||||||
|
|
||||||
// Example events:
|
await brain.init()
|
||||||
// - Index created for frequently queried field
|
|
||||||
// - Cache strategy changed due to low hit rate
|
|
||||||
// - Batch size increased due to high throughput
|
|
||||||
// - Storage migrated due to space constraints
|
|
||||||
// - Model switched due to multilingual content
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Conclusion
|
See the [API Reference](../api/README.md#configuration) for the complete option
|
||||||
|
list.
|
||||||
Brainy's zero-configuration and auto-adaptation capabilities mean you can focus on your application logic while Brainy handles:
|
|
||||||
|
|
||||||
- Environment detection and optimization
|
|
||||||
- Storage selection and migration
|
|
||||||
- Performance tuning and scaling
|
|
||||||
- Resource management
|
|
||||||
- Error recovery
|
|
||||||
- Workload optimization
|
|
||||||
|
|
||||||
Just create a Brainy instance and start using it. Brainy will learn, adapt, and optimize itself for your specific use case—no configuration required.
|
|
||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Architecture Overview](./overview.md)
|
- [Architecture Overview](./overview.md)
|
||||||
- [Storage Architecture](./storage.md)
|
- [Storage Adapters](../concepts/storage-adapters.md)
|
||||||
- [Performance Guide](../guides/performance.md)
|
- [Scaling Guide](../SCALING.md)
|
||||||
- [Augmentations System](./augmentations.md)
|
- [API Reference](../api/README.md)
|
||||||
|
|
|
||||||
|
|
@ -1,549 +0,0 @@
|
||||||
# 🌐 Brainy API Exposure Architecture
|
|
||||||
|
|
||||||
## 📡 Current State: Built-in MCP Support
|
|
||||||
|
|
||||||
Brainy **already has** Model Context Protocol (MCP) support built-in:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Already exists in Brainy!
|
|
||||||
import { BrainyMCPService } from 'brainy/mcp'
|
|
||||||
|
|
||||||
const brain = new BrainyData()
|
|
||||||
const mcpService = new BrainyMCPService(brain)
|
|
||||||
|
|
||||||
// Handles MCP requests
|
|
||||||
const response = await mcpService.handleRequest({
|
|
||||||
type: 'data_access',
|
|
||||||
operation: 'search',
|
|
||||||
parameters: { query: 'find documents' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### What's Already Built:
|
|
||||||
- **BrainyMCPAdapter** - Exposes data operations
|
|
||||||
- **MCPAugmentationToolset** - Exposes augmentations as MCP tools
|
|
||||||
- **BrainyMCPService** - Unified service layer
|
|
||||||
- **ServerSearchConduitAugmentation** - Connect to remote Brainy instances
|
|
||||||
|
|
||||||
## 🚀 The Missing Piece: API Server Augmentation
|
|
||||||
|
|
||||||
What's **NOT** built yet is a unified API server that exposes REST, WebSocket, and MCP over network. This SHOULD be an augmentation!
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
/**
|
|
||||||
* Universal API Server Augmentation
|
|
||||||
* Exposes Brainy through REST, WebSocket, MCP, and GraphQL
|
|
||||||
*/
|
|
||||||
export class APIServerAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'api-server'
|
|
||||||
readonly timing = 'after' as const
|
|
||||||
readonly operations = ['all'] as const // Monitor all operations
|
|
||||||
readonly priority = 5 // Low priority, runs last
|
|
||||||
|
|
||||||
private httpServer?: any
|
|
||||||
private wsServer?: any
|
|
||||||
private mcpService?: BrainyMCPService
|
|
||||||
private clients = new Set<any>()
|
|
||||||
private apiKeys = new Map<string, any>()
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
const config = this.context.config.apiServer || {}
|
|
||||||
|
|
||||||
if (!config.enabled) {
|
|
||||||
this.log('API Server disabled in config')
|
|
||||||
return
|
|
||||||
}
|
|
||||||
|
|
||||||
// Initialize MCP service
|
|
||||||
this.mcpService = new BrainyMCPService(this.context.brain, {
|
|
||||||
enableAuth: config.requireAuth
|
|
||||||
})
|
|
||||||
|
|
||||||
// Start servers based on environment
|
|
||||||
if (typeof process !== 'undefined' && process.versions?.node) {
|
|
||||||
await this.startNodeServers(config)
|
|
||||||
} else if (typeof Deno !== 'undefined') {
|
|
||||||
await this.startDenoServer(config)
|
|
||||||
} else if (typeof self !== 'undefined') {
|
|
||||||
await this.startServiceWorker(config)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async startNodeServers(config: any) {
|
|
||||||
const express = await import('express')
|
|
||||||
const { WebSocketServer } = await import('ws')
|
|
||||||
const cors = await import('cors')
|
|
||||||
|
|
||||||
const app = express.default()
|
|
||||||
|
|
||||||
// Middleware
|
|
||||||
app.use(cors.default(config.cors))
|
|
||||||
app.use(express.json())
|
|
||||||
app.use(this.authMiddleware.bind(this))
|
|
||||||
app.use(this.rateLimitMiddleware.bind(this))
|
|
||||||
|
|
||||||
// REST API Routes
|
|
||||||
this.setupRESTRoutes(app)
|
|
||||||
|
|
||||||
// Start HTTP server
|
|
||||||
this.httpServer = app.listen(config.port || 3000, () => {
|
|
||||||
this.log(`REST API listening on port ${config.port || 3000}`)
|
|
||||||
})
|
|
||||||
|
|
||||||
// WebSocket server for real-time
|
|
||||||
this.wsServer = new WebSocketServer({
|
|
||||||
server: this.httpServer,
|
|
||||||
path: '/ws'
|
|
||||||
})
|
|
||||||
|
|
||||||
this.setupWebSocketServer()
|
|
||||||
|
|
||||||
// MCP over WebSocket
|
|
||||||
this.setupMCPWebSocket()
|
|
||||||
}
|
|
||||||
|
|
||||||
private setupRESTRoutes(app: any) {
|
|
||||||
// Health check
|
|
||||||
app.get('/health', (req: any, res: any) => {
|
|
||||||
res.json({ status: 'healthy', version: '2.0.0' })
|
|
||||||
})
|
|
||||||
|
|
||||||
// Search endpoint
|
|
||||||
app.post('/api/search', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const { query, limit = 10, options = {} } = req.body
|
|
||||||
const results = await this.context.brain.search(query, limit, options)
|
|
||||||
res.json({ success: true, results })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Add data endpoint
|
|
||||||
app.post('/api/add', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const { content, metadata } = req.body
|
|
||||||
const id = await this.context.brain.add(content, metadata)
|
|
||||||
res.json({ success: true, id })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Get endpoint
|
|
||||||
app.get('/api/get/:id', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const data = await this.context.brain.get(req.params.id)
|
|
||||||
res.json({ success: true, data })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(404).json({
|
|
||||||
success: false,
|
|
||||||
error: 'Not found'
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Delete endpoint
|
|
||||||
app.delete('/api/delete/:id', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
await this.context.brain.delete(req.params.id)
|
|
||||||
res.json({ success: true })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Relate endpoint
|
|
||||||
app.post('/api/relate', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const { source, target, verb, metadata } = req.body
|
|
||||||
await this.context.brain.relate(source, target, verb, metadata)
|
|
||||||
res.json({ success: true })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Find endpoint (complex queries)
|
|
||||||
app.post('/api/find', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const results = await this.context.brain.find(req.body)
|
|
||||||
res.json({ success: true, results })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Cluster endpoint
|
|
||||||
app.post('/api/cluster', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const { algorithm = 'kmeans', options = {} } = req.body
|
|
||||||
const clusters = await this.context.brain.cluster(algorithm, options)
|
|
||||||
res.json({ success: true, clusters })
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// MCP endpoint (for non-WebSocket MCP)
|
|
||||||
app.post('/api/mcp', async (req: any, res: any) => {
|
|
||||||
try {
|
|
||||||
const response = await this.mcpService.handleRequest(req.body)
|
|
||||||
res.json(response)
|
|
||||||
} catch (error) {
|
|
||||||
res.status(500).json({
|
|
||||||
success: false,
|
|
||||||
error: error.message
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// GraphQL endpoint (optional)
|
|
||||||
if (this.context.config.apiServer?.enableGraphQL) {
|
|
||||||
this.setupGraphQL(app)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private setupWebSocketServer() {
|
|
||||||
this.wsServer.on('connection', (ws: any) => {
|
|
||||||
this.clients.add(ws)
|
|
||||||
|
|
||||||
ws.on('message', async (message: string) => {
|
|
||||||
try {
|
|
||||||
const msg = JSON.parse(message)
|
|
||||||
await this.handleWebSocketMessage(msg, ws)
|
|
||||||
} catch (error) {
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'error',
|
|
||||||
error: error.message
|
|
||||||
}))
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
ws.on('close', () => {
|
|
||||||
this.clients.delete(ws)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
private async handleWebSocketMessage(msg: any, ws: any) {
|
|
||||||
switch (msg.type) {
|
|
||||||
case 'subscribe':
|
|
||||||
// Subscribe to operations
|
|
||||||
ws.subscriptions = msg.operations || ['all']
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'subscribed',
|
|
||||||
operations: ws.subscriptions
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'search':
|
|
||||||
const results = await this.context.brain.search(msg.query, msg.limit)
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'searchResults',
|
|
||||||
results
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'mcp':
|
|
||||||
// Handle MCP over WebSocket
|
|
||||||
const response = await this.mcpService.handleRequest(msg.request)
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'mcpResponse',
|
|
||||||
response
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private setupMCPWebSocket() {
|
|
||||||
// Dedicated MCP WebSocket endpoint
|
|
||||||
const { WebSocketServer } = require('ws')
|
|
||||||
const mcpWs = new WebSocketServer({
|
|
||||||
port: (this.context.config.apiServer?.mcpPort || 3001),
|
|
||||||
path: '/mcp'
|
|
||||||
})
|
|
||||||
|
|
||||||
mcpWs.on('connection', (ws: any) => {
|
|
||||||
ws.on('message', async (message: string) => {
|
|
||||||
try {
|
|
||||||
const request = JSON.parse(message)
|
|
||||||
const response = await this.mcpService.handleRequest(request)
|
|
||||||
ws.send(JSON.stringify(response))
|
|
||||||
} catch (error) {
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
error: error.message,
|
|
||||||
type: 'error'
|
|
||||||
}))
|
|
||||||
}
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
this.log(`MCP WebSocket listening on port ${this.context.config.apiServer?.mcpPort || 3001}`)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Broadcast changes to all connected clients
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Broadcast to WebSocket clients
|
|
||||||
if (this.clients.size > 0) {
|
|
||||||
const message = JSON.stringify({
|
|
||||||
type: 'operation',
|
|
||||||
operation,
|
|
||||||
params: this.sanitizeParams(params),
|
|
||||||
timestamp: Date.now()
|
|
||||||
})
|
|
||||||
|
|
||||||
for (const client of this.clients) {
|
|
||||||
if (client.subscriptions?.includes('all') ||
|
|
||||||
client.subscriptions?.includes(operation)) {
|
|
||||||
client.send(message)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
private authMiddleware(req: any, res: any, next: any) {
|
|
||||||
if (!this.context.config.apiServer?.requireAuth) {
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
|
|
||||||
const apiKey = req.headers['x-api-key']
|
|
||||||
if (!apiKey || !this.apiKeys.has(apiKey)) {
|
|
||||||
return res.status(401).json({ error: 'Unauthorized' })
|
|
||||||
}
|
|
||||||
|
|
||||||
req.user = this.apiKeys.get(apiKey)
|
|
||||||
next()
|
|
||||||
}
|
|
||||||
|
|
||||||
private rateLimitMiddleware(req: any, res: any, next: any) {
|
|
||||||
// Simple rate limiting
|
|
||||||
const ip = req.ip
|
|
||||||
const limit = this.context.config.apiServer?.rateLimit || 100
|
|
||||||
|
|
||||||
// Implementation details...
|
|
||||||
next()
|
|
||||||
}
|
|
||||||
|
|
||||||
private sanitizeParams(params: any) {
|
|
||||||
// Remove sensitive data before broadcasting
|
|
||||||
const safe = { ...params }
|
|
||||||
delete safe.apiKey
|
|
||||||
delete safe.password
|
|
||||||
return safe
|
|
||||||
}
|
|
||||||
|
|
||||||
protected async onShutdown() {
|
|
||||||
// Close all connections
|
|
||||||
for (const client of this.clients) {
|
|
||||||
client.close()
|
|
||||||
}
|
|
||||||
|
|
||||||
// Close servers
|
|
||||||
if (this.httpServer) {
|
|
||||||
await new Promise(resolve => this.httpServer.close(resolve))
|
|
||||||
}
|
|
||||||
if (this.wsServer) {
|
|
||||||
this.wsServer.close()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Usage: Deploy Brainy as a Server
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
import { APIServerAugmentation } from 'brainy/augmentations'
|
|
||||||
|
|
||||||
const brain = new BrainyData({
|
|
||||||
apiServer: {
|
|
||||||
enabled: true,
|
|
||||||
port: 3000,
|
|
||||||
mcpPort: 3001,
|
|
||||||
requireAuth: true,
|
|
||||||
rateLimit: 100,
|
|
||||||
cors: { origin: '*' },
|
|
||||||
enableGraphQL: false
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Register the API server augmentation
|
|
||||||
brain.augmentations.register(new APIServerAugmentation())
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
console.log('Brainy API Server running!')
|
|
||||||
console.log('REST API: http://localhost:3000')
|
|
||||||
console.log('WebSocket: ws://localhost:3000/ws')
|
|
||||||
console.log('MCP: ws://localhost:3001/mcp')
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔌 Client Usage
|
|
||||||
|
|
||||||
### REST API
|
|
||||||
```javascript
|
|
||||||
// Search
|
|
||||||
const response = await fetch('http://localhost:3000/api/search', {
|
|
||||||
method: 'POST',
|
|
||||||
headers: {
|
|
||||||
'Content-Type': 'application/json',
|
|
||||||
'X-API-Key': 'your-key'
|
|
||||||
},
|
|
||||||
body: JSON.stringify({
|
|
||||||
query: 'find documents about AI',
|
|
||||||
limit: 10
|
|
||||||
})
|
|
||||||
})
|
|
||||||
const { results } = await response.json()
|
|
||||||
```
|
|
||||||
|
|
||||||
### WebSocket (Real-time)
|
|
||||||
```javascript
|
|
||||||
const ws = new WebSocket('ws://localhost:3000/ws')
|
|
||||||
|
|
||||||
ws.onopen = () => {
|
|
||||||
// Subscribe to operations
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'subscribe',
|
|
||||||
operations: ['add', 'delete', 'relate']
|
|
||||||
}))
|
|
||||||
}
|
|
||||||
|
|
||||||
ws.onmessage = (event) => {
|
|
||||||
const msg = JSON.parse(event.data)
|
|
||||||
console.log('Operation:', msg.operation, msg.params)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### MCP (AI Agents)
|
|
||||||
```javascript
|
|
||||||
const mcpWs = new WebSocket('ws://localhost:3001/mcp')
|
|
||||||
|
|
||||||
mcpWs.send(JSON.stringify({
|
|
||||||
type: 'data_access',
|
|
||||||
operation: 'search',
|
|
||||||
requestId: '123',
|
|
||||||
parameters: { query: 'test' }
|
|
||||||
}))
|
|
||||||
|
|
||||||
mcpWs.onmessage = (event) => {
|
|
||||||
const response = JSON.parse(event.data)
|
|
||||||
console.log('MCP Response:', response)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🏗️ Architecture Benefits
|
|
||||||
|
|
||||||
### Why as an Augmentation?
|
|
||||||
|
|
||||||
1. **Optional** - Not everyone needs a server
|
|
||||||
2. **Configurable** - Easy to enable/disable
|
|
||||||
3. **Extensible** - Add custom endpoints
|
|
||||||
4. **Integrated** - Hooks into all operations
|
|
||||||
5. **Real-time** - Broadcasts changes automatically
|
|
||||||
|
|
||||||
### Security Features
|
|
||||||
|
|
||||||
- **API Key Authentication**
|
|
||||||
- **Rate Limiting**
|
|
||||||
- **CORS Configuration**
|
|
||||||
- **Parameter Sanitization**
|
|
||||||
- **SSL/TLS Support** (with proper certs)
|
|
||||||
|
|
||||||
## 🌍 Deployment Options
|
|
||||||
|
|
||||||
### Local Development
|
|
||||||
```bash
|
|
||||||
npm install brainy
|
|
||||||
node server.js # Your server file with APIServerAugmentation
|
|
||||||
```
|
|
||||||
|
|
||||||
### Docker Container
|
|
||||||
```dockerfile
|
|
||||||
FROM node:18-alpine
|
|
||||||
WORKDIR /app
|
|
||||||
COPY package*.json ./
|
|
||||||
RUN npm install brainy
|
|
||||||
COPY server.js .
|
|
||||||
EXPOSE 3000 3001
|
|
||||||
CMD ["node", "server.js"]
|
|
||||||
```
|
|
||||||
|
|
||||||
### Cloud Deployment
|
|
||||||
Deploy to any Node.js hosting:
|
|
||||||
- Vercel Edge Functions
|
|
||||||
- Cloudflare Workers (with adapter)
|
|
||||||
- AWS Lambda (with adapter)
|
|
||||||
- Google Cloud Run
|
|
||||||
- Traditional VPS
|
|
||||||
|
|
||||||
### Browser Service Worker
|
|
||||||
```javascript
|
|
||||||
// In browser, use Service Worker for local API
|
|
||||||
if ('serviceWorker' in navigator) {
|
|
||||||
// APIServerAugmentation can create a Service Worker
|
|
||||||
// that intercepts fetch() calls and handles them locally
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 The Complete Picture
|
|
||||||
|
|
||||||
```
|
|
||||||
┌─────────────────────────────────────────────┐
|
|
||||||
│ Client Applications │
|
|
||||||
├─────────────┬───────────┬──────────────────┤
|
|
||||||
│ REST │ WebSocket │ MCP │
|
|
||||||
└──────┬──────┴─────┬─────┴──────┬───────────┘
|
|
||||||
│ │ │
|
|
||||||
▼ ▼ ▼
|
|
||||||
┌─────────────────────────────────────────────┐
|
|
||||||
│ APIServerAugmentation │
|
|
||||||
│ (Unified API exposure as augmentation) │
|
|
||||||
└─────────────────┬───────────────────────────┘
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌─────────────────────────────────────────────┐
|
|
||||||
│ BrainyData Core │
|
|
||||||
│ (with all augmentations in pipeline) │
|
|
||||||
└─────────────────────────────────────────────┘
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌─────────────────────────────────────────────┐
|
|
||||||
│ Storage Layer │
|
|
||||||
│ (FileSystem, S3, OPFS, Memory) │
|
|
||||||
└─────────────────────────────────────────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 Summary
|
|
||||||
|
|
||||||
1. **MCP is built-in** - Already in Brainy core
|
|
||||||
2. **API Server should be an augmentation** - Optional, configurable
|
|
||||||
3. **Exposes everything** - REST, WebSocket, MCP, GraphQL
|
|
||||||
4. **Real-time by default** - Broadcasts all operations
|
|
||||||
5. **Secure** - Auth, rate limiting, CORS
|
|
||||||
6. **Deploy anywhere** - Node, Deno, Browser, Cloud
|
|
||||||
|
|
||||||
The beauty is that the API server is just another augmentation - it hooks into the pipeline like everything else and exposes Brainy's capabilities to the world!
|
|
||||||
|
|
@ -1,774 +0,0 @@
|
||||||
# 🚀 Real-World Augmentation Examples
|
|
||||||
|
|
||||||
## 1. 💬 Chat Interface Augmentation
|
|
||||||
**"Talk to your data through natural language"**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from './brainyAugmentation.js'
|
|
||||||
|
|
||||||
export class ChatInterfaceAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'chat-interface'
|
|
||||||
readonly timing = 'after' as const // Process after operations
|
|
||||||
readonly operations = ['search', 'add', 'delete'] as const
|
|
||||||
readonly priority = 30 // Medium priority
|
|
||||||
|
|
||||||
private chatHistory: Array<{role: string, content: string}> = []
|
|
||||||
private llmClient: any // User's chosen LLM
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
// User provides their own LLM
|
|
||||||
this.llmClient = this.context.config.llmClient || null
|
|
||||||
if (!this.llmClient) {
|
|
||||||
this.log('Chat augmentation needs LLM client in config')
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
// If params include natural language query
|
|
||||||
if (params.chatQuery) {
|
|
||||||
// Convert natural language to Brainy operations
|
|
||||||
const intent = await this.parseIntent(params.chatQuery)
|
|
||||||
|
|
||||||
// Transform params based on intent
|
|
||||||
if (intent.type === 'search') {
|
|
||||||
params.query = intent.query
|
|
||||||
params.k = intent.limit || 10
|
|
||||||
} else if (intent.type === 'add') {
|
|
||||||
params.content = intent.content
|
|
||||||
params.metadata = { ...params.metadata, source: 'chat' }
|
|
||||||
}
|
|
||||||
|
|
||||||
// Store in chat history
|
|
||||||
this.chatHistory.push({
|
|
||||||
role: 'user',
|
|
||||||
content: params.chatQuery
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
// Execute the operation
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Generate conversational response
|
|
||||||
if (params.chatQuery && this.llmClient) {
|
|
||||||
const response = await this.generateResponse(operation, result)
|
|
||||||
|
|
||||||
this.chatHistory.push({
|
|
||||||
role: 'assistant',
|
|
||||||
content: response
|
|
||||||
})
|
|
||||||
|
|
||||||
// Enhance result with chat response
|
|
||||||
return {
|
|
||||||
...result,
|
|
||||||
chatResponse: response,
|
|
||||||
chatHistory: this.chatHistory
|
|
||||||
} as T
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
private async parseIntent(query: string) {
|
|
||||||
// Use Brainy's NLP patterns + LLM to understand intent
|
|
||||||
const prompt = `Parse this query into a Brainy operation:
|
|
||||||
Query: ${query}
|
|
||||||
|
|
||||||
Return JSON with:
|
|
||||||
- type: 'search' | 'add' | 'delete' | 'relate'
|
|
||||||
- query: search terms or content
|
|
||||||
- filters: any metadata filters
|
|
||||||
- limit: number of results`
|
|
||||||
|
|
||||||
const response = await this.llmClient.complete(prompt)
|
|
||||||
return JSON.parse(response)
|
|
||||||
}
|
|
||||||
|
|
||||||
private async generateResponse(operation: string, result: any) {
|
|
||||||
const prompt = `Generate a friendly response for this operation:
|
|
||||||
Operation: ${operation}
|
|
||||||
Result: ${JSON.stringify(result).slice(0, 500)}
|
|
||||||
Chat History: ${JSON.stringify(this.chatHistory.slice(-3))}
|
|
||||||
|
|
||||||
Be conversational and helpful.`
|
|
||||||
|
|
||||||
return await this.llmClient.complete(prompt)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage:
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
new ChatInterfaceAugmentation()
|
|
||||||
],
|
|
||||||
llmClient: openai // Bring your own LLM
|
|
||||||
})
|
|
||||||
|
|
||||||
// Now you can chat!
|
|
||||||
const result = await brain.search({
|
|
||||||
chatQuery: "Show me all documents about project roadmap from last week"
|
|
||||||
})
|
|
||||||
console.log(result.chatResponse) // "I found 5 documents about the project roadmap..."
|
|
||||||
```
|
|
||||||
|
|
||||||
## 2. 🤖 MCP Agent Memory Augmentation
|
|
||||||
**"Provide persistent memory for AI agents through MCP"**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from './brainyAugmentation.js'
|
|
||||||
import { Server } from '@modelcontextprotocol/sdk'
|
|
||||||
|
|
||||||
export class MCPAgentMemoryAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'mcp-agent-memory'
|
|
||||||
readonly timing = 'around' as const // Wrap operations
|
|
||||||
readonly operations = ['all'] as const // Monitor everything
|
|
||||||
readonly priority = 70 // High priority
|
|
||||||
|
|
||||||
private mcpServer: Server
|
|
||||||
private agentSessions: Map<string, any> = new Map()
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
// Initialize MCP server
|
|
||||||
this.mcpServer = new Server({
|
|
||||||
name: 'brainy-memory',
|
|
||||||
version: '1.0.0'
|
|
||||||
})
|
|
||||||
|
|
||||||
// Register MCP tools for agents
|
|
||||||
this.mcpServer.setRequestHandler('tools/list', () => ({
|
|
||||||
tools: [
|
|
||||||
{
|
|
||||||
name: 'remember',
|
|
||||||
description: 'Store information in long-term memory',
|
|
||||||
inputSchema: {
|
|
||||||
type: 'object',
|
|
||||||
properties: {
|
|
||||||
content: { type: 'string' },
|
|
||||||
category: { type: 'string' },
|
|
||||||
importance: { type: 'number' }
|
|
||||||
}
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
|
||||||
name: 'recall',
|
|
||||||
description: 'Retrieve information from memory',
|
|
||||||
inputSchema: {
|
|
||||||
type: 'object',
|
|
||||||
properties: {
|
|
||||||
query: { type: 'string' },
|
|
||||||
category: { type: 'string' },
|
|
||||||
limit: { type: 'number' }
|
|
||||||
}
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
|
||||||
name: 'forget',
|
|
||||||
description: 'Remove information from memory',
|
|
||||||
inputSchema: {
|
|
||||||
type: 'object',
|
|
||||||
properties: {
|
|
||||||
query: { type: 'string' },
|
|
||||||
category: { type: 'string' }
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}))
|
|
||||||
|
|
||||||
// Handle tool calls from agents
|
|
||||||
this.mcpServer.setRequestHandler('tools/call', async (request) => {
|
|
||||||
const { name, arguments: args } = request.params
|
|
||||||
|
|
||||||
switch (name) {
|
|
||||||
case 'remember':
|
|
||||||
return await this.rememberForAgent(args)
|
|
||||||
case 'recall':
|
|
||||||
return await this.recallForAgent(args)
|
|
||||||
case 'forget':
|
|
||||||
return await this.forgetForAgent(args)
|
|
||||||
default:
|
|
||||||
throw new Error(`Unknown tool: ${name}`)
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Start MCP server
|
|
||||||
await this.mcpServer.connect(process.stdin, process.stdout)
|
|
||||||
this.log('MCP Agent Memory server started')
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
// Extract agent context if present
|
|
||||||
const agentId = params.metadata?._agentId || 'default'
|
|
||||||
const sessionId = params.metadata?._sessionId
|
|
||||||
|
|
||||||
// Track agent operations
|
|
||||||
if (agentId && sessionId) {
|
|
||||||
if (!this.agentSessions.has(sessionId)) {
|
|
||||||
this.agentSessions.set(sessionId, {
|
|
||||||
agentId,
|
|
||||||
startTime: Date.now(),
|
|
||||||
operations: []
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
const session = this.agentSessions.get(sessionId)
|
|
||||||
session.operations.push({
|
|
||||||
operation,
|
|
||||||
params: { ...params },
|
|
||||||
timestamp: Date.now()
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
// Execute with agent context
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Auto-remember important operations
|
|
||||||
if (operation === 'add' && agentId) {
|
|
||||||
await this.autoRemember(agentId, params, result)
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
private async rememberForAgent(args: any) {
|
|
||||||
// Store in Brainy with agent-specific metadata
|
|
||||||
const id = await this.context.brain.add(args.content, {
|
|
||||||
_agentMemory: true,
|
|
||||||
_agentId: args.agentId || 'default',
|
|
||||||
category: args.category,
|
|
||||||
importance: args.importance || 0.5,
|
|
||||||
timestamp: new Date().toISOString()
|
|
||||||
})
|
|
||||||
|
|
||||||
return {
|
|
||||||
content: [
|
|
||||||
{
|
|
||||||
type: 'text',
|
|
||||||
text: `Remembered with ID: ${id}`
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async recallForAgent(args: any) {
|
|
||||||
// Search agent's memories
|
|
||||||
const results = await this.context.brain.search(args.query, args.limit || 10, {
|
|
||||||
where: {
|
|
||||||
_agentMemory: true,
|
|
||||||
_agentId: args.agentId || 'default',
|
|
||||||
category: args.category
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
return {
|
|
||||||
content: [
|
|
||||||
{
|
|
||||||
type: 'text',
|
|
||||||
text: JSON.stringify(results, null, 2)
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async forgetForAgent(args: any) {
|
|
||||||
// Remove specific memories
|
|
||||||
const results = await this.context.brain.find({
|
|
||||||
where: {
|
|
||||||
_agentMemory: true,
|
|
||||||
_agentId: args.agentId || 'default',
|
|
||||||
category: args.category
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
for (const item of results) {
|
|
||||||
await this.context.brain.delete(item.id)
|
|
||||||
}
|
|
||||||
|
|
||||||
return {
|
|
||||||
content: [
|
|
||||||
{
|
|
||||||
type: 'text',
|
|
||||||
text: `Forgot ${results.length} memories`
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async autoRemember(agentId: string, params: any, result: any) {
|
|
||||||
// Automatically remember important information
|
|
||||||
if (params.metadata?.important) {
|
|
||||||
await this.context.brain.add(params.content, {
|
|
||||||
...params.metadata,
|
|
||||||
_agentMemory: true,
|
|
||||||
_agentId: agentId,
|
|
||||||
_autoRemembered: true,
|
|
||||||
_originalOperation: 'add',
|
|
||||||
_resultId: result
|
|
||||||
})
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage:
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
new MCPAgentMemoryAugmentation()
|
|
||||||
]
|
|
||||||
})
|
|
||||||
|
|
||||||
// Now AI agents can use Brainy as memory through MCP!
|
|
||||||
// Agents connect via MCP and use remember/recall/forget tools
|
|
||||||
```
|
|
||||||
|
|
||||||
## 3. 🌐 API Server Augmentation
|
|
||||||
**"Expose Brainy through REST, WebSocket, and MCP APIs"**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from './brainyAugmentation.js'
|
|
||||||
import { BrainyMCPService } from '../mcp/brainyMCPService.js'
|
|
||||||
|
|
||||||
export class APIServerAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'api-server'
|
|
||||||
readonly timing = 'after' as const
|
|
||||||
readonly operations = ['all'] as ('all')[]
|
|
||||||
readonly priority = 5 // Low priority, runs after other augmentations
|
|
||||||
|
|
||||||
private httpServer: any
|
|
||||||
private wsServer: any
|
|
||||||
private mcpService: BrainyMCPService
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
// Initialize MCP service
|
|
||||||
this.mcpService = new BrainyMCPService(this.context.brain)
|
|
||||||
|
|
||||||
// Start HTTP server with REST endpoints
|
|
||||||
await this.startHTTPServer()
|
|
||||||
|
|
||||||
// Start WebSocket server for real-time
|
|
||||||
await this.startWebSocketServer()
|
|
||||||
|
|
||||||
this.log(`API Server running on port ${this.config.port || 3000}`)
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Broadcast operation to WebSocket clients
|
|
||||||
this.broadcast({
|
|
||||||
type: 'operation',
|
|
||||||
operation,
|
|
||||||
params: this.sanitizeParams(params),
|
|
||||||
timestamp: Date.now()
|
|
||||||
})
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
private async startHTTPServer() {
|
|
||||||
// REST endpoints: /api/search, /api/add, /api/get/:id, etc.
|
|
||||||
// MCP endpoint: /api/mcp
|
|
||||||
// Health check: /health
|
|
||||||
}
|
|
||||||
|
|
||||||
private async startWebSocketServer() {
|
|
||||||
// WebSocket for real-time subscriptions
|
|
||||||
// Clients can subscribe to specific operations
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage:
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new APIServerAugmentation({ port: 3000 }))
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Now access Brainy via:
|
|
||||||
// - REST: http://localhost:3000/api/*
|
|
||||||
// - WebSocket: ws://localhost:3000/ws
|
|
||||||
// - MCP: http://localhost:3000/api/mcp
|
|
||||||
```
|
|
||||||
|
|
||||||
## 4. 📊 Graph Visualization Augmentation
|
|
||||||
**"Real-time graph visualization with clustering"**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from './brainyAugmentation.js'
|
|
||||||
import { WebSocketServer } from 'ws'
|
|
||||||
|
|
||||||
export class GraphVisualizationAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'graph-visualization'
|
|
||||||
readonly timing = 'after' as const
|
|
||||||
readonly operations = ['all'] as const // Monitor all changes
|
|
||||||
readonly priority = 20
|
|
||||||
|
|
||||||
private wsServer: WebSocketServer
|
|
||||||
private graphState: {
|
|
||||||
nodes: Map<string, any>
|
|
||||||
edges: Map<string, any>
|
|
||||||
clusters: Map<string, Set<string>>
|
|
||||||
}
|
|
||||||
private clients: Set<any> = new Set()
|
|
||||||
|
|
||||||
protected async onInitialize(): Promise<void> {
|
|
||||||
// Initialize WebSocket server for real-time updates
|
|
||||||
this.wsServer = new WebSocketServer({
|
|
||||||
port: this.context.config.visualizationPort || 8080
|
|
||||||
})
|
|
||||||
|
|
||||||
this.graphState = {
|
|
||||||
nodes: new Map(),
|
|
||||||
edges: new Map(),
|
|
||||||
clusters: new Map()
|
|
||||||
}
|
|
||||||
|
|
||||||
// Load initial graph state
|
|
||||||
await this.loadGraphState()
|
|
||||||
|
|
||||||
// Handle client connections
|
|
||||||
this.wsServer.on('connection', (ws) => {
|
|
||||||
this.clients.add(ws)
|
|
||||||
|
|
||||||
// Send initial state
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'init',
|
|
||||||
data: this.serializeGraphState()
|
|
||||||
}))
|
|
||||||
|
|
||||||
// Handle client messages
|
|
||||||
ws.on('message', async (message) => {
|
|
||||||
const msg = JSON.parse(message.toString())
|
|
||||||
await this.handleClientMessage(msg, ws)
|
|
||||||
})
|
|
||||||
|
|
||||||
ws.on('close', () => {
|
|
||||||
this.clients.delete(ws)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
|
|
||||||
// Start clustering in background
|
|
||||||
this.startClusteringWorker()
|
|
||||||
|
|
||||||
this.log('Graph visualization server started on port ' +
|
|
||||||
(this.context.config.visualizationPort || 8080))
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Update graph state based on operation
|
|
||||||
switch (operation) {
|
|
||||||
case 'add':
|
|
||||||
case 'addNoun':
|
|
||||||
await this.handleNodeAdded(result, params)
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'relate':
|
|
||||||
case 'addVerb':
|
|
||||||
await this.handleEdgeAdded(params)
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'delete':
|
|
||||||
await this.handleNodeDeleted(params)
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'search':
|
|
||||||
await this.handleSearchPerformed(params, result)
|
|
||||||
break
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
|
|
||||||
private async handleNodeAdded(id: string, data: any) {
|
|
||||||
// Add node to graph
|
|
||||||
const node = {
|
|
||||||
id,
|
|
||||||
label: data.content?.slice(0, 50) || id,
|
|
||||||
type: data.metadata?.type || 'default',
|
|
||||||
metadata: data.metadata,
|
|
||||||
position: this.calculatePosition(id),
|
|
||||||
clusterId: null
|
|
||||||
}
|
|
||||||
|
|
||||||
this.graphState.nodes.set(id, node)
|
|
||||||
|
|
||||||
// Broadcast to clients
|
|
||||||
this.broadcast({
|
|
||||||
type: 'nodeAdded',
|
|
||||||
data: node
|
|
||||||
})
|
|
||||||
|
|
||||||
// Trigger re-clustering
|
|
||||||
this.scheduleReClustering()
|
|
||||||
}
|
|
||||||
|
|
||||||
private async handleEdgeAdded(params: any) {
|
|
||||||
const edge = {
|
|
||||||
id: `${params.source}-${params.verb}-${params.target}`,
|
|
||||||
source: params.source,
|
|
||||||
target: params.target,
|
|
||||||
label: params.verb,
|
|
||||||
weight: params.weight || 1
|
|
||||||
}
|
|
||||||
|
|
||||||
this.graphState.edges.set(edge.id, edge)
|
|
||||||
|
|
||||||
this.broadcast({
|
|
||||||
type: 'edgeAdded',
|
|
||||||
data: edge
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
private async handleSearchPerformed(params: any, results: any) {
|
|
||||||
// Highlight search results in visualization
|
|
||||||
const highlightNodes = results.map((r: any) => r.id)
|
|
||||||
|
|
||||||
this.broadcast({
|
|
||||||
type: 'highlight',
|
|
||||||
data: {
|
|
||||||
nodes: highlightNodes,
|
|
||||||
query: params.query,
|
|
||||||
duration: 5000 // Highlight for 5 seconds
|
|
||||||
}
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
private async loadGraphState() {
|
|
||||||
// Load all nodes (nouns)
|
|
||||||
const nouns = await this.context.brain.getAllNouns()
|
|
||||||
for (const noun of nouns) {
|
|
||||||
this.graphState.nodes.set(noun.id, {
|
|
||||||
id: noun.id,
|
|
||||||
label: noun.content?.slice(0, 50) || noun.id,
|
|
||||||
type: noun.type,
|
|
||||||
metadata: noun.metadata,
|
|
||||||
position: this.calculatePosition(noun.id)
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
// Load all edges (verbs/relationships)
|
|
||||||
const verbs = await this.context.brain.getAllVerbs()
|
|
||||||
for (const verb of verbs) {
|
|
||||||
this.graphState.edges.set(verb.id, {
|
|
||||||
id: verb.id,
|
|
||||||
source: verb.source,
|
|
||||||
target: verb.target,
|
|
||||||
label: verb.type,
|
|
||||||
weight: verb.weight
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
// Initial clustering
|
|
||||||
await this.performClustering()
|
|
||||||
}
|
|
||||||
|
|
||||||
private async performClustering() {
|
|
||||||
// Use Brainy's clustering capabilities
|
|
||||||
const clusteringResult = await this.context.brain.cluster({
|
|
||||||
algorithm: 'hierarchical',
|
|
||||||
threshold: 0.7
|
|
||||||
})
|
|
||||||
|
|
||||||
// Update cluster state
|
|
||||||
this.graphState.clusters.clear()
|
|
||||||
for (const [clusterId, nodeIds] of Object.entries(clusteringResult)) {
|
|
||||||
this.graphState.clusters.set(clusterId, new Set(nodeIds as string[]))
|
|
||||||
|
|
||||||
// Update nodes with cluster IDs
|
|
||||||
for (const nodeId of nodeIds as string[]) {
|
|
||||||
const node = this.graphState.nodes.get(nodeId)
|
|
||||||
if (node) {
|
|
||||||
node.clusterId = clusterId
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Broadcast cluster update
|
|
||||||
this.broadcast({
|
|
||||||
type: 'clustersUpdated',
|
|
||||||
data: this.serializeClusters()
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
private startClusteringWorker() {
|
|
||||||
// Re-cluster periodically or when graph changes significantly
|
|
||||||
setInterval(async () => {
|
|
||||||
if (this.graphState.nodes.size > 0) {
|
|
||||||
await this.performClustering()
|
|
||||||
}
|
|
||||||
}, 30000) // Every 30 seconds
|
|
||||||
}
|
|
||||||
|
|
||||||
private scheduleReClustering = (() => {
|
|
||||||
let timeout: NodeJS.Timeout
|
|
||||||
return () => {
|
|
||||||
clearTimeout(timeout)
|
|
||||||
timeout = setTimeout(() => this.performClustering(), 5000)
|
|
||||||
}
|
|
||||||
})()
|
|
||||||
|
|
||||||
private calculatePosition(id: string) {
|
|
||||||
// Simple force-directed layout position
|
|
||||||
const hash = id.split('').reduce((a, b) => {
|
|
||||||
a = ((a << 5) - a) + b.charCodeAt(0)
|
|
||||||
return a & a
|
|
||||||
}, 0)
|
|
||||||
|
|
||||||
return {
|
|
||||||
x: (hash % 1000) - 500,
|
|
||||||
y: ((hash * 7) % 1000) - 500
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private broadcast(message: any) {
|
|
||||||
const data = JSON.stringify(message)
|
|
||||||
for (const client of this.clients) {
|
|
||||||
client.send(data)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async handleClientMessage(msg: any, ws: any) {
|
|
||||||
switch (msg.type) {
|
|
||||||
case 'requestClustering':
|
|
||||||
await this.performClustering()
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'search':
|
|
||||||
const results = await this.context.brain.search(msg.query)
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'searchResults',
|
|
||||||
data: results
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'getNodeDetails':
|
|
||||||
const node = await this.context.brain.get(msg.nodeId)
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'nodeDetails',
|
|
||||||
data: node
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
|
|
||||||
case 'expandNode':
|
|
||||||
const connections = await this.context.brain.getConnections(msg.nodeId)
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'nodeConnections',
|
|
||||||
data: connections
|
|
||||||
}))
|
|
||||||
break
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private serializeGraphState() {
|
|
||||||
return {
|
|
||||||
nodes: Array.from(this.graphState.nodes.values()),
|
|
||||||
edges: Array.from(this.graphState.edges.values()),
|
|
||||||
clusters: this.serializeClusters()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private serializeClusters() {
|
|
||||||
const clusters: any = {}
|
|
||||||
for (const [id, nodeIds] of this.graphState.clusters) {
|
|
||||||
clusters[id] = Array.from(nodeIds)
|
|
||||||
}
|
|
||||||
return clusters
|
|
||||||
}
|
|
||||||
|
|
||||||
protected async onShutdown() {
|
|
||||||
this.wsServer.close()
|
|
||||||
this.clients.clear()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage:
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
new GraphVisualizationAugmentation()
|
|
||||||
],
|
|
||||||
visualizationPort: 8080
|
|
||||||
})
|
|
||||||
|
|
||||||
// Now connect a web-based graph viz tool to ws://localhost:8080
|
|
||||||
// It receives real-time updates as data changes!
|
|
||||||
```
|
|
||||||
|
|
||||||
## 4. 🌐 Multi-Agent Team Coordination
|
|
||||||
**"Multiple AI agents sharing knowledge and coordinating tasks"**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
export class TeamCoordinationAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'team-coordination'
|
|
||||||
readonly timing = 'around' as const
|
|
||||||
readonly operations = ['all'] as const
|
|
||||||
readonly priority = 85
|
|
||||||
|
|
||||||
private agents: Map<string, AgentState> = new Map()
|
|
||||||
private tasks: Map<string, Task> = new Map()
|
|
||||||
private sharedMemory: Map<string, any> = new Map()
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
const agentId = params.metadata?._agentId
|
|
||||||
|
|
||||||
if (agentId) {
|
|
||||||
// Track agent activity
|
|
||||||
this.updateAgentState(agentId, operation, params)
|
|
||||||
|
|
||||||
// Check if operation needs coordination
|
|
||||||
if (await this.needsCoordination(operation, params)) {
|
|
||||||
return await this.coordinatedExecute(agentId, operation, params, next)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
|
|
||||||
private async coordinatedExecute<T>(
|
|
||||||
agentId: string,
|
|
||||||
operation: string,
|
|
||||||
params: any,
|
|
||||||
next: () => Promise<T>
|
|
||||||
): Promise<T> {
|
|
||||||
// Acquire distributed lock
|
|
||||||
const lockId = await this.acquireLock(operation, params)
|
|
||||||
|
|
||||||
try {
|
|
||||||
// Check shared memory for related work
|
|
||||||
const relatedWork = await this.findRelatedWork(params)
|
|
||||||
if (relatedWork) {
|
|
||||||
params.metadata._relatedWork = relatedWork
|
|
||||||
}
|
|
||||||
|
|
||||||
// Execute with team context
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Update shared memory
|
|
||||||
await this.updateSharedMemory(agentId, operation, params, result)
|
|
||||||
|
|
||||||
// Notify other agents
|
|
||||||
await this.notifyTeam(agentId, operation, result)
|
|
||||||
|
|
||||||
return result
|
|
||||||
} finally {
|
|
||||||
await this.releaseLock(lockId)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Key Patterns
|
|
||||||
|
|
||||||
All these augmentations follow the same pattern:
|
|
||||||
|
|
||||||
1. **Extend BaseAugmentation**
|
|
||||||
2. **Define timing & operations**
|
|
||||||
3. **Initialize resources** in `onInitialize()`
|
|
||||||
4. **Intercept operations** in `execute()`
|
|
||||||
5. **Clean up** in `onShutdown()`
|
|
||||||
|
|
||||||
They can:
|
|
||||||
- **Add APIs** (REST, WebSocket, MCP)
|
|
||||||
- **Transform data** (chat queries → operations)
|
|
||||||
- **Coordinate agents** (distributed locking, shared memory)
|
|
||||||
- **Visualize in real-time** (WebSocket broadcasts)
|
|
||||||
- **Integrate any service** (LLMs, databases, APIs)
|
|
||||||
|
|
||||||
The beauty is they all use the **same simple interface** but achieve vastly different goals!
|
|
||||||
|
|
@ -1,306 +0,0 @@
|
||||||
# 🔄 How Augmentations Hook Into Brainy
|
|
||||||
|
|
||||||
## The Complete Pipeline Architecture
|
|
||||||
|
|
||||||
```
|
|
||||||
User Code → BrainyData Method → Augmentation Pipeline → Storage/Operations
|
|
||||||
↑ ↓
|
|
||||||
└────── Augmentations Execute Here ──┘
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 How Augmentations Register & Execute
|
|
||||||
|
|
||||||
### 1. **Registration During Initialization**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// In BrainyData constructor/init
|
|
||||||
class BrainyData {
|
|
||||||
private augmentations = new AugmentationRegistry()
|
|
||||||
|
|
||||||
async init() {
|
|
||||||
// Register built-in augmentations in priority order
|
|
||||||
this.augmentations.register(new WALAugmentation()) // Priority: 100
|
|
||||||
this.augmentations.register(new EntityRegistryAugmentation()) // Priority: 90
|
|
||||||
this.augmentations.register(new NeuralImportAugmentation()) // Priority: 80
|
|
||||||
this.augmentations.register(new BatchProcessingAugmentation()) // Priority: 50
|
|
||||||
|
|
||||||
// Initialize all with context
|
|
||||||
const context: AugmentationContext = {
|
|
||||||
brain: this,
|
|
||||||
storage: this.storage,
|
|
||||||
config: this.config,
|
|
||||||
log: (msg, level) => console.log(msg)
|
|
||||||
}
|
|
||||||
|
|
||||||
await this.augmentations.initialize(context)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Execution Through Method Interception**
|
|
||||||
|
|
||||||
Every BrainyData operation wraps its core logic with augmentation execution:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Example: The add() method
|
|
||||||
async add(content: string, metadata?: any): Promise<string> {
|
|
||||||
// Augmentations wrap the core operation
|
|
||||||
return this.augmentations.execute(
|
|
||||||
'add', // Operation name
|
|
||||||
{ content, metadata }, // Parameters
|
|
||||||
async () => { // Core operation
|
|
||||||
// Actual add logic here
|
|
||||||
const id = generateId()
|
|
||||||
await this.storage.set(id, { content, metadata })
|
|
||||||
return id
|
|
||||||
}
|
|
||||||
)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **The Execution Chain**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// In AugmentationRegistry
|
|
||||||
async execute<T>(operation: string, params: any, mainOperation: () => Promise<T>): Promise<T> {
|
|
||||||
// 1. Filter augmentations that should run for this operation
|
|
||||||
const applicable = this.augmentations.filter(aug =>
|
|
||||||
aug.shouldExecute(operation, params)
|
|
||||||
)
|
|
||||||
|
|
||||||
// 2. Sort by priority (already sorted during registration)
|
|
||||||
// Priority 100 runs first, then 90, 80, etc.
|
|
||||||
|
|
||||||
// 3. Create middleware chain
|
|
||||||
let index = 0
|
|
||||||
const executeNext = async (): Promise<T> => {
|
|
||||||
if (index >= applicable.length) {
|
|
||||||
// All augmentations processed, run main operation
|
|
||||||
return mainOperation()
|
|
||||||
}
|
|
||||||
|
|
||||||
const augmentation = applicable[index++]
|
|
||||||
// Each augmentation decides what to do with the operation
|
|
||||||
return augmentation.execute(operation, params, executeNext)
|
|
||||||
}
|
|
||||||
|
|
||||||
return executeNext()
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎭 The Four Timing Modes in Action
|
|
||||||
|
|
||||||
### **`timing: 'before'`** - Pre-processing
|
|
||||||
```typescript
|
|
||||||
class NeuralImportAugmentation {
|
|
||||||
timing = 'before'
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Analyze data BEFORE storage
|
|
||||||
const analysis = await this.analyzeWithAI(params.content)
|
|
||||||
params.metadata._neural = analysis
|
|
||||||
|
|
||||||
// Continue with enhanced params
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### **`timing: 'after'`** - Post-processing
|
|
||||||
```typescript
|
|
||||||
class NotionSynapse {
|
|
||||||
timing = 'after'
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Let operation complete first
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Then sync to Notion
|
|
||||||
await this.syncToNotion(op, params, result)
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### **`timing: 'around'`** - Wrapping
|
|
||||||
```typescript
|
|
||||||
class WALAugmentation {
|
|
||||||
timing = 'around'
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Write to WAL before
|
|
||||||
await this.wal.write({ op, params, timestamp: Date.now() })
|
|
||||||
|
|
||||||
try {
|
|
||||||
// Execute operation
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Mark as committed
|
|
||||||
await this.wal.commit()
|
|
||||||
|
|
||||||
return result
|
|
||||||
} catch (error) {
|
|
||||||
// Rollback on failure
|
|
||||||
await this.wal.rollback()
|
|
||||||
throw error
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### **`timing: 'replace'`** - Complete replacement
|
|
||||||
```typescript
|
|
||||||
class S3StorageAugmentation {
|
|
||||||
timing = 'replace'
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
if (op === 'storage.get') {
|
|
||||||
// Don't call next() - completely replace
|
|
||||||
return await this.s3.getObject(params.key)
|
|
||||||
}
|
|
||||||
// For other operations, pass through
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📊 Real Example: How `brain.add()` Works
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// User calls:
|
|
||||||
await brain.add("John is a developer", { type: "person" })
|
|
||||||
|
|
||||||
// This triggers the chain:
|
|
||||||
|
|
||||||
1. BrainyData.add() calls augmentations.execute('add', params, coreLogic)
|
|
||||||
|
|
||||||
2. AugmentationRegistry filters applicable augmentations:
|
|
||||||
- WALAugmentation (priority: 100, operations: ['all'])
|
|
||||||
- EntityRegistryAugmentation (priority: 90, operations: ['add'])
|
|
||||||
- NeuralImportAugmentation (priority: 80, operations: ['add'])
|
|
||||||
- BatchProcessingAugmentation (priority: 50, operations: ['add'])
|
|
||||||
|
|
||||||
3. Execution chain (highest priority first):
|
|
||||||
|
|
||||||
WALAugmentation.execute() {
|
|
||||||
await wal.write(operation) // Log to WAL
|
|
||||||
const result = await next() // Call next in chain
|
|
||||||
await wal.commit() // Commit WAL
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
↓
|
|
||||||
EntityRegistryAugmentation.execute() {
|
|
||||||
const hash = computeHash(params.content)
|
|
||||||
if (registry.has(hash)) {
|
|
||||||
return registry.get(hash) // Return existing ID
|
|
||||||
}
|
|
||||||
const result = await next() // Continue chain
|
|
||||||
registry.set(hash, result) // Register new entity
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
↓
|
|
||||||
NeuralImportAugmentation.execute() {
|
|
||||||
const analysis = await analyzeWithAI(params)
|
|
||||||
params.metadata._neural = analysis // Add AI insights
|
|
||||||
return next() // Continue with enhanced data
|
|
||||||
}
|
|
||||||
↓
|
|
||||||
BatchProcessingAugmentation.execute() {
|
|
||||||
batch.add(params) // Add to batch
|
|
||||||
if (batch.isFull()) {
|
|
||||||
await batch.flush() // Process batch if full
|
|
||||||
}
|
|
||||||
return next() // Continue
|
|
||||||
}
|
|
||||||
↓
|
|
||||||
Core add() logic {
|
|
||||||
// Finally, the actual storage operation
|
|
||||||
const id = generateId()
|
|
||||||
await storage.set(id, params)
|
|
||||||
await index.add(id, vector)
|
|
||||||
return id
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔌 Dynamic Registration
|
|
||||||
|
|
||||||
Augmentations can be registered at any time:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// During initialization
|
|
||||||
brain.augmentations.register(new CustomAugmentation())
|
|
||||||
|
|
||||||
// Or later, dynamically
|
|
||||||
const synapse = new NotionSynapse({ apiKey: 'xxx' })
|
|
||||||
brain.augmentations.register(synapse)
|
|
||||||
|
|
||||||
// From Brain Cloud marketplace
|
|
||||||
import { EmotionalIntelligence } from '@brain-cloud/empathy'
|
|
||||||
brain.augmentations.register(new EmotionalIntelligence())
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Operation Targeting
|
|
||||||
|
|
||||||
Augmentations declare which operations they care about:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class SearchOptimizer {
|
|
||||||
operations = ['search', 'searchText', 'findSimilar'] // Only search ops
|
|
||||||
}
|
|
||||||
|
|
||||||
class GlobalLogger {
|
|
||||||
operations = ['all'] // Every operation
|
|
||||||
}
|
|
||||||
|
|
||||||
class StorageReplacer {
|
|
||||||
operations = ['storage'] // Storage operations only
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔍 On-Demand Execution
|
|
||||||
|
|
||||||
Some augmentations can be triggered manually:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Get specific augmentation
|
|
||||||
const neuralImport = brain.augmentations.get('neural-import')
|
|
||||||
|
|
||||||
// Use its public API directly
|
|
||||||
const analysis = await neuralImport.getNeuralAnalysis(data, 'json')
|
|
||||||
|
|
||||||
// Or trigger through operations
|
|
||||||
await brain.add(data) // Automatically uses neural import if registered
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📈 Priority System
|
|
||||||
|
|
||||||
```
|
|
||||||
100: Critical Infrastructure (WAL, Transactions)
|
|
||||||
90: Data Integrity (Entity Registry, Deduplication)
|
|
||||||
80: Data Processing (Neural Import, Transformation)
|
|
||||||
50: Performance (Batching, Caching)
|
|
||||||
10: Features (Scoring, Analytics)
|
|
||||||
1: Monitoring (Logging, Metrics)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🌊 The Flow
|
|
||||||
|
|
||||||
1. **User Action** → `brain.add()`, `brain.search()`, etc.
|
|
||||||
2. **Method Wraps** → Core logic wrapped with `augmentations.execute()`
|
|
||||||
3. **Filter** → Find augmentations for this operation
|
|
||||||
4. **Sort** → Order by priority
|
|
||||||
5. **Chain** → Each augmentation calls next() or not
|
|
||||||
6. **Core** → Eventually hits actual implementation
|
|
||||||
7. **Unwind** → Results flow back through chain
|
|
||||||
8. **Return** → Enhanced result to user
|
|
||||||
|
|
||||||
## 💡 Key Insights
|
|
||||||
|
|
||||||
1. **Everything is interceptable** - All operations go through the pipeline
|
|
||||||
2. **Augmentations compose** - They stack like middleware
|
|
||||||
3. **Priority matters** - Higher priority runs first
|
|
||||||
4. **Timing is flexible** - before/after/around/replace covers all needs
|
|
||||||
5. **Simple but powerful** - One interface, infinite possibilities
|
|
||||||
|
|
||||||
This is why the single `BrainyAugmentation` interface works for EVERYTHING - it's just middleware with superpowers! 🚀
|
|
||||||
|
|
@ -1,288 +0,0 @@
|
||||||
# Simple Guide: Creating Augmentations
|
|
||||||
|
|
||||||
## The One Interface That Rules Them All
|
|
||||||
|
|
||||||
**EVERY augmentation is a `BrainyAugmentation`:**
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface BrainyAugmentation {
|
|
||||||
name: string // Unique name
|
|
||||||
timing: 'before' | 'after' | 'around' | 'replace' // When to run
|
|
||||||
operations: string[] // What to intercept
|
|
||||||
priority: number // Order (higher = first)
|
|
||||||
|
|
||||||
initialize(context): Promise<void> // Setup
|
|
||||||
execute(op, params, next): Promise<T> // Do work
|
|
||||||
shutdown?(): Promise<void> // Cleanup (optional)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
That's it! Every augmentation implements this interface.
|
|
||||||
|
|
||||||
## Creating Different Types of Augmentations
|
|
||||||
|
|
||||||
### 1. Basic Feature Augmentation
|
|
||||||
|
|
||||||
**Use Case:** Add logging, caching, validation, etc.
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from 'brainy'
|
|
||||||
|
|
||||||
export class LoggingAugmentation extends BaseAugmentation {
|
|
||||||
name = 'logging'
|
|
||||||
timing = 'around' // Wrap operations
|
|
||||||
operations = ['add', 'delete'] // What to log
|
|
||||||
priority = 10 // Low priority
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
console.log(`Starting ${op}`)
|
|
||||||
const result = await next()
|
|
||||||
console.log(`Completed ${op}`)
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage
|
|
||||||
brain.augmentations.register(new LoggingAugmentation())
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Storage Augmentation
|
|
||||||
|
|
||||||
**Use Case:** Provide a storage backend (special: has `provideStorage()` method)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { StorageAugmentation } from 'brainy'
|
|
||||||
|
|
||||||
export class RedisStorageAugmentation extends StorageAugmentation {
|
|
||||||
constructor(config) {
|
|
||||||
super('redis-storage') // Pass name to parent
|
|
||||||
this.config = config
|
|
||||||
}
|
|
||||||
|
|
||||||
// Special method for storage only!
|
|
||||||
async provideStorage() {
|
|
||||||
return new RedisAdapter(this.config)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage (BEFORE init!)
|
|
||||||
brain.augmentations.register(new RedisStorageAugmentation({
|
|
||||||
host: 'localhost',
|
|
||||||
port: 6379
|
|
||||||
}))
|
|
||||||
await brain.init() // Will use Redis!
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Data Processing Augmentation
|
|
||||||
|
|
||||||
**Use Case:** Transform or validate data before storage
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
export class ValidationAugmentation extends BaseAugmentation {
|
|
||||||
name = 'validator'
|
|
||||||
timing = 'before' // Run before operation
|
|
||||||
operations = ['add'] // Validate on add
|
|
||||||
priority = 50
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Validate data
|
|
||||||
if (!params.data || !params.data.title) {
|
|
||||||
throw new Error('Title is required')
|
|
||||||
}
|
|
||||||
|
|
||||||
// Add timestamp
|
|
||||||
params.data.createdAt = new Date()
|
|
||||||
|
|
||||||
// Continue with modified params
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. External System Augmentation (Synapse)
|
|
||||||
|
|
||||||
**Use Case:** Sync with external systems like Notion, Slack, etc.
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
export class NotionSyncAugmentation extends BaseAugmentation {
|
|
||||||
name = 'notion-sync'
|
|
||||||
timing = 'after' // Sync after local operation
|
|
||||||
operations = ['add', 'update', 'delete']
|
|
||||||
priority = 30
|
|
||||||
|
|
||||||
private notion: NotionClient
|
|
||||||
|
|
||||||
async initialize(context) {
|
|
||||||
await super.initialize(context)
|
|
||||||
this.notion = new NotionClient(this.apiKey)
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Do local operation first
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// Then sync to Notion
|
|
||||||
if (op === 'add') {
|
|
||||||
await this.notion.createPage({
|
|
||||||
title: params.data.title,
|
|
||||||
content: params.data.content
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. Performance Optimization Augmentation
|
|
||||||
|
|
||||||
**Use Case:** Add caching, batching, deduplication
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
export class CacheAugmentation extends BaseAugmentation {
|
|
||||||
name = 'smart-cache'
|
|
||||||
timing = 'around' // Wrap to check cache
|
|
||||||
operations = ['search'] // Cache searches only
|
|
||||||
priority = 60
|
|
||||||
|
|
||||||
private cache = new Map()
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const key = JSON.stringify(params)
|
|
||||||
|
|
||||||
// Check cache
|
|
||||||
if (this.cache.has(key)) {
|
|
||||||
this.log('Cache hit!')
|
|
||||||
return this.cache.get(key)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Miss - execute and cache
|
|
||||||
const result = await next()
|
|
||||||
this.cache.set(key, result)
|
|
||||||
|
|
||||||
// Clear old entries if too many
|
|
||||||
if (this.cache.size > 1000) {
|
|
||||||
const firstKey = this.cache.keys().next().value
|
|
||||||
this.cache.delete(firstKey)
|
|
||||||
}
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Quick Reference: When to Use Each Timing
|
|
||||||
|
|
||||||
| Timing | Use For | Example |
|
|
||||||
|--------|---------|---------|
|
|
||||||
| `before` | Validation, transformation | Check required fields |
|
|
||||||
| `after` | Logging, syncing, analytics | Send to external API |
|
|
||||||
| `around` | Caching, error handling, timing | Wrap with try/catch |
|
|
||||||
| `replace` | Complete replacement | Storage backends |
|
|
||||||
|
|
||||||
## Quick Reference: Common Operations
|
|
||||||
|
|
||||||
| Operation | Description |
|
|
||||||
|-----------|-------------|
|
|
||||||
| `'add'` | Adding data to brain |
|
|
||||||
| `'search'` | Searching/querying |
|
|
||||||
| `'update'` | Updating existing data |
|
|
||||||
| `'delete'` | Removing data |
|
|
||||||
| `'storage'` | Storage resolution (special) |
|
|
||||||
| `'all'` | Intercept everything |
|
|
||||||
|
|
||||||
## The Context Object
|
|
||||||
|
|
||||||
Every augmentation gets this during `initialize()`:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
{
|
|
||||||
brain: BrainyData, // The brain instance
|
|
||||||
storage: StorageAdapter, // Storage backend
|
|
||||||
config: BrainyDataConfig, // Configuration
|
|
||||||
log: (msg, level) => void // Logger
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Priority Guidelines
|
|
||||||
|
|
||||||
| Priority | Use For |
|
|
||||||
|----------|---------|
|
|
||||||
| 100 | Storage (critical infrastructure) |
|
|
||||||
| 80-99 | System operations (WAL, connections) |
|
|
||||||
| 50-79 | Performance (caching, batching) |
|
|
||||||
| 20-49 | Features (validation, transformation) |
|
|
||||||
| 1-19 | Logging, analytics |
|
|
||||||
|
|
||||||
## Complete Working Example
|
|
||||||
|
|
||||||
Here's a full augmentation that adds word count to all documents:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from 'brainy'
|
|
||||||
|
|
||||||
export class WordCountAugmentation extends BaseAugmentation {
|
|
||||||
name = 'word-counter'
|
|
||||||
timing = 'before'
|
|
||||||
operations = ['add', 'update']
|
|
||||||
priority = 40
|
|
||||||
|
|
||||||
async execute(operation, params, next) {
|
|
||||||
// Add word count to metadata
|
|
||||||
if (params.data && params.data.content) {
|
|
||||||
const wordCount = params.data.content.split(/\s+/).length
|
|
||||||
params.metadata = params.metadata || {}
|
|
||||||
params.metadata.wordCount = wordCount
|
|
||||||
|
|
||||||
this.log(`Added word count: ${wordCount}`)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Continue with enhanced params
|
|
||||||
return next()
|
|
||||||
}
|
|
||||||
|
|
||||||
async initialize(context) {
|
|
||||||
await super.initialize(context)
|
|
||||||
this.log('Word counter ready!')
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// Usage
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new WordCountAugmentation())
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Now all adds include word count
|
|
||||||
await brain.add('Hello world', {
|
|
||||||
content: 'This is a test document with nine words here'
|
|
||||||
})
|
|
||||||
// Automatically adds: metadata.wordCount = 9
|
|
||||||
```
|
|
||||||
|
|
||||||
## Key Points to Remember
|
|
||||||
|
|
||||||
1. **All augmentations are `BrainyAugmentation`** - One interface
|
|
||||||
2. **Storage augmentations** add `provideStorage()` method
|
|
||||||
3. **Register before `init()`** for storage, anytime for others
|
|
||||||
4. **Use `BaseAugmentation`** for convenience (has helpers)
|
|
||||||
5. **`next()` is crucial** - Always call it (unless `replace`)
|
|
||||||
6. **Order matters** - Use priority to control execution order
|
|
||||||
|
|
||||||
## Testing Your Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
describe('MyAugmentation', () => {
|
|
||||||
it('should enhance data', async () => {
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new MyAugmentation())
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
await brain.add('test', { data: 'test' })
|
|
||||||
const result = await brain.search('test')
|
|
||||||
|
|
||||||
expect(result[0].metadata.enhanced).toBe(true)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
That's it! Augmentations are simple middleware that intercept operations. Pick your timing, operations, and priority, then implement `execute()`!
|
|
||||||
|
|
@ -1,292 +0,0 @@
|
||||||
# 🧠 The Complete Brainy Architecture Vision
|
|
||||||
|
|
||||||
## 🎯 The Genius: Everything is an Augmentation
|
|
||||||
|
|
||||||
```
|
|
||||||
🧠 BRAINY CORE
|
|
||||||
│
|
|
||||||
┌────────┴────────┐
|
|
||||||
│ Augmentations │
|
|
||||||
│ Pipeline │
|
|
||||||
└────────┬────────┘
|
|
||||||
│
|
|
||||||
┌────────────────────┼────────────────────┐
|
|
||||||
│ │ │
|
|
||||||
Data Processing External API Exposure
|
|
||||||
Augmentations Connections Augmentations
|
|
||||||
│ │ │
|
|
||||||
NeuralImport Synapses APIServer
|
|
||||||
EntityRegistry (Notion,etc) (REST/WS)
|
|
||||||
BatchProcessing │ MCPServer
|
|
||||||
IntelligentScoring │ GraphQLServer
|
|
||||||
│ │ │
|
|
||||||
└────────────────────┴────────────────────┘
|
|
||||||
│
|
|
||||||
Storage Layer
|
|
||||||
(FS, S3, OPFS, Memory)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 How It All Works Together
|
|
||||||
|
|
||||||
### 1. **Core Pipeline**
|
|
||||||
Every operation flows through the augmentation pipeline:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
User Action → BrainyData Method → Augmentation Pipeline → Storage
|
|
||||||
↑
|
|
||||||
All Augmentations Execute Here
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **Augmentation Categories (All Using Same Interface!)**
|
|
||||||
|
|
||||||
#### 🧬 **Data Processing** (timing: 'before')
|
|
||||||
- **NeuralImport** - AI understands data before storage
|
|
||||||
- **EntityRegistry** - Deduplicates entities
|
|
||||||
- **BatchProcessing** - Optimizes bulk operations
|
|
||||||
|
|
||||||
#### 🌐 **External Connections** (timing: 'after')
|
|
||||||
- **Synapses** - Sync with Notion, Salesforce, etc.
|
|
||||||
- **WebSocketBroadcast** - Real-time updates to clients
|
|
||||||
- **TeamCoordination** - Multi-agent synchronization
|
|
||||||
|
|
||||||
#### 📡 **API Exposure** (timing: 'after' or separate process)
|
|
||||||
- **APIServerAugmentation** - REST/WebSocket/MCP server
|
|
||||||
- **GraphQLAugmentation** - GraphQL endpoint
|
|
||||||
- **ServiceWorkerAugmentation** - Browser local API
|
|
||||||
|
|
||||||
#### 💾 **Storage Backends** (timing: 'replace')
|
|
||||||
- **S3StorageAugmentation** - Use S3 instead of local
|
|
||||||
- **RedisAugmentation** - Use Redis for caching
|
|
||||||
- **PostgresAugmentation** - Use Postgres for persistence
|
|
||||||
|
|
||||||
#### 🛡️ **Infrastructure** (timing: 'around')
|
|
||||||
- **WALAugmentation** - Write-ahead logging
|
|
||||||
- **TransactionAugmentation** - ACID transactions
|
|
||||||
- **CacheAugmentation** - Multi-level caching
|
|
||||||
|
|
||||||
## 🌟 The Beautiful Simplicity
|
|
||||||
|
|
||||||
### One Interface Rules All
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface BrainyAugmentation {
|
|
||||||
name: string
|
|
||||||
timing: 'before' | 'after' | 'around' | 'replace'
|
|
||||||
operations: string[]
|
|
||||||
priority: number
|
|
||||||
initialize(context): Promise<void>
|
|
||||||
execute(operation, params, next): Promise<any>
|
|
||||||
shutdown?(): Promise<void>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
This single interface can:
|
|
||||||
- **Process data** with AI
|
|
||||||
- **Connect** to any external service
|
|
||||||
- **Expose** APIs (REST, WebSocket, MCP, GraphQL)
|
|
||||||
- **Replace** storage backends
|
|
||||||
- **Add** infrastructure (WAL, transactions, caching)
|
|
||||||
- **Coordinate** distributed systems
|
|
||||||
- **Visualize** data in real-time
|
|
||||||
- Literally **ANYTHING**
|
|
||||||
|
|
||||||
## 🏗️ Real-World Deployment Architecture
|
|
||||||
|
|
||||||
### Scenario 1: Local Development
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
new NeuralImportAugmentation(), // AI processing
|
|
||||||
new EntityRegistryAugmentation(), // Deduplication
|
|
||||||
new WALAugmentation() // Durability
|
|
||||||
]
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Scenario 2: Production Server
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
// Infrastructure
|
|
||||||
new WALAugmentation(),
|
|
||||||
new ConnectionPoolAugmentation(),
|
|
||||||
new RequestDeduplicatorAugmentation(),
|
|
||||||
|
|
||||||
// Data Processing
|
|
||||||
new NeuralImportAugmentation(),
|
|
||||||
new EntityRegistryAugmentation(),
|
|
||||||
new BatchProcessingAugmentation(),
|
|
||||||
|
|
||||||
// External Connections
|
|
||||||
new NotionSynapse({ apiKey: 'xxx' }),
|
|
||||||
new SlackSynapse({ token: 'xxx' }),
|
|
||||||
|
|
||||||
// API Exposure
|
|
||||||
new APIServerAugmentation({ port: 3000 }),
|
|
||||||
new MCPServerAugmentation({ port: 3001 }),
|
|
||||||
|
|
||||||
// Monitoring
|
|
||||||
new MetricsAugmentation(),
|
|
||||||
new LoggingAugmentation()
|
|
||||||
]
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Scenario 3: Distributed AI Agent System
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
// Agent Coordination
|
|
||||||
new TeamCoordinationAugmentation(),
|
|
||||||
new DistributedLockAugmentation(),
|
|
||||||
new SharedMemoryAugmentation(),
|
|
||||||
|
|
||||||
// Agent Memory
|
|
||||||
new MCPAgentMemoryAugmentation(),
|
|
||||||
new ConversationHistoryAugmentation(),
|
|
||||||
|
|
||||||
// Real-time Communication
|
|
||||||
new WebSocketBroadcastAugmentation(),
|
|
||||||
new PubSubAugmentation(),
|
|
||||||
|
|
||||||
// Visualization
|
|
||||||
new GraphVisualizationAugmentation()
|
|
||||||
]
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔌 How API Exposure Works
|
|
||||||
|
|
||||||
The **APIServerAugmentation** is special - it can run in two modes:
|
|
||||||
|
|
||||||
### Mode 1: Embedded (Same Process)
|
|
||||||
```typescript
|
|
||||||
brain.augmentations.register(new APIServerAugmentation())
|
|
||||||
// API server runs in same process, hooks into pipeline
|
|
||||||
```
|
|
||||||
|
|
||||||
### Mode 2: Standalone (Separate Process)
|
|
||||||
```typescript
|
|
||||||
// server.js - separate file
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
import { APIServerAugmentation } from 'brainy/augmentations'
|
|
||||||
|
|
||||||
const brain = new BrainyData()
|
|
||||||
const apiServer = new APIServerAugmentation()
|
|
||||||
|
|
||||||
// Can also run as standalone server connecting to remote Brainy
|
|
||||||
apiServer.connectToRemoteBrainy('ws://brainy-host:8080')
|
|
||||||
apiServer.listen(3000)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎭 The Four Timing Modes in Practice
|
|
||||||
|
|
||||||
### System Startup Sequence
|
|
||||||
```
|
|
||||||
1. INITIALIZE Phase
|
|
||||||
└─> All augmentations initialize (storage, connections, servers)
|
|
||||||
|
|
||||||
2. OPERATION Phase (for each operation)
|
|
||||||
├─> 'before' augmentations (NeuralImport, Validation)
|
|
||||||
├─> 'around' augmentations start (WAL, Transactions)
|
|
||||||
├─> 'replace' augmentations (if any, skip core)
|
|
||||||
├─> Core operation (or replaced operation)
|
|
||||||
├─> 'around' augmentations complete (Commit/Rollback)
|
|
||||||
└─> 'after' augmentations (Sync, Broadcast, Log)
|
|
||||||
|
|
||||||
3. SHUTDOWN Phase
|
|
||||||
└─> All augmentations cleanup (close connections, flush buffers)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🌍 Deployment Patterns
|
|
||||||
|
|
||||||
### Pattern 1: Monolithic
|
|
||||||
Everything in one process:
|
|
||||||
```
|
|
||||||
┌─────────────────────────────┐
|
|
||||||
│ Single Node.js Process │
|
|
||||||
│ ┌───────────────────────┐ │
|
|
||||||
│ │ BrainyData Core │ │
|
|
||||||
│ ├───────────────────────┤ │
|
|
||||||
│ │ All Augmentations │ │
|
|
||||||
│ ├───────────────────────┤ │
|
|
||||||
│ │ API Server │ │
|
|
||||||
│ └───────────────────────┘ │
|
|
||||||
└─────────────────────────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
### Pattern 2: Microservices
|
|
||||||
Distributed across services:
|
|
||||||
```
|
|
||||||
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
|
|
||||||
│ Brainy Core │────▶│ API Gateway │────▶│ Clients │
|
|
||||||
└──────┬───────┘ └──────────────┘ └──────────────┘
|
|
||||||
│
|
|
||||||
├──────────────┐
|
|
||||||
▼ ▼
|
|
||||||
┌──────────────┐ ┌──────────────┐
|
|
||||||
│ Synapses │ │ AI Agents │
|
|
||||||
│ Service │ │ Service │
|
|
||||||
└──────────────┘ └──────────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
### Pattern 3: Edge Computing
|
|
||||||
Brainy at the edge:
|
|
||||||
```
|
|
||||||
┌─────────────────────────────────────┐
|
|
||||||
│ CloudFlare Worker │
|
|
||||||
│ ┌───────────────────────────────┐ │
|
|
||||||
│ │ Brainy (Memory Storage) │ │
|
|
||||||
│ │ + API Server Augmentation │ │
|
|
||||||
│ └───────────────────────────────┘ │
|
|
||||||
└─────────────────────────────────────┘
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
┌───────────────┐
|
|
||||||
│ S3 Storage │
|
|
||||||
└───────────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 The Power of Composition
|
|
||||||
|
|
||||||
Any combination works because everything uses the same interface:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Local AI Assistant
|
|
||||||
[NeuralImport, ChatInterface, LocalStorage]
|
|
||||||
|
|
||||||
// Production API
|
|
||||||
[WAL, S3Storage, APIServer, RateLimiting]
|
|
||||||
|
|
||||||
// Multi-Agent System
|
|
||||||
[TeamCoordination, MCPServer, GraphVisualization]
|
|
||||||
|
|
||||||
// Data Pipeline
|
|
||||||
[KafkaConsumer, NeuralImport, PostgresStorage]
|
|
||||||
|
|
||||||
// Real-time Analytics
|
|
||||||
[StreamProcessing, Clustering, WebSocketBroadcast]
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Key Insights
|
|
||||||
|
|
||||||
1. **No Special Cases** - Everything is an augmentation
|
|
||||||
2. **Complete Flexibility** - Mix and match any combination
|
|
||||||
3. **Environment Agnostic** - Works in browser, Node, Deno, edge
|
|
||||||
4. **Protocol Agnostic** - REST, WebSocket, MCP, GraphQL, gRPC
|
|
||||||
5. **Storage Agnostic** - Local, S3, Redis, Postgres, anything
|
|
||||||
6. **Infinitely Extensible** - Just add more augmentations
|
|
||||||
|
|
||||||
## 🧠 The Philosophy
|
|
||||||
|
|
||||||
> "Make everything an augmentation, and the system becomes infinitely flexible while remaining dead simple."
|
|
||||||
|
|
||||||
This is why Brainy can be:
|
|
||||||
- A local embedded database
|
|
||||||
- A distributed knowledge graph
|
|
||||||
- An AI agent memory system
|
|
||||||
- A real-time collaboration platform
|
|
||||||
- A data pipeline processor
|
|
||||||
- All of the above simultaneously
|
|
||||||
|
|
||||||
**One interface. Infinite possibilities. That's the Brainy way.** 🚀
|
|
||||||
|
|
@ -1,18 +0,0 @@
|
||||||
# Augmentations Documentation Archive
|
|
||||||
|
|
||||||
This directory contains historical augmentation documentation from the Brainy 2.0 development process.
|
|
||||||
|
|
||||||
## Current Documentation
|
|
||||||
The definitive augmentation documentation is now located at:
|
|
||||||
- **`/docs/augmentations/README.md`** - Main augmentation guide
|
|
||||||
- **`/docs/architecture/augmentations.md`** - Architecture details
|
|
||||||
|
|
||||||
## Archived Documents
|
|
||||||
These documents represent the evolution of the augmentation system design:
|
|
||||||
- Various implementation approaches
|
|
||||||
- Pipeline architecture exploration
|
|
||||||
- Storage augmentation patterns
|
|
||||||
- Example implementations
|
|
||||||
|
|
||||||
## Note
|
|
||||||
These documents are archived to preserve development history while maintaining a clean documentation structure.
|
|
||||||
|
|
@ -1,276 +0,0 @@
|
||||||
# Storage Augmentations Guide
|
|
||||||
|
|
||||||
## Overview
|
|
||||||
|
|
||||||
Brainy uses a unified augmentation system for storage backends. This guide explains the difference between built-in storage augmentations and how to create custom ones.
|
|
||||||
|
|
||||||
## Built-in Storage Augmentations
|
|
||||||
|
|
||||||
These wrap existing, battle-tested storage adapters from `/storage/adapters/`:
|
|
||||||
|
|
||||||
| Augmentation | Underlying Adapter | Environment | Description |
|
|
||||||
|--------------|-------------------|-------------|-------------|
|
|
||||||
| `MemoryStorageAugmentation` | `MemoryStorage` | Universal | Fast in-memory storage (not persistent) |
|
|
||||||
| `FileSystemStorageAugmentation` | `FileSystemStorage` | Node.js | Persistent file-based storage |
|
|
||||||
| `OPFSStorageAugmentation` | `OPFSStorage` | Browser | Browser persistent storage |
|
|
||||||
| `S3StorageAugmentation` | `S3CompatibleStorage` | Universal | Amazon S3 with throttling & caching |
|
|
||||||
| `R2StorageAugmentation` | `R2Storage` | Universal | Cloudflare R2 storage |
|
|
||||||
| `GCSStorageAugmentation` | `S3CompatibleStorage` | Universal | Google Cloud Storage |
|
|
||||||
|
|
||||||
### Architecture of Built-in Storage
|
|
||||||
|
|
||||||
```
|
|
||||||
BrainyData
|
|
||||||
↓
|
|
||||||
StorageAugmentation (thin wrapper)
|
|
||||||
↓
|
|
||||||
StorageAdapter (actual implementation in /storage/adapters/)
|
|
||||||
↓
|
|
||||||
Actual Storage (filesystem, S3, memory, etc.)
|
|
||||||
```
|
|
||||||
|
|
||||||
### Why This Design?
|
|
||||||
|
|
||||||
1. **Preserve existing code** - Storage adapters have years of bug fixes
|
|
||||||
2. **Complex features intact** - S3 throttling, caching, retry logic preserved
|
|
||||||
3. **Minimal wrapper** - Augmentations are just 20-30 lines
|
|
||||||
4. **Zero feature loss** - All 30+ StorageAdapter methods work unchanged
|
|
||||||
|
|
||||||
## Using Built-in Storage
|
|
||||||
|
|
||||||
### 1. Zero-Config (Auto-Selection)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
// Automatically selects:
|
|
||||||
// - Node.js → FileSystemStorage
|
|
||||||
// - Browser → OPFSStorage (or Memory fallback)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Configuration-Based
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
s3Storage: {
|
|
||||||
bucketName: 'my-bucket',
|
|
||||||
accessKeyId: 'xxx',
|
|
||||||
secretAccessKey: 'yyy'
|
|
||||||
}
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Augmentation Override
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new S3StorageAugmentation({
|
|
||||||
bucketName: 'my-bucket',
|
|
||||||
region: 'us-east-1',
|
|
||||||
accessKeyId: 'xxx',
|
|
||||||
secretAccessKey: 'yyy'
|
|
||||||
}))
|
|
||||||
await brain.init()
|
|
||||||
```
|
|
||||||
|
|
||||||
## Creating Custom Storage Augmentations
|
|
||||||
|
|
||||||
Custom storage augmentations can either:
|
|
||||||
1. Wrap an existing adapter (like built-ins do)
|
|
||||||
2. Implement the StorageAdapter interface directly
|
|
||||||
|
|
||||||
### Option 1: Wrapping an Existing Adapter
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { StorageAugmentation } from 'brainy'
|
|
||||||
import { CustomAdapter } from './my-custom-adapter'
|
|
||||||
|
|
||||||
export class CustomStorageAugmentation extends StorageAugmentation {
|
|
||||||
private config: CustomConfig
|
|
||||||
|
|
||||||
constructor(config: CustomConfig) {
|
|
||||||
super('custom-storage')
|
|
||||||
this.config = config
|
|
||||||
}
|
|
||||||
|
|
||||||
async provideStorage(): Promise<StorageAdapter> {
|
|
||||||
// Create and return your adapter
|
|
||||||
const adapter = new CustomAdapter(this.config)
|
|
||||||
return adapter
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Option 2: Self-Contained Implementation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { StorageAugmentation, StorageAdapter } from 'brainy'
|
|
||||||
import Redis from 'ioredis'
|
|
||||||
|
|
||||||
export class RedisStorageAugmentation extends StorageAugmentation {
|
|
||||||
private redis: Redis
|
|
||||||
|
|
||||||
constructor(config: RedisConfig) {
|
|
||||||
super('redis-storage')
|
|
||||||
this.redis = new Redis(config)
|
|
||||||
}
|
|
||||||
|
|
||||||
async provideStorage(): Promise<StorageAdapter> {
|
|
||||||
// Return an object implementing StorageAdapter
|
|
||||||
return {
|
|
||||||
async init() {
|
|
||||||
await this.redis.ping()
|
|
||||||
},
|
|
||||||
|
|
||||||
async saveNoun(noun) {
|
|
||||||
await this.redis.set(
|
|
||||||
`noun:${noun.id}`,
|
|
||||||
JSON.stringify(noun)
|
|
||||||
)
|
|
||||||
},
|
|
||||||
|
|
||||||
async getNoun(id) {
|
|
||||||
const data = await this.redis.get(`noun:${id}`)
|
|
||||||
return data ? JSON.parse(data) : null
|
|
||||||
},
|
|
||||||
|
|
||||||
async deleteNoun(id) {
|
|
||||||
await this.redis.del(`noun:${id}`)
|
|
||||||
},
|
|
||||||
|
|
||||||
// ... implement all 30+ required methods
|
|
||||||
// See StorageAdapter interface in coreTypes.ts
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## StorageAdapter Interface Requirements
|
|
||||||
|
|
||||||
Your custom storage must implement these core methods:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface StorageAdapter {
|
|
||||||
// Initialization
|
|
||||||
init(): Promise<void>
|
|
||||||
|
|
||||||
// Noun operations
|
|
||||||
saveNoun(noun: HNSWNoun): Promise<void>
|
|
||||||
getNoun(id: string): Promise<HNSWNoun | null>
|
|
||||||
deleteNoun(id: string): Promise<void>
|
|
||||||
getNounsByNounType(type: string): Promise<HNSWNoun[]>
|
|
||||||
|
|
||||||
// Verb operations
|
|
||||||
saveVerb(verb: HNSWVerb): Promise<void>
|
|
||||||
getVerb(id: string): Promise<HNSWVerb | null>
|
|
||||||
deleteVerb(id: string): Promise<void>
|
|
||||||
getVerbsBySource(sourceId: string): Promise<HNSWVerb[]>
|
|
||||||
getVerbsByTarget(targetId: string): Promise<HNSWVerb[]>
|
|
||||||
|
|
||||||
// Metadata operations
|
|
||||||
saveMetadata(id: string, metadata: any): Promise<void>
|
|
||||||
getMetadata(id: string): Promise<any | null>
|
|
||||||
saveVerbMetadata(id: string, metadata: any): Promise<void>
|
|
||||||
getVerbMetadata(id: string): Promise<any | null>
|
|
||||||
|
|
||||||
// Pagination
|
|
||||||
getNouns(options?: PaginationOptions): Promise<PaginatedResult>
|
|
||||||
getVerbs(options?: PaginationOptions): Promise<PaginatedResult>
|
|
||||||
|
|
||||||
// Statistics
|
|
||||||
getStatistics(): Promise<StatisticsData | null>
|
|
||||||
saveStatistics(stats: StatisticsData): Promise<void>
|
|
||||||
incrementStatistic(type: string, service: string): Promise<void>
|
|
||||||
|
|
||||||
// Utility
|
|
||||||
clear(): Promise<void>
|
|
||||||
getStorageStatus(): Promise<StorageStatus>
|
|
||||||
|
|
||||||
// ... plus ~10 more methods
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Publishing to Brain Cloud (Future)
|
|
||||||
|
|
||||||
Custom storage augmentations can be published to the Brain Cloud marketplace:
|
|
||||||
|
|
||||||
```json
|
|
||||||
// package.json
|
|
||||||
{
|
|
||||||
"name": "@brain-cloud/redis-storage",
|
|
||||||
"version": "1.0.0",
|
|
||||||
"brainy": {
|
|
||||||
"type": "augmentation",
|
|
||||||
"category": "storage",
|
|
||||||
"implements": "StorageAdapter"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
Users will be able to install via:
|
|
||||||
```bash
|
|
||||||
brainy augment install redis-storage
|
|
||||||
```
|
|
||||||
|
|
||||||
## Best Practices
|
|
||||||
|
|
||||||
1. **Use existing adapters when possible** - They're well-tested
|
|
||||||
2. **Implement all methods** - StorageAdapter has 30+ required methods
|
|
||||||
3. **Handle errors gracefully** - Storage is critical infrastructure
|
|
||||||
4. **Include connection pooling** - For network-based storage
|
|
||||||
5. **Add retry logic** - Network operations can fail
|
|
||||||
6. **Implement caching** - Reduce latency for hot data
|
|
||||||
7. **Track statistics** - Use BaseStorageAdapter if possible
|
|
||||||
8. **Document configuration** - Make it easy for users
|
|
||||||
|
|
||||||
## Examples in the Wild
|
|
||||||
|
|
||||||
### MongoDB Storage (Community)
|
|
||||||
```typescript
|
|
||||||
class MongoStorageAugmentation extends StorageAugmentation {
|
|
||||||
async provideStorage() {
|
|
||||||
const client = new MongoClient(this.uri)
|
|
||||||
const db = client.db('brainy')
|
|
||||||
|
|
||||||
return {
|
|
||||||
async saveNoun(noun) {
|
|
||||||
await db.collection('nouns').replaceOne(
|
|
||||||
{ _id: noun.id },
|
|
||||||
noun,
|
|
||||||
{ upsert: true }
|
|
||||||
)
|
|
||||||
},
|
|
||||||
// ... full implementation
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### PostgreSQL Storage (Premium)
|
|
||||||
```typescript
|
|
||||||
class PostgreSQLStorageAugmentation extends StorageAugmentation {
|
|
||||||
async provideStorage() {
|
|
||||||
const pool = new Pool(this.config)
|
|
||||||
|
|
||||||
return {
|
|
||||||
async saveNoun(noun) {
|
|
||||||
await pool.query(
|
|
||||||
'INSERT INTO nouns (id, data) VALUES ($1, $2) ON CONFLICT (id) DO UPDATE SET data = $2',
|
|
||||||
[noun.id, JSON.stringify(noun)]
|
|
||||||
)
|
|
||||||
},
|
|
||||||
// ... full implementation
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Summary
|
|
||||||
|
|
||||||
- **Built-in augmentations** wrap existing adapters (thin layer)
|
|
||||||
- **Custom augmentations** can wrap OR implement directly
|
|
||||||
- **Storage adapters** in `/storage/adapters/` are for core only
|
|
||||||
- **Premium storage** comes as self-contained augmentations
|
|
||||||
- **Everything uses** the same StorageAdapter interface
|
|
||||||
- **Zero-config** still works perfectly
|
|
||||||
|
|
||||||
This design provides maximum flexibility while preserving all existing functionality!
|
|
||||||
|
|
@ -1,239 +0,0 @@
|
||||||
# 🧠 Unified Augmentation System
|
|
||||||
|
|
||||||
## The Single Interface That Rules Them All
|
|
||||||
|
|
||||||
Brainy uses ONE elegant interface for ALL augmentations:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface BrainyAugmentation {
|
|
||||||
name: string
|
|
||||||
timing: 'before' | 'after' | 'around' | 'replace'
|
|
||||||
operations: string[]
|
|
||||||
priority: number
|
|
||||||
initialize(context): Promise<void>
|
|
||||||
execute<T>(operation, params, next): Promise<T>
|
|
||||||
shutdown?(): Promise<void>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Why This Works for EVERYTHING
|
|
||||||
|
|
||||||
### 🎭 The Four Timing Modes
|
|
||||||
|
|
||||||
1. **`before`**: Pre-process data
|
|
||||||
- Data validation
|
|
||||||
- Authentication checks
|
|
||||||
- Input transformation
|
|
||||||
|
|
||||||
2. **`after`**: Post-process results
|
|
||||||
- Logging
|
|
||||||
- Analytics
|
|
||||||
- Cache updates
|
|
||||||
|
|
||||||
3. **`around`**: Wrap operations (middleware)
|
|
||||||
- Error handling
|
|
||||||
- Performance monitoring
|
|
||||||
- Transaction management
|
|
||||||
|
|
||||||
4. **`replace`**: Complete replacement
|
|
||||||
- Alternative storage backends
|
|
||||||
- Mock implementations
|
|
||||||
- Custom algorithms
|
|
||||||
|
|
||||||
### 🎯 Operation Targeting
|
|
||||||
|
|
||||||
Augmentations can target:
|
|
||||||
- Specific operations: `['add', 'search']`
|
|
||||||
- All operations: `['all']`
|
|
||||||
- Pattern matching: Operations containing certain strings
|
|
||||||
|
|
||||||
### 🔄 The Execute Chain
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
async execute<T>(operation, params, next): Promise<T> {
|
|
||||||
// Before logic
|
|
||||||
console.log(`Starting ${operation}`)
|
|
||||||
|
|
||||||
// Call next (or don't!)
|
|
||||||
const result = await next()
|
|
||||||
|
|
||||||
// After logic
|
|
||||||
console.log(`Completed ${operation}`)
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📦 Categories of Augmentations
|
|
||||||
|
|
||||||
While using the same interface, augmentations naturally fall into categories:
|
|
||||||
|
|
||||||
### 1. **Data Processing**
|
|
||||||
```typescript
|
|
||||||
class NeuralImportAugmentation {
|
|
||||||
timing = 'before'
|
|
||||||
operations = ['add', 'addNoun']
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Analyze data with AI
|
|
||||||
const enhanced = await this.processWithAI(params)
|
|
||||||
// Continue with enhanced data
|
|
||||||
return next(enhanced)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. **External Connections (Synapses)**
|
|
||||||
```typescript
|
|
||||||
class NotionSynapse {
|
|
||||||
timing = 'after'
|
|
||||||
operations = ['add', 'update', 'delete']
|
|
||||||
|
|
||||||
async initialize(context) {
|
|
||||||
await this.connectToNotion()
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const result = await next()
|
|
||||||
// Sync to Notion after local operation
|
|
||||||
await this.syncToNotion(op, params)
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. **Storage Backends**
|
|
||||||
```typescript
|
|
||||||
class S3StorageAugmentation {
|
|
||||||
timing = 'replace'
|
|
||||||
operations = ['storage']
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Don't call next() - replace entirely
|
|
||||||
return await this.s3Client.store(params)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. **Real-time Communication**
|
|
||||||
```typescript
|
|
||||||
class WebSocketBroadcast {
|
|
||||||
timing = 'after'
|
|
||||||
operations = ['all']
|
|
||||||
|
|
||||||
async initialize(context) {
|
|
||||||
this.ws = new WebSocket(url)
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const result = await next()
|
|
||||||
// Broadcast changes
|
|
||||||
this.ws.send({ op, params, result })
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. **AI Agent Coordination**
|
|
||||||
```typescript
|
|
||||||
class TeamMemoryAugmentation {
|
|
||||||
timing = 'around'
|
|
||||||
operations = ['add', 'search']
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
// Acquire distributed lock
|
|
||||||
await this.acquireLock(op)
|
|
||||||
try {
|
|
||||||
// Synchronize with team
|
|
||||||
const teamData = await this.syncWithTeam(params)
|
|
||||||
const result = await next(teamData)
|
|
||||||
// Broadcast result to team
|
|
||||||
await this.broadcastToTeam(result)
|
|
||||||
return result
|
|
||||||
} finally {
|
|
||||||
await this.releaseLock(op)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 6. **Analytics & Prediction**
|
|
||||||
```typescript
|
|
||||||
class PredictiveAnalytics {
|
|
||||||
timing = 'after'
|
|
||||||
operations = ['search']
|
|
||||||
|
|
||||||
async execute(op, params, next) {
|
|
||||||
const results = await next()
|
|
||||||
// Analyze search patterns
|
|
||||||
this.recordPattern(params, results)
|
|
||||||
// Add predictions
|
|
||||||
results.predictions = await this.predict(params)
|
|
||||||
return results
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔌 How Augmentations Connect
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// In BrainyData initialization
|
|
||||||
const brain = new BrainyData({
|
|
||||||
augmentations: [
|
|
||||||
new NeuralImportAugmentation(),
|
|
||||||
new NotionSynapse({ apiKey: 'xxx' }),
|
|
||||||
new TeamMemoryAugmentation(),
|
|
||||||
new PredictiveAnalytics()
|
|
||||||
]
|
|
||||||
})
|
|
||||||
|
|
||||||
// Or dynamically
|
|
||||||
brain.augmentations.register(new CustomAugmentation())
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Priority System
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Execution order (highest first)
|
|
||||||
100: Critical (WAL, Storage)
|
|
||||||
50: Performance (Cache, Dedup)
|
|
||||||
10: Features (Scoring, Analytics)
|
|
||||||
1: Optional (Logging)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🌍 Brain Cloud Integration
|
|
||||||
|
|
||||||
All augmentations (free, community, premium) use this SAME interface:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// From Brain Cloud marketplace
|
|
||||||
import { EmotionalIntelligence } from '@brainy-cloud/empathy'
|
|
||||||
|
|
||||||
const empathy = new EmotionalIntelligence()
|
|
||||||
// It's just a BrainyAugmentation!
|
|
||||||
brain.augmentations.register(empathy)
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💡 Why This Design Wins
|
|
||||||
|
|
||||||
1. **Simplicity**: One interface to learn
|
|
||||||
2. **Flexibility**: Can do literally anything
|
|
||||||
3. **Composability**: Stack augmentations like middleware
|
|
||||||
4. **Extensibility**: Easy to add new augmentations
|
|
||||||
5. **Marketplace Ready**: All augmentations compatible
|
|
||||||
|
|
||||||
## 🚀 The Future is Unified
|
|
||||||
|
|
||||||
No more complex type hierarchies. No more ISenseAugmentation, IConduitAugmentation, etc.
|
|
||||||
|
|
||||||
Just one beautiful, simple interface that can:
|
|
||||||
- Process data with AI
|
|
||||||
- Connect to any platform
|
|
||||||
- Coordinate AI teams
|
|
||||||
- Provide predictive analytics
|
|
||||||
- Add empathy to AI
|
|
||||||
- Store anywhere
|
|
||||||
- Communicate in real-time
|
|
||||||
- And literally anything else you can imagine
|
|
||||||
|
|
||||||
**One interface. Infinite possibilities. That's the Brainy way.** 🧠✨
|
|
||||||
|
|
@ -1,421 +0,0 @@
|
||||||
# 🔌 Brainy 2.0 Augmentations Complete Reference
|
|
||||||
|
|
||||||
> **All 27 augmentations that power Brainy's extensibility - with locations, usage, and examples**
|
|
||||||
|
|
||||||
## Quick Start
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from '@soulcraft/brainy'
|
|
||||||
|
|
||||||
const brain = new BrainyData({
|
|
||||||
// Augmentations auto-configure based on environment
|
|
||||||
storage: 'auto', // Storage augmentation
|
|
||||||
cache: true, // Cache augmentation
|
|
||||||
index: true // Index augmentation
|
|
||||||
})
|
|
||||||
|
|
||||||
await brain.init() // Augmentations initialize automatically
|
|
||||||
```
|
|
||||||
|
|
||||||
## Core Concepts
|
|
||||||
|
|
||||||
### What are Augmentations?
|
|
||||||
Augmentations are modular extensions that add functionality to Brainy without cluttering the core API. They follow a unified interface and can be:
|
|
||||||
- **Auto-enabled**: Based on configuration (cache, index, storage)
|
|
||||||
- **Manually registered**: For custom functionality
|
|
||||||
- **Chained**: Multiple augmentations work together seamlessly
|
|
||||||
|
|
||||||
### Augmentation Lifecycle
|
|
||||||
1. **Registration**: Augmentations register before init()
|
|
||||||
2. **Initialization**: Two-phase init (storage first, then others)
|
|
||||||
3. **Execution**: Hook into operations (before/after/both)
|
|
||||||
4. **Shutdown**: Clean teardown on brain.shutdown()
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Storage Augmentations (8 total)
|
|
||||||
|
|
||||||
### MemoryStorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Auto-enabled**: When `storage: 'memory'` or in test environments
|
|
||||||
**Purpose**: In-memory storage for testing and temporary data
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ storage: 'memory' })
|
|
||||||
```
|
|
||||||
|
|
||||||
### FileSystemStorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Auto-enabled**: When `storage: 'filesystem'` or Node.js detected
|
|
||||||
**Purpose**: Persistent file-based storage for Node.js applications
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: { type: 'filesystem', path: './data' }
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### OPFSStorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Auto-enabled**: When `storage: 'opfs'` or browser with OPFS support
|
|
||||||
**Purpose**: Browser-based persistent storage using Origin Private File System
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ storage: 'opfs' })
|
|
||||||
```
|
|
||||||
|
|
||||||
### S3StorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Manual**: Requires AWS credentials
|
|
||||||
**Purpose**: AWS S3-compatible cloud storage
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 's3',
|
|
||||||
bucket: 'my-bucket',
|
|
||||||
region: 'us-east-1',
|
|
||||||
credentials: { accessKeyId, secretAccessKey }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### R2StorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Manual**: Requires Cloudflare credentials
|
|
||||||
**Purpose**: Cloudflare R2 storage (S3-compatible)
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'r2',
|
|
||||||
accountId: 'xxx',
|
|
||||||
bucket: 'my-bucket',
|
|
||||||
credentials: { accessKeyId, secretAccessKey }
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### GCSStorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentations.ts`
|
|
||||||
**Manual**: Requires Google Cloud credentials
|
|
||||||
**Purpose**: Google Cloud Storage
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
storage: {
|
|
||||||
type: 'gcs',
|
|
||||||
bucket: 'my-bucket',
|
|
||||||
projectId: 'my-project'
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### StorageAugmentation (base)
|
|
||||||
**Location**: `src/augmentations/storageAugmentation.ts`
|
|
||||||
**Purpose**: Base class for custom storage implementations
|
|
||||||
|
|
||||||
### DynamicStorageAugmentation
|
|
||||||
**Location**: `src/augmentations/storageAugmentation.ts`
|
|
||||||
**Purpose**: Runtime storage adapter switching
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Performance Augmentations (7 total)
|
|
||||||
|
|
||||||
### CacheAugmentation
|
|
||||||
**Location**: `src/augmentations/cacheAugmentation.ts`
|
|
||||||
**Auto-enabled**: When `cache: true` (default)
|
|
||||||
**Purpose**: LRU cache for search results and frequent queries
|
|
||||||
```typescript
|
|
||||||
brain.clearCache() // Exposed via API
|
|
||||||
brain.getCacheStats() // Cache hit/miss statistics
|
|
||||||
```
|
|
||||||
|
|
||||||
### IndexAugmentation
|
|
||||||
**Location**: `src/augmentations/indexAugmentation.ts`
|
|
||||||
**Auto-enabled**: When `index: true` (default)
|
|
||||||
**Purpose**: Metadata indexing for O(1) field lookups
|
|
||||||
```typescript
|
|
||||||
brain.rebuildMetadataIndex() // Exposed via API
|
|
||||||
// Enables fast where queries:
|
|
||||||
brain.find({ where: { category: 'tech' } })
|
|
||||||
```
|
|
||||||
|
|
||||||
### MetricsAugmentation
|
|
||||||
**Location**: `src/augmentations/metricsAugmentation.ts`
|
|
||||||
**Auto-enabled**: Always active
|
|
||||||
**Purpose**: Performance metrics and statistics collection
|
|
||||||
```typescript
|
|
||||||
brain.getStatistics() // Comprehensive metrics
|
|
||||||
```
|
|
||||||
|
|
||||||
### MonitoringAugmentation
|
|
||||||
**Location**: `src/augmentations/monitoringAugmentation.ts`
|
|
||||||
**Manual**: Register for detailed monitoring
|
|
||||||
**Purpose**: Real-time performance monitoring and alerts
|
|
||||||
|
|
||||||
### BatchProcessingAugmentation
|
|
||||||
**Location**: `src/augmentations/batchProcessingAugmentation.ts`
|
|
||||||
**Auto-enabled**: For batch operations
|
|
||||||
**Purpose**: Optimizes bulk add/update/delete operations
|
|
||||||
```typescript
|
|
||||||
brain.addNouns([...]) // Automatically batched
|
|
||||||
```
|
|
||||||
|
|
||||||
### RequestDeduplicatorAugmentation
|
|
||||||
**Location**: `src/augmentations/requestDeduplicatorAugmentation.ts`
|
|
||||||
**Auto-enabled**: Always active
|
|
||||||
**Purpose**: Prevents duplicate concurrent operations
|
|
||||||
|
|
||||||
### ConnectionPoolAugmentation
|
|
||||||
**Location**: `src/augmentations/connectionPoolAugmentation.ts`
|
|
||||||
**Auto-enabled**: For network storage
|
|
||||||
**Purpose**: Connection pooling for cloud storage adapters
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Data Integrity Augmentations (3 total)
|
|
||||||
|
|
||||||
### WALAugmentation
|
|
||||||
**Location**: `src/augmentations/walAugmentation.ts`
|
|
||||||
**Auto-enabled**: When `wal: true`
|
|
||||||
**Purpose**: Write-ahead logging for crash recovery
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ wal: true })
|
|
||||||
// Automatic recovery on restart after crash
|
|
||||||
```
|
|
||||||
|
|
||||||
### EntityRegistryAugmentation
|
|
||||||
**Location**: `src/augmentations/entityRegistryAugmentation.ts`
|
|
||||||
**Auto-enabled**: For streaming operations
|
|
||||||
**Purpose**: High-speed deduplication for real-time data
|
|
||||||
```typescript
|
|
||||||
// Prevents duplicate entities in streaming scenarios
|
|
||||||
brain.addNoun(data) // Automatically deduplicated
|
|
||||||
```
|
|
||||||
|
|
||||||
### AutoRegisterEntitiesAugmentation
|
|
||||||
**Location**: `src/augmentations/entityRegistryAugmentation.ts`
|
|
||||||
**Manual**: For automatic entity discovery
|
|
||||||
**Purpose**: Auto-discovers and registers entities from data
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Intelligence Augmentations (2 total)
|
|
||||||
|
|
||||||
### NeuralImportAugmentation
|
|
||||||
**Location**: `src/augmentations/neuralImport.ts`
|
|
||||||
**Manual**: Via `brain.neuralImport()`
|
|
||||||
**Purpose**: AI-powered smart data import
|
|
||||||
```typescript
|
|
||||||
const result = await brain.neuralImport(data, {
|
|
||||||
confidenceThreshold: 0.7,
|
|
||||||
autoApply: true
|
|
||||||
})
|
|
||||||
// Automatically detects entities and relationships
|
|
||||||
```
|
|
||||||
|
|
||||||
### IntelligentVerbScoringAugmentation
|
|
||||||
**Location**: `src/augmentations/intelligentVerbScoringAugmentation.ts`
|
|
||||||
**Auto-enabled**: When verbs are used
|
|
||||||
**Purpose**: ML-based relationship strength scoring
|
|
||||||
```typescript
|
|
||||||
brain.verbScoring.train(feedback)
|
|
||||||
brain.verbScoring.getScore(verbId)
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Communication Augmentations (4 total)
|
|
||||||
|
|
||||||
### APIServerAugmentation
|
|
||||||
**Location**: `src/augmentations/apiServerAugmentation.ts`
|
|
||||||
**Manual**: For server deployments
|
|
||||||
**Purpose**: REST/WebSocket/MCP API server
|
|
||||||
```typescript
|
|
||||||
const augmentation = new APIServerAugmentation()
|
|
||||||
await brain.registerAugmentation(augmentation)
|
|
||||||
// Exposes full Brainy API over network
|
|
||||||
```
|
|
||||||
|
|
||||||
### WebSocketConduitAugmentation
|
|
||||||
**Location**: `src/augmentations/conduitAugmentations.ts`
|
|
||||||
**Manual**: For Brainy-to-Brainy sync
|
|
||||||
**Purpose**: Real-time sync between Brainy instances
|
|
||||||
```typescript
|
|
||||||
const conduit = new WebSocketConduitAugmentation()
|
|
||||||
await conduit.establishConnection('ws://other-brain')
|
|
||||||
```
|
|
||||||
|
|
||||||
### ServerSearchConduitAugmentation
|
|
||||||
**Location**: `src/augmentations/serverSearchAugmentations.ts`
|
|
||||||
**Manual**: For client-server search
|
|
||||||
**Purpose**: Search remote Brainy instance, cache locally
|
|
||||||
|
|
||||||
### ServerSearchActivationAugmentation
|
|
||||||
**Location**: `src/augmentations/serverSearchAugmentations.ts`
|
|
||||||
**Manual**: Works with ServerSearchConduit
|
|
||||||
**Purpose**: Triggers and manages server search operations
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## External Integration (2 total)
|
|
||||||
|
|
||||||
### SynapseAugmentation (base)
|
|
||||||
**Location**: `src/augmentations/synapseAugmentation.ts`
|
|
||||||
**Purpose**: Base class for external platform integrations
|
|
||||||
```typescript
|
|
||||||
// Example: NotionSynapse, SlackSynapse, etc.
|
|
||||||
class NotionSynapse extends SynapseAugmentation {
|
|
||||||
async fetchData() { /* Notion API calls */ }
|
|
||||||
async pushData() { /* Sync to Notion */ }
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### ExampleFileSystemSynapse
|
|
||||||
**Location**: `src/augmentations/synapseAugmentation.ts`
|
|
||||||
**Purpose**: Example implementation for file system sync
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Augmentation Configuration
|
|
||||||
|
|
||||||
### Auto-Configuration
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({
|
|
||||||
// These auto-register augmentations:
|
|
||||||
storage: 'auto', // Storage augmentation
|
|
||||||
cache: true, // Cache augmentation
|
|
||||||
index: true, // Index augmentation
|
|
||||||
wal: true, // WAL augmentation
|
|
||||||
metrics: true // Metrics augmentation
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Manual Registration
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
// Register before init()
|
|
||||||
const customAug = new MyCustomAugmentation()
|
|
||||||
await brain.registerAugmentation(customAug)
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
```
|
|
||||||
|
|
||||||
### Creating Custom Augmentations
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation } from '@soulcraft/brainy'
|
|
||||||
|
|
||||||
class MyAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'my-augmentation'
|
|
||||||
readonly timing = 'after' // before | after | both
|
|
||||||
readonly operations = ['addNoun', 'search'] // Which ops to hook
|
|
||||||
readonly priority = 10 // Execution order (lower = earlier)
|
|
||||||
|
|
||||||
protected async onInit(): Promise<void> {
|
|
||||||
// Initialize your augmentation
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T>(
|
|
||||||
operation: string,
|
|
||||||
params: any,
|
|
||||||
context?: AugmentationContext
|
|
||||||
): Promise<T | void> {
|
|
||||||
// Your augmentation logic
|
|
||||||
if (operation === 'addNoun') {
|
|
||||||
console.log('Noun added:', params)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
protected async onShutdown(): Promise<void> {
|
|
||||||
// Cleanup
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Augmentation Timing & Priority
|
|
||||||
|
|
||||||
### Timing Options
|
|
||||||
- **`before`**: Runs before the operation (can modify params)
|
|
||||||
- **`after`**: Runs after the operation (can see results)
|
|
||||||
- **`both`**: Runs before AND after
|
|
||||||
|
|
||||||
### Priority (lower = earlier)
|
|
||||||
1. Storage augmentations (priority: 0)
|
|
||||||
2. Cache/Index augmentations (priority: 5-10)
|
|
||||||
3. Monitoring/Metrics (priority: 15-20)
|
|
||||||
4. Conduits/Synapses (priority: 20-30)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Key Integration Points
|
|
||||||
|
|
||||||
### Where Augmentations Hook In
|
|
||||||
|
|
||||||
**BrainyData Constructor**:
|
|
||||||
- Storage augmentations register based on config
|
|
||||||
- Cache/Index augmentations auto-register if enabled
|
|
||||||
|
|
||||||
**brain.init()**:
|
|
||||||
- Two-phase initialization (storage first, then others)
|
|
||||||
- Augmentations can access brain instance via context
|
|
||||||
|
|
||||||
**Operations** (addNoun, search, etc.):
|
|
||||||
- Augmentations execute based on timing and operations filter
|
|
||||||
- Can modify params (before) or see results (after)
|
|
||||||
|
|
||||||
**brain.shutdown()**:
|
|
||||||
- All augmentations cleaned up in reverse order
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Performance Impact
|
|
||||||
|
|
||||||
Most augmentations have minimal overhead:
|
|
||||||
- **Cache**: ~1ms per search (saves 10-100ms on hits)
|
|
||||||
- **Index**: ~1ms per operation (saves 100ms+ on queries)
|
|
||||||
- **Metrics**: <1ms per operation
|
|
||||||
- **Storage**: Varies by adapter (memory: 0ms, S3: 50-200ms)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Best Practices
|
|
||||||
|
|
||||||
1. **Let auto-configuration work**: Most apps need zero manual config
|
|
||||||
2. **Storage first**: Always configure storage before other augmentations
|
|
||||||
3. **Use built-in augmentations**: They're optimized and battle-tested
|
|
||||||
4. **Custom augmentations**: Extend BaseAugmentation for consistency
|
|
||||||
5. **Respect timing**: Use 'before' to modify, 'after' to observe
|
|
||||||
6. **Mind priority**: Lower numbers execute first
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Troubleshooting
|
|
||||||
|
|
||||||
### Augmentation not working?
|
|
||||||
```typescript
|
|
||||||
// Check if registered
|
|
||||||
brain.listAugmentations()
|
|
||||||
|
|
||||||
// Check if enabled
|
|
||||||
brain.isAugmentationEnabled('cache')
|
|
||||||
|
|
||||||
// Enable/disable at runtime
|
|
||||||
brain.enableAugmentation('cache')
|
|
||||||
brain.disableAugmentation('cache')
|
|
||||||
```
|
|
||||||
|
|
||||||
### Performance issues?
|
|
||||||
```typescript
|
|
||||||
// Check augmentation overhead
|
|
||||||
const stats = brain.getStatistics()
|
|
||||||
console.log(stats.augmentations)
|
|
||||||
|
|
||||||
// Disable non-critical augmentations
|
|
||||||
brain.disableAugmentation('monitoring')
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
*Augmentations make Brainy infinitely extensible while keeping the core API clean and simple!*
|
|
||||||
|
|
@ -1,477 +0,0 @@
|
||||||
# 🛠️ Brainy Augmentation Developer Guide
|
|
||||||
|
|
||||||
> **How to create, test, and use augmentations in Brainy 2.0**
|
|
||||||
|
|
||||||
## Quick Start: Your First Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BaseAugmentation, BrainyAugmentation, AugmentationContext } from '@soulcraft/brainy'
|
|
||||||
|
|
||||||
export class MyFirstAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'my-first-augmentation'
|
|
||||||
readonly timing = 'after' as const // When to run: before | after | both
|
|
||||||
readonly operations = ['addNoun'] as const // Which operations to hook
|
|
||||||
readonly priority = 10 // Execution order (lower = first)
|
|
||||||
|
|
||||||
protected async onInit(): Promise<void> {
|
|
||||||
// Initialize your augmentation
|
|
||||||
console.log('MyFirstAugmentation initialized!')
|
|
||||||
}
|
|
||||||
|
|
||||||
async execute<T = any>(
|
|
||||||
operation: string,
|
|
||||||
params: any,
|
|
||||||
context?: AugmentationContext
|
|
||||||
): Promise<T | void> {
|
|
||||||
// Your augmentation logic
|
|
||||||
if (operation === 'addNoun') {
|
|
||||||
console.log('Noun added:', params.noun)
|
|
||||||
// You can access the brain instance
|
|
||||||
const stats = await context?.brain.getStatistics()
|
|
||||||
console.log('Total nouns:', stats.totalNouns)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
protected async onShutdown(): Promise<void> {
|
|
||||||
// Cleanup
|
|
||||||
console.log('MyFirstAugmentation shutting down')
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Using Your Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from '@soulcraft/brainy'
|
|
||||||
import { MyFirstAugmentation } from './my-first-augmentation'
|
|
||||||
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
// Register before init()
|
|
||||||
brain.augmentations.register(new MyFirstAugmentation())
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Now your augmentation runs automatically!
|
|
||||||
await brain.addNoun('Hello World')
|
|
||||||
// Console: "Noun added: { id: '...', vector: [...], metadata: {} }"
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Augmentation Lifecycle
|
|
||||||
|
|
||||||
### 1. Registration Phase
|
|
||||||
```typescript
|
|
||||||
const aug = new MyAugmentation()
|
|
||||||
brain.augmentations.register(aug) // Before brain.init()!
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Initialization Phase
|
|
||||||
```typescript
|
|
||||||
await brain.init() // Calls aug.initialize() internally
|
|
||||||
// Your onInit() method runs here
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Execution Phase
|
|
||||||
```typescript
|
|
||||||
await brain.addNoun('data') // Your execute() method runs
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. Shutdown Phase
|
|
||||||
```typescript
|
|
||||||
await brain.shutdown() // Your onShutdown() method runs
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Timing Options
|
|
||||||
|
|
||||||
### `before` - Modify Input
|
|
||||||
```typescript
|
|
||||||
class ValidationAugmentation extends BaseAugmentation {
|
|
||||||
readonly timing = 'before' as const
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<any> {
|
|
||||||
if (operation === 'addNoun') {
|
|
||||||
// Validate and/or modify params
|
|
||||||
if (!params.content) {
|
|
||||||
throw new Error('Content required')
|
|
||||||
}
|
|
||||||
// Return modified params
|
|
||||||
return { ...params, validated: true }
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### `after` - React to Results
|
|
||||||
```typescript
|
|
||||||
class LoggingAugmentation extends BaseAugmentation {
|
|
||||||
readonly timing = 'after' as const
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<void> {
|
|
||||||
if (operation === 'search') {
|
|
||||||
console.log(`Search for "${params.query}" returned ${params.result.length} results`)
|
|
||||||
}
|
|
||||||
// Don't return anything - just observe
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### `both` - Before AND After
|
|
||||||
```typescript
|
|
||||||
class TimingAugmentation extends BaseAugmentation {
|
|
||||||
readonly timing = 'both' as const
|
|
||||||
private startTime?: number
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, context?: AugmentationContext): Promise<void> {
|
|
||||||
if (!this.startTime) {
|
|
||||||
// Before execution
|
|
||||||
this.startTime = Date.now()
|
|
||||||
} else {
|
|
||||||
// After execution
|
|
||||||
const duration = Date.now() - this.startTime
|
|
||||||
console.log(`${operation} took ${duration}ms`)
|
|
||||||
this.startTime = undefined
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Operation Hooks
|
|
||||||
|
|
||||||
### Core Operations You Can Hook
|
|
||||||
```typescript
|
|
||||||
readonly operations = [
|
|
||||||
'addNoun', // Adding data
|
|
||||||
'updateNoun', // Updating data
|
|
||||||
'deleteNoun', // Deleting data
|
|
||||||
'getNoun', // Retrieving data
|
|
||||||
'search', // Searching
|
|
||||||
'find', // Triple Intelligence queries
|
|
||||||
'addVerb', // Adding relationships
|
|
||||||
'deleteVerb', // Removing relationships
|
|
||||||
'clear', // Clearing data
|
|
||||||
'all' // Hook ALL operations
|
|
||||||
] as const
|
|
||||||
```
|
|
||||||
|
|
||||||
### Example: Multi-Operation Hook
|
|
||||||
```typescript
|
|
||||||
class AuditAugmentation extends BaseAugmentation {
|
|
||||||
readonly operations = ['addNoun', 'updateNoun', 'deleteNoun'] as const
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<void> {
|
|
||||||
// Log all data modifications
|
|
||||||
await this.logToAuditTrail(operation, params)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Accessing Brain Context
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class ContextAwareAugmentation extends BaseAugmentation {
|
|
||||||
async execute<T>(
|
|
||||||
operation: string,
|
|
||||||
params: any,
|
|
||||||
context?: AugmentationContext
|
|
||||||
): Promise<void> {
|
|
||||||
// Access the brain instance
|
|
||||||
const brain = context?.brain
|
|
||||||
if (!brain) return
|
|
||||||
|
|
||||||
// Use any brain method
|
|
||||||
const stats = await brain.getStatistics()
|
|
||||||
const size = await brain.size()
|
|
||||||
const results = await brain.search('query')
|
|
||||||
|
|
||||||
// Access other augmentations
|
|
||||||
const cache = brain.augmentations.get('cache')
|
|
||||||
if (cache) {
|
|
||||||
await cache.clear()
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Real-World Examples
|
|
||||||
|
|
||||||
### 1. Backup Augmentation
|
|
||||||
```typescript
|
|
||||||
class BackupAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'backup'
|
|
||||||
readonly timing = 'after' as const
|
|
||||||
readonly operations = ['addNoun', 'updateNoun', 'deleteNoun'] as const
|
|
||||||
readonly priority = 5
|
|
||||||
|
|
||||||
private changes = 0
|
|
||||||
private readonly backupThreshold = 100
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, context?: AugmentationContext): Promise<void> {
|
|
||||||
this.changes++
|
|
||||||
|
|
||||||
if (this.changes >= this.backupThreshold) {
|
|
||||||
await this.performBackup(context?.brain)
|
|
||||||
this.changes = 0
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
private async performBackup(brain?: any): Promise<void> {
|
|
||||||
if (!brain) return
|
|
||||||
const backup = await brain.backup()
|
|
||||||
await this.saveToCloud(backup)
|
|
||||||
console.log('Automatic backup completed')
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Rate Limiting Augmentation
|
|
||||||
```typescript
|
|
||||||
class RateLimitAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'rate-limit'
|
|
||||||
readonly timing = 'before' as const
|
|
||||||
readonly operations = ['search', 'find'] as const
|
|
||||||
readonly priority = 100 // High priority - run first
|
|
||||||
|
|
||||||
private requests = new Map<string, number[]>()
|
|
||||||
private readonly limit = 100 // 100 requests
|
|
||||||
private readonly window = 60000 // per minute
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<void> {
|
|
||||||
const now = Date.now()
|
|
||||||
const key = params.userId || 'anonymous'
|
|
||||||
|
|
||||||
// Get request timestamps
|
|
||||||
const timestamps = this.requests.get(key) || []
|
|
||||||
|
|
||||||
// Remove old timestamps
|
|
||||||
const recent = timestamps.filter(t => now - t < this.window)
|
|
||||||
|
|
||||||
// Check limit
|
|
||||||
if (recent.length >= this.limit) {
|
|
||||||
throw new Error('Rate limit exceeded')
|
|
||||||
}
|
|
||||||
|
|
||||||
// Add current request
|
|
||||||
recent.push(now)
|
|
||||||
this.requests.set(key, recent)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Encryption Augmentation
|
|
||||||
```typescript
|
|
||||||
class EncryptionAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'encryption'
|
|
||||||
readonly timing = 'both' as const
|
|
||||||
readonly operations = ['addNoun', 'getNoun'] as const
|
|
||||||
readonly priority = 90 // Run early
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<any> {
|
|
||||||
if (operation === 'addNoun') {
|
|
||||||
// Encrypt before storing
|
|
||||||
if (params.metadata?.sensitive) {
|
|
||||||
params.content = await this.encrypt(params.content)
|
|
||||||
params.encrypted = true
|
|
||||||
}
|
|
||||||
return params
|
|
||||||
}
|
|
||||||
|
|
||||||
if (operation === 'getNoun' && params.result?.encrypted) {
|
|
||||||
// Decrypt after retrieval
|
|
||||||
params.result.content = await this.decrypt(params.result.content)
|
|
||||||
delete params.result.encrypted
|
|
||||||
return params.result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Testing Your Augmentation
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { describe, it, expect } from 'vitest'
|
|
||||||
import { BrainyData } from '@soulcraft/brainy'
|
|
||||||
import { MyAugmentation } from './my-augmentation'
|
|
||||||
|
|
||||||
describe('MyAugmentation', () => {
|
|
||||||
it('should hook into addNoun', async () => {
|
|
||||||
const brain = new BrainyData({ storage: 'memory' })
|
|
||||||
const aug = new MyAugmentation()
|
|
||||||
|
|
||||||
// Spy on the execute method
|
|
||||||
const executeSpy = vi.spyOn(aug, 'execute')
|
|
||||||
|
|
||||||
brain.augmentations.register(aug)
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Trigger the augmentation
|
|
||||||
await brain.addNoun('test data')
|
|
||||||
|
|
||||||
// Verify it was called
|
|
||||||
expect(executeSpy).toHaveBeenCalledWith(
|
|
||||||
'addNoun',
|
|
||||||
expect.objectContaining({ content: 'test data' }),
|
|
||||||
expect.any(Object)
|
|
||||||
)
|
|
||||||
})
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Best Practices
|
|
||||||
|
|
||||||
### 1. Use Proper Timing
|
|
||||||
- `before`: Validation, modification, rate limiting
|
|
||||||
- `after`: Logging, metrics, side effects
|
|
||||||
- `both`: Timing, tracing, wrapping
|
|
||||||
|
|
||||||
### 2. Set Appropriate Priority
|
|
||||||
```typescript
|
|
||||||
// Priority guidelines
|
|
||||||
100: Critical (auth, rate limiting)
|
|
||||||
50: Important (validation, transformation)
|
|
||||||
10: Normal (logging, metrics)
|
|
||||||
1: Optional (debugging, tracing)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Handle Errors Gracefully
|
|
||||||
```typescript
|
|
||||||
async execute<T>(operation: string, params: any): Promise<void> {
|
|
||||||
try {
|
|
||||||
await this.riskyOperation()
|
|
||||||
} catch (error) {
|
|
||||||
// Log but don't break the main operation
|
|
||||||
console.error(`Augmentation error in ${this.name}:`, error)
|
|
||||||
// Optionally report to monitoring
|
|
||||||
this.reportError(error)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. Be Performance Conscious
|
|
||||||
```typescript
|
|
||||||
class CachedAugmentation extends BaseAugmentation {
|
|
||||||
private cache = new Map<string, any>()
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any): Promise<any> {
|
|
||||||
const key = this.getCacheKey(params)
|
|
||||||
|
|
||||||
// Check cache first
|
|
||||||
if (this.cache.has(key)) {
|
|
||||||
return this.cache.get(key)
|
|
||||||
}
|
|
||||||
|
|
||||||
// Expensive operation
|
|
||||||
const result = await this.expensiveOperation(params)
|
|
||||||
this.cache.set(key, result)
|
|
||||||
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. Clean Up Resources
|
|
||||||
```typescript
|
|
||||||
protected async onShutdown(): Promise<void> {
|
|
||||||
// Close connections
|
|
||||||
await this.connection?.close()
|
|
||||||
|
|
||||||
// Clear intervals
|
|
||||||
clearInterval(this.interval)
|
|
||||||
|
|
||||||
// Flush buffers
|
|
||||||
await this.flush()
|
|
||||||
|
|
||||||
// Clear caches
|
|
||||||
this.cache.clear()
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Publishing Your Augmentation (Future)
|
|
||||||
|
|
||||||
### Package Structure
|
|
||||||
```
|
|
||||||
my-augmentation/
|
|
||||||
├── src/
|
|
||||||
│ └── index.ts # Your augmentation
|
|
||||||
├── dist/ # Built output
|
|
||||||
├── tests/
|
|
||||||
│ └── augmentation.test.ts
|
|
||||||
├── package.json
|
|
||||||
├── tsconfig.json
|
|
||||||
└── README.md
|
|
||||||
```
|
|
||||||
|
|
||||||
### package.json
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"name": "@mycompany/brainy-custom-augmentation",
|
|
||||||
"version": "1.0.0",
|
|
||||||
"main": "dist/index.js",
|
|
||||||
"types": "dist/index.d.ts",
|
|
||||||
"keywords": ["brainy-augmentation"],
|
|
||||||
"peerDependencies": {
|
|
||||||
"@soulcraft/brainy": ">=2.0.0"
|
|
||||||
},
|
|
||||||
"brainy": {
|
|
||||||
"type": "augmentation",
|
|
||||||
"class": "CustomAugmentation",
|
|
||||||
"timing": "after",
|
|
||||||
"operations": ["addNoun"],
|
|
||||||
"priority": 10
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Future: Brain Cloud Registry
|
|
||||||
```bash
|
|
||||||
# Coming in 2.1+
|
|
||||||
npm run build
|
|
||||||
npm test
|
|
||||||
brainy publish # Publishes to brain-cloud registry
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## FAQ
|
|
||||||
|
|
||||||
### Q: Can I modify the operation result?
|
|
||||||
**A**: Yes, if `timing: 'before'`, return modified params. If `timing: 'after'`, you can see but not modify results.
|
|
||||||
|
|
||||||
### Q: Can augmentations communicate?
|
|
||||||
**A**: Yes, through the context: `context.brain.augmentations.get('other-augmentation')`
|
|
||||||
|
|
||||||
### Q: What if my augmentation fails?
|
|
||||||
**A**: Handle errors internally. Don't break the main operation unless critical.
|
|
||||||
|
|
||||||
### Q: Can I use async operations?
|
|
||||||
**A**: Yes, everything is async-friendly.
|
|
||||||
|
|
||||||
### Q: How do I access storage directly?
|
|
||||||
**A**: Through context: `context.brain.storage` (but prefer using brain methods)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Get Help
|
|
||||||
|
|
||||||
- **GitHub**: [github.com/soulcraft/brainy](https://github.com/soulcraft/brainy)
|
|
||||||
- **Discord**: [discord.gg/brainy](https://discord.gg/brainy)
|
|
||||||
- **Examples**: See `/examples/augmentations/` in the repo
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
*Start building your augmentation today! The marketplace is coming in 2.1 🚀*
|
|
||||||
|
|
@ -1,206 +0,0 @@
|
||||||
# Brainy Augmentations
|
|
||||||
|
|
||||||
Augmentations are the core extensibility mechanism in Brainy. They allow you to modify, enhance, and extend Brainy's behavior without changing the core code.
|
|
||||||
|
|
||||||
## Core Principle: One Interface, Infinite Possibilities
|
|
||||||
|
|
||||||
Every augmentation implements the same simple `BrainyAugmentation` interface:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
interface BrainyAugmentation {
|
|
||||||
name: string
|
|
||||||
timing: 'before' | 'after' | 'around' | 'replace'
|
|
||||||
operations: string[]
|
|
||||||
priority: number
|
|
||||||
initialize(context): Promise<void>
|
|
||||||
execute(operation, params, next): Promise<any>
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
This single interface can handle EVERYTHING - from adding AI capabilities to exposing APIs to replacing storage backends.
|
|
||||||
|
|
||||||
## Available Augmentations
|
|
||||||
|
|
||||||
### 🧠 Data Processing
|
|
||||||
Augmentations that enhance how data is processed and stored.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **NeuralImportAugmentation** | AI-powered entity and relationship extraction | `before` | ✅ Production |
|
|
||||||
| **EntityRegistryAugmentation** | High-performance entity deduplication | `before` | ✅ Production |
|
|
||||||
| **BatchProcessingAugmentation** | Optimizes bulk operations | `around` | ✅ Production |
|
|
||||||
| **IntelligentVerbScoringAugmentation** | Learns relationship importance over time | `after` | ✅ Production |
|
|
||||||
|
|
||||||
### 🔌 External Connections (Synapses)
|
|
||||||
Connect Brainy to external services and data sources.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **NotionSynapse** | Sync with Notion databases | `after` | 📝 Example |
|
|
||||||
| **SalesforceSynapse** | Connect to Salesforce CRM | `after` | 📝 Example |
|
|
||||||
| **SlackSynapse** | Import Slack conversations | `after` | 📝 Example |
|
|
||||||
| **GoogleDriveSynapse** | Sync Google Drive documents | `after` | 📝 Example |
|
|
||||||
|
|
||||||
### 🌐 API Exposure
|
|
||||||
Expose Brainy through various protocols.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **APIServerAugmentation** | REST, WebSocket, and MCP server | `after` | ✅ Production |
|
|
||||||
| **GraphQLAugmentation** | GraphQL API endpoint | `after` | 🚧 Planned |
|
|
||||||
| **gRPCAugmentation** | gRPC service | `after` | 🚧 Planned |
|
|
||||||
|
|
||||||
### 💾 Storage Backends
|
|
||||||
Replace or enhance the storage layer.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **WALAugmentation** | Write-ahead logging for durability | `around` | ✅ Production |
|
|
||||||
| **S3StorageAugmentation** | Use S3 as storage backend | `replace` | 📝 Example |
|
|
||||||
| **RedisAugmentation** | Redis caching layer | `around` | 📝 Example |
|
|
||||||
| **PostgresAugmentation** | PostgreSQL persistence | `replace` | 📝 Example |
|
|
||||||
|
|
||||||
### 🔄 Real-time & Sync
|
|
||||||
Handle real-time updates and synchronization.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **WebSocketConduitAugmentation** | WebSocket client connections | `after` | ⚠️ Legacy |
|
|
||||||
| **ServerSearchAugmentation** | Connect to remote Brainy servers | `after` | ⚠️ Legacy |
|
|
||||||
| **TeamCoordinationAugmentation** | Multi-agent synchronization | `after` | 📝 Example |
|
|
||||||
|
|
||||||
### 🛡️ Infrastructure
|
|
||||||
Core infrastructure and reliability features.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **ConnectionPoolAugmentation** | Optimize cloud storage connections | `before` | ✅ Production |
|
|
||||||
| **RequestDeduplicatorAugmentation** | Prevent duplicate concurrent requests | `before` | ✅ Production |
|
|
||||||
| **TransactionAugmentation** | ACID transaction support | `around` | 🚧 Planned |
|
|
||||||
| **CacheAugmentation** | Multi-level caching | `around` | ✅ Production |
|
|
||||||
|
|
||||||
### 📊 Monitoring & Analytics
|
|
||||||
Track and analyze Brainy's behavior.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **MetricsAugmentation** | Prometheus metrics | `after` | 📝 Example |
|
|
||||||
| **LoggingAugmentation** | Structured logging | `after` | 📝 Example |
|
|
||||||
| **TracingAugmentation** | Distributed tracing | `around` | 🚧 Planned |
|
|
||||||
|
|
||||||
### 🤖 AI & Chat
|
|
||||||
AI-powered interfaces and chat capabilities.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **ChatInterfaceAugmentation** | Natural language interface | `before` | 📝 Example |
|
|
||||||
| **MCPAgentMemoryAugmentation** | AI agent memory via MCP | `after` | 📝 Example |
|
|
||||||
| **LLMQueryAugmentation** | LLM-enhanced queries | `before` | 📝 Example |
|
|
||||||
|
|
||||||
### 📈 Visualization
|
|
||||||
Visual representations of data.
|
|
||||||
|
|
||||||
| Augmentation | Description | Timing | Status |
|
|
||||||
|-------------|-------------|--------|--------|
|
|
||||||
| **GraphVisualizationAugmentation** | Real-time graph visualization | `after` | 📝 Example |
|
|
||||||
| **DashboardAugmentation** | Web-based dashboard | `after` | 🚧 Planned |
|
|
||||||
|
|
||||||
## Status Legend
|
|
||||||
|
|
||||||
- ✅ **Production**: Fully implemented and tested
|
|
||||||
- 📝 **Example**: Example implementation available
|
|
||||||
- 🚧 **Planned**: On the roadmap
|
|
||||||
- ⚠️ **Legacy**: Being replaced by newer augmentations
|
|
||||||
|
|
||||||
## Using Augmentations
|
|
||||||
|
|
||||||
### Zero-Config Approach
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
// Just register augmentations - they work automatically!
|
|
||||||
brain.augmentations.register(new WALAugmentation())
|
|
||||||
brain.augmentations.register(new EntityRegistryAugmentation())
|
|
||||||
brain.augmentations.register(new APIServerAugmentation())
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
```
|
|
||||||
|
|
||||||
### With Configuration
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
brain.augmentations.register(
|
|
||||||
new APIServerAugmentation({
|
|
||||||
port: 8080,
|
|
||||||
auth: { required: true }
|
|
||||||
})
|
|
||||||
)
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
```
|
|
||||||
|
|
||||||
## Creating Custom Augmentations
|
|
||||||
|
|
||||||
See [Creating Custom Augmentations](./creating-augmentations.md) for a complete guide.
|
|
||||||
|
|
||||||
Quick example:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class MyAugmentation extends BaseAugmentation {
|
|
||||||
readonly name = 'my-augmentation'
|
|
||||||
readonly timing = 'after'
|
|
||||||
readonly operations = ['add', 'search']
|
|
||||||
readonly priority = 50
|
|
||||||
|
|
||||||
async execute<T>(operation: string, params: any, next: () => Promise<T>): Promise<T> {
|
|
||||||
console.log(`Before ${operation}`)
|
|
||||||
const result = await next()
|
|
||||||
console.log(`After ${operation}`)
|
|
||||||
return result
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Augmentation Timing
|
|
||||||
|
|
||||||
### `before`
|
|
||||||
Executes before the main operation. Used for:
|
|
||||||
- Input validation
|
|
||||||
- Data transformation
|
|
||||||
- Authentication checks
|
|
||||||
|
|
||||||
### `after`
|
|
||||||
Executes after the main operation. Used for:
|
|
||||||
- Broadcasting updates
|
|
||||||
- Syncing to external services
|
|
||||||
- Logging and metrics
|
|
||||||
|
|
||||||
### `around`
|
|
||||||
Wraps the main operation. Used for:
|
|
||||||
- Transactions
|
|
||||||
- Caching
|
|
||||||
- Error handling
|
|
||||||
|
|
||||||
### `replace`
|
|
||||||
Completely replaces the main operation. Used for:
|
|
||||||
- Alternative storage backends
|
|
||||||
- Mock implementations
|
|
||||||
- Proxy operations
|
|
||||||
|
|
||||||
## Priority System
|
|
||||||
|
|
||||||
Higher numbers execute first:
|
|
||||||
- **100**: Critical system operations
|
|
||||||
- **50**: Performance optimizations
|
|
||||||
- **10**: Enhancement features
|
|
||||||
- **1**: Optional features
|
|
||||||
|
|
||||||
## Related Documentation
|
|
||||||
|
|
||||||
- [API Server Augmentation](./api-server.md) - Complete API server documentation
|
|
||||||
- [Creating Augmentations](./creating-augmentations.md) - How to build your own
|
|
||||||
- [Augmentation Examples](../AUGMENTATION-EXAMPLES.md) - Real-world examples
|
|
||||||
- [Architecture Overview](../COMPLETE-ARCHITECTURE-VISION.md) - System architecture
|
|
||||||
|
|
@ -1,404 +0,0 @@
|
||||||
# API Server Augmentation
|
|
||||||
|
|
||||||
## Overview
|
|
||||||
|
|
||||||
The `APIServerAugmentation` is a powerful augmentation that exposes your Brainy instance through REST, WebSocket, and MCP (Model Context Protocol) APIs. It transforms Brainy into a full-featured API server with zero configuration required.
|
|
||||||
|
|
||||||
## Features
|
|
||||||
|
|
||||||
### 🌐 REST API
|
|
||||||
Complete CRUD operations and advanced queries through HTTP endpoints.
|
|
||||||
|
|
||||||
### 🔌 WebSocket Server
|
|
||||||
Real-time bidirectional communication with automatic operation broadcasting.
|
|
||||||
|
|
||||||
### 🧠 MCP Integration
|
|
||||||
Built-in Model Context Protocol support for AI agent communication.
|
|
||||||
|
|
||||||
### 📊 Operation Broadcasting
|
|
||||||
Automatically broadcasts all Brainy operations to subscribed WebSocket clients.
|
|
||||||
|
|
||||||
### 🔒 Optional Security
|
|
||||||
Built-in authentication and rate limiting when needed.
|
|
||||||
|
|
||||||
## Installation
|
|
||||||
|
|
||||||
The APIServerAugmentation is included in Brainy core. No additional installation required.
|
|
||||||
|
|
||||||
For Node.js environments, you may want to install optional dependencies:
|
|
||||||
```bash
|
|
||||||
npm install express cors ws
|
|
||||||
```
|
|
||||||
|
|
||||||
## Zero-Config Usage
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
import { APIServerAugmentation } from 'brainy/augmentations'
|
|
||||||
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
// Register the API server augmentation
|
|
||||||
brain.augmentations.register(new APIServerAugmentation())
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Server is now running at http://localhost:3000
|
|
||||||
console.log('API Server ready!')
|
|
||||||
console.log('REST: http://localhost:3000/api/*')
|
|
||||||
console.log('WebSocket: ws://localhost:3000/ws')
|
|
||||||
console.log('MCP: http://localhost:3000/api/mcp')
|
|
||||||
```
|
|
||||||
|
|
||||||
## Configuration Options
|
|
||||||
|
|
||||||
While zero-config works great, you can customize the server:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const apiServer = new APIServerAugmentation({
|
|
||||||
enabled: true, // Enable/disable the server
|
|
||||||
port: 3000, // HTTP port
|
|
||||||
host: '0.0.0.0', // Bind address
|
|
||||||
|
|
||||||
cors: {
|
|
||||||
origin: '*', // CORS allowed origins
|
|
||||||
credentials: true // Allow credentials
|
|
||||||
},
|
|
||||||
|
|
||||||
auth: {
|
|
||||||
required: false, // Require authentication
|
|
||||||
apiKeys: [], // Valid API keys
|
|
||||||
bearerTokens: [] // Valid bearer tokens
|
|
||||||
},
|
|
||||||
|
|
||||||
rateLimit: {
|
|
||||||
windowMs: 60000, // Rate limit window (ms)
|
|
||||||
max: 100 // Max requests per window
|
|
||||||
}
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
## REST API Endpoints
|
|
||||||
|
|
||||||
### Health Check
|
|
||||||
```http
|
|
||||||
GET /health
|
|
||||||
```
|
|
||||||
Returns server status and basic metrics.
|
|
||||||
|
|
||||||
### Search
|
|
||||||
```http
|
|
||||||
POST /api/search
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"query": "search text",
|
|
||||||
"limit": 10,
|
|
||||||
"options": {}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Add Data
|
|
||||||
```http
|
|
||||||
POST /api/add
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"content": "data to add",
|
|
||||||
"metadata": {
|
|
||||||
"key": "value"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Get by ID
|
|
||||||
```http
|
|
||||||
GET /api/get/:id
|
|
||||||
```
|
|
||||||
|
|
||||||
### Delete
|
|
||||||
```http
|
|
||||||
DELETE /api/delete/:id
|
|
||||||
```
|
|
||||||
|
|
||||||
### Create Relationship
|
|
||||||
```http
|
|
||||||
POST /api/relate
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"source": "id1",
|
|
||||||
"target": "id2",
|
|
||||||
"verb": "relates_to",
|
|
||||||
"metadata": {}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Complex Queries
|
|
||||||
```http
|
|
||||||
POST /api/find
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"where": { "type": "document" },
|
|
||||||
"like": "machine learning",
|
|
||||||
"limit": 10
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Clustering
|
|
||||||
```http
|
|
||||||
POST /api/cluster
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"algorithm": "kmeans",
|
|
||||||
"options": {
|
|
||||||
"k": 5
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Statistics
|
|
||||||
```http
|
|
||||||
GET /api/stats
|
|
||||||
```
|
|
||||||
|
|
||||||
### Operation History
|
|
||||||
```http
|
|
||||||
GET /api/history
|
|
||||||
```
|
|
||||||
|
|
||||||
## WebSocket API
|
|
||||||
|
|
||||||
### Connection
|
|
||||||
```javascript
|
|
||||||
const ws = new WebSocket('ws://localhost:3000/ws')
|
|
||||||
|
|
||||||
ws.onopen = () => {
|
|
||||||
console.log('Connected to Brainy WebSocket')
|
|
||||||
}
|
|
||||||
|
|
||||||
ws.onmessage = (event) => {
|
|
||||||
const msg = JSON.parse(event.data)
|
|
||||||
console.log('Received:', msg)
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Subscribe to Operations
|
|
||||||
```javascript
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'subscribe',
|
|
||||||
operations: ['all'] // or specific: ['add', 'search', 'delete']
|
|
||||||
}))
|
|
||||||
```
|
|
||||||
|
|
||||||
### Search via WebSocket
|
|
||||||
```javascript
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'search',
|
|
||||||
query: 'your search',
|
|
||||||
limit: 10,
|
|
||||||
requestId: 'unique-id'
|
|
||||||
}))
|
|
||||||
```
|
|
||||||
|
|
||||||
### Add Data via WebSocket
|
|
||||||
```javascript
|
|
||||||
ws.send(JSON.stringify({
|
|
||||||
type: 'add',
|
|
||||||
content: 'data to add',
|
|
||||||
metadata: {},
|
|
||||||
requestId: 'unique-id'
|
|
||||||
}))
|
|
||||||
```
|
|
||||||
|
|
||||||
### Operation Broadcasts
|
|
||||||
When subscribed, you'll receive real-time updates:
|
|
||||||
```javascript
|
|
||||||
{
|
|
||||||
"type": "operation",
|
|
||||||
"operation": "add",
|
|
||||||
"params": { /* sanitized parameters */ },
|
|
||||||
"timestamp": 1234567890,
|
|
||||||
"duration": 15
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## MCP (Model Context Protocol)
|
|
||||||
|
|
||||||
The MCP endpoint allows AI agents to interact with Brainy:
|
|
||||||
|
|
||||||
```http
|
|
||||||
POST /api/mcp
|
|
||||||
Content-Type: application/json
|
|
||||||
|
|
||||||
{
|
|
||||||
"method": "search",
|
|
||||||
"params": {
|
|
||||||
"query": "find documents about AI"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## Authentication
|
|
||||||
|
|
||||||
When authentication is enabled:
|
|
||||||
|
|
||||||
### API Key
|
|
||||||
```http
|
|
||||||
GET /api/stats
|
|
||||||
X-API-Key: your-api-key
|
|
||||||
```
|
|
||||||
|
|
||||||
### Bearer Token
|
|
||||||
```http
|
|
||||||
GET /api/stats
|
|
||||||
Authorization: Bearer your-token
|
|
||||||
```
|
|
||||||
|
|
||||||
## Environment Support
|
|
||||||
|
|
||||||
### Node.js ✅
|
|
||||||
Full support with Express, WebSocket, and all features.
|
|
||||||
|
|
||||||
### Deno 🚧
|
|
||||||
Planned support using Deno.serve() or oak framework.
|
|
||||||
|
|
||||||
### Browser/Service Worker 🚧
|
|
||||||
Planned support for intercepting fetch() calls locally.
|
|
||||||
|
|
||||||
## How It Works
|
|
||||||
|
|
||||||
The APIServerAugmentation hooks into Brainy's augmentation pipeline:
|
|
||||||
|
|
||||||
1. **Timing**: Executes `after` operations complete
|
|
||||||
2. **Operations**: Monitors `all` operations
|
|
||||||
3. **Broadcasting**: Sends operation details to subscribed clients
|
|
||||||
4. **History**: Maintains operation history (last 1000 operations)
|
|
||||||
|
|
||||||
## Example: Multi-Client Sync
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Server
|
|
||||||
const brain = new BrainyData()
|
|
||||||
brain.augmentations.register(new APIServerAugmentation())
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Client 1 - WebSocket subscriber
|
|
||||||
const ws1 = new WebSocket('ws://localhost:3000/ws')
|
|
||||||
ws1.onopen = () => {
|
|
||||||
ws1.send(JSON.stringify({
|
|
||||||
type: 'subscribe',
|
|
||||||
operations: ['add', 'delete']
|
|
||||||
}))
|
|
||||||
}
|
|
||||||
ws1.onmessage = (e) => {
|
|
||||||
console.log('Client 1 received update:', JSON.parse(e.data))
|
|
||||||
}
|
|
||||||
|
|
||||||
// Client 2 - REST API user
|
|
||||||
fetch('http://localhost:3000/api/add', {
|
|
||||||
method: 'POST',
|
|
||||||
headers: { 'Content-Type': 'application/json' },
|
|
||||||
body: JSON.stringify({
|
|
||||||
content: 'New data',
|
|
||||||
metadata: { source: 'client2' }
|
|
||||||
})
|
|
||||||
})
|
|
||||||
// Client 1 automatically receives notification!
|
|
||||||
```
|
|
||||||
|
|
||||||
## Performance Considerations
|
|
||||||
|
|
||||||
- **Operation History**: Limited to last 1000 operations
|
|
||||||
- **WebSocket Heartbeat**: Every 30 seconds
|
|
||||||
- **Client Timeout**: 60 seconds of inactivity
|
|
||||||
- **Parameter Sanitization**: Sensitive fields removed, large content truncated
|
|
||||||
- **Rate Limiting**: In-memory tracking (use Redis in production)
|
|
||||||
|
|
||||||
## Security Notes
|
|
||||||
|
|
||||||
1. **Default Configuration**: No auth, open CORS - suitable for development
|
|
||||||
2. **Production**: Enable auth, configure CORS, use HTTPS
|
|
||||||
3. **Sensitive Data**: Parameters are sanitized before broadcasting
|
|
||||||
4. **Rate Limiting**: Basic in-memory implementation included
|
|
||||||
|
|
||||||
## Comparison with Previous Implementations
|
|
||||||
|
|
||||||
The APIServerAugmentation unifies and replaces:
|
|
||||||
- `BrainyMCPBroadcast` - Node-specific WebSocket/HTTP server
|
|
||||||
- `WebSocketConduitAugmentation` - WebSocket client functionality
|
|
||||||
- `ServerSearchAugmentations` - Remote Brainy connections
|
|
||||||
|
|
||||||
Benefits of the unified approach:
|
|
||||||
- Single augmentation for all API needs
|
|
||||||
- Consistent interface across protocols
|
|
||||||
- Automatic operation broadcasting
|
|
||||||
- Environment-aware implementation
|
|
||||||
- Zero-configuration philosophy
|
|
||||||
|
|
||||||
## Advanced Usage
|
|
||||||
|
|
||||||
### Custom Operation Filtering
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
class FilteredAPIServer extends APIServerAugmentation {
|
|
||||||
shouldExecute(operation: string, params: any): boolean {
|
|
||||||
// Don't broadcast sensitive operations
|
|
||||||
if (operation === 'delete' && params.sensitive) {
|
|
||||||
return false
|
|
||||||
}
|
|
||||||
return true
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Integration with Other Augmentations
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
|
|
||||||
// Stack augmentations for complete system
|
|
||||||
brain.augmentations.register(new WALAugmentation()) // Durability
|
|
||||||
brain.augmentations.register(new EntityRegistryAugmentation()) // Dedup
|
|
||||||
brain.augmentations.register(new APIServerAugmentation()) // API
|
|
||||||
|
|
||||||
await brain.init()
|
|
||||||
// All augmentations work together seamlessly!
|
|
||||||
```
|
|
||||||
|
|
||||||
## Troubleshooting
|
|
||||||
|
|
||||||
### Server won't start
|
|
||||||
- Check if port is already in use
|
|
||||||
- Verify Node.js dependencies are installed: `npm install express cors ws`
|
|
||||||
- Check console for error messages
|
|
||||||
|
|
||||||
### WebSocket connections drop
|
|
||||||
- Ensure heartbeat responses are handled
|
|
||||||
- Check for proxy/firewall issues
|
|
||||||
- Verify CORS configuration
|
|
||||||
|
|
||||||
### Authentication not working
|
|
||||||
- Ensure `auth.required` is set to `true`
|
|
||||||
- Verify API keys or bearer tokens are correctly configured
|
|
||||||
- Check request headers are properly formatted
|
|
||||||
|
|
||||||
## Future Enhancements
|
|
||||||
|
|
||||||
- [ ] Deno server implementation
|
|
||||||
- [ ] Service Worker implementation
|
|
||||||
- [ ] GraphQL endpoint
|
|
||||||
- [ ] gRPC support
|
|
||||||
- [ ] Built-in SSL/TLS
|
|
||||||
- [ ] Redis-based rate limiting
|
|
||||||
- [ ] Prometheus metrics endpoint
|
|
||||||
- [ ] OpenAPI/Swagger documentation
|
|
||||||
|
|
||||||
## Related Documentation
|
|
||||||
|
|
||||||
- [Augmentation System Overview](../AUGMENTATION-SYSTEM.md)
|
|
||||||
- [BrainyAugmentation Interface](./brainy-augmentation.md)
|
|
||||||
- [MCP Integration](../mcp/README.md)
|
|
||||||
- [Zero-Config Philosophy](../ZERO-CONFIG.md)
|
|
||||||
386
docs/concepts/consistency-model.md
Normal file
386
docs/concepts/consistency-model.md
Normal file
|
|
@ -0,0 +1,386 @@
|
||||||
|
---
|
||||||
|
title: Consistency Model
|
||||||
|
slug: concepts/consistency-model
|
||||||
|
public: true
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 4
|
||||||
|
description: The exact guarantees behind Brainy's Db API — snapshot isolation, atomic transactions, two levels of compare-and-swap, time travel, retention, snapshots, crash recovery, and the reserved-field contract.
|
||||||
|
next:
|
||||||
|
- guides/snapshots-and-time-travel
|
||||||
|
- guides/optimistic-concurrency
|
||||||
|
---
|
||||||
|
|
||||||
|
# Consistency Model
|
||||||
|
|
||||||
|
Brainy 8.0's consistency story rests on one mechanism: **generational MVCC**
|
||||||
|
— multi-version concurrency control over immutable, generation-stamped
|
||||||
|
records. It is exposed through a single value type, the **`Db`**: an
|
||||||
|
immutable, point-in-time view of the whole store that you query like the
|
||||||
|
live brain.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const db = brain.now() // pin the current state — O(1), no I/O
|
||||||
|
|
||||||
|
await brain.transact([
|
||||||
|
{ op: 'update', id: invoiceId, metadata: { status: 'paid' } }
|
||||||
|
])
|
||||||
|
|
||||||
|
await db.get(invoiceId) // still 'pending' — pinned, forever
|
||||||
|
await brain.get(invoiceId) // 'paid' — live
|
||||||
|
await db.release() // unpin when done
|
||||||
|
```
|
||||||
|
|
||||||
|
This page states the guarantees precisely — what is promised, what it costs,
|
||||||
|
and where the honest limits are. The design record is
|
||||||
|
[ADR-001](../ADR-001-generational-mvcc.md); every guarantee below is proven
|
||||||
|
by a dedicated test in `tests/integration/db-mvcc.test.ts`.
|
||||||
|
|
||||||
|
## The generation clock
|
||||||
|
|
||||||
|
A **monotonic generation counter** is the store's logical clock:
|
||||||
|
|
||||||
|
- It advances **once per committed `transact()` batch** and once per
|
||||||
|
single-operation write (`add`/`update`/`remove`/`relate`/…).
|
||||||
|
- `brain.generation()` reads it; it is persisted in the data directory and
|
||||||
|
**never reissued** — not across restarts, and not across `restore()`
|
||||||
|
(the counter is floored at its pre-restore value).
|
||||||
|
|
||||||
|
Every `Db` is pinned at one generation. `db.generation` and `db.timestamp`
|
||||||
|
identify the view; `newerDb.since(olderDb)` returns exactly the entity and
|
||||||
|
relationship ids that committed transactions touched between two views.
|
||||||
|
|
||||||
|
## Snapshot isolation for reads
|
||||||
|
|
||||||
|
**Guarantee:** a `Db` reads exactly the state at its pinned generation, no
|
||||||
|
matter what commits afterwards — including deletes. There are no torn reads,
|
||||||
|
no partially applied batches, and no drift over time.
|
||||||
|
|
||||||
|
- `brain.now()` pins the current generation in O(1).
|
||||||
|
- `brain.transact()` returns a `Db` pinned at the freshly committed
|
||||||
|
generation.
|
||||||
|
- `brain.asOf(generation | Date | snapshotPath)` pins past state.
|
||||||
|
|
||||||
|
While nothing has committed past the pin, reads delegate to the live fast
|
||||||
|
paths — pinning is free until history actually moves. Once later
|
||||||
|
transactions commit, the view keeps serving the **full query surface** at
|
||||||
|
its generation (see "Reading the past" below).
|
||||||
|
|
||||||
|
Writers are never blocked by readers and readers never block writers: a
|
||||||
|
pinned view stays valid because nothing overwrites the immutable records it
|
||||||
|
resolves from (the LMDB reader-pin model).
|
||||||
|
|
||||||
|
## Transaction atomicity
|
||||||
|
|
||||||
|
`brain.transact(ops)` executes a declarative batch — `add`, `update`,
|
||||||
|
`remove`, `relate`, `unrelate` — **atomically as exactly one generation**:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const db = await brain.transact([
|
||||||
|
{ op: 'add', id: orderId, type: NounType.Document, subtype: 'order', data: 'Order #1042' },
|
||||||
|
{ op: 'add', id: itemId, type: NounType.Thing, subtype: 'line-item', data: 'Widget x3' },
|
||||||
|
{ op: 'relate', from: orderId, to: itemId, type: VerbType.Contains, subtype: 'order-line' }
|
||||||
|
], { meta: { author: 'order-service', requestId: 'req-9f2' } })
|
||||||
|
|
||||||
|
db.receipt.ids // resolved id per operation, in input order
|
||||||
|
```
|
||||||
|
|
||||||
|
Either every operation applies, or none do and the store is byte-identical
|
||||||
|
to its pre-transaction state. Operation semantics mirror the corresponding
|
||||||
|
single-operation methods — validation, subtype enforcement, relationship
|
||||||
|
deduplication, delete cascades — and later operations may reference ids
|
||||||
|
created earlier in the same batch.
|
||||||
|
|
||||||
|
**The commit point is one atomic rename.** The durability protocol:
|
||||||
|
|
||||||
|
1. Before-images of every touched id are staged into an immutable
|
||||||
|
generation directory and **fsynced**.
|
||||||
|
2. The batch executes through the transaction manager (which has its own
|
||||||
|
operation-level rollback for non-crash failures).
|
||||||
|
3. The store manifest is replaced via atomic temp-file rename and fsynced.
|
||||||
|
**The rename is the commit** — a generation is committed if and only if
|
||||||
|
the manifest says so.
|
||||||
|
|
||||||
|
**Crash recovery:** on the next open, any staged generation above the
|
||||||
|
manifest watermark is an uncommitted transaction; its before-images are
|
||||||
|
restored (idempotently — recovery can itself crash and rerun) and derived
|
||||||
|
indexes never observe the rolled-back state. A crash anywhere before the
|
||||||
|
rename rolls back to the exact pre-transaction bytes; a crash after it keeps
|
||||||
|
the transaction.
|
||||||
|
|
||||||
|
Transaction metadata (`meta`) is reified Datomic-style: recorded in an
|
||||||
|
append-only transaction log readable via `brain.transactionLog()` — audit
|
||||||
|
fields live in the database, not in commit messages.
|
||||||
|
|
||||||
|
## Two levels of compare-and-swap
|
||||||
|
|
||||||
|
Concurrent `transact()` calls commit serially (snapshot-isolated batches).
|
||||||
|
For lost-update protection across a read–modify–write cycle, Brainy offers
|
||||||
|
CAS at two granularities:
|
||||||
|
|
||||||
|
| Granularity | Mechanism | Conflict error | Use when |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Per entity** | `_rev` + `{ op: 'update', ifRev }` (also on `brain.update()`) | `RevisionConflictError` | "This entity must not have changed since I read it." |
|
||||||
|
| **Whole store** | `transact(ops, { ifAtGeneration })` | `GenerationConflictError` | "*Nothing* may have committed since I read." |
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const view = brain.now()
|
||||||
|
const order = await view.get(orderId)
|
||||||
|
|
||||||
|
try {
|
||||||
|
await brain.transact(
|
||||||
|
[{ op: 'update', id: orderId, metadata: { total: recompute(order) }, ifRev: order._rev }],
|
||||||
|
{ ifAtGeneration: view.generation }
|
||||||
|
)
|
||||||
|
} catch (err) {
|
||||||
|
if (err instanceof GenerationConflictError) {
|
||||||
|
// Something committed since the pin — re-read and retry.
|
||||||
|
}
|
||||||
|
} finally {
|
||||||
|
await view.release()
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
An `ifRev` conflict on any operation rejects the **whole batch**; an
|
||||||
|
`ifAtGeneration` conflict is detected before anything is staged. Both leave
|
||||||
|
the store untouched and the generation counter unchanged. See
|
||||||
|
[Optimistic concurrency with `_rev`](../guides/optimistic-concurrency.md)
|
||||||
|
for the per-entity pattern in depth.
|
||||||
|
|
||||||
|
## Reading the past
|
||||||
|
|
||||||
|
`brain.asOf()` accepts a generation number, a `Date` (resolved through the
|
||||||
|
transaction log to the newest generation committed at or before it), or a
|
||||||
|
snapshot directory path. Historical views serve the **full query surface**
|
||||||
|
— `get()`, `find()` in every mode, semantic search, graph traversal,
|
||||||
|
cursors, aggregation — through two complementary paths:
|
||||||
|
|
||||||
|
- **Record path** (free): `get()`, metadata-level `find()`, and
|
||||||
|
filter-based `related()` resolve directly through the immutable record
|
||||||
|
layer. Ids untouched since the pin still ride the live fast paths.
|
||||||
|
- **Index path** (paid once): index-accelerated queries — semantic/vector
|
||||||
|
search, graph traversal, cursors, aggregation — are served by an
|
||||||
|
**at-generation index materialization** built lazily on first use:
|
||||||
|
Brainy reconstructs in-memory indexes over the exact record set at that
|
||||||
|
generation. This costs O(n at the pinned generation) time and memory,
|
||||||
|
**once per `Db`**, cached until `release()`. That is the open-core price
|
||||||
|
of historical index queries, stated plainly.
|
||||||
|
|
||||||
|
A native index provider implementing the optional
|
||||||
|
`VersionedIndexProvider` plugin capability serves the same historical reads
|
||||||
|
from its retained index segments **without any rebuild** — the materializer
|
||||||
|
is the correctness baseline, the provider is the accelerator. Semantics are
|
||||||
|
identical on both paths.
|
||||||
|
|
||||||
|
### History granularity — the honest limit
|
||||||
|
|
||||||
|
Generation *records* are written per `transact()` batch only.
|
||||||
|
Single-operation writes (`add`/`update`/`remove`/`relate`/… outside
|
||||||
|
`transact()`) advance the generation counter — so watermarks and CAS stay
|
||||||
|
sound — but do **not** stage before-images: they remain visible through
|
||||||
|
earlier pins and are not reported by `db.since()`. Code that needs pinned
|
||||||
|
isolation across its own writes uses `transact()`. This is the documented
|
||||||
|
8.0 contract, not an accident.
|
||||||
|
|
||||||
|
## Speculative writes: `db.with()`
|
||||||
|
|
||||||
|
`db.with(ops)` returns a new `Db` whose reads see the operations applied
|
||||||
|
**in memory, on top of the view** — Datomic's `with`. Nothing touches disk,
|
||||||
|
the generation counter, or index providers:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const current = brain.now()
|
||||||
|
const whatIf = await current.with([
|
||||||
|
{ op: 'update', id: employeeId, metadata: { team: 'platform' } }
|
||||||
|
])
|
||||||
|
|
||||||
|
await whatIf.find({ where: { team: 'platform' } }) // sees the change
|
||||||
|
await brain.get(employeeId) // unchanged — nothing committed
|
||||||
|
```
|
||||||
|
|
||||||
|
**The one boundary:** overlay entities carry no embeddings (`with()` never
|
||||||
|
invokes the embedder), so index-accelerated queries and `persist()` on a
|
||||||
|
speculative view throw `SpeculativeOverlayError` rather than returning
|
||||||
|
silently incomplete results. `get()`, metadata-filter `find()`, and
|
||||||
|
filter-based `related()` work fully on overlays. To get the full surface,
|
||||||
|
commit the same operations with `brain.transact()`.
|
||||||
|
|
||||||
|
## Retention and compaction
|
||||||
|
|
||||||
|
Historical records cost disk space, so retention is explicit:
|
||||||
|
|
||||||
|
- Every live `Db` holds a refcounted **pin**; a record-set is never
|
||||||
|
reclaimed while any pin could need it — pinned reads stay correct across
|
||||||
|
compaction, always.
|
||||||
|
- The **`retention`** knob governs auto-compaction (on every `flush()`/
|
||||||
|
`close()`): unset → ADAPTIVE (disk/RAM-pressure, zero-config) · `'all'` →
|
||||||
|
unbounded · `{ maxGenerations?, maxAge?, maxBytes? }` → explicit caps.
|
||||||
|
`brain.compactHistory({ maxGenerations?, maxAge?, maxBytes? })` reclaims
|
||||||
|
manually on the same caps, and records the **horizon** — `asOf()` below it
|
||||||
|
throws `GenerationCompactedError`, explicitly, never partial data.
|
||||||
|
- To keep a state readable forever, `persist()` it first: snapshots are
|
||||||
|
self-contained and unaffected by compaction of the source store.
|
||||||
|
|
||||||
|
Release `Db` values you do not keep (including the ones `transact()`
|
||||||
|
returns). A `FinalizationRegistry` backstop releases leaked pins at garbage
|
||||||
|
collection, but explicit `release()` is what makes compaction
|
||||||
|
deterministic.
|
||||||
|
|
||||||
|
## Durability: snapshots and restore
|
||||||
|
|
||||||
|
`db.persist(path)` cuts a **self-contained snapshot** under the store's
|
||||||
|
commit mutex, so no commit or compaction can interleave. On filesystem
|
||||||
|
storage it is built from **hard links**: because every data file is
|
||||||
|
immutable-by-rename, linking is safe — the snapshot is created without
|
||||||
|
copying entity data, shares disk space with the source, and later writes to
|
||||||
|
the source can never alter it (rewrites swap inodes; the snapshot keeps the
|
||||||
|
old bytes). Cross-device targets fall back to byte copies; in-memory stores
|
||||||
|
serialize to the same directory layout, producing a real, durable store.
|
||||||
|
|
||||||
|
Two rules keep snapshots honest:
|
||||||
|
|
||||||
|
- `persist()` requires the view to still be the store's **latest**
|
||||||
|
generation (a snapshot captures current bytes); a view that history has
|
||||||
|
moved past throws `GenerationConflictError` instead of persisting the
|
||||||
|
wrong state.
|
||||||
|
- `brain.restore(path, { confirm: true })` replaces the store's entire
|
||||||
|
state from a snapshot via byte copy (never links — the snapshot stays
|
||||||
|
independent), rebuilds all indexes, and floors the generation counter so
|
||||||
|
observed generation numbers are never reissued. Live pins do not survive
|
||||||
|
a restore — release them first (a warning is logged when any exist).
|
||||||
|
|
||||||
|
`Brainy.load(path)` (or `brain.asOf(path)`) opens a snapshot as a
|
||||||
|
self-contained **read-only** store with the full query surface, including
|
||||||
|
vector search.
|
||||||
|
|
||||||
|
## Reserved fields
|
||||||
|
|
||||||
|
Some field names belong to Brainy, not to your metadata. They live at **top
|
||||||
|
level** on every entity and relationship, have dedicated write paths, and may
|
||||||
|
never appear inside a `metadata` bag:
|
||||||
|
|
||||||
|
| Entities (nouns) | Relationships (verbs) | Canonical write path |
|
||||||
|
|---|---|---|
|
||||||
|
| `noun` | `verb` | the `type` param of `add()` / `relate()` |
|
||||||
|
| `subtype` | `subtype` | the `subtype` param |
|
||||||
|
| `visibility` | `visibility` | the `visibility` param (`'public'` \| `'internal'`) |
|
||||||
|
| `confidence` | `confidence` | the `confidence` param |
|
||||||
|
| `weight` | `weight` | the `weight` param |
|
||||||
|
| `service` | `service` | the `service` param (fixed at create time) |
|
||||||
|
| `data` | `data` | the `data` param |
|
||||||
|
| `createdBy` | `createdBy` | the `createdBy` param of `add()` (system-managed on verbs) |
|
||||||
|
| `createdAt`, `updatedAt`, `_rev` | `createdAt`, `updatedAt`, `_rev` | system-managed (`ifRev` for CAS) |
|
||||||
|
|
||||||
|
The canonical machine-readable lists are exported as
|
||||||
|
`RESERVED_ENTITY_FIELDS` and `RESERVED_RELATION_FIELDS` (defined in
|
||||||
|
`src/types/reservedFields.ts`, the single source of truth). Three layers
|
||||||
|
enforce the contract:
|
||||||
|
|
||||||
|
1. **Compile time** — every `metadata` param (`add`, `update`, `relate`,
|
||||||
|
`updateRelation`, and the matching `transact()` operations) rejects a
|
||||||
|
literal reserved key as a TypeScript error.
|
||||||
|
2. **Write time** — untyped (JavaScript) callers that pass one anyway are
|
||||||
|
normalized: user-settable fields (`confidence`, `weight`, `subtype`, and
|
||||||
|
`service`/`createdBy` at create time) are remapped to their dedicated
|
||||||
|
param — **top-level wins** when both are supplied — and system-managed
|
||||||
|
fields are dropped with a one-shot warning naming the correct write path.
|
||||||
|
`update({ metadata: { confidence: 0.9 } })` therefore behaves exactly
|
||||||
|
like `update({ confidence: 0.9 })`.
|
||||||
|
3. **Read time** — every read path (`get`, `find`, `search`,
|
||||||
|
`related`, batch reads, and historical `asOf()` materialization)
|
||||||
|
surfaces reserved fields **only at top level**: `entity.metadata` and
|
||||||
|
`relation.metadata` contain only your custom fields, always.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const id = await brain.add({
|
||||||
|
type: 'document', subtype: 'invoice',
|
||||||
|
data: 'Invoice #42', confidence: 0.95, // reserved → top-level params
|
||||||
|
metadata: { customer: 'acme', total: 129.5 } // custom fields only
|
||||||
|
})
|
||||||
|
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
entity.confidence // 0.95 — top level
|
||||||
|
entity.metadata // { customer: 'acme', total: 129.5 }
|
||||||
|
```
|
||||||
|
|
||||||
|
### Visibility — `public` / `internal` / `system`
|
||||||
|
|
||||||
|
`visibility` is a reserved tier that controls whether an entity or relationship
|
||||||
|
surfaces on Brainy's **default** user-facing reads. The absence of the field is
|
||||||
|
exactly equivalent to `'public'`.
|
||||||
|
|
||||||
|
| Tier | Counted in `getNounCount()` / `stats()`? | Returned by default `find()` / `related()`? | Opt-in |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `'public'` (default, or field absent) | yes | yes | — |
|
||||||
|
| `'internal'` | no | no | `find({ includeInternal: true })` / `related({ includeInternal: true })` |
|
||||||
|
| `'system'` | no | no | `find({ includeSystem: true })` / `related({ includeSystem: true })` |
|
||||||
|
|
||||||
|
- **`'public'`** — normal data. Counted and returned everywhere. Stored lean:
|
||||||
|
the field is omitted on disk for public records, so existing data needs no
|
||||||
|
migration.
|
||||||
|
- **`'internal'`** — your app's own bookkeeping (audit trails, derived caches,
|
||||||
|
scratch entities) that should not pollute default queries, counts, or
|
||||||
|
`stats()`, yet must stay retrievable on demand. Set it via the `visibility`
|
||||||
|
param; read it back with the `includeInternal` opt-in.
|
||||||
|
- **`'system'`** — Brainy's own plumbing (for example the Virtual File System
|
||||||
|
root entity). Hidden everywhere by default — even when `includeInternal` is
|
||||||
|
set — and surfaced only with the explicit `includeSystem` opt-in. The
|
||||||
|
`'system'` tier is **not** part of the public `add()` / `relate()` param type
|
||||||
|
(`'public' | 'internal'`); only internal Brainy code assigns it.
|
||||||
|
|
||||||
|
The opt-ins are applied as a **hard candidate filter** — hidden entities are
|
||||||
|
removed before `limit` / `offset` are applied, so a default `find({ limit: 10 })`
|
||||||
|
always returns ten *visible* results when that many exist, never a short page.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// App-internal scratch entity: present, retrievable, but out of the way.
|
||||||
|
await brain.add({ type: 'task', data: 'reindex job', visibility: 'internal' })
|
||||||
|
|
||||||
|
await brain.getNounCount() // unchanged — internal not counted
|
||||||
|
await brain.find({ type: 'task' }) // [] — hidden by default
|
||||||
|
await brain.find({ type: 'task', includeInternal: true }) // includes it
|
||||||
|
|
||||||
|
// A brand-new brain reports zero user entities even though the VFS root exists:
|
||||||
|
const fresh = new Brainy()
|
||||||
|
await fresh.init()
|
||||||
|
await fresh.getNounCount() // 0 — the root is visibility:'system'
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Note (8.0):** the structural `Contains` edges the VFS creates between
|
||||||
|
> directories and files are left at the default (`public`) visibility for now —
|
||||||
|
> only the VFS *root entity* is `'system'`. Marking those edges system requires
|
||||||
|
> companion changes to VFS traversal and is out of scope for this change.
|
||||||
|
|
||||||
|
## What is not guaranteed
|
||||||
|
|
||||||
|
Stated plainly, so nothing surprises you in production:
|
||||||
|
|
||||||
|
- **Single-writer.** Brainy is a single-writer, many-reader database
|
||||||
|
([multi-process model](./multi-process.md)). Transactions are atomic
|
||||||
|
within one writer process — there is no distributed or cross-process
|
||||||
|
transaction coordination.
|
||||||
|
- **History granularity.** Every write is its own immutable generation —
|
||||||
|
`transact()` batches AND single-operation `add`/`update`/`remove`/`relate`.
|
||||||
|
A pin always freezes against later writes, and every write is addressable
|
||||||
|
via `asOf()`. (`transact()` groups several operations into ONE atomic
|
||||||
|
generation; durability of single-op history is batched via async
|
||||||
|
group-commit — a hard crash can lose only the last un-flushed window's
|
||||||
|
*history*, never live data.)
|
||||||
|
- **Compacted history is gone.** `asOf()` below the compaction horizon
|
||||||
|
fails explicitly; persist what you must keep.
|
||||||
|
- **Counter persistence is coalesced for single-operation writes.** Durable
|
||||||
|
artifacts (records, manifests, snapshots) always persist the counter at
|
||||||
|
their own commit points, so a crash inside the coalescing window can lose
|
||||||
|
only counter values nothing durable ever referenced.
|
||||||
|
- **Speculative overlays are metadata-only readers.** Index-accelerated
|
||||||
|
queries on `with()` views throw rather than guess.
|
||||||
|
|
||||||
|
## Where to go next
|
||||||
|
|
||||||
|
- [Snapshots & Time Travel](../guides/snapshots-and-time-travel.md) — the
|
||||||
|
recipes: backup, restore, time-travel debugging, what-if analysis, audit
|
||||||
|
trails.
|
||||||
|
- [Optimistic concurrency with `_rev`](../guides/optimistic-concurrency.md)
|
||||||
|
— the per-entity CAS pattern.
|
||||||
|
- [ADR-001: Generational MVCC](../ADR-001-generational-mvcc.md) — the full
|
||||||
|
design record, including the persisted layout and the proof table.
|
||||||
117
docs/concepts/generation-fact-log.md
Normal file
117
docs/concepts/generation-fact-log.md
Normal file
|
|
@ -0,0 +1,117 @@
|
||||||
|
---
|
||||||
|
title: The Generation Fact Log
|
||||||
|
slug: concepts/generation-fact-log
|
||||||
|
public: true
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 6
|
||||||
|
description: Every committed write also appends a self-verifying "fact" — an after-image commit record — to an append-only log. What facts are, the crash-safety model, the scanFacts() streaming surface, family stamps, and how index providers consume the log for sequential heals.
|
||||||
|
next:
|
||||||
|
- concepts/consistency-model
|
||||||
|
- guides/snapshots-and-time-travel
|
||||||
|
---
|
||||||
|
|
||||||
|
# The Generation Fact Log
|
||||||
|
|
||||||
|
Since 8.4.0, every committed generation also appends a **fact** — a compact record of what each
|
||||||
|
touched entity or relationship *became* — to an append-only, checksummed log under
|
||||||
|
`_generations/facts/`. Where the generational history answers *"what did things look like
|
||||||
|
before?"* (before-images, powering `asOf()` and rollback), the fact log answers *"what happened,
|
||||||
|
in order?"* — one sequential, self-verifying stream of the store's present being written.
|
||||||
|
|
||||||
|
Nothing about querying changes. The fact log exists for three consumers:
|
||||||
|
|
||||||
|
1. **Index heals and rebuilds** — one sequential read in commit order replaces a per-entity
|
||||||
|
directory walk over millions of files.
|
||||||
|
2. **Incremental catch-up** — a derived index that knows which generation it reflects reads *just
|
||||||
|
the gap*, instead of rebuilding from scratch.
|
||||||
|
3. **Replay and audit tooling** — anything that wants the store's committed timeline as a stream.
|
||||||
|
|
||||||
|
## What a fact is
|
||||||
|
|
||||||
|
One fact per committed generation:
|
||||||
|
|
||||||
|
- **`generation`** and **`timestamp`** — which commit, when.
|
||||||
|
- **`ops`** — every write in that commit: `{ kind: 'noun' | 'verb', id, record }` where `record`
|
||||||
|
holds the entity's full after-image (both stored legs), or **`null` for a tombstone** — a
|
||||||
|
removal carries no body, by design.
|
||||||
|
- **`meta`** — the transaction metadata `transact()` was submitted with, when present.
|
||||||
|
- **`blobHashes`** — content-blob references, for exact reclamation accounting.
|
||||||
|
|
||||||
|
Facts accumulate **from the first write after upgrading** — pre-existing history is not
|
||||||
|
retroactively converted, and consumers fall back to the enumeration walk when no log exists.
|
||||||
|
|
||||||
|
## Crash safety, in one paragraph
|
||||||
|
|
||||||
|
Facts are appended and fsynced **inside the same durability window as the commit itself**, before
|
||||||
|
the commit point — so after a crash, the log can only ever be *ahead* of committed truth, never
|
||||||
|
behind it with a hole. On open, the store reconciles the log back to the committed watermark:
|
||||||
|
torn tails are detected by per-record checksums and cut; whole records beyond the watermark are
|
||||||
|
truncated. The invariant every reader can rely on: **an absent generation was never committed; a
|
||||||
|
present fact was.** `transact()` facts are durable the moment `transact()` returns; single-op
|
||||||
|
facts share the same group-commit flush as the rest of their generation, so a hard kill loses the
|
||||||
|
fact and the generation *together* — never a torn state.
|
||||||
|
|
||||||
|
## Reading the log
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const scan = brain.scanFacts({ fromGeneration: 1 })
|
||||||
|
if (scan) {
|
||||||
|
// Telemetry up front — progress bars get a denominator from second zero.
|
||||||
|
console.log(scan.headGeneration, scan.segmentCount, scan.approxFactCount)
|
||||||
|
|
||||||
|
for await (const batch of scan.batches()) {
|
||||||
|
// Each batch: { facts, firstGeneration, lastGeneration, factCount, byteSize, segmentId }
|
||||||
|
for (const fact of batch.facts) {
|
||||||
|
for (const op of fact.ops) {
|
||||||
|
if (op.record === null) {
|
||||||
|
// a tombstone: op.id was removed in this generation
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log(scan.summary()) // { factsYielded, segmentsRead } — the cross-check
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `scanFacts()` returns `null` when the store hosts no fact log (older store, or a storage adapter
|
||||||
|
without binary append support) — fall back to enumerating entities.
|
||||||
|
- Scans run against a **snapshot**: facts appended after the scan opens never bleed in, each fact
|
||||||
|
is yielded exactly once, and a detected gap aborts loudly — never a silent skip.
|
||||||
|
- `brain.factSegmentPaths()` returns the immutable, *sealed* segment files for zero-copy consumers
|
||||||
|
(the append-mutable tail is excluded — read it through `scanFacts()`).
|
||||||
|
|
||||||
|
## Family stamps: how a projection proves it's current
|
||||||
|
|
||||||
|
Anything derived from the store — an index, the entity file tree itself — carries a **family
|
||||||
|
stamp**: a small JSON record of *which committed generation the projection reflects*
|
||||||
|
(`sourceGeneration`) plus the invariants that verify it whole (exact per-file byte sizes for
|
||||||
|
bounded families; rollup invariants like entity counts for unbounded ones). At open, coherence is
|
||||||
|
a **comparison**, not a walk:
|
||||||
|
|
||||||
|
- stamp equals the committed watermark and invariants hold → serve;
|
||||||
|
- stamp is behind → the projection reads just the gap from the fact log;
|
||||||
|
- invariants fail → loud, named divergence — `brain.repairIndex()` rebuilds from canonical and
|
||||||
|
re-stamps.
|
||||||
|
|
||||||
|
The verifier is exported (`verifyFamilyStamp`) so every projection — TypeScript or native — runs
|
||||||
|
literally the same check.
|
||||||
|
|
||||||
|
## For plugin authors: the storage capability
|
||||||
|
|
||||||
|
Index providers receive the storage adapter, not the brain — so the host wires the log onto it.
|
||||||
|
Feature-detect and prefer the stream; fall back to enumeration:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const scan = storage.scanFacts?.({ fromGeneration: stamp.sourceGeneration + 1 })
|
||||||
|
if (scan) {
|
||||||
|
// sequential catch-up from the log
|
||||||
|
} else {
|
||||||
|
// enumeration walk (older store or adapter)
|
||||||
|
}
|
||||||
|
const committed = storage.committedGeneration?.() // the watermark stamps compare against
|
||||||
|
```
|
||||||
|
|
||||||
|
Providers must never construct their own reader over the log's files — the open path belongs to
|
||||||
|
the single writer (it reconciles the log at open); the capability is the sanctioned seam.
|
||||||
153
docs/concepts/multi-process.md
Normal file
153
docs/concepts/multi-process.md
Normal file
|
|
@ -0,0 +1,153 @@
|
||||||
|
---
|
||||||
|
title: Multi-Process Model
|
||||||
|
slug: concepts/multi-process
|
||||||
|
public: true
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 5
|
||||||
|
description: How Brainy coordinates a single writer with any number of readers on a filesystem data directory — and how to safely inspect a live store.
|
||||||
|
next:
|
||||||
|
- guides/inspection
|
||||||
|
---
|
||||||
|
|
||||||
|
# Multi-Process Model
|
||||||
|
|
||||||
|
Brainy is a **single-writer, many-reader** database when backed by filesystem
|
||||||
|
storage. This page explains the model, the guarantees, and the safe ways to
|
||||||
|
inspect a live store from a second process.
|
||||||
|
|
||||||
|
## The rule
|
||||||
|
|
||||||
|
For one data directory:
|
||||||
|
|
||||||
|
- **One writer** at a time. The writer acquires an exclusive lock on the
|
||||||
|
directory at `init()` and releases it on `close()`.
|
||||||
|
- **Any number of readers**, concurrent with each other and with the writer.
|
||||||
|
Readers open via `Brainy.openReadOnly()` — they never touch the writer
|
||||||
|
lock.
|
||||||
|
|
||||||
|
Any attempt to open a second writer on the same directory throws:
|
||||||
|
|
||||||
|
```
|
||||||
|
BrainyError: Another writer holds this Brainy directory.
|
||||||
|
PID: 1774431 on host app-host-1
|
||||||
|
Started: 2026-05-15T14:22:11Z
|
||||||
|
Heartbeat: 2026-05-15T14:22:34Z
|
||||||
|
Version: 7.21.0
|
||||||
|
Directory: /data/brain
|
||||||
|
```
|
||||||
|
|
||||||
|
This is intentional. Two writers sharing a directory would silently corrupt
|
||||||
|
in-memory indexes and produce wrong query results — the worst possible default
|
||||||
|
for an operations tool.
|
||||||
|
|
||||||
|
## Why a lock?
|
||||||
|
|
||||||
|
Brainy keeps its primary indexes (HNSW, metadata, graph adjacency) in memory.
|
||||||
|
On disk, those indexes are persisted incrementally as writes flush. A second
|
||||||
|
process opening the same directory:
|
||||||
|
|
||||||
|
- Loads the *persisted* state into a fresh in-memory copy.
|
||||||
|
- Has no awareness of writes the first process buffered but hasn't flushed.
|
||||||
|
- Will overwrite the persisted state on its own next flush, racing the first
|
||||||
|
process and corrupting whichever wins.
|
||||||
|
|
||||||
|
The fix is the lock: refuse to open a second writer. SQLite has done the same
|
||||||
|
since the late 1990s (`SQLITE_BUSY`).
|
||||||
|
|
||||||
|
## What about Cor?
|
||||||
|
|
||||||
|
Brainy + Cor compose cleanly under this model:
|
||||||
|
|
||||||
|
- Cor stores its column-index segments inside the same `rootDir` (under
|
||||||
|
`indexes/_column_index/{field}/`).
|
||||||
|
- Segments (`*.cidx` files) are **immutable** once written. Cor mmaps them
|
||||||
|
read-only.
|
||||||
|
- The `MANIFEST.json` per field is updated via atomic rename — readers see
|
||||||
|
either the old or new manifest, never a torn file.
|
||||||
|
|
||||||
|
A reader process can safely mmap Cor segments alongside a live writer
|
||||||
|
without coordination. The single Brainy writer lock at
|
||||||
|
`<rootDir>/locks/_writer.lock` covers Cor too, because Cor segment
|
||||||
|
writes happen on the writer's side.
|
||||||
|
|
||||||
|
## Stale-lock detection
|
||||||
|
|
||||||
|
If a writer crashes or is forcibly killed, its lock file is left behind. To
|
||||||
|
avoid a permanently-jammed directory, Brainy treats a lock as stale when:
|
||||||
|
|
||||||
|
1. The recorded `hostname` equals the current host (cross-host PID checks
|
||||||
|
are unsafe), AND
|
||||||
|
2. The recorded `pid` is no longer alive (`process.kill(pid, 0)` returns
|
||||||
|
`ESRCH`), OR the `lastHeartbeat` field is older than 60 seconds.
|
||||||
|
|
||||||
|
A live writer rewrites `lastHeartbeat` every 10 seconds, so a hung writer
|
||||||
|
that's missed several heartbeats is treated as dead. Stale locks are
|
||||||
|
overwritten with a warning.
|
||||||
|
|
||||||
|
If stale detection cannot prove the existing lock is dead — for example, a
|
||||||
|
crashed writer on a different host writing to a shared filesystem — pass
|
||||||
|
`{ force: true }` to override. A warning is logged either way.
|
||||||
|
|
||||||
|
## Heartbeat and shutdown
|
||||||
|
|
||||||
|
The heartbeat interval rewrites the lock file every 10 seconds. The timer
|
||||||
|
is unref'd, so it does not keep the event loop alive on its own.
|
||||||
|
|
||||||
|
On normal shutdown the writer releases the lock in `close()`. The shutdown
|
||||||
|
hooks Brainy registers for `SIGTERM`, `SIGINT`, and `beforeExit` also
|
||||||
|
release the lock so a container restart doesn't strand the directory.
|
||||||
|
|
||||||
|
## How to inspect a live writer
|
||||||
|
|
||||||
|
Use `Brainy.openReadOnly()`. It does not acquire the writer lock, so it
|
||||||
|
coexists with whatever the writer is doing:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const reader = await Brainy.openReadOnly({
|
||||||
|
storage: { type: 'filesystem', path: '/data/brain' }
|
||||||
|
})
|
||||||
|
|
||||||
|
const stats = await reader.stats()
|
||||||
|
const bookings = await reader.find({ where: { entityType: 'booking' } })
|
||||||
|
|
||||||
|
await reader.close()
|
||||||
|
```
|
||||||
|
|
||||||
|
What the reader sees reflects the writer's most recent **flush** to disk. If
|
||||||
|
you need fresher state, ask the writer to flush before opening:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const reader = await Brainy.openReadOnly({
|
||||||
|
storage: { type: 'filesystem', path: '/data/brain' }
|
||||||
|
})
|
||||||
|
|
||||||
|
const acked = await reader.requestFlush({ timeoutMs: 5000 })
|
||||||
|
if (!acked) {
|
||||||
|
console.warn('Writer did not respond; results reflect last natural flush.')
|
||||||
|
}
|
||||||
|
|
||||||
|
const fresh = await reader.find({ where: { entityType: 'booking' } })
|
||||||
|
```
|
||||||
|
|
||||||
|
The CLI `brainy inspect` subcommands all do this for you by default
|
||||||
|
(`--no-fresh` to opt out).
|
||||||
|
|
||||||
|
## What's not enforced (yet)
|
||||||
|
|
||||||
|
- **Non-filesystem backends** are out of scope in 8.0, which ships only the
|
||||||
|
filesystem and memory adapters. A custom `BaseStorage` subclass that is not
|
||||||
|
filesystem-backed does not enforce multi-process locking by default: two
|
||||||
|
processes can both succeed at `init()` in writer mode and clobber each
|
||||||
|
other's writes. A best-effort warning is logged in writer mode against a
|
||||||
|
non-filesystem backend.
|
||||||
|
- **Long-running readers** do not automatically pick up new Cor segments
|
||||||
|
the writer publishes. One-shot inspector calls re-open the store and see
|
||||||
|
fresh segments; a reader that stays open for hours sees its column store
|
||||||
|
as-of the time it opened.
|
||||||
|
|
||||||
|
## Reading material
|
||||||
|
|
||||||
|
- `Brainy.openReadOnly()` — [API reference](../api/brainy.md)
|
||||||
|
- `brainy inspect` — [inspection guide](../guides/inspection.md)
|
||||||
|
- Cor columnar storage — see `node_modules/@soulcraft/cor/README.md`
|
||||||
189
docs/concepts/storage-adapters.md
Normal file
189
docs/concepts/storage-adapters.md
Normal file
|
|
@ -0,0 +1,189 @@
|
||||||
|
---
|
||||||
|
title: Storage Adapter Inheritance Contract
|
||||||
|
slug: concepts/storage-adapters
|
||||||
|
public: true
|
||||||
|
category: concepts
|
||||||
|
template: concept
|
||||||
|
order: 6
|
||||||
|
description: How storage adapters extend Brainy's BaseStorage and FileSystemStorage to inherit multi-process safety, what plugin authors need to override, and the contract Brainy promises to keep stable.
|
||||||
|
next:
|
||||||
|
- concepts/multi-process
|
||||||
|
- guides/inspection
|
||||||
|
---
|
||||||
|
|
||||||
|
# Storage Adapter Inheritance Contract
|
||||||
|
|
||||||
|
Brainy's storage layer is designed for plugins to extend cleanly. A plugin
|
||||||
|
that subclasses `BaseStorage` or `FileSystemStorage` inherits new behavior
|
||||||
|
Brainy adds over time without code changes — provided the import resolution
|
||||||
|
brings in the right Brainy version at runtime.
|
||||||
|
|
||||||
|
This page documents what plugin authors can rely on, what they need to
|
||||||
|
override, and the install-time failure modes that the defensive
|
||||||
|
`hasStorageMethod()` guard exists to handle.
|
||||||
|
|
||||||
|
## The class hierarchy
|
||||||
|
|
||||||
|
```
|
||||||
|
BaseStorageAdapter (counts, batch ops, multi-tenancy hooks)
|
||||||
|
↑
|
||||||
|
BaseStorage (type-statistics, lifecycle helpers, generation
|
||||||
|
hooks, default no-op multi-process methods)
|
||||||
|
↑
|
||||||
|
FileSystemStorage (real filesystem I/O, writer-lock implementation,
|
||||||
|
flush-request watcher, atomic writes)
|
||||||
|
↑
|
||||||
|
<your plugin's storage> (Cor's MmapFileSystemStorage, etc.)
|
||||||
|
```
|
||||||
|
|
||||||
|
When you `extend FileSystemStorage`, your adapter inherits every method on
|
||||||
|
the chain — including the ones Brainy adds in a later release — for free.
|
||||||
|
JavaScript's prototype chain resolves method lookups dynamically; nothing
|
||||||
|
about the inheritance is baked in at class-definition time.
|
||||||
|
|
||||||
|
## What you inherit for free
|
||||||
|
|
||||||
|
A plugin whose storage class extends `FileSystemStorage` automatically gets
|
||||||
|
the full multi-process safety surface:
|
||||||
|
|
||||||
|
| Method | What it does | Override? |
|
||||||
|
|---|---|---|
|
||||||
|
| `acquireWriterLock(opts)` | Write `_writer.lock`, start heartbeat, throw on conflict | No |
|
||||||
|
| `releaseWriterLock()` | Clean up lock file + heartbeat timer | No |
|
||||||
|
| `readWriterLock()` | Read lock file as `WriterLockInfo` | No |
|
||||||
|
| `startFlushRequestWatcher(cb)` | Poll `_flush_requests/` and invoke `cb` | No |
|
||||||
|
| `stopFlushRequestWatcher()` | Stop the polling timer | No |
|
||||||
|
| `requestFlushOverFilesystem(timeoutMs)` | Drop a `.req`, await `.ack` | No |
|
||||||
|
| `supportsMultiProcessLocking()` | Return `false` (default) | **Yes — override to `true`** |
|
||||||
|
|
||||||
|
The only required override is the capability flag. Returning `true` from
|
||||||
|
`supportsMultiProcessLocking()` is the signal Brainy uses to decide whether
|
||||||
|
to call `acquireWriterLock()` at init.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { FileSystemStorage } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
export class MmapFileSystemStorage extends FileSystemStorage {
|
||||||
|
public supportsMultiProcessLocking(): boolean {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
// ... your mmap-specific overrides ...
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
That's the full ceremony for inheriting multi-process safety.
|
||||||
|
|
||||||
|
## When NOT to extend FileSystemStorage
|
||||||
|
|
||||||
|
If your storage is **not filesystem-backed** (a custom
|
||||||
|
network backend), extend `BaseStorage` directly:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { BaseStorage } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
export class MyCloudStorage extends BaseStorage {
|
||||||
|
// BaseStorage's default no-op implementations of the multi-process
|
||||||
|
// methods stay in effect. `supportsMultiProcessLocking()` returns false
|
||||||
|
// by default — keep it that way unless you've implemented an object-
|
||||||
|
// versioned lease or similar cross-process synchronization for your
|
||||||
|
// backend.
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Brainy treats cloud backends as not-multi-process-safe by default and logs a
|
||||||
|
one-line warning at init. That's the correct behavior until cloud locking
|
||||||
|
ships (currently out of scope — see
|
||||||
|
[`concepts/multi-process`](./multi-process.md)).
|
||||||
|
|
||||||
|
## What `hasStorageMethod()` actually guards against
|
||||||
|
|
||||||
|
The defensive check at every new-storage-method call site (`brainy.ts`,
|
||||||
|
`hasStorageMethod(name)`) does **not** exist to handle "plugin bundles a
|
||||||
|
stale BaseStorage." Plugins ship a dist that preserves the dynamic ESM
|
||||||
|
import (verify in your plugin's `dist/`: `import { FileSystemStorage } from
|
||||||
|
'@soulcraft/brainy'` is not rewritten to a vendored copy). The prototype
|
||||||
|
chain at runtime resolves to whatever Brainy version your consumer has
|
||||||
|
installed.
|
||||||
|
|
||||||
|
`hasStorageMethod()` protects against **build/install artifacts** that break
|
||||||
|
the prototype chain at the consumer-app level:
|
||||||
|
|
||||||
|
- **Stale `node_modules`** — a lingering install from before the consumer
|
||||||
|
upgraded Brainy. The package.json says `@soulcraft/brainy@7.22.0` but
|
||||||
|
`node_modules/@soulcraft/brainy` is still 7.20.x.
|
||||||
|
- **Lockfile drift** — `bun.lockb` / `package-lock.json` pins a brainy
|
||||||
|
version older than the package.json range, and `bun install` honors the
|
||||||
|
lockfile.
|
||||||
|
- **Docker layer cache** — the image reuses a `node_modules` from an
|
||||||
|
earlier build that predates the brainy bump.
|
||||||
|
- **Bundler quirks** — some bundlers (esbuild, webpack) flatten the
|
||||||
|
prototype chain at build time and lose later prototype mutations. Brainy
|
||||||
|
doesn't mutate prototypes at runtime, but bundler behavior can still
|
||||||
|
cause method lookups to fail in non-Node environments.
|
||||||
|
|
||||||
|
In any of those, calling `storage.acquireWriterLock(...)` unconditionally
|
||||||
|
throws `TypeError: storage.acquireWriterLock is not a function`. The guard
|
||||||
|
turns that into a logged warning + graceful no-op so the app still boots,
|
||||||
|
and the warning names the adapter class plus a remediation hint:
|
||||||
|
|
||||||
|
```
|
||||||
|
[brainy] Storage adapter `MmapFileSystemStorage` is missing the multi-process
|
||||||
|
methods on its prototype chain. Writer locking and the flush-request RPC are
|
||||||
|
disabled for this directory. Likely fix: clean install (`rm -rf node_modules
|
||||||
|
bun.lockb && bun install`) or rebuild your container image to refresh
|
||||||
|
`@soulcraft/brainy` to ≥7.21. See docs/concepts/storage-adapters.md.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Authoring a new storage adapter — minimum checklist
|
||||||
|
|
||||||
|
1. **Extend the right base class.**
|
||||||
|
- Filesystem-backed → `FileSystemStorage`.
|
||||||
|
- Cloud / network / custom → `BaseStorage`.
|
||||||
|
|
||||||
|
2. **Override the capability flag.** If filesystem-backed:
|
||||||
|
```typescript
|
||||||
|
public supportsMultiProcessLocking(): boolean { return true }
|
||||||
|
```
|
||||||
|
|
||||||
|
3. **Don't `super.X()`-wrap the multi-process methods**. They're inherited;
|
||||||
|
leaving them inherited means `hasStorageMethod()` finds them on the
|
||||||
|
prototype chain. Re-declaring them as `super.X()` wrappers makes the
|
||||||
|
helper resolve to your wrapper, which can fool the guard if your
|
||||||
|
constructor runs before the super class initializes.
|
||||||
|
|
||||||
|
4. **Do override `init()` / `flush()` / `close()`** as needed. Always call
|
||||||
|
`super.init()` / `super.flush()` / `super.close()` first so the
|
||||||
|
filesystem prep, writer-lock acquisition, and lock release happen in the
|
||||||
|
expected order.
|
||||||
|
|
||||||
|
5. **Verify the inheritance.** A one-line smoke test in your plugin's
|
||||||
|
`__tests__/`:
|
||||||
|
```typescript
|
||||||
|
const s = new MyStorage(rootDir)
|
||||||
|
await s.init()
|
||||||
|
assert(typeof s.acquireWriterLock === 'function')
|
||||||
|
assert(s.supportsMultiProcessLocking() === true)
|
||||||
|
```
|
||||||
|
If `acquireWriterLock` is undefined the prototype chain is broken at
|
||||||
|
install time — fix install, not your plugin.
|
||||||
|
|
||||||
|
6. **Pin your peer dep generously.** `"peerDependencies": {
|
||||||
|
"@soulcraft/brainy": "^7.21.0" }` accepts any compatible 7.x. Don't pin
|
||||||
|
to an exact patch unless you're tracking a known regression.
|
||||||
|
|
||||||
|
## Future direction
|
||||||
|
|
||||||
|
The 7 multi-process methods are currently defaults on `BaseStorage`. A
|
||||||
|
future refactor may extract them into a `MultiProcessSafeStorage`
|
||||||
|
interface/mixin for cleaner separation — only adapters that opt in would
|
||||||
|
expose them. This would require a minor bump and is tracked as an internal
|
||||||
|
follow-up; consumers don't need to anticipate the change.
|
||||||
|
|
||||||
|
## Reading material
|
||||||
|
|
||||||
|
- [`concepts/multi-process`](./multi-process.md) — the writer-lock model,
|
||||||
|
heartbeat semantics, what the lock protects.
|
||||||
|
- [`guides/inspection`](../guides/inspection.md) — `brainy inspect` and the
|
||||||
|
read-only mode.
|
||||||
|
- `node_modules/@soulcraft/brainy/dist/storage/baseStorage.d.ts` — the
|
||||||
|
authoritative type signatures for every method this page references.
|
||||||
155
docs/eli5.md
Normal file
155
docs/eli5.md
Normal file
|
|
@ -0,0 +1,155 @@
|
||||||
|
---
|
||||||
|
title: What is Brainy?
|
||||||
|
slug: getting-started/what-is-brainy
|
||||||
|
public: true
|
||||||
|
category: getting-started
|
||||||
|
template: guide
|
||||||
|
order: 0
|
||||||
|
description: Plain-language guide covering what Brainy does, how it compares to other tools, and what you can build with it. No jargon, no code — just clear analogies.
|
||||||
|
next:
|
||||||
|
- getting-started/installation
|
||||||
|
- getting-started/quick-start
|
||||||
|
---
|
||||||
|
|
||||||
|
# Brainy and Cor — Explained Simply
|
||||||
|
|
||||||
|
*A plain-language guide for anyone who wants to understand what this thing actually does.*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What is Brainy?
|
||||||
|
|
||||||
|
Imagine you have the world's smartest librarian.
|
||||||
|
|
||||||
|
You walk up and say *"I'm looking for something about climate change — but only books published after 2020, and only ones written by authors I've already read."* A normal library would make you dig through a card catalogue, then cross-reference a list of authors, then scan the shelves yourself. That takes a while.
|
||||||
|
|
||||||
|
Your smart librarian does all three at the same time — in less than the time it takes to blink.
|
||||||
|
|
||||||
|
That's Brainy. It's a knowledge database that can search by **meaning**, follow **connections**, and filter by **labels** — all at once, in a single question.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Three Superpowers
|
||||||
|
|
||||||
|
### 1. Meaning Search (the "fuzzy" superpower)
|
||||||
|
|
||||||
|
When you search for "automobile," Brainy also finds results about "car," "vehicle," and "sedan" — because it understands what words *mean*, not just how they're spelled. It reads your data the way a person would, not the way a search box does.
|
||||||
|
|
||||||
|
Think of it like the librarian who finds books on "heartbreak" when you ask for something about "loneliness."
|
||||||
|
|
||||||
|
### 2. Relationship Walking (the "follow the thread" superpower)
|
||||||
|
|
||||||
|
Every piece of information can be connected to other pieces. A Person *works at* a Company. A Project *depends on* a Tool. A Recipe *contains* Ingredients.
|
||||||
|
|
||||||
|
Brainy can follow these connections across many hops in one step. Ask for "everything connected to this author, two steps out" and Brainy returns the author's books, the books' publishers, the publishers' other authors — without you needing to chain four separate lookups yourself.
|
||||||
|
|
||||||
|
Think of it like the librarian who not only hands you the book you asked for, but also knows which shelf it came from, who donated it, and what other books arrived in the same donation.
|
||||||
|
|
||||||
|
### 3. Label Filtering (the "narrow it down" superpower)
|
||||||
|
|
||||||
|
Sometimes meaning and connections aren't enough — you need precision. "Only recipes with fewer than 500 calories." "Only events from last week." "Only documents tagged as urgent."
|
||||||
|
|
||||||
|
Brainy can narrow any result set down by exact labels or ranges in the same breath as the other two searches. No extra steps.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What Else Can It Do?
|
||||||
|
|
||||||
|
- **Virtual file cabinet.** Brainy includes a full filesystem you can use to store, organize, and semantically search files — PDFs, documents, anything — the same way you search everything else.
|
||||||
|
|
||||||
|
- **Live dashboards.** You can define running totals that Brainy keeps updated automatically — things like "total sales this month by region" or "average response time per service." Every time new data comes in, the numbers stay current with no manual recalculation.
|
||||||
|
|
||||||
|
- **Time travel.** Every committed change becomes part of the database's history. You can pin the current state as a frozen view, see the whole knowledge base exactly as it was last week, try out changes in a scratch copy that never touches the real data, and take instant backups.
|
||||||
|
|
||||||
|
- **Universal vocabulary.** Brainy ships with a shared language of 42 kinds of things (Person, Document, Task, Concept, Event…) and 127 kinds of connections (Contains, DependsOn, Creates, RelatedTo…). This means data from different sources speaks the same language without you having to translate.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What is Cor?
|
||||||
|
|
||||||
|
Cor is a turbocharger for Brainy.
|
||||||
|
|
||||||
|
Same car. Same controls. Same fuel. You just swap in a faster engine under the hood, and everything that used to take a moment now happens instantly.
|
||||||
|
|
||||||
|
Technically, Cor is an optional plugin written in Rust — a lower-level language that runs much closer to the raw metal of your processor. It plugs into Brainy and takes over the most compute-intensive work: the distance calculations that power meaning search, the number-crunching behind live aggregates, and the set operations that drive label filtering.
|
||||||
|
|
||||||
|
You install it with one line, register it with one call, and Brainy automatically uses it everywhere it can help.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## How Much Faster?
|
||||||
|
|
||||||
|
Plain language:
|
||||||
|
|
||||||
|
- **Searches** go from "the blink of an eye" to "faster than a blink." The overall speedup is **5.2× on average** across all operations.
|
||||||
|
- **Live aggregates** are rebuilt using all CPU cores in parallel, so re-indexing large datasets takes a fraction of the time.
|
||||||
|
- **Analytics** that aren't even possible in pure JavaScript — real-time anomaly detection, streaming percentile estimates, approximate unique counts — become available because Cor brings the native capabilities required to run them efficiently.
|
||||||
|
|
||||||
|
If Brainy is what makes knowledge fast, Cor is what makes Brainy feel instant.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What Does Brainy Replace?
|
||||||
|
|
||||||
|
Most applications that need to store and search knowledge end up stitching together several specialized tools. Brainy replaces all of them with one — a single free, open-source library in place of multiple paid services.
|
||||||
|
|
||||||
|
### Before and After
|
||||||
|
|
||||||
|
**Before Brainy** — a pile of services:
|
||||||
|
- Pinecone (vectors) + Neo4j (graph) + MongoDB (docs)
|
||||||
|
- Algolia (search) + Redis (cache) + PostgreSQL + pgvector
|
||||||
|
- Plus glue code, sync jobs, ETL pipelines, and 3am incidents
|
||||||
|
|
||||||
|
**After Brainy** — one thing:
|
||||||
|
Search, graph, filter, files, time travel, and imports — unified in a single library.
|
||||||
|
|
||||||
|
### What Each Tool Is Missing
|
||||||
|
|
||||||
|
| Tool | Search | Graph | Filter | VFS | Time travel | Import |
|
||||||
|
|---|:---:|:---:|:---:|:---:|:---:|:---:|
|
||||||
|
| **Brainy** | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||||
|
| *— Vector databases —* | | | | | | |
|
||||||
|
| Pinecone | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| Weaviate | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| Qdrant | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| Chroma | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| *— Graph databases —* | | | | | | |
|
||||||
|
| Neo4j | ❌ | ✅ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| *— Document stores —* | | | | | | |
|
||||||
|
| MongoDB | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| Firestore | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| DynamoDB | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| *— Relational + vector —* | | | | | | |
|
||||||
|
| PostgreSQL + pgvector | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| MySQL | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| SQLite | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| *— Search engines —* | | | | | | |
|
||||||
|
| Elasticsearch | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| Algolia | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ |
|
||||||
|
| *— Cache —* | | | | | | |
|
||||||
|
| Redis | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
||||||
|
|
||||||
|
Brainy is the only row with every box checked. And it runs all of them in a single query — no stitching services together.
|
||||||
|
|
||||||
|
### One library, any scale
|
||||||
|
|
||||||
|
Brainy scales from a quick experiment to serious production datasets without changing a line of code. Small datasets live entirely in memory. Larger ones spill to disk, where Brainy shards and compresses everything automatically. Need a backup or a copy? Snapshots are instant — the same API the whole way.
|
||||||
|
|
||||||
|
Add Cor and you also unlock memory-mapped storage — aggregate state lives directly in the operating system's memory with zero serialization overhead, as fast as the hardware allows.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What Can You Build?
|
||||||
|
|
||||||
|
### Common applications
|
||||||
|
|
||||||
|
- **AI agents with persistent memory** — Give any AI an always-on, self-organizing knowledge graph that persists between sessions and across agents.
|
||||||
|
- **Searchable knowledge bases** — Build institutional memory that links documents automatically and surfaces answers across the full web of related information.
|
||||||
|
- **Semantic document search** — Index PDFs, code, or media and find them by meaning, not just keywords.
|
||||||
|
- **Relationship-aware recommendations** — Power product catalogs or content platforms where every recommendation understands what connects to what.
|
||||||
|
- **Safe experiments** — Test risky changes against a scratch copy of the knowledge base, audit exactly what changed and when, and roll back to any snapshot instantly.
|
||||||
|
- **Unified business platforms** — Combine booking, CRM, inventory, and analytics in one queryable knowledge graph with no sync pipeline.
|
||||||
|
|
||||||
|
### What Brainy is good at
|
||||||
|
|
||||||
|
Brainy is the engine underneath production systems that need to combine semantic search, structured filtering, and graph traversal in a single query — agent memory, knowledge-base platforms, business operations consoles, multi-agent coordination, and more. The combination of vector + graph + metadata search in one indexed call is what differentiates it from running three engines side by side.
|
||||||
|
|
@ -1,415 +0,0 @@
|
||||||
# 🚀 Brainy 2.0 - Complete Feature List
|
|
||||||
|
|
||||||
> **The Truth**: Brainy is MORE powerful than previously documented! This is the complete list of ALL implemented features.
|
|
||||||
|
|
||||||
## 🧠 Core Intelligence Engine
|
|
||||||
|
|
||||||
### Triple Intelligence System ✅
|
|
||||||
Unified query system that automatically combines:
|
|
||||||
- **Vector Search**: HNSW-indexed semantic similarity (O(log n) performance)
|
|
||||||
- **Graph Traversal**: Relationship-based discovery
|
|
||||||
- **Field Filtering**: Metadata and attribute queries
|
|
||||||
- **Auto-optimization**: Queries are automatically optimized based on data patterns
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// All three intelligences work together automatically
|
|
||||||
const results = await brain.find({
|
|
||||||
like: 'AI research', // Vector search
|
|
||||||
where: { year: 2024 }, // Metadata filtering
|
|
||||||
connected: { to: authorId } // Graph traversal
|
|
||||||
})
|
|
||||||
```
|
|
||||||
|
|
||||||
### Neural Query Understanding ✅
|
|
||||||
- **220+ embedded patterns** for query intent detection
|
|
||||||
- Natural language query processing
|
|
||||||
- Automatic query type detection
|
|
||||||
- Query rewriting and optimization
|
|
||||||
|
|
||||||
## 🔧 12+ Production Augmentations
|
|
||||||
|
|
||||||
### 1. WAL (Write-Ahead Logging) ✅
|
|
||||||
```typescript
|
|
||||||
import { WALAugmentation } from 'brainy'
|
|
||||||
// Full crash recovery, checkpointing, replay
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Entity Registry ✅
|
|
||||||
```typescript
|
|
||||||
import { EntityRegistryAugmentation } from 'brainy'
|
|
||||||
// Bloom filter-based deduplication for streaming data
|
|
||||||
// Handles millions of entities with minimal memory
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Auto-Register Entities ✅
|
|
||||||
```typescript
|
|
||||||
import { AutoRegisterEntitiesAugmentation } from 'brainy'
|
|
||||||
// Automatically extracts and registers entities from text
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. Intelligent Verb Scoring ✅
|
|
||||||
```typescript
|
|
||||||
import { IntelligentVerbScoringAugmentation } from 'brainy'
|
|
||||||
// Multi-factor relationship strength:
|
|
||||||
// - Semantic similarity
|
|
||||||
// - Temporal decay
|
|
||||||
// - Frequency amplification
|
|
||||||
// - Context awareness
|
|
||||||
```
|
|
||||||
|
|
||||||
### 5. Batch Processing ✅
|
|
||||||
```typescript
|
|
||||||
import { BatchProcessingAugmentation } from 'brainy'
|
|
||||||
// Adaptive batching with backpressure
|
|
||||||
// Dynamically adjusts batch size based on load
|
|
||||||
```
|
|
||||||
|
|
||||||
### 6. Connection Pool ✅
|
|
||||||
```typescript
|
|
||||||
import { ConnectionPoolAugmentation } from 'brainy'
|
|
||||||
// Auto-scaling connection management
|
|
||||||
// Optimized for distributed operations
|
|
||||||
```
|
|
||||||
|
|
||||||
### 7. Request Deduplicator ✅
|
|
||||||
```typescript
|
|
||||||
import { RequestDeduplicatorAugmentation } from 'brainy'
|
|
||||||
// In-flight request deduplication
|
|
||||||
// 3x performance boost for concurrent operations
|
|
||||||
```
|
|
||||||
|
|
||||||
### 8. WebSocket Conduit ✅
|
|
||||||
```typescript
|
|
||||||
import { WebSocketConduitAugmentation } from 'brainy'
|
|
||||||
// Real-time bidirectional streaming
|
|
||||||
// Auto-reconnection and heartbeat
|
|
||||||
```
|
|
||||||
|
|
||||||
### 9. WebRTC Conduit ✅
|
|
||||||
```typescript
|
|
||||||
import { WebRTCConduitAugmentation } from 'brainy'
|
|
||||||
// Peer-to-peer data channels
|
|
||||||
// Direct browser-to-browser communication
|
|
||||||
```
|
|
||||||
|
|
||||||
### 10. Memory Storage Optimization ✅
|
|
||||||
```typescript
|
|
||||||
import { MemoryStorageAugmentation } from 'brainy'
|
|
||||||
// Memory-specific optimizations
|
|
||||||
// Circular buffers, compression
|
|
||||||
```
|
|
||||||
|
|
||||||
### 11. Server Search Conduit ✅
|
|
||||||
```typescript
|
|
||||||
import { ServerSearchConduitAugmentation } from 'brainy'
|
|
||||||
// Distributed query execution
|
|
||||||
// Load balancing across nodes
|
|
||||||
```
|
|
||||||
|
|
||||||
### 12. Neural Import ✅
|
|
||||||
```typescript
|
|
||||||
import { NeuralImportAugmentation } from 'brainy'
|
|
||||||
// AI-powered data understanding
|
|
||||||
// Automatic entity detection and classification
|
|
||||||
// Relationship discovery
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🤖 Neural Import Capabilities (FULLY IMPLEMENTED!)
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const neuralImport = new NeuralImport(brain)
|
|
||||||
|
|
||||||
// ALL of these work TODAY:
|
|
||||||
await neuralImport.neuralImport('data.csv')
|
|
||||||
await neuralImport.detectEntitiesWithNeuralAnalysis(data)
|
|
||||||
await neuralImport.detectNounType(entity)
|
|
||||||
await neuralImport.detectRelationships(entities)
|
|
||||||
await neuralImport.generateInsights(data)
|
|
||||||
```
|
|
||||||
|
|
||||||
### Features:
|
|
||||||
- **Auto-detects file format** (CSV, JSON, XML, etc.)
|
|
||||||
- **Identifies entity types** using AI
|
|
||||||
- **Discovers relationships** between entities
|
|
||||||
- **Generates insights** about the data
|
|
||||||
- **Creates optimal graph structure** automatically
|
|
||||||
|
|
||||||
## 🎯 Zero-Config Model Loading Cascade
|
|
||||||
|
|
||||||
Brainy automatically loads models with ZERO configuration required:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData() // That's it!
|
|
||||||
await brain.init()
|
|
||||||
// Models load automatically from best available source
|
|
||||||
```
|
|
||||||
|
|
||||||
### Loading Priority:
|
|
||||||
1. **Local Cache** (./models) - Instant, no network
|
|
||||||
2. **CDN** (models.soulcraft.com) - Fast, global [Coming Soon]
|
|
||||||
3. **GitHub Releases** - Reliable backup
|
|
||||||
4. **HuggingFace** - Ultimate fallback
|
|
||||||
|
|
||||||
### Key Features:
|
|
||||||
- **Automatic fallback** if sources fail
|
|
||||||
- **Model verification** with checksums
|
|
||||||
- **Offline support** with bundled models
|
|
||||||
- **No environment variables needed**
|
|
||||||
- **Works in all environments** (Node, Browser, Workers)
|
|
||||||
|
|
||||||
## 🏢 Distributed Operation Modes
|
|
||||||
|
|
||||||
### Reader Mode ✅
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ mode: 'reader' })
|
|
||||||
// Optimized for read-heavy workloads
|
|
||||||
// 80% cache ratio, aggressive prefetch
|
|
||||||
// 1 hour TTL, minimal writes
|
|
||||||
```
|
|
||||||
|
|
||||||
### Writer Mode ✅
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ mode: 'writer' })
|
|
||||||
// Optimized for write-heavy workloads
|
|
||||||
// Large write buffers, batch writes
|
|
||||||
// Minimal caching, fast ingestion
|
|
||||||
```
|
|
||||||
|
|
||||||
### Hybrid Mode ✅
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData({ mode: 'hybrid' })
|
|
||||||
// Balanced for mixed workloads
|
|
||||||
// Adaptive caching and batching
|
|
||||||
```
|
|
||||||
|
|
||||||
## 💾 Advanced Caching System
|
|
||||||
|
|
||||||
### 3-Level Cache Architecture ✅
|
|
||||||
```typescript
|
|
||||||
const cacheConfig = {
|
|
||||||
hotCache: {
|
|
||||||
size: 1000, // L1 - RAM
|
|
||||||
ttl: 60000 // 1 minute
|
|
||||||
},
|
|
||||||
warmCache: {
|
|
||||||
size: 10000, // L2 - Fast storage
|
|
||||||
ttl: 300000 // 5 minutes
|
|
||||||
},
|
|
||||||
coldCache: {
|
|
||||||
size: 100000, // L3 - Persistent
|
|
||||||
ttl: null // No expiry
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Cache Features:
|
|
||||||
- **Automatic promotion/demotion** between levels
|
|
||||||
- **LRU eviction** within each level
|
|
||||||
- **Compression** for cold cache
|
|
||||||
- **Statistics tracking** for optimization
|
|
||||||
|
|
||||||
## 📊 Comprehensive Statistics
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
const stats = await brain.getStatistics()
|
|
||||||
// Returns detailed metrics:
|
|
||||||
{
|
|
||||||
nouns: {
|
|
||||||
count, created, updated, deleted,
|
|
||||||
size, avgSize
|
|
||||||
},
|
|
||||||
verbs: {
|
|
||||||
count, created, types,
|
|
||||||
weights: { min, max, avg }
|
|
||||||
},
|
|
||||||
vectors: {
|
|
||||||
dimensions: 384,
|
|
||||||
indexSize, partitions,
|
|
||||||
avgSearchTime
|
|
||||||
},
|
|
||||||
cache: {
|
|
||||||
hits, misses, evictions,
|
|
||||||
hitRate, sizes
|
|
||||||
},
|
|
||||||
performance: {
|
|
||||||
operations, avgTimes,
|
|
||||||
p95Latency, p99Latency
|
|
||||||
},
|
|
||||||
storage: {
|
|
||||||
used, available,
|
|
||||||
compression, files
|
|
||||||
},
|
|
||||||
throttling: {
|
|
||||||
delays, rateLimited,
|
|
||||||
backoffMs, retries
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🚀 GPU Acceleration Support
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// Automatic GPU detection
|
|
||||||
const device = await detectBestDevice()
|
|
||||||
// Returns: 'cpu' | 'webgpu' | 'cuda'
|
|
||||||
|
|
||||||
// WebGPU in browser (when available)
|
|
||||||
if (device === 'webgpu') {
|
|
||||||
// Transformer models use WebGPU automatically
|
|
||||||
}
|
|
||||||
|
|
||||||
// CUDA in Node.js (requires ONNX Runtime GPU)
|
|
||||||
if (device === 'cuda') {
|
|
||||||
// Automatically uses GPU for embeddings
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🔄 Adaptive Systems
|
|
||||||
|
|
||||||
### Adaptive Backpressure ✅
|
|
||||||
```typescript
|
|
||||||
// Automatically adjusts flow based on system load
|
|
||||||
// Prevents OOM and maintains throughput
|
|
||||||
```
|
|
||||||
|
|
||||||
### Adaptive Socket Manager ✅
|
|
||||||
```typescript
|
|
||||||
// Dynamic connection pooling
|
|
||||||
// Scales connections based on traffic patterns
|
|
||||||
```
|
|
||||||
|
|
||||||
### Cache Auto-Configuration ✅
|
|
||||||
```typescript
|
|
||||||
// Sizes cache based on available memory
|
|
||||||
// Adjusts strategies based on usage patterns
|
|
||||||
```
|
|
||||||
|
|
||||||
### S3 Throttling Protection ✅
|
|
||||||
```typescript
|
|
||||||
// Built-in exponential backoff
|
|
||||||
// Rate limit detection and adaptation
|
|
||||||
// Automatic retry with jitter
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🛠️ Storage Adapters
|
|
||||||
|
|
||||||
All included, auto-selected based on environment:
|
|
||||||
|
|
||||||
### FileSystem Storage ✅
|
|
||||||
- Default for Node.js
|
|
||||||
- Efficient file-based storage
|
|
||||||
- Automatic directory management
|
|
||||||
|
|
||||||
### Memory Storage ✅
|
|
||||||
- Ultra-fast in-memory operations
|
|
||||||
- Perfect for testing and temporary data
|
|
||||||
- Circular buffer support
|
|
||||||
|
|
||||||
### OPFS Storage ✅
|
|
||||||
- Browser persistent storage
|
|
||||||
- Survives page refreshes
|
|
||||||
- Quota management
|
|
||||||
|
|
||||||
### S3 Storage ✅
|
|
||||||
- AWS S3 compatible
|
|
||||||
- Automatic multipart uploads
|
|
||||||
- Throttling protection
|
|
||||||
- Batch operations
|
|
||||||
|
|
||||||
## 🎨 Natural Language Processing
|
|
||||||
|
|
||||||
### Built-in Patterns (220+)
|
|
||||||
- Question types (what, why, how, when, where)
|
|
||||||
- Temporal queries (yesterday, last week, 2024)
|
|
||||||
- Comparative queries (better than, similar to)
|
|
||||||
- Aggregations (count, sum, average)
|
|
||||||
- Filters (only, except, without)
|
|
||||||
- Relationships (related to, connected with)
|
|
||||||
|
|
||||||
### Coverage: 94-98% of typical queries!
|
|
||||||
|
|
||||||
## 🔐 Security Features
|
|
||||||
|
|
||||||
### Built-in Security ✅
|
|
||||||
- Automatic input sanitization
|
|
||||||
- SQL injection prevention
|
|
||||||
- XSS protection for web contexts
|
|
||||||
- Rate limiting support
|
|
||||||
|
|
||||||
### Encryption Ready ✅
|
|
||||||
```typescript
|
|
||||||
import { crypto } from 'brainy/utils'
|
|
||||||
// AES-256-GCM encryption utilities
|
|
||||||
// Key derivation functions
|
|
||||||
// Secure random generation
|
|
||||||
```
|
|
||||||
|
|
||||||
## 🎯 Key Design Principles
|
|
||||||
|
|
||||||
### 1. Zero Configuration
|
|
||||||
```typescript
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
// Everything else is automatic!
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. Fixed Dimensions (384)
|
|
||||||
- **ALWAYS** uses all-MiniLM-L6-v2 model
|
|
||||||
- **ALWAYS** 384 dimensions
|
|
||||||
- **NOT** configurable (by design)
|
|
||||||
- Ensures everything works together
|
|
||||||
|
|
||||||
### 3. Progressive Enhancement
|
|
||||||
- Starts simple, scales automatically
|
|
||||||
- Adapts to workload patterns
|
|
||||||
- Optimizes based on usage
|
|
||||||
|
|
||||||
### 4. Universal Compatibility
|
|
||||||
- Works in Node.js 18+
|
|
||||||
- Works in modern browsers
|
|
||||||
- Works in Web Workers
|
|
||||||
- Works in Edge environments
|
|
||||||
|
|
||||||
## 📦 What Ships in Core (MIT Licensed)
|
|
||||||
|
|
||||||
**EVERYTHING** is included in the core package:
|
|
||||||
- ✅ All engines (vector, graph, field, neural)
|
|
||||||
- ✅ All augmentations (12+)
|
|
||||||
- ✅ All storage adapters
|
|
||||||
- ✅ All distributed modes
|
|
||||||
- ✅ Complete statistics
|
|
||||||
- ✅ GPU support
|
|
||||||
- ✅ No feature limitations
|
|
||||||
- ✅ No premium tiers
|
|
||||||
- ✅ 100% MIT licensed
|
|
||||||
|
|
||||||
## 🚀 Quick Start
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
import { BrainyData } from 'brainy'
|
|
||||||
|
|
||||||
// Zero config required!
|
|
||||||
const brain = new BrainyData()
|
|
||||||
await brain.init()
|
|
||||||
|
|
||||||
// Add data (auto-detects type)
|
|
||||||
await brain.addNoun('Content here')
|
|
||||||
|
|
||||||
// Search with natural language
|
|
||||||
const results = await brain.find('related content from last week')
|
|
||||||
|
|
||||||
// Everything else is automatic!
|
|
||||||
```
|
|
||||||
|
|
||||||
## 📈 Performance Characteristics
|
|
||||||
|
|
||||||
- **Vector Search**: O(log n) with HNSW indexing
|
|
||||||
- **Graph Traversal**: O(k) for k-hop queries
|
|
||||||
- **Field Filtering**: O(1) with metadata index
|
|
||||||
- **Memory Usage**: ~100MB base + data
|
|
||||||
- **Embedding Speed**: ~100ms for batch of 10
|
|
||||||
- **Query Speed**: <10ms for most queries
|
|
||||||
|
|
||||||
## 🎉 Summary
|
|
||||||
|
|
||||||
Brainy 2.0 is a **complete**, **production-ready** AI database that requires **ZERO configuration**. Every feature listed here is **implemented and working** today. No configuration, no setup, no complexity - just powerful AI capabilities that work out of the box!
|
|
||||||
230
docs/guides/MIGRATING_TO_V5.11.md
Normal file
230
docs/guides/MIGRATING_TO_V5.11.md
Normal file
|
|
@ -0,0 +1,230 @@
|
||||||
|
# Migrating to v5.11.1
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
v5.11.1 introduces a **breaking change** with **massive performance benefits**:
|
||||||
|
|
||||||
|
- `brain.get()` now loads **metadata-only by default** (76-81% faster!)
|
||||||
|
- Vector embeddings require **explicit opt-in**: `{ includeVectors: true }`
|
||||||
|
|
||||||
|
**Impact**: Only ~6% of codebases need changes (code that computes similarity on retrieved entities).
|
||||||
|
|
||||||
|
## What Changed
|
||||||
|
|
||||||
|
### Before (v5.11.0 and earlier)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
// entity.vector was ALWAYS loaded (384 dimensions, 6KB)
|
||||||
|
console.log(entity.vector.length) // 384
|
||||||
|
```
|
||||||
|
|
||||||
|
### After (v5.11.1)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// DEFAULT: Metadata-only (76-81% faster)
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
console.log(entity.vector) // [] (empty array - not loaded)
|
||||||
|
|
||||||
|
// EXPLICIT: Full entity with vectors
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
console.log(entity.vector.length) // 384
|
||||||
|
```
|
||||||
|
|
||||||
|
## Who Needs to Update?
|
||||||
|
|
||||||
|
### ✅ NO CHANGES NEEDED (94% of code)
|
||||||
|
|
||||||
|
If you use `brain.get()` for:
|
||||||
|
- **VFS operations** (readFile, stat, readdir)
|
||||||
|
- **Existence checks**: `if (await brain.get(id))`
|
||||||
|
- **Metadata access**: `entity.data`, `entity.type`, `entity.metadata`
|
||||||
|
- **Relationship traversal**
|
||||||
|
- **Admin tools**, import utilities, data APIs
|
||||||
|
|
||||||
|
→ **Zero changes needed, automatic 76-81% speedup!**
|
||||||
|
|
||||||
|
### ⚠️ REQUIRES UPDATE (~6% of code)
|
||||||
|
|
||||||
|
If you use `brain.get()` AND then compute similarity on the returned entity:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ❌ BEFORE (v5.11.0) - will break in v5.11.1
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
const similar = await brain.similar({ to: entity.vector }) // entity.vector is [] !
|
||||||
|
|
||||||
|
// ✅ AFTER (v5.11.1) - add includeVectors
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
const similar = await brain.similar({ to: entity.vector }) // Works!
|
||||||
|
```
|
||||||
|
|
||||||
|
**Note**: `brain.similar({ to: entityId })` (using ID) still works - no changes needed!
|
||||||
|
|
||||||
|
## Migration Steps
|
||||||
|
|
||||||
|
### Step 1: Find Affected Code
|
||||||
|
|
||||||
|
Search your codebase for patterns that use vectors from `brain.get()`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Find brain.get() calls that access .vector
|
||||||
|
grep -r "await brain.get(" --include="*.ts" --include="*.js" | \
|
||||||
|
grep -E "(\.vector|entity\.vector)"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 2: Update Pattern-by-Pattern
|
||||||
|
|
||||||
|
#### Pattern 1: Similarity Using Retrieved Entity Vector
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ❌ BEFORE
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
const similar = await brain.similar({ to: entity.vector })
|
||||||
|
|
||||||
|
// ✅ AFTER - Option A: Add includeVectors
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
const similar = await brain.similar({ to: entity.vector })
|
||||||
|
|
||||||
|
// ✅ AFTER - Option B: Use ID directly (recommended)
|
||||||
|
const similar = await brain.similar({ to: id })
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Pattern 2: Manual Vector Operations
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ❌ BEFORE
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
const magnitude = Math.sqrt(entity.vector.reduce((sum, v) => sum + v*v, 0))
|
||||||
|
|
||||||
|
// ✅ AFTER
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
const magnitude = Math.sqrt(entity.vector.reduce((sum, v) => sum + v*v, 0))
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Pattern 3: Vector Assertions in Tests
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// ❌ BEFORE
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
expect(entity.vector).toBeDefined()
|
||||||
|
expect(entity.vector.length).toBe(384)
|
||||||
|
|
||||||
|
// ✅ AFTER
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
expect(entity.vector).toBeDefined()
|
||||||
|
expect(entity.vector.length).toBe(384)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 3: Verify Migration
|
||||||
|
|
||||||
|
Run your test suite to catch any remaining issues:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm test
|
||||||
|
```
|
||||||
|
|
||||||
|
Look for errors like:
|
||||||
|
- `entity.vector is empty` or `entity.vector.length is 0`
|
||||||
|
- `Cannot compute similarity on empty vector`
|
||||||
|
|
||||||
|
Add `{ includeVectors: true }` wherever these errors occur.
|
||||||
|
|
||||||
|
## Performance Impact
|
||||||
|
|
||||||
|
### Before Migration
|
||||||
|
```
|
||||||
|
brain.get(): 43ms, 6KB per call
|
||||||
|
VFS readFile(): 53ms per file
|
||||||
|
VFS readdir(100 files): 5.3s
|
||||||
|
```
|
||||||
|
|
||||||
|
### After Migration
|
||||||
|
```
|
||||||
|
brain.get(): 10ms, 300 bytes per call (76-81% faster) ✨
|
||||||
|
brain.get({ includeVectors: true }): 43ms, 6KB (unchanged)
|
||||||
|
VFS readFile(): ~13ms per file (75% faster) ✨
|
||||||
|
VFS readdir(100 files): ~1.3s (75% faster) ✨
|
||||||
|
```
|
||||||
|
|
||||||
|
**Result**:
|
||||||
|
- VFS operations: **75% faster**
|
||||||
|
- Metadata access: **76-81% faster**
|
||||||
|
- Vector similarity: **Unchanged** (still fast when needed)
|
||||||
|
|
||||||
|
## TypeScript Support
|
||||||
|
|
||||||
|
The new `GetOptions` interface is fully typed:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
interface GetOptions {
|
||||||
|
/**
|
||||||
|
* Include 384-dimensional vector embeddings in the response
|
||||||
|
*
|
||||||
|
* Default: false (metadata-only for 76-81% speedup)
|
||||||
|
*/
|
||||||
|
includeVectors?: boolean
|
||||||
|
}
|
||||||
|
|
||||||
|
// TypeScript will autocomplete and validate
|
||||||
|
const entity = await brain.get(id, { includeVectors: true })
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rollback Plan
|
||||||
|
|
||||||
|
If you encounter issues, you can temporarily force full entity loading everywhere:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Temporary wrapper (NOT RECOMMENDED - defeats optimization)
|
||||||
|
async function getLegacy(id: string) {
|
||||||
|
return brain.get(id, { includeVectors: true })
|
||||||
|
}
|
||||||
|
|
||||||
|
// Use throughout codebase while migrating
|
||||||
|
const entity = await getLegacy(id)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Important**: This defeats the 76-81% performance improvement. Only use temporarily while fixing affected code.
|
||||||
|
|
||||||
|
## FAQ
|
||||||
|
|
||||||
|
### Q: Why did you make this a breaking change?
|
||||||
|
|
||||||
|
**A**: The performance gains are massive (76-81% speedup, 95% less bandwidth) and affect 94% of code positively. Only ~6% of code needs updates. The net benefit is enormous.
|
||||||
|
|
||||||
|
### Q: Do I need to update my VFS code?
|
||||||
|
|
||||||
|
**A**: No! VFS automatically benefits from the optimization with zero code changes. Your VFS operations are now 75% faster automatically.
|
||||||
|
|
||||||
|
### Q: Will brain.similar() still work?
|
||||||
|
|
||||||
|
**A**: Yes! `brain.similar({ to: entityId })` works exactly as before. Only `brain.similar({ to: entity.vector })` requires the entity to be loaded with `{ includeVectors: true }`.
|
||||||
|
|
||||||
|
### Q: What about backward compatibility?
|
||||||
|
|
||||||
|
**A**: Entities returned without vectors have `vector: []` (empty array), which is type-safe. Code that doesn't use vectors continues to work. Only code that explicitly uses `entity.vector` needs updating.
|
||||||
|
|
||||||
|
### Q: Can I check if vectors are loaded?
|
||||||
|
|
||||||
|
**A**: Yes! Check `entity.vector.length > 0` to detect if vectors were loaded.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const entity = await brain.get(id)
|
||||||
|
if (entity.vector.length > 0) {
|
||||||
|
// Vectors are loaded
|
||||||
|
} else {
|
||||||
|
// Metadata-only
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Support
|
||||||
|
|
||||||
|
If you encounter migration issues:
|
||||||
|
|
||||||
|
1. Check the [VFS Performance Guide](../vfs/VFS_PERFORMANCE.md)
|
||||||
|
2. Review [API Reference](../api/README.md)
|
||||||
|
3. See [Performance Documentation](../PERFORMANCE.md)
|
||||||
|
4. File an issue: https://github.com/soulcraft/brainy/issues
|
||||||
|
|
||||||
|
## Changelog
|
||||||
|
|
||||||
|
See [CHANGELOG.md](../../CHANGELOG.md) for complete v5.11.1 release notes.
|
||||||
593
docs/guides/aggregation.md
Normal file
593
docs/guides/aggregation.md
Normal file
|
|
@ -0,0 +1,593 @@
|
||||||
|
# Aggregation Guide
|
||||||
|
|
||||||
|
> Real-time analytics on your entity data with incremental running totals
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Brainy's aggregation engine computes running totals at write time, so reading aggregate results is always O(1) regardless of dataset size. Define an aggregate once, and every `add()`, `update()`, and `delete()` automatically updates the running metrics.
|
||||||
|
|
||||||
|
No batch jobs. No scheduled recalculations. Aggregates stay current with every write.
|
||||||
|
|
||||||
|
**Defining over existing data:** if you define an aggregate on a store that already holds
|
||||||
|
matching entities, Brainy backfills it from those entities on the first query (a one-time scan,
|
||||||
|
then purely incremental). So `defineAggregate()` behaves the same whether you define it before
|
||||||
|
or after the data exists.
|
||||||
|
|
||||||
|
**Reopening a persisted brain:** aggregate state persists across restarts. Re-defining the
|
||||||
|
same aggregate at boot (the normal declarative pattern) adopts the persisted state directly —
|
||||||
|
no rescan. A backfill scan runs only when the definition actually changed, when no persisted
|
||||||
|
state exists, or when the state failed to load; and however many aggregates need backfilling,
|
||||||
|
they share a single scan.
|
||||||
|
|
||||||
|
## Quick Start
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { Brainy, NounType } from '@soulcraft/brainy'
|
||||||
|
|
||||||
|
const brain = new Brainy()
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// 1. Define an aggregate
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'sales_by_category',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: ['category'],
|
||||||
|
metrics: {
|
||||||
|
revenue: { op: 'sum', field: 'amount' },
|
||||||
|
count: { op: 'count' },
|
||||||
|
average: { op: 'avg', field: 'amount' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// 2. Add entities — aggregates update automatically
|
||||||
|
await brain.add({
|
||||||
|
data: 'Coffee purchase',
|
||||||
|
type: NounType.Event,
|
||||||
|
metadata: { category: 'food', amount: 5.50 }
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.add({
|
||||||
|
data: 'Laptop purchase',
|
||||||
|
type: NounType.Event,
|
||||||
|
metadata: { category: 'electronics', amount: 1200 }
|
||||||
|
})
|
||||||
|
|
||||||
|
await brain.add({
|
||||||
|
data: 'Lunch purchase',
|
||||||
|
type: NounType.Event,
|
||||||
|
metadata: { category: 'food', amount: 12.00 }
|
||||||
|
})
|
||||||
|
|
||||||
|
// 3. Query results
|
||||||
|
const results = await brain.find({ aggregate: 'sales_by_category' })
|
||||||
|
|
||||||
|
// Results:
|
||||||
|
// [
|
||||||
|
// { groupKey: { category: 'food' }, metrics: { revenue: 17.50, count: 2, average: 8.75 } },
|
||||||
|
// { groupKey: { category: 'electronics' }, metrics: { revenue: 1200, count: 1, average: 1200 } }
|
||||||
|
// ]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Aggregation Operations
|
||||||
|
|
||||||
|
Brainy supports 7 aggregation operations:
|
||||||
|
|
||||||
|
### `sum` — Running Total
|
||||||
|
|
||||||
|
Adds up all values of a numeric field.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
total_revenue: { op: 'sum', field: 'amount' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `count` — Entity Count
|
||||||
|
|
||||||
|
Counts the number of matching entities. No `field` required.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
order_count: { op: 'count' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `avg` — Running Average
|
||||||
|
|
||||||
|
Computes `sum / count` incrementally.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
average_price: { op: 'avg', field: 'price' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `min` — Minimum Value
|
||||||
|
|
||||||
|
Tracks the minimum value across all entities in each group.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
lowest_price: { op: 'min', field: 'price' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `max` — Maximum Value
|
||||||
|
|
||||||
|
Tracks the maximum value across all entities in each group.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
highest_price: { op: 'max', field: 'price' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `stddev` — Sample Standard Deviation
|
||||||
|
|
||||||
|
Computes the sample standard deviation using Welford's numerically stable online algorithm. Updates incrementally without storing individual values.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
price_spread: { op: 'stddev', field: 'price' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `variance` — Sample Variance
|
||||||
|
|
||||||
|
Computes the sample variance (square of standard deviation) using Welford's online algorithm.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
metrics: {
|
||||||
|
price_variance: { op: 'variance', field: 'price' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## GROUP BY Dimensions
|
||||||
|
|
||||||
|
Every aggregate requires at least one `groupBy` dimension. Results are grouped by the unique combinations of dimension values.
|
||||||
|
|
||||||
|
### Plain Fields
|
||||||
|
|
||||||
|
Group by a metadata field value:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
groupBy: ['category']
|
||||||
|
// Produces groups: { category: 'food' }, { category: 'electronics' }, ...
|
||||||
|
```
|
||||||
|
|
||||||
|
### Multiple Fields
|
||||||
|
|
||||||
|
Group by multiple fields for composite keys:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
groupBy: ['category', 'region']
|
||||||
|
// Produces groups: { category: 'food', region: 'US' }, { category: 'food', region: 'EU' }, ...
|
||||||
|
```
|
||||||
|
|
||||||
|
### Time Windows
|
||||||
|
|
||||||
|
Group by a timestamp field bucketed into time periods:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
groupBy: [{ field: 'date', window: 'month' }]
|
||||||
|
// Produces groups: { date: '2024-01' }, { date: '2024-02' }, ...
|
||||||
|
```
|
||||||
|
|
||||||
|
Available time window granularities:
|
||||||
|
|
||||||
|
| Window | Format | Example |
|
||||||
|
|--------|--------|---------|
|
||||||
|
| `hour` | `YYYY-MM-DDThh` | `2024-01-15T14` |
|
||||||
|
| `day` | `YYYY-MM-DD` | `2024-01-15` |
|
||||||
|
| `week` | `YYYY-Wnn` | `2024-W03` |
|
||||||
|
| `month` | `YYYY-MM` | `2024-01` |
|
||||||
|
| `quarter` | `YYYY-Qn` | `2024-Q1` |
|
||||||
|
| `year` | `YYYY` | `2024` |
|
||||||
|
| `{ seconds: N }` | ISO 8601 | Custom interval |
|
||||||
|
|
||||||
|
### Combined Dimensions
|
||||||
|
|
||||||
|
Mix plain fields and time windows:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'monthly_sales',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: ['region', { field: 'date', window: 'month' }],
|
||||||
|
metrics: {
|
||||||
|
revenue: { op: 'sum', field: 'amount' },
|
||||||
|
count: { op: 'count' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Produces groups like:
|
||||||
|
// { region: 'US', date: '2024-01' }
|
||||||
|
// { region: 'US', date: '2024-02' }
|
||||||
|
// { region: 'EU', date: '2024-01' }
|
||||||
|
```
|
||||||
|
|
||||||
|
### Array Fields (Unnest)
|
||||||
|
|
||||||
|
Group by **each element** of an array-valued field — for tag frequencies, label counts, and
|
||||||
|
faceted breakdowns. Mark the dimension `{ field, unnest: true }`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'tag_frequency',
|
||||||
|
source: { type: NounType.Document },
|
||||||
|
groupBy: [{ field: 'tags', unnest: true }],
|
||||||
|
metrics: { count: { op: 'count' } }
|
||||||
|
})
|
||||||
|
|
||||||
|
// A document tagged ['ml', 'ai'] contributes once to the 'ml' group and once to 'ai'.
|
||||||
|
// Duplicate tags on one entity count once; an entity with no tags joins no group.
|
||||||
|
const top = await brain.queryAggregate('tag_frequency', { orderBy: 'count', order: 'desc' })
|
||||||
|
// [ { groupKey: { tags: 'ai' }, metrics: { count: 3 }, count: 3 }, ... ]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Querying Aggregates
|
||||||
|
|
||||||
|
Aggregate results are queried through the standard `find()` method.
|
||||||
|
|
||||||
|
### Basic Query
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const results = await brain.find({ aggregate: 'sales_by_category' })
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by Group Key
|
||||||
|
|
||||||
|
Use `where` to filter on group key values:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const foodOnly = await brain.find({
|
||||||
|
aggregate: 'sales_by_category',
|
||||||
|
where: { category: 'food' }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by Metric Value (HAVING)
|
||||||
|
|
||||||
|
Use `having` to filter groups by their **computed metric values** — the analytics equivalent of
|
||||||
|
SQL `HAVING`. (`where` filters group *keys*; `having` filters *metrics*.)
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const bigCategories = await brain.find({
|
||||||
|
aggregate: 'sales_by_category',
|
||||||
|
having: { revenue: { greaterThan: 1000 } }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
`having` accepts the same operators as `where`, applied to each group's metric results plus
|
||||||
|
`count`. It is evaluated per group — **O(groups), independent of entity count** — before sorting
|
||||||
|
and pagination, so it stays cheap even over billions of entities.
|
||||||
|
|
||||||
|
### Sort and Paginate
|
||||||
|
|
||||||
|
Sort by any metric or group key field:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const topCategories = await brain.find({
|
||||||
|
aggregate: {
|
||||||
|
name: 'sales_by_category',
|
||||||
|
orderBy: 'revenue',
|
||||||
|
order: 'desc',
|
||||||
|
limit: 10
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Combined Parameters
|
||||||
|
|
||||||
|
`where`, `orderBy`, `limit`, and `offset` from the outer `find()` call merge automatically with the aggregate query:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const recentTopSpenders = await brain.find({
|
||||||
|
aggregate: 'monthly_sales',
|
||||||
|
where: { region: 'US' },
|
||||||
|
orderBy: 'revenue',
|
||||||
|
order: 'desc',
|
||||||
|
limit: 12,
|
||||||
|
offset: 0
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Result Format
|
||||||
|
|
||||||
|
`find({ aggregate })` returns `Result<T>` rows (for uniformity with the rest of `find()`),
|
||||||
|
with the aggregate fields surfaced **both** at the top level and, for backward compatibility,
|
||||||
|
flattened into `metadata`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
{
|
||||||
|
id: string,
|
||||||
|
score: 1.0,
|
||||||
|
type: NounType.Measurement,
|
||||||
|
groupKey: { category: 'food' }, // top-level — the group key values
|
||||||
|
metrics: { revenue: 17.50, count: 2, average: 8.75 }, // top-level — computed metrics
|
||||||
|
count: 2, // top-level — entities in the group
|
||||||
|
metadata: { // legacy mirror of the same data
|
||||||
|
__aggregate: 'sales_by_category',
|
||||||
|
category: 'food',
|
||||||
|
revenue: 17.50, count: 2, average: 8.75
|
||||||
|
},
|
||||||
|
entity: Entity
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### `queryAggregate()` — the report-friendly view
|
||||||
|
|
||||||
|
For dashboards and reports, prefer `brain.queryAggregate(name, params)`. It returns the clean
|
||||||
|
`AggregateResult[]` shape directly — no search-result wrapper:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const rows = await brain.queryAggregate('sales_by_category', {
|
||||||
|
orderBy: 'revenue',
|
||||||
|
order: 'desc',
|
||||||
|
limit: 10
|
||||||
|
})
|
||||||
|
// [
|
||||||
|
// { groupKey: { category: 'electronics' }, metrics: { revenue: 1200, count: 1, average: 1200 }, count: 1 },
|
||||||
|
// { groupKey: { category: 'food' }, metrics: { revenue: 17.50, count: 2, average: 8.75 }, count: 2 }
|
||||||
|
// ]
|
||||||
|
```
|
||||||
|
|
||||||
|
It accepts the same `where` / `having` / `orderBy` / `order` / `limit` / `offset` params as the
|
||||||
|
`find({ aggregate })` form.
|
||||||
|
|
||||||
|
## Source Filtering
|
||||||
|
|
||||||
|
Control which entities feed into an aggregate with the `source` property.
|
||||||
|
|
||||||
|
### Filter by Entity Type
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'event_stats',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: ['category'],
|
||||||
|
metrics: { count: { op: 'count' } }
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by Multiple Types
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
source: { type: [NounType.Event, NounType.Document] }
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by Metadata
|
||||||
|
|
||||||
|
Use the same `where` syntax as `find()`:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
source: {
|
||||||
|
type: NounType.Event,
|
||||||
|
where: { domain: 'financial', subtype: 'transaction' }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Filter by Service
|
||||||
|
|
||||||
|
For multi-tenant deployments:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
source: { service: 'tenant-123' }
|
||||||
|
```
|
||||||
|
|
||||||
|
Entities that don't match the source filter are silently skipped during incremental updates.
|
||||||
|
|
||||||
|
## Incremental Updates
|
||||||
|
|
||||||
|
The aggregation engine hooks into every write operation:
|
||||||
|
|
||||||
|
### On `add()`
|
||||||
|
|
||||||
|
When a new entity matches an aggregate's source filter:
|
||||||
|
1. The group key is computed from the entity's metadata
|
||||||
|
2. Each metric in the matching group is incremented
|
||||||
|
3. New groups are created automatically
|
||||||
|
|
||||||
|
### On `update()`
|
||||||
|
|
||||||
|
When an existing entity is updated:
|
||||||
|
1. The old entity's contribution is reversed from its group
|
||||||
|
2. The new entity's contribution is applied to its (potentially different) group
|
||||||
|
3. Handles group key changes — an entity moving from category "food" to "drink" updates both groups
|
||||||
|
|
||||||
|
### On `delete()`
|
||||||
|
|
||||||
|
When an entity is deleted:
|
||||||
|
1. The entity's contribution is reversed from its group
|
||||||
|
2. If a group becomes empty (all metric counts reach zero), it's removed
|
||||||
|
|
||||||
|
### Aggregate Entity Exclusion
|
||||||
|
|
||||||
|
Materialized `NounType.Measurement` entities are automatically excluded from all source matching, preventing infinite feedback loops. Entities with `service: 'brainy:aggregation'` or `metadata.__aggregate` are always skipped.
|
||||||
|
|
||||||
|
## Materialization
|
||||||
|
|
||||||
|
Materialization writes aggregate results as `NounType.Measurement` entities, making them automatically available through OData, Google Sheets, SSE, and webhook integrations.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'daily_metrics',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: [{ field: 'date', window: 'day' }],
|
||||||
|
metrics: {
|
||||||
|
total: { op: 'sum', field: 'amount' },
|
||||||
|
count: { op: 'count' }
|
||||||
|
},
|
||||||
|
materialize: true
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Debounce Configuration
|
||||||
|
|
||||||
|
During high-throughput ingestion, materialization is debounced to avoid excessive writes:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
materialize: {
|
||||||
|
debounceMs: 2000, // Wait 2 seconds after last update before writing
|
||||||
|
trackSources: true // Track which entities contributed
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The default debounce interval is 1000ms.
|
||||||
|
|
||||||
|
## Multiple Aggregates
|
||||||
|
|
||||||
|
Define multiple aggregates that process the same entities:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Revenue by category
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'category_revenue',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: ['category'],
|
||||||
|
metrics: { total: { op: 'sum', field: 'amount' } }
|
||||||
|
})
|
||||||
|
|
||||||
|
// Monthly trends
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'monthly_trends',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: [{ field: 'date', window: 'month' }],
|
||||||
|
metrics: {
|
||||||
|
revenue: { op: 'sum', field: 'amount' },
|
||||||
|
count: { op: 'count' },
|
||||||
|
avg_order: { op: 'avg', field: 'amount' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// Regional breakdown with statistical analysis
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'regional_analysis',
|
||||||
|
source: { type: NounType.Event },
|
||||||
|
groupBy: ['region'],
|
||||||
|
metrics: {
|
||||||
|
revenue: { op: 'sum', field: 'amount' },
|
||||||
|
spread: { op: 'stddev', field: 'amount' },
|
||||||
|
variance: { op: 'variance', field: 'amount' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
Each `add()` call updates all matching aggregates automatically.
|
||||||
|
|
||||||
|
## Removing Aggregates
|
||||||
|
|
||||||
|
Remove an aggregate and clean up its state:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.removeAggregate('category_revenue')
|
||||||
|
```
|
||||||
|
|
||||||
|
## Persistence
|
||||||
|
|
||||||
|
Aggregate definitions and running state are automatically persisted:
|
||||||
|
|
||||||
|
- **On `flush()`/`close()`**: All dirty aggregate state is written to storage
|
||||||
|
- **On `init()`**: Definitions and state are restored from storage
|
||||||
|
- **Change detection**: Definition changes are detected via FNV-1a hashing — only changed aggregates reset their state on restart
|
||||||
|
|
||||||
|
## Native Acceleration
|
||||||
|
|
||||||
|
When [Cor](https://www.npmjs.com/package/@soulcraft/cor) is installed as a plugin, the aggregation engine automatically uses Rust-accelerated computation:
|
||||||
|
|
||||||
|
- Incremental updates run in Rust with BTreeMap-backed precise MIN/MAX
|
||||||
|
- Welford's online stddev/variance computed natively
|
||||||
|
- Rebuild uses Rayon parallel iterators across CPU cores (above 1,000 entities)
|
||||||
|
- Time window bucketing uses integer arithmetic without `Date` object allocation
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
plugins: ['@soulcraft/cor']
|
||||||
|
})
|
||||||
|
await brain.init()
|
||||||
|
|
||||||
|
// Aggregation automatically uses native engine
|
||||||
|
brain.defineAggregate({ ... })
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify native acceleration is active:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const diag = brain.diagnostics()
|
||||||
|
console.log(diag.providers.aggregation)
|
||||||
|
// { source: 'plugin' }
|
||||||
|
```
|
||||||
|
|
||||||
|
## Common Patterns
|
||||||
|
|
||||||
|
### Financial Analytics
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'monthly_spending',
|
||||||
|
source: {
|
||||||
|
type: NounType.Event,
|
||||||
|
where: { domain: 'financial', subtype: 'transaction' }
|
||||||
|
},
|
||||||
|
groupBy: [
|
||||||
|
'category',
|
||||||
|
{ field: 'date', window: 'month' }
|
||||||
|
],
|
||||||
|
metrics: {
|
||||||
|
total: { op: 'sum', field: 'amount' },
|
||||||
|
count: { op: 'count' },
|
||||||
|
average: { op: 'avg', field: 'amount' },
|
||||||
|
highest: { op: 'max', field: 'amount' },
|
||||||
|
lowest: { op: 'min', field: 'amount' }
|
||||||
|
},
|
||||||
|
materialize: true
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Time-Series Monitoring
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'hourly_metrics',
|
||||||
|
source: { type: NounType.Event, where: { domain: 'monitoring' } },
|
||||||
|
groupBy: [
|
||||||
|
'service',
|
||||||
|
{ field: 'timestamp', window: 'hour' }
|
||||||
|
],
|
||||||
|
metrics: {
|
||||||
|
request_count: { op: 'count' },
|
||||||
|
avg_latency: { op: 'avg', field: 'latency_ms' },
|
||||||
|
max_latency: { op: 'max', field: 'latency_ms' },
|
||||||
|
error_count: { op: 'sum', field: 'is_error' },
|
||||||
|
latency_spread: { op: 'stddev', field: 'latency_ms' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
### Content Analytics
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.defineAggregate({
|
||||||
|
name: 'content_stats',
|
||||||
|
source: { type: NounType.Document },
|
||||||
|
groupBy: ['author', { field: 'publishedAt', window: 'month' }],
|
||||||
|
metrics: {
|
||||||
|
articles: { op: 'count' },
|
||||||
|
total_words: { op: 'sum', field: 'wordCount' },
|
||||||
|
avg_words: { op: 'avg', field: 'wordCount' }
|
||||||
|
}
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
## Performance
|
||||||
|
|
||||||
|
Aggregation complexity per write is O(A x G x M) where A = matching aggregates, G = groupBy dimensions, M = metrics. For typical configurations (2-5 aggregates, 1-3 dimensions, 3-5 metrics), this is effectively O(1).
|
||||||
|
|
||||||
|
With Cor native acceleration:
|
||||||
|
|
||||||
|
| Operation | Throughput | Latency |
|
||||||
|
|-----------|-----------|---------|
|
||||||
|
| Incremental update (1K entities) | 809 ops/s | 1.2 ms |
|
||||||
|
| Rebuild (10K entities) | 475 ops/s | 2.1 ms |
|
||||||
|
| Rebuild (100K entities, Rayon) | 66 ops/s | 15.2 ms |
|
||||||
|
| Query (1K groups, sort + paginate) | 986 ops/s | 1.0 ms |
|
||||||
|
|
@ -21,7 +21,7 @@ Enterprise features on our roadmap.
|
||||||
**Everyone gets bank-level security features:**
|
**Everyone gets bank-level security features:**
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
security: {
|
security: {
|
||||||
encryption: 'aes-256-gcm', // Military-grade encryption
|
encryption: 'aes-256-gcm', // Military-grade encryption
|
||||||
keyRotation: true, // Automatic key rotation
|
keyRotation: true, // Automatic key rotation
|
||||||
|
|
@ -46,11 +46,9 @@ const brain = new BrainyData({
|
||||||
**Everyone gets mission-critical reliability:**
|
**Everyone gets mission-critical reliability:**
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
import { WALAugmentation } from 'brainy'
|
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new WALAugmentation({
|
|
||||||
enabled: true, // Write-ahead logging
|
enabled: true, // Write-ahead logging
|
||||||
redundancy: 3, // Triple redundancy
|
redundancy: 3, // Triple redundancy
|
||||||
checkpointInterval: 1000, // Frequent checkpoints
|
checkpointInterval: 1000, // Frequent checkpoints
|
||||||
|
|
@ -104,7 +102,7 @@ const performance = {
|
||||||
```typescript
|
```typescript
|
||||||
import { MonitoringAugmentation } from 'brainy'
|
import { MonitoringAugmentation } from 'brainy'
|
||||||
|
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
augmentations: [
|
augmentations: [
|
||||||
new MonitoringAugmentation({
|
new MonitoringAugmentation({
|
||||||
metrics: 'all', // Complete metrics
|
metrics: 'all', // Complete metrics
|
||||||
|
|
@ -167,39 +165,35 @@ await brain.syncWith({
|
||||||
- **Webhook support**: React to changes
|
- **Webhook support**: React to changes
|
||||||
- **API generation**: Auto-generate REST/GraphQL APIs
|
- **API generation**: Auto-generate REST/GraphQL APIs
|
||||||
|
|
||||||
### 🌍 Enterprise Scale 🚧 Coming Soon
|
### 🌍 Scale
|
||||||
|
|
||||||
**Everyone gets planetary scale:**
|
**Everyone gets the same scale model:**
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Same architecture Netflix uses, free for you
|
// Pure JS by default; install the optional native provider for billions of vectors
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy()
|
||||||
clustering: {
|
|
||||||
enabled: true, // Distributed mode
|
|
||||||
sharding: 'automatic', // Auto-sharding
|
|
||||||
replication: 3, // Triple replication
|
|
||||||
consensus: 'raft', // Strong consistency
|
|
||||||
geoDistribution: true // Multi-region support
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
// Handles everything from 1 to 1 billion entities
|
// 1 → ~1M vectors: pure-JS HNSW, zero extra setup
|
||||||
|
// 1M → 10B+ vectors: install @soulcraft/cor for the native DiskANN provider
|
||||||
```
|
```
|
||||||
|
|
||||||
**Scaling features:**
|
**Scaling model:**
|
||||||
- **Horizontal scaling**: Add nodes as needed
|
- **Single process, no cluster**: Brainy runs in one process — no coordinator,
|
||||||
- **Auto-sharding**: Distributes data automatically
|
no peer discovery, no consensus to operate
|
||||||
- **Multi-region**: Global distribution
|
- **Optional native provider**: install `@soulcraft/cor` to back the index with
|
||||||
- **Load balancing**: Automatic request distribution
|
on-disk DiskANN that scales to 10B+ vectors on one machine
|
||||||
- **Zero-downtime upgrades**: Rolling updates
|
- **Per-tenant pools**: isolate tenants by giving each its own Brainy instance and
|
||||||
- **Infinite scale**: No upper limits
|
storage directory
|
||||||
|
- **Horizontal read scaling**: run many reader processes against one shared on-disk
|
||||||
|
store (single writer, many readers); replicate the artifact with your operator
|
||||||
|
tooling
|
||||||
|
|
||||||
### 🛡️ Enterprise Compliance 🚧 Coming Soon
|
### 🛡️ Enterprise Compliance 🚧 Coming Soon
|
||||||
|
|
||||||
**Everyone gets compliance tools:**
|
**Everyone gets compliance tools:**
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
compliance: {
|
compliance: {
|
||||||
gdpr: {
|
gdpr: {
|
||||||
rightToDelete: true, // Automatic PII deletion
|
rightToDelete: true, // Automatic PII deletion
|
||||||
|
|
@ -237,7 +231,7 @@ const brain = new BrainyData({
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// Advanced AI capabilities for everyone
|
// Advanced AI capabilities for everyone
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
ai: {
|
ai: {
|
||||||
embeddings: 'state-of-the-art', // Best models available
|
embeddings: 'state-of-the-art', // Best models available
|
||||||
dimensions: 1536, // High-precision vectors
|
dimensions: 1536, // High-precision vectors
|
||||||
|
|
@ -269,7 +263,7 @@ const anomalies = await brain.detectAnomalies()
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
// CI/CD and DevOps features
|
// CI/CD and DevOps features
|
||||||
const brain = new BrainyData({
|
const brain = new Brainy({
|
||||||
operations: {
|
operations: {
|
||||||
blueGreen: true, // Zero-downtime deployments
|
blueGreen: true, // Zero-downtime deployments
|
||||||
canary: true, // Gradual rollouts
|
canary: true, // Gradual rollouts
|
||||||
|
|
@ -318,7 +312,7 @@ Your hobby project today might be tomorrow's unicorn startup. With Brainy, you w
|
||||||
### Startups
|
### Startups
|
||||||
```typescript
|
```typescript
|
||||||
// A 2-person startup gets the same features as Amazon
|
// A 2-person startup gets the same features as Amazon
|
||||||
const startup = new BrainyData()
|
const startup = new Brainy()
|
||||||
// ✓ Full durability
|
// ✓ Full durability
|
||||||
// ✓ Complete security
|
// ✓ Complete security
|
||||||
// ✓ Unlimited scale
|
// ✓ Unlimited scale
|
||||||
|
|
@ -328,7 +322,7 @@ const startup = new BrainyData()
|
||||||
### Education
|
### Education
|
||||||
```typescript
|
```typescript
|
||||||
// Students learn with production-grade tools
|
// Students learn with production-grade tools
|
||||||
const classroom = new BrainyData()
|
const classroom = new Brainy()
|
||||||
// ✓ No feature restrictions
|
// ✓ No feature restrictions
|
||||||
// ✓ Real enterprise experience
|
// ✓ Real enterprise experience
|
||||||
// ✓ Free forever
|
// ✓ Free forever
|
||||||
|
|
@ -337,7 +331,7 @@ const classroom = new BrainyData()
|
||||||
### Non-Profits
|
### Non-Profits
|
||||||
```typescript
|
```typescript
|
||||||
// NGOs get enterprise features without enterprise costs
|
// NGOs get enterprise features without enterprise costs
|
||||||
const nonprofit = new BrainyData()
|
const nonprofit = new Brainy()
|
||||||
// ✓ Compliance tools
|
// ✓ Compliance tools
|
||||||
// ✓ Security features
|
// ✓ Security features
|
||||||
// ✓ Scale for impact
|
// ✓ Scale for impact
|
||||||
|
|
@ -347,7 +341,7 @@ const nonprofit = new BrainyData()
|
||||||
### Enterprises
|
### Enterprises
|
||||||
```typescript
|
```typescript
|
||||||
// Enterprises get everything plus peace of mind
|
// Enterprises get everything plus peace of mind
|
||||||
const enterprise = new BrainyData()
|
const enterprise = new Brainy()
|
||||||
// ✓ Proven at scale
|
// ✓ Proven at scale
|
||||||
// ✓ Community tested
|
// ✓ Community tested
|
||||||
// ✓ No vendor lock-in
|
// ✓ No vendor lock-in
|
||||||
|
|
@ -399,14 +393,14 @@ npm install brainy
|
||||||
```
|
```
|
||||||
|
|
||||||
```typescript
|
```typescript
|
||||||
import { BrainyData } from 'brainy'
|
import { Brainy } from 'brainy'
|
||||||
|
|
||||||
// Create your enterprise-grade database
|
// Create your enterprise-grade database
|
||||||
const brain = new BrainyData()
|
const brain = new Brainy()
|
||||||
await brain.init()
|
await brain.init()
|
||||||
|
|
||||||
// You're now running the same tech as Fortune 500 companies
|
// You're now running the same tech as Fortune 500 companies
|
||||||
await brain.addNoun("Your data is enterprise-grade", {
|
await brain.add("Your data is enterprise-grade", {
|
||||||
secure: true,
|
secure: true,
|
||||||
durable: true,
|
durable: true,
|
||||||
scalable: true,
|
scalable: true,
|
||||||
|
|
@ -444,4 +438,4 @@ Brainy is more than software—it's a movement to democratize enterprise technol
|
||||||
- [Zero Configuration](../architecture/zero-config.md)
|
- [Zero Configuration](../architecture/zero-config.md)
|
||||||
- [Augmentations System](../architecture/augmentations.md)
|
- [Augmentations System](../architecture/augmentations.md)
|
||||||
- [Architecture Overview](../architecture/overview.md)
|
- [Architecture Overview](../architecture/overview.md)
|
||||||
- [Getting Started](./getting-started.md)
|
- [API Reference](../api/README.md)
|
||||||
181
docs/guides/export-and-import.md
Normal file
181
docs/guides/export-and-import.md
Normal file
|
|
@ -0,0 +1,181 @@
|
||||||
|
---
|
||||||
|
title: Export & Import (portable graph)
|
||||||
|
slug: guides/export-and-import
|
||||||
|
public: true
|
||||||
|
category: guides
|
||||||
|
template: guide
|
||||||
|
order: 9
|
||||||
|
description: Export part or all of a brain to a portable, versioned PortableGraph document and import it back — by id, collection, connected neighbourhood, VFS subtree, predicate, or the whole brain. export() lives on the immutable Db, so asOf()/with() give time-travel and what-if exports.
|
||||||
|
next:
|
||||||
|
- guides/subtypes-and-facets
|
||||||
|
- api/README
|
||||||
|
---
|
||||||
|
|
||||||
|
# Export & Import (portable graph)
|
||||||
|
|
||||||
|
Brainy serializes part or all of a brain — an item, a collection, a connected
|
||||||
|
neighbourhood, a VFS subtree, a predicate match, or the whole brain — into a single
|
||||||
|
versioned JSON document (`PortableGraph`), and restores it.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const graph = await brain.export() // whole brain → PortableGraph
|
||||||
|
await brain.import(graph) // restore (merge by id, re-embed if no vectors)
|
||||||
|
```
|
||||||
|
|
||||||
|
It is **portable** (human-readable JSON), **versioned** (`formatVersion`, so a document
|
||||||
|
written by 7.x imports cleanly into 8.0), and **current-state** (the entities and edges as
|
||||||
|
they are now — no generation history). Use it for portable artifacts, partial exports,
|
||||||
|
cross-environment moves, and version upgrades.
|
||||||
|
|
||||||
|
`export()` is a method on the **immutable `Db` value**, so it composes with every way of
|
||||||
|
obtaining one:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.export(sel) // = brain.now().export(sel)
|
||||||
|
;(await brain.asOf(gen)).export(sel) // time-travel export (a past generation)
|
||||||
|
brain.now().with(ops).export(sel) // what-if export (a speculative state)
|
||||||
|
```
|
||||||
|
|
||||||
|
## When to use which
|
||||||
|
|
||||||
|
| You want… | Use |
|
||||||
|
|-----------|-----|
|
||||||
|
| A portable, partial-or-whole, cross-version graph document | **`brain.export()` / `brain.import()`** (this guide) |
|
||||||
|
| A whole-brain snapshot **with generation history** | `brain.now().persist(path)` / `Brainy.load(path)` (native) |
|
||||||
|
| To ingest a CSV / PDF / Excel / JSON **file** as new entities | `brain.import(file)` — see [Import Anything](./import-anything.md) |
|
||||||
|
|
||||||
|
`import()` is **polymorphic**: hand it a `PortableGraph` and it does the graph round-trip;
|
||||||
|
hand it a file/buffer and it does foreign-file ingestion (dispatched on the document's
|
||||||
|
`format: 'brainy-portable-graph'` tag).
|
||||||
|
|
||||||
|
## Exporting
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.export(selector?, options?): Promise<PortableGraph>
|
||||||
|
// (also on any Db: brain.now().export(...), (await brain.asOf(g)).export(...))
|
||||||
|
```
|
||||||
|
|
||||||
|
### Selectors — *what* to export
|
||||||
|
|
||||||
|
Omit the selector to export the whole brain. Otherwise pick a node set:
|
||||||
|
|
||||||
|
| Scenario | Selector |
|
||||||
|
|----------|----------|
|
||||||
|
| Just an item (or items) | `{ ids: ['a', 'b'] }` |
|
||||||
|
| A collection + its children | `{ collection: collectionId }` (alias `memberOf`) |
|
||||||
|
| A connected neighbourhood | `{ connected: { from: id, depth: 2, verbs?, direction? } }` |
|
||||||
|
| A VFS directory / file (+ subtree) | `{ vfsPath: '/docs', recursive?: true }` |
|
||||||
|
| Everything matching a predicate | `{ type, subtype, where, service, visibility }` |
|
||||||
|
| The whole brain | *(omit)* |
|
||||||
|
|
||||||
|
The selector reuses `find()`'s grammar — *"export what `find()` would match, minus ranking
|
||||||
|
and limit."* Structural and predicate selectors **compose**:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
// Members of a collection whose status is "open"
|
||||||
|
await brain.export({ collection: collectionId, where: { status: 'open' } })
|
||||||
|
|
||||||
|
// Already have find() results? Export exactly those with the ids selector
|
||||||
|
const hits = await brain.find({ type: NounType.Document })
|
||||||
|
await brain.export({ ids: hits.map(r => r.id) })
|
||||||
|
```
|
||||||
|
|
||||||
|
### Options — *how* to serialize
|
||||||
|
|
||||||
|
| Option | Default | Effect |
|
||||||
|
|--------|---------|--------|
|
||||||
|
| `includeVectors` | `false` | Carry embedding vectors verbatim. Off ⇒ `import()` re-embeds from `data`. |
|
||||||
|
| `includeContent` | `false` | Include VFS file bytes in `blobs` so files round-trip byte-identically. |
|
||||||
|
| `includeSystem` | `false` | Include `visibility:'system'` entities such as the VFS root. |
|
||||||
|
| `edges` | `'induced'` | `'induced'` (both endpoints in the set), `'incident'` (also dangling edges, recorded in `danglingIds`), or `'none'` (nodes only). |
|
||||||
|
|
||||||
|
## Importing
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
brain.import(graph, options?): Promise<ImportResult>
|
||||||
|
```
|
||||||
|
|
||||||
|
The whole graph is applied as **one atomic transaction** — it advances the brain exactly
|
||||||
|
one generation, or none on failure.
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const result = await brain.import(graph, { onConflict: 'merge' })
|
||||||
|
// → { imported, merged, skipped, reembedded, blobsWritten, errors }
|
||||||
|
```
|
||||||
|
|
||||||
|
| Option | Default | Effect |
|
||||||
|
|--------|---------|--------|
|
||||||
|
| `onConflict` | `'merge'` | `'merge'` (update existing id in place — assemble many exported graphs), `'replace'` (delete + recreate), or `'skip'`. |
|
||||||
|
| `reembed` | `'auto'` | `'auto'` (use the carried vector, else re-embed from `data`) or `'never'` (require a carried vector; record an error if absent). |
|
||||||
|
| `remapIds` | — | Rewrite every id on the way in, e.g. to clone a template subgraph under fresh ids. |
|
||||||
|
| `meta` | — | Transaction metadata recorded in the tx-log alongside the new generation. |
|
||||||
|
|
||||||
|
The default `onConflict: 'merge'` lets you assemble one working graph from many exported
|
||||||
|
documents that share entity ids — re-importing an id merges rather than duplicates.
|
||||||
|
|
||||||
|
## The `PortableGraph` format
|
||||||
|
|
||||||
|
```jsonc
|
||||||
|
{
|
||||||
|
"format": "brainy-portable-graph", // identifies the document type
|
||||||
|
"formatVersion": 1, // import gates on this (cross-version migration)
|
||||||
|
"brainyVersion": "8.0.0",
|
||||||
|
"createdAt": "2026-06-16T…Z",
|
||||||
|
"embedding": { "model": "all-MiniLM-L6-v2", "dimensions": 384 },
|
||||||
|
"selector": { … }, // echoes what was exported (provenance)
|
||||||
|
"entities": [
|
||||||
|
{
|
||||||
|
"id": "…", "type": "Document", "subtype": "invoice", "visibility": "public",
|
||||||
|
"data": "…", // the embedding source
|
||||||
|
"confidence": 1, "weight": 1, "service": "…",
|
||||||
|
"vector": [ … ], // only with includeVectors
|
||||||
|
"metadata": { … } // custom fields only (reserved fields are top-level)
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"relations": [
|
||||||
|
{ "id":"…", "from":"…", "to":"…", "type":"Contains", "subtype":"…",
|
||||||
|
"weight":1, "confidence":1, "metadata": { … } }
|
||||||
|
],
|
||||||
|
"blobs": { "<sha256>": "<base64>" }, // only with includeContent
|
||||||
|
"danglingIds": [ "…" ], // only with edges:'incident'
|
||||||
|
"stats": { "entityCount": 0, "relationCount": 0, "blobCount": 0, "vectorDimensions": 384 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Standard fields (`subtype`, `visibility`, `data`, `confidence`, `weight`, `service`) sit at
|
||||||
|
the top level of each entity; `metadata` holds **only** custom user fields — mirroring the
|
||||||
|
in-memory `Entity` shape, so `import()` maps each field to its dedicated parameter. The
|
||||||
|
TypeScript types (`PortableGraph`, `PortableGraphEntity`, `PortableGraphRelation`, `ExportSelector`,
|
||||||
|
`ExportOptions`, `ImportOptions`, `ImportResult`) are exported from the package root.
|
||||||
|
|
||||||
|
## Generations & time-travel
|
||||||
|
|
||||||
|
The portable document is **current-state** — it never embeds generation history (that keeps
|
||||||
|
it cross-version-portable). History lives where it's queryable:
|
||||||
|
|
||||||
|
- **During a session:** `brain.asOf(g)` / `brain.now().with(ops)` on the live brain. Because
|
||||||
|
`export()` is on the `Db`, `(await brain.asOf(g)).export()` serializes a *past* generation
|
||||||
|
and `brain.now().with(ops).export()` serializes a *speculative* one.
|
||||||
|
- **A whole-brain snapshot with history:** `brain.now().persist(path)` / `Brainy.load(path)`
|
||||||
|
(native, generation-preserving) — a separate facility from this portable format.
|
||||||
|
|
||||||
|
Note: only `transact()` (and the write shortcuts that commit through it) advances a
|
||||||
|
generation, so time-travel export differs across transaction boundaries.
|
||||||
|
|
||||||
|
## Cross-version (7.x → 8.0)
|
||||||
|
|
||||||
|
Because the document is shared and versioned, a PortableGraph written by 7.x imports into 8.0:
|
||||||
|
`formatVersion` is read forward, `subtype` is carried so 8.0 re-types correctly, and the
|
||||||
|
same 384-dimension model on both lines means `includeVectors:false` re-embeds identically
|
||||||
|
(or `true` carries vectors verbatim).
|
||||||
|
|
||||||
|
## VFS
|
||||||
|
|
||||||
|
VFS directories are `Collection` entities and files are entities linked by `Contains`, so
|
||||||
|
the whole filesystem (or any subtree) exports through the `vfsPath` selector:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
await brain.export({ vfsPath: '/' }, { includeContent: true }) // all VFS + bytes
|
||||||
|
await brain.export({ vfsPath: '/docs' }, { includeContent: true }) // one directory
|
||||||
|
await brain.export({ vfsPath: '/a/b.txt' }, { includeContent: true }) // one file
|
||||||
|
```
|
||||||
99
docs/guides/external-backups-and-sparse-storage.md
Normal file
99
docs/guides/external-backups-and-sparse-storage.md
Normal file
|
|
@ -0,0 +1,99 @@
|
||||||
|
---
|
||||||
|
title: External Backups & Sparse Storage
|
||||||
|
slug: guides/external-backups
|
||||||
|
public: true
|
||||||
|
category: guides
|
||||||
|
template: guide
|
||||||
|
order: 10
|
||||||
|
description: How to back up a brain directory with external tools (tar, rsync, cp) without exploding sparse files — why a store can show 100+ GB "apparent" size on a small disk, which files are sparse, and how persist()/restore() handle it for you.
|
||||||
|
next:
|
||||||
|
- guides/snapshots-and-time-travel
|
||||||
|
- concepts/storage-adapters
|
||||||
|
---
|
||||||
|
|
||||||
|
# External Backups & Sparse Storage
|
||||||
|
|
||||||
|
The built-in snapshot path — [`db.persist()` and `brain.restore()`](/docs/guides/snapshots-and-time-travel) —
|
||||||
|
already handles everything on this page for you. Read this when you back up a brain directory with
|
||||||
|
**external tools**: `tar`, `rsync`, `cp`, `scp`, or a filesystem-level backup agent.
|
||||||
|
|
||||||
|
## The one-sentence rule
|
||||||
|
|
||||||
|
> **Always use the sparse-aware flag**: `tar czSf` (capital `S`), `rsync --sparse`,
|
||||||
|
> `cp --sparse=always`. A naive copy can turn a 2 GB store into a 100+ GB one — or fail
|
||||||
|
> the disk entirely.
|
||||||
|
|
||||||
|
## Why: some files are sparse
|
||||||
|
|
||||||
|
When a native accelerator plugin is active, parts of the index live in **memory-mapped files**
|
||||||
|
created at a large fixed virtual size — the file's *apparent* size — while the filesystem only
|
||||||
|
allocates blocks that were actually written. A brand-new id-mapper file can report tens of
|
||||||
|
gigabytes in `ls -l` while occupying a few megabytes on disk.
|
||||||
|
|
||||||
|
Check the difference yourself:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ls -lh brain-data/_id_mapper/ # APPARENT size (can be huge)
|
||||||
|
du -sh brain-data/ # ALLOCATED size (the real footprint)
|
||||||
|
```
|
||||||
|
|
||||||
|
The sparse candidates in a brain directory:
|
||||||
|
|
||||||
|
| Path | What it is |
|
||||||
|
|---|---|
|
||||||
|
| `_id_mapper/` | The native id-mapper's mmap files (large fixed virtual size) |
|
||||||
|
| `_blobs/` | Native index files (vector base, segments) — may be mmap-backed |
|
||||||
|
|
||||||
|
Everything else (entities, `_system`, `_generations`, `_cas` content blobs) is ordinary dense data.
|
||||||
|
|
||||||
|
## Doing it right
|
||||||
|
|
||||||
|
**tar** — the `S` flag detects holes and stores only real data:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
tar czSf brain-backup.tgz /data/brain
|
||||||
|
# restore preserves the holes:
|
||||||
|
tar xzSf brain-backup.tgz -C /data/
|
||||||
|
```
|
||||||
|
|
||||||
|
**rsync**:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
rsync -a --sparse /data/brain/ backup-host:/backups/brain/
|
||||||
|
```
|
||||||
|
|
||||||
|
**cp**:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cp -a --sparse=always /data/brain /backups/brain
|
||||||
|
```
|
||||||
|
|
||||||
|
**What goes wrong without the flag:** the copy *materializes* every hole as real zero bytes.
|
||||||
|
A store whose apparent size exceeds the target disk fails with `ENOSPC` partway through — and a
|
||||||
|
copy that *does* fit silently costs the full apparent size in storage and transfer time.
|
||||||
|
|
||||||
|
## What the built-in paths do (so you don't have to)
|
||||||
|
|
||||||
|
- **`db.persist(path)`** snapshots via **hard links** — instant and space-shared, since every data
|
||||||
|
file is immutable-by-rename. The handful of append-in-place files (the transaction log, the
|
||||||
|
commit fact log's tail segment) and mmap-mutated directories (`_id_mapper/`) are **byte-copied**
|
||||||
|
instead, so a post-snapshot write can never reach through a shared inode into your backup.
|
||||||
|
- **`brain.restore(path, { confirm: true })`** is **non-destructive and sparse-aware**: the snapshot
|
||||||
|
is copied into a staging area *before* any live data is touched (all-zero blocks stay holes), and
|
||||||
|
only after the copy fully succeeds does an atomic swap move it into place. A failed copy —
|
||||||
|
including `ENOSPC` — leaves the live store exactly as it was. A crash mid-swap completes forward
|
||||||
|
on the next open.
|
||||||
|
|
||||||
|
## Live-store caveats for external tools
|
||||||
|
|
||||||
|
1. **Prefer snapshotting a `persist()` output, not the live directory.** `persist()` produces a
|
||||||
|
crash-consistent, immutable snapshot; running `tar` against a live, actively-written directory
|
||||||
|
can capture a torn mid-write state. If you must archive live, stop writes first (or accept that
|
||||||
|
the archive is only as consistent as the moment's flush state).
|
||||||
|
2. **Never prune or "clean up" files inside a brain directory.** Index files that look stale or
|
||||||
|
redundant are load-bearing; the store protects its declared index families from in-process
|
||||||
|
deletion, but an external `rm` bypasses that fence. If space is the concern, `du -sh` first —
|
||||||
|
the allocated size is usually far smaller than it looks.
|
||||||
|
3. **Verify restores by opening them.** `Brainy.load(path)` opens any snapshot or restored
|
||||||
|
directory read-only — the store verifies its own coherence at open and reports loudly if
|
||||||
|
anything is missing or torn.
|
||||||
153
docs/guides/find-limits.md
Normal file
153
docs/guides/find-limits.md
Normal file
|
|
@ -0,0 +1,153 @@
|
||||||
|
---
|
||||||
|
title: Query Limits & Pagination
|
||||||
|
slug: guides/find-limits
|
||||||
|
public: true
|
||||||
|
category: guides
|
||||||
|
template: guide
|
||||||
|
order: 8
|
||||||
|
description: How Brainy caps `find({ limit })` to prevent OOM, the three escape valves when the cap is too tight, and why pagination is the future-proof pattern.
|
||||||
|
next:
|
||||||
|
- guides/aggregation
|
||||||
|
- api/reference
|
||||||
|
---
|
||||||
|
|
||||||
|
# Query Limits & Pagination
|
||||||
|
|
||||||
|
Brainy's `find()` returns entities into a JavaScript array. The size of that array is bounded by an auto-configured cap so a single query can never run the host out of memory. This guide explains the cap, the three ways to raise it when your use case justifies it, and the one pattern that scales no matter what cap is in effect: pagination.
|
||||||
|
|
||||||
|
## Why the cap exists
|
||||||
|
|
||||||
|
Every entity Brainy returns carries:
|
||||||
|
|
||||||
|
- A 384-dim float32 embedding vector (1.5 KB)
|
||||||
|
- Standard fields: `id`, `type`, `subtype`, timestamps, confidence, weight (~200 bytes)
|
||||||
|
- User metadata (variable — typical 5-10 KB, can spike to 20+ KB)
|
||||||
|
|
||||||
|
Conservative budget: **25 KB per result**. A `find({ limit: 100_000 })` against a brain with rich metadata can claim ~2.5 GB before Brainy's iteration starts. JavaScript's GC + V8's heap targets can't absorb that swing without paging or OOM in production.
|
||||||
|
|
||||||
|
The cap is a safety net. It's not the only reason your query might be slow — graph traversal and HNSW search have their own perf characteristics — but it's the one that turns a slow query into a sudden runtime error.
|
||||||
|
|
||||||
|
## The auto-configured cap (7.30.2+)
|
||||||
|
|
||||||
|
Brainy picks `maxLimit` from the first of these that's available:
|
||||||
|
|
||||||
|
| Priority | Source | Formula |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | Constructor option `maxQueryLimit` | Hard cap at supplied value, max 100 000 |
|
||||||
|
| 2 | Constructor option `reservedQueryMemory` | `floor(reservedQueryMemory / 25 KB)` capped at 100 000 |
|
||||||
|
| 3 | Detected container memory limit (Cloud Run, Kubernetes, cgroups v1/v2) | `floor(containerLimit × 0.25 / 25 KB)` capped at 100 000 |
|
||||||
|
| 4 | Free system memory | `floor(availableMemory / 25 KB)` capped at 100 000 |
|
||||||
|
|
||||||
|
Worked example: a 4 GB Cloud Run container picks priority 3 → `floor(4 GB × 0.25 / 25 KB) = floor(40 960) = 40 000` results. A 900 MB free-memory box on priority 4 gets `floor(900 MB / 25 KB) = ~36 000`.
|
||||||
|
|
||||||
|
The cap is fixed at construction and never changes at runtime. Query timing is recorded
|
||||||
|
for diagnostics only — a burst of slow queries cannot silently shrink the cap, and the
|
||||||
|
auto-detected tiers (3 and 4) never go below a floor of 10 000.
|
||||||
|
|
||||||
|
> **Calibration note.** Pre-7.30.2 used 100 KB per result instead of 25 KB, which produced caps that were 4× too tight for typical workloads (an 8 KB / result reality). 7.30.2 recalibrated to match observed entity sizes; existing `limit: 10_000` safety patterns now pass silently on any reasonably-sized box.
|
||||||
|
|
||||||
|
## What happens when you exceed the cap
|
||||||
|
|
||||||
|
`find({ limit })` enforces in **two tiers**:
|
||||||
|
|
||||||
|
### Soft tier: `maxLimit < limit ≤ 2 × maxLimit`
|
||||||
|
|
||||||
|
You get a one-time warning per call site:
|
||||||
|
|
||||||
|
```
|
||||||
|
[Brainy] find({ limit: 50000 }) exceeds the auto-configured query limit of
|
||||||
|
40000 (basis: detected container memory limit). Choose one:
|
||||||
|
• Increase the cap: new Brainy({ maxQueryLimit: 50000 })
|
||||||
|
• Reserve more memory: new Brainy({ reservedQueryMemory: 1310720000 })
|
||||||
|
• Paginate: split the query with { limit, offset } pages
|
||||||
|
at YourService.loadDashboard (/app/src/dashboard.ts:142:18)
|
||||||
|
Docs: https://soulcraft.com/docs/guides/find-limits
|
||||||
|
```
|
||||||
|
|
||||||
|
**The query proceeds.** Brainy returns the result set you asked for; the warning is a teaching signal, not a block. Existing code that relied on the cap silently allowing safety-cap limits (`limit: 10_000` against a 9 K-cap box) keeps working — the warning shows you the recipe so you can fix it intentionally.
|
||||||
|
|
||||||
|
### Hard tier: `limit > 2 × maxLimit`
|
||||||
|
|
||||||
|
Same message, but thrown as an error. This is real OOM territory; the cap stops being a recommendation and becomes a guardrail.
|
||||||
|
|
||||||
|
## The three escape valves
|
||||||
|
|
||||||
|
### 1. Raise the cap at construction — `maxQueryLimit`
|
||||||
|
|
||||||
|
When the auto-config is wrong for your workload (e.g. you know your entities are smaller than 25 KB average and you need bigger result sets), set an explicit cap:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
storage: { type: 'filesystem', path: './data' },
|
||||||
|
maxQueryLimit: 50_000 // raises the cap; still hard-clamped at 100 000
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This is the right answer when:
|
||||||
|
- Your entity metadata is genuinely small (e.g. 1-2 KB) and 25 KB per result is over-conservative
|
||||||
|
- You're running on a box with lots of headroom and 25% of memory underestimates what you can spare for queries
|
||||||
|
- You need a known-good limit that doesn't change when the box's free-memory wiggles at startup
|
||||||
|
|
||||||
|
### 2. Reserve more memory for queries — `reservedQueryMemory`
|
||||||
|
|
||||||
|
When you want the cap to be memory-derived but more generous than the default 25% slice:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const brain = new Brainy({
|
||||||
|
reservedQueryMemory: 1024 * 1024 * 1024 // 1 GB → ~40 000 result cap
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This is the right answer when:
|
||||||
|
- Your host's memory budget for queries is known and stable, regardless of free-memory at startup
|
||||||
|
- You want the formula to scale with the documented per-result size (25 KB) instead of a hard number
|
||||||
|
|
||||||
|
### 3. Paginate — the future-proof pattern
|
||||||
|
|
||||||
|
If your query genuinely needs to walk all matches in a category, don't fight the cap — walk in pages:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
async function findAll<T>(params: FindParams<T>, pageSize = 1000): Promise<Result<T>[]> {
|
||||||
|
const all: Result<T>[] = []
|
||||||
|
let offset = 0
|
||||||
|
while (true) {
|
||||||
|
const page = await brain.find({ ...params, limit: pageSize, offset })
|
||||||
|
all.push(...page)
|
||||||
|
if (page.length < pageSize) break
|
||||||
|
offset += page.length
|
||||||
|
}
|
||||||
|
return all
|
||||||
|
}
|
||||||
|
|
||||||
|
// Use it just like find():
|
||||||
|
const allEvents = await findAll({ type: NounType.Event, where: { status: 'open' } })
|
||||||
|
```
|
||||||
|
|
||||||
|
For very large brains, prefer the streaming API which avoids holding the full result set in memory at all:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
for await (const entity of brain.streaming.entities({ type: NounType.Event })) {
|
||||||
|
// process one entity at a time
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## When to use which
|
||||||
|
|
||||||
|
| Situation | Recommended valve |
|
||||||
|
|---|---|
|
||||||
|
| The cap is unreasonably low for your known entity size | `maxQueryLimit` |
|
||||||
|
| You want a memory-derived cap but more generous than 25% | `reservedQueryMemory` |
|
||||||
|
| Your query needs ALL matches in a category | Pagination or `brain.streaming.entities()` |
|
||||||
|
| You hit the cap once during a one-off migration | `maxQueryLimit` or `migrateField` (which already paginates internally) |
|
||||||
|
| You're hitting the cap on a recurring user-facing query | Pagination — the cap will get tighter in 8.0, not looser |
|
||||||
|
|
||||||
|
## A note on Brainy 8.0
|
||||||
|
|
||||||
|
8.0's Datomic-style `Db` API may make per-call limits stricter to keep snapshot semantics cheap. **Pagination is the only pattern that's guaranteed to keep working unchanged.** Code that paginates today doesn't need to revisit when 8.0 ships.
|
||||||
|
|
||||||
|
## Reference
|
||||||
|
|
||||||
|
- `BrainyConfig.maxQueryLimit?: number` — explicit cap override (max 100 000)
|
||||||
|
- `BrainyConfig.reservedQueryMemory?: number` — memory budget for queries (bytes)
|
||||||
|
- `find({ limit, offset })` — paginated find
|
||||||
|
- `brain.streaming.entities(filter)` — streaming alternative for very large traversals
|
||||||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Add a link
Reference in a new issue