2026-01-27 15:38:21 -08:00
# Storage Architecture
2025-08-26 12:32:21 -07:00
2026-01-27 15:38:21 -08:00
> **Updated**: Metadata/vector separation, UUID-based sharding, lifecycle management
2025-08-26 12:32:21 -07:00
## Storage Structure
2026-01-27 15:38:21 -08:00
### Architecture: Metadata/Vector Separation
2025-10-17 14:47:53 -07:00
2026-01-27 15:38:21 -08:00
entities and relationships are split into **2 separate files** for optimal performance at billion-entity scale:
2025-10-17 14:47:53 -07:00
2025-08-26 12:32:21 -07:00
```
brainy-data/
2026-01-27 15:38:21 -08:00
├── _system/ # System metadata (not sharded)
│ ├── statistics.json # Performance metrics
│ ├── __metadata_field_index__ *.json # Field indexes
│ └── __metadata_sorted_index__ *.json # Sorted indexes
2025-10-17 14:47:53 -07:00
│
├── entities/
2026-01-27 15:38:21 -08:00
│ ├── nouns/
│ │ ├── vectors/ # HNSW graph data (sharded by UUID)
│ │ │ ├── 00/ # Shard 00 (first 2 hex digits)
│ │ │ │ ├── 00123456-....json # Vector + HNSW connections
│ │ │ │ └── 00abcdef-....json
│ │ │ ├── 01/ ... ff/ # 256 shards total
│ │ │
│ │ └── metadata/ # Business data (sharded by UUID)
│ │ ├── 00/
│ │ │ ├── 00123456-....json # Entity metadata only
│ │ │ └── 00abcdef-....json
│ │ ├── 01/ ... ff/
│ │
│ └── verbs/
│ ├── vectors/ # Relationship vectors (sharded)
│ │ ├── 00/ ... ff/
│ │
│ └── metadata/ # Relationship data (sharded)
│ ├── 00/ ... ff/
2025-10-17 14:47:53 -07:00
```
### Why Split Metadata and Vectors?
**Performance at scale:**
- **HNSW operations**: Only load vectors (4KB) during search, not metadata (2-10KB)
- **Filtering**: Only load metadata during filtering, not vectors
- **Pagination**: Load metadata IDs first, fetch vectors/metadata on-demand
- **Result**: 60-70% reduction in I/O for typical queries at million-entity scale
### UUID-Based Sharding (256 Shards)
**How it works:**
```typescript
const uuid = "3fa85f64-5717-4562-b3fc-2c963f66afa6"
2026-01-27 15:38:21 -08:00
const shard = uuid.substring(0, 2) // "3f"
2025-10-17 14:47:53 -07:00
2026-01-27 15:38:21 -08:00
// Vector path: entities/nouns/vectors/3f/3fa85f64-....json
2025-10-17 14:47:53 -07:00
// Metadata path: entities/nouns/metadata/3f/3fa85f64-....json
2025-08-26 12:32:21 -07:00
```
2025-10-17 14:47:53 -07:00
**Benefits:**
- **Uniform distribution**: ~3,900 entities per shard (at 1M scale)
- **Cloud storage optimization**: 200x faster than unsharded (30s → 150ms)
- **Parallel operations**: Load 256 shards in parallel
- **Predictable**: Deterministic shard assignment
2025-08-26 12:32:21 -07:00
## Storage Adapters
2026-01-27 15:38:21 -08:00
Brainy provides multiple storage adapters with identical APIs and production features:
2025-08-26 12:32:21 -07:00
### FileSystem Storage (Node.js)
```typescript
2025-09-30 17:09:15 -07:00
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'filesystem',
path: './data',
compression: true // Gzip compression (60-80% space savings)
}
2025-08-26 12:32:21 -07:00
})
```
- **Use case**: Server applications, CLI tools
2025-10-17 14:47:53 -07:00
- **Performance**: Direct file I/O with optional compression
2025-08-26 12:32:21 -07:00
- **Persistence**: Permanent on disk
2026-01-27 15:38:21 -08:00
- **Features**:
- **Gzip Compression**: 60-80% storage savings with minimal CPU overhead
- **Batch Delete**: Efficient bulk deletion with retries
- **UUID Sharding**: Automatic 256-shard distribution
2025-08-26 12:32:21 -07:00
2025-10-17 14:47:53 -07:00
### S3 Compatible Storage (AWS, MinIO, R2)
2025-08-26 12:32:21 -07:00
```typescript
2025-09-30 17:09:15 -07:00
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 's3',
bucket: 'my-brainy-data',
region: 'us-east-1',
credentials: {
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY
}
}
2025-08-26 12:32:21 -07:00
})
```
- **Use case**: Distributed applications, cloud deployments
- **Performance**: Network dependent, with intelligent caching
2025-10-17 14:47:53 -07:00
- **Persistence**: Cloud storage durability (99.999999999%)
2026-01-27 15:38:21 -08:00
- **Features**:
- **Lifecycle Policies**: Automatic tier transitions (Standard → IA → Glacier → Deep Archive)
- **Intelligent-Tiering**: Automatic optimization based on access patterns (up to 95% savings)
- **Batch Delete**: Efficient bulk deletion (1000 objects per request)
- **Cost Impact**: $138k/year → $5.9k/year at 500TB (96% savings!)
2025-10-17 14:47:53 -07:00
### Google Cloud Storage (GCS)
```typescript
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'gcs',
bucketName: 'my-brainy-data',
keyFilename: './service-account.json' // Or use ADC
}
2025-10-17 14:47:53 -07:00
})
```
- **Use case**: Google Cloud deployments
- **Performance**: Global CDN with edge caching
- **Persistence**: 99.999999999% durability
2026-01-27 15:38:21 -08:00
- **Features**:
- **Lifecycle Policies**: Automatic tier transitions (Standard → Nearline → Coldline → Archive)
- **Autoclass**: Intelligent automatic tier optimization
- **Batch Delete**: Efficient bulk operations
- **Cost Impact**: $138k/year → $8.3k/year at 500TB (94% savings!)
2025-10-17 14:47:53 -07:00
### Azure Blob Storage
```typescript
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'azure',
connectionString: process.env.AZURE_STORAGE_CONNECTION_STRING,
containerName: 'brainy-data'
}
2025-10-17 14:47:53 -07:00
})
```
- **Use case**: Azure cloud deployments
- **Performance**: Global replication with CDN
- **Persistence**: LRS, ZRS, GRS, RA-GRS options
2026-01-27 15:38:21 -08:00
- **Features**:
- **Blob Tier Management**: Hot/Cool/Archive tiers (99% cost savings)
- **Lifecycle Policies**: Automatic tier transitions and deletions
- **Batch Delete**: BlobBatchClient for efficient bulk operations
- **Batch Tier Changes**: Move thousands of blobs efficiently
- **Archive Rehydration**: Smart rehydration with priority options
2025-08-26 12:32:21 -07:00
### Origin Private File System (Browser)
```typescript
2025-09-30 17:09:15 -07:00
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'opfs'
}
2025-08-26 12:32:21 -07:00
})
```
- **Use case**: Browser applications, PWAs
- **Performance**: Near-native file system speed
- **Persistence**: Permanent in browser (with quota limits)
2026-01-27 15:38:21 -08:00
- **Features**:
- **Quota Monitoring**: Real-time quota tracking and warnings
- **Batch Delete**: Efficient bulk deletion
- **Storage Status**: Detailed usage/available reporting
2025-08-26 12:32:21 -07:00
## Metadata Indexing System
### Field Discovery Index
Tracks all unique values for each field:
```json
// __metadata_field_index__field_category.json
{
2026-01-27 15:38:21 -08:00
"values": {
"technology": 45,
"science": 32,
"business": 28
},
"lastUpdated": 1699564234567
2025-08-26 12:32:21 -07:00
}
```
### Value-Based Indexes
Maps field+value combinations to entity IDs:
```json
// __metadata_index__category_technology_chunk0.json
{
2026-01-27 15:38:21 -08:00
"field": "category",
"value": "technology",
"ids": ["uuid1", "uuid2", "uuid3", ...],
"chunk": 0,
"total": 45
2025-08-26 12:32:21 -07:00
}
```
### Index Chunking
Large indexes automatically chunk for performance:
- **Chunk size**: 10,000 IDs per chunk
- **Auto-splitting**: Transparent to queries
- **Parallel loading**: Chunks load on demand
## Entity Registry
High-performance deduplication system for streaming data:
### Registry Structure
```json
// __entity_registry__ .json
{
2026-01-27 15:38:21 -08:00
"mappings": {
"did:plc:alice123": "550e8400-e29b-41d4-a716-446655440000",
"handle:alice.bsky.social": "550e8400-e29b-41d4-a716-446655440000"
},
"stats": {
"totalMappings": 10000,
"lastSync": 1699564234567
}
2025-08-26 12:32:21 -07:00
}
```
### Performance Characteristics
- **Lookup**: O(1) in-memory hash map
- **Persistence**: Configurable (memory/storage/hybrid)
- **Cache**: LRU with configurable TTL
- **Sync**: Periodic or on-demand
Ensures durability and enables recovery:
```json
{
2026-01-27 15:38:21 -08:00
"timestamp": 1699564234567,
"operation": "add",
"data": {
"id": "550e8400-e29b-41d4-a716-446655440000",
"content": "...",
"metadata": {}
},
"checksum": "sha256:..."
2025-08-26 12:32:21 -07:00
}
```
### Recovery Process
2. Replay operations from last checkpoint
3. Verify checksums for integrity
2026-01-27 15:38:21 -08:00
## Storage Optimization
2025-10-17 14:47:53 -07:00
### 1. Lifecycle Policies (Cloud Storage)
2025-08-26 12:32:21 -07:00
2025-10-17 14:47:53 -07:00
**Automatic cost optimization through tier transitions:**
2025-08-26 12:32:21 -07:00
```typescript
2025-10-17 14:47:53 -07:00
// S3: Set lifecycle policy for automatic archival
await storage.setLifecyclePolicy({
2026-01-27 15:38:21 -08:00
rules: [{
id: 'archive-old-data',
prefix: 'entities/',
status: 'Enabled',
transitions: [
{ days: 30, storageClass: 'STANDARD_IA' }, // Move to IA after 30 days
{ days: 90, storageClass: 'GLACIER' }, // Archive after 90 days
{ days: 365, storageClass: 'DEEP_ARCHIVE' } // Deep archive after 1 year
]
}]
2025-10-17 14:47:53 -07:00
})
// GCS: Set lifecycle policy
await storage.setLifecyclePolicy({
2026-01-27 15:38:21 -08:00
rules: [{
condition: { age: 30 },
action: { type: 'SetStorageClass', storageClass: 'NEARLINE' }
}, {
condition: { age: 90 },
action: { type: 'SetStorageClass', storageClass: 'COLDLINE' }
}, {
condition: { age: 365 },
action: { type: 'SetStorageClass', storageClass: 'ARCHIVE' }
}]
2025-10-17 14:47:53 -07:00
})
// Azure: Set lifecycle policy
await storage.setLifecyclePolicy({
2026-01-27 15:38:21 -08:00
rules: [{
name: 'archiveOldData',
enabled: true,
type: 'Lifecycle',
definition: {
filters: { blobTypes: ['blockBlob'] },
actions: {
baseBlob: {
tierToCool: { daysAfterModificationGreaterThan: 30 },
tierToArchive: { daysAfterModificationGreaterThan: 90 }
}
}
}
}]
2025-10-17 14:47:53 -07:00
})
```
**Cost Impact (500TB dataset):**
| Storage | Before | After | Savings |
|---------|--------|-------|---------|
| **AWS S3** | $138,000/yr | $5,940/yr | **96%** |
| **GCS** | $138,000/yr | $8,300/yr | **94%** |
| **Azure** | $107,520/yr | $5,016/yr | **95%** |
### 2. Intelligent-Tiering (S3)
**Automatic optimization without retrieval fees:**
```typescript
// Enable S3 Intelligent-Tiering
await storage.enableIntelligentTiering('entities/', 'auto-optimize')
// Benefits:
// - Automatic tier transitions based on access patterns
// - No retrieval fees (unlike Glacier)
// - Up to 95% cost savings
// - No performance impact on frequently accessed data
```
### 3. Autoclass (GCS)
**Google Cloud's intelligent automatic optimization:**
```typescript
// Enable GCS Autoclass
await storage.enableAutoclass({
2026-01-27 15:38:21 -08:00
terminalStorageClass: 'ARCHIVE' // Optional: Set lowest tier
2025-10-17 14:47:53 -07:00
})
// Benefits:
// - Automatic optimization based on access patterns
// - No data retrieval delays
// - Transparent tier transitions
// - Up to 94% cost savings
```
### 4. Compression (FileSystem)
```typescript
// Enable gzip compression for local storage
2025-09-30 17:09:15 -07:00
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'filesystem',
path: './data',
compression: true // 60-80% space savings
}
2025-08-26 12:32:21 -07:00
})
2025-10-17 14:47:53 -07:00
// Performance impact:
// - Write: +10-20ms per file (gzip compression)
// - Read: +5-10ms per file (gzip decompression)
// - Space savings: 60-80% for typical JSON data
// - CPU overhead: Minimal (~5% CPU)
2025-08-26 12:32:21 -07:00
```
2025-10-17 14:47:53 -07:00
### 5. Batch Operations
2025-08-26 12:32:21 -07:00
```typescript
2026-01-27 15:38:21 -08:00
// Efficient batch delete
2025-10-17 14:47:53 -07:00
await storage.batchDelete([
2026-01-27 15:38:21 -08:00
'entities/nouns/vectors/00/00123456-....json',
'entities/nouns/metadata/00/00123456-....json',
// ... up to 1000 objects
2025-10-17 14:47:53 -07:00
])
// Benefits:
// - S3: 1000 objects per request (vs 1 per request)
// - GCS: 100 objects per request
// - Azure: 256 objects per batch
// - Automatic retry logic with exponential backoff
// - Throttling protection
2025-08-26 12:32:21 -07:00
// Batch writes for performance
await brain.addBatch([
2026-01-27 15:38:21 -08:00
{ content: "item1", metadata: {} },
{ content: "item2", metadata: {} },
{ content: "item3", metadata: {} }
2025-08-26 12:32:21 -07:00
])
// Single transaction, optimized I/O
```
2025-10-17 14:47:53 -07:00
### 6. Quota Monitoring (OPFS)
```typescript
// Get quota status for browser storage
const status = await storage.getStorageStatus()
console.log(status)
// {
2026-01-27 15:38:21 -08:00
// type: 'opfs',
// available: true,
// details: {
// usage: 45829120, // 43.7 MB used
// quota: 536870912, // 512 MB available
// usagePercent: 8.5,
// quotaExceeded: false
// }
2025-10-17 14:47:53 -07:00
// }
// Proactive quota management:
// - Monitor usage before writes
// - Warn users when approaching quota
// - Automatically clean up old data
```
### 7. Tier Management (Azure)
```typescript
// Change blob tier for cost optimization
2026-01-27 15:38:21 -08:00
await storage.changeBlobTier(blobPath, 'Cool') // Hot → Cool (50% savings)
await storage.changeBlobTier(blobPath, 'Archive') // Cool → Archive (99% savings)
2025-10-17 14:47:53 -07:00
// Batch tier changes (efficient)
await storage.batchChangeTier([blob1, blob2, blob3], 'Cool')
// Rehydrate from Archive when needed
2026-01-27 15:38:21 -08:00
await storage.rehydrateBlob(blobPath, 'Standard') // Standard or High priority
2025-10-17 14:47:53 -07:00
```
### 8. Caching Strategy
```typescript
// Configure caching per storage type
const brain = new Brainy({
2026-01-27 15:38:21 -08:00
storage: {
type: 'filesystem',
cache: {
enabled: true,
maxSize: 1000, // Maximum cached items
ttl: 300000, // 5 minutes
strategy: 'lru' // Least recently used
}
}
2025-10-17 14:47:53 -07:00
})
```
2025-08-26 12:32:21 -07:00
## Concurrent Access
### Locking Mechanism
```typescript
// Automatic locking for write operations
await brain.storage.withLock('resource-id', async () => {
2026-01-27 15:38:21 -08:00
// Exclusive access to resource
await brain.storage.saveNoun(id, data)
2025-08-26 12:32:21 -07:00
})
```
### Read-Write Separation
- **Reads**: Non-blocking, parallel
- **Writes**: Serialized with locks
- **Hybrid**: Read-heavy optimization
## Migration and Backup
### Export Data
```typescript
// Export entire database
const backup = await brain.export({
2026-01-27 15:38:21 -08:00
format: 'json',
includeVectors: true,
includeIndexes: false
2025-08-26 12:32:21 -07:00
})
```
### Import Data
```typescript
// Import from backup
await brain.import(backup, {
2026-01-27 15:38:21 -08:00
mode: 'merge', // or 'replace'
validateSchema: true
2025-08-26 12:32:21 -07:00
})
```
### Storage Migration
```typescript
// Migrate between storage types
2025-09-30 17:09:15 -07:00
const oldBrain = new Brainy({ storage: { type: 'filesystem' } })
const newBrain = new Brainy({ storage: { type: 's3' } })
2025-08-26 12:32:21 -07:00
await oldBrain.init()
await newBrain.init()
// Transfer all data
const data = await oldBrain.export()
await newBrain.import(data)
```
## Performance Tuning
### Storage-Specific Optimizations
#### FileSystem
- **Directory sharding**: Split files across subdirectories
- **Async I/O**: Non-blocking file operations
- **Buffer pooling**: Reuse buffers for efficiency
#### S3
- **Multipart uploads**: For large objects
- **Request batching**: Combine small operations
- **CDN integration**: Edge caching for reads
#### OPFS
- **Quota management**: Monitor and request increases
- **Worker offloading**: Heavy operations in workers
- **Transaction batching**: Group operations
### Monitoring
```typescript
// Get storage statistics
const stats = await brain.storage.getStatistics()
console.log(stats)
// {
2026-01-27 15:38:21 -08:00
// totalSize: 1048576,
// entityCount: 1000,
// indexSize: 204800,
// walSize: 10240,
// cacheHitRate: 0.85
2025-08-26 12:32:21 -07:00
// }
```
2026-01-27 15:38:21 -08:00
## Best Practices
2025-08-26 12:32:21 -07:00
### Choose the Right Adapter
2025-10-17 14:47:53 -07:00
1. **Development** : FileSystem with compression (local persistence, small storage footprint)
2. **Production Server** : FileSystem with compression or cloud storage with lifecycle policies
3. **Browser Apps** : OPFS with quota monitoring
4. **Distributed** : S3/GCS/Azure with Intelligent-Tiering/Autoclass
2025-08-26 12:32:21 -07:00
### Optimize for Your Use Case
2025-10-17 14:47:53 -07:00
1. **Read-heavy** : Enable aggressive caching + cloud CDN
2. **Write-heavy** : Batch operations + async writes
2025-10-10 14:09:30 -07:00
3. **Real-time** : FileSystem with periodic snapshots
2025-10-17 14:47:53 -07:00
4. **Archival** : Cloud storage with lifecycle policies (96% cost savings!)
5. **Large-scale** : Metadata/vector separation + UUID sharding + lifecycle policies
2026-01-27 15:38:21 -08:00
### Cost Optimization
2025-10-17 14:47:53 -07:00
1. **Enable lifecycle policies** for cloud storage (automated cost reduction)
2. **Use Intelligent-Tiering (S3)** or Autoclass (GCS) for automatic optimization
3. **Enable compression** for FileSystem storage (60-80% space savings)
4. **Monitor quota** for OPFS (prevent quota exceeded errors)
5. **Use batch operations** for bulk deletions (efficient API usage)
6. **Consider tier management** for Azure (Hot/Cool/Archive tiers)
**Example Cost Savings (500TB dataset):**
- Without lifecycle policies: ** $138,000/year**
2026-01-27 15:38:21 -08:00
- With lifecycle policies: ** $5,940/year**
2025-10-17 14:47:53 -07:00
- **Savings: $132,060/year (96%)**
2025-08-26 12:32:21 -07:00
### Monitor and Maintain
1. Regular statistics collection
2025-10-17 14:47:53 -07:00
2. Monitor lifecycle policy effectiveness
2025-08-26 12:32:21 -07:00
3. Index optimization
4. Cache tuning based on hit rates
2025-10-17 14:47:53 -07:00
5. Track storage costs and tier distribution
6. Review quota usage (OPFS) and storage growth patterns
### Production Deployment Checklist
- ✅ Enable lifecycle policies on cloud storage
- ✅ Configure batch delete for cleanup operations
- ✅ Enable compression for FileSystem storage
- ✅ Set up quota monitoring for OPFS
- ✅ Configure appropriate tier transitions
- ✅ Enable Intelligent-Tiering (S3) or Autoclass (GCS)
- ✅ Monitor storage costs and optimize regularly
2025-08-26 12:32:21 -07:00
## API Reference
2025-10-17 14:47:53 -07:00
See the [Storage API ](../api/storage.md ) for complete method documentation.
---
**Version**: 4.0.0
**Last Updated**: 2025-10-17
**Key Features**: Metadata/vector separation, UUID sharding, lifecycle management, tier optimization