brainy/docs/architecture/storage-architecture.md
David Snelling 89036deb20 refactor: rename BackupData → PortableGraph (the type is interchange, not a backup)
Parity with the 8.0 rename, backported to the 7.x line. The brain.data()
export()/import() document type was BackupData, but it is the portable, versioned
interchange representation of a graph (entities + relations + optional vectors),
NOT a backup — exported for transport between instances, versions, and products.
The actual backup is the native snapshot, so "Backup*" mis-signalled.

Rename every developer-visible symbol, JSDoc, comment and doc:
- BackupData→PortableGraph, BackupEntity→PortableGraphEntity,
  BackupRelation→PortableGraphRelation, BACKUP_FORMAT[_VERSION]→
  PORTABLE_GRAPH_FORMAT[_VERSION], internal toBackup* helpers→toPortableGraph*.
- src/api/DataAPI.ts, src/index.ts, src/cli/commands/core.ts, docs, and the test
  (renamed data-backup.test.ts → data-portable-graph.test.ts).

The on-the-wire `format` tag is also renamed 'brainy-backup' → 'brainy-portable-graph':
the export/import format was introduced in 7.32.0 and has not been adopted by any
consumer, so there are no stored documents to stay compatible with — a clean rename
beats carrying a legacy tag forward. No deprecated aliases (nothing to alias).

1505 unit green; build clean; data-portable-graph.test.ts 20/20.
2026-06-19 12:30:06 -07:00

16 KiB

Storage Architecture

Updated: Metadata/vector separation, UUID-based sharding, lifecycle management

Storage Structure

Architecture: Metadata/Vector Separation

entities and relationships are split into 2 separate files for optimal performance at billion-entity scale:

brainy-data/
├── _system/ # System metadata (not sharded)
│ ├── statistics.json # Performance metrics
│ ├── __metadata_field_index__*.json # Field indexes
│ └── __metadata_sorted_index__*.json # Sorted indexes
│
├── entities/
│ ├── nouns/
│ │ ├── vectors/ # HNSW graph data (sharded by UUID)
│ │ │ ├── 00/ # Shard 00 (first 2 hex digits)
│ │ │ │ ├── 00123456-....json # Vector + HNSW connections
│ │ │ │ └── 00abcdef-....json
│ │ │ ├── 01/ ... ff/ # 256 shards total
│ │ │
│ │ └── metadata/ # Business data (sharded by UUID)
│ │ ├── 00/
│ │ │ ├── 00123456-....json # Entity metadata only
│ │ │ └── 00abcdef-....json
│ │ ├── 01/ ... ff/
│ │
│ └── verbs/
│ ├── vectors/ # Relationship vectors (sharded)
│ │ ├── 00/ ... ff/
│ │
│ └── metadata/ # Relationship data (sharded)
│ ├── 00/ ... ff/

Why Split Metadata and Vectors?

Performance at scale:

  • HNSW operations: Only load vectors (4KB) during search, not metadata (2-10KB)
  • Filtering: Only load metadata during filtering, not vectors
  • Pagination: Load metadata IDs first, fetch vectors/metadata on-demand
  • Result: 60-70% reduction in I/O for typical queries at million-entity scale

UUID-Based Sharding (256 Shards)

How it works:

const uuid = "3fa85f64-5717-4562-b3fc-2c963f66afa6"
const shard = uuid.substring(0, 2) // "3f"

// Vector path: entities/nouns/vectors/3f/3fa85f64-....json
// Metadata path: entities/nouns/metadata/3f/3fa85f64-....json

Benefits:

  • Uniform distribution: ~3,900 entities per shard (at 1M scale)
  • Cloud storage optimization: 200x faster than unsharded (30s → 150ms)
  • Parallel operations: Load 256 shards in parallel
  • Predictable: Deterministic shard assignment

Storage Adapters

Brainy provides multiple storage adapters with identical APIs and production features:

FileSystem Storage (Node.js)

const brain = new Brainy({
 storage: {
 type: 'filesystem',
 path: './data',
 compression: true // Gzip compression (60-80% space savings)
 }
})
  • Use case: Server applications, CLI tools
  • Performance: Direct file I/O with optional compression
  • Persistence: Permanent on disk
  • Features:
  • Gzip Compression: 60-80% storage savings with minimal CPU overhead
  • Batch Delete: Efficient bulk deletion with retries
  • UUID Sharding: Automatic 256-shard distribution

S3 Compatible Storage (AWS, MinIO, R2)

const brain = new Brainy({
 storage: {
 type: 's3',
 bucket: 'my-brainy-data',
 region: 'us-east-1',
 credentials: {
 accessKeyId: process.env.AWS_ACCESS_KEY_ID,
 secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY
 }
 }
})
  • Use case: Distributed applications, cloud deployments
  • Performance: Network dependent, with intelligent caching
  • Persistence: Cloud storage durability (99.999999999%)
  • Features:
  • Lifecycle Policies: Automatic tier transitions (Standard → IA → Glacier → Deep Archive)
  • Intelligent-Tiering: Automatic optimization based on access patterns (up to 95% savings)
  • Batch Delete: Efficient bulk deletion (1000 objects per request)
  • Cost Impact: $138k/year → $5.9k/year at 500TB (96% savings!)

Google Cloud Storage (GCS)

const brain = new Brainy({
 storage: {
 type: 'gcs',
 bucketName: 'my-brainy-data',
 keyFilename: './service-account.json' // Or use ADC
 }
})
  • Use case: Google Cloud deployments
  • Performance: Global CDN with edge caching
  • Persistence: 99.999999999% durability
  • Features:
  • Lifecycle Policies: Automatic tier transitions (Standard → Nearline → Coldline → Archive)
  • Autoclass: Intelligent automatic tier optimization
  • Batch Delete: Efficient bulk operations
  • Cost Impact: $138k/year → $8.3k/year at 500TB (94% savings!)

Azure Blob Storage

const brain = new Brainy({
 storage: {
 type: 'azure',
 connectionString: process.env.AZURE_STORAGE_CONNECTION_STRING,
 containerName: 'brainy-data'
 }
})
  • Use case: Azure cloud deployments
  • Performance: Global replication with CDN
  • Persistence: LRS, ZRS, GRS, RA-GRS options
  • Features:
  • Blob Tier Management: Hot/Cool/Archive tiers (99% cost savings)
  • Lifecycle Policies: Automatic tier transitions and deletions
  • Batch Delete: BlobBatchClient for efficient bulk operations
  • Batch Tier Changes: Move thousands of blobs efficiently
  • Archive Rehydration: Smart rehydration with priority options

Origin Private File System (Browser)

const brain = new Brainy({
 storage: {
 type: 'opfs'
 }
})
  • Use case: Browser applications, PWAs
  • Performance: Near-native file system speed
  • Persistence: Permanent in browser (with quota limits)
  • Features:
  • Quota Monitoring: Real-time quota tracking and warnings
  • Batch Delete: Efficient bulk deletion
  • Storage Status: Detailed usage/available reporting

Metadata Indexing System

Field Discovery Index

Tracks all unique values for each field:

// __metadata_field_index__field_category.json
{
 "values": {
 "technology": 45,
 "science": 32,
 "business": 28
 },
 "lastUpdated": 1699564234567
}

Value-Based Indexes

Maps field+value combinations to entity IDs:

// __metadata_index__category_technology_chunk0.json
{
 "field": "category",
 "value": "technology",
 "ids": ["uuid1", "uuid2", "uuid3", ...],
 "chunk": 0,
 "total": 45
}

Index Chunking

Large indexes automatically chunk for performance:

  • Chunk size: 10,000 IDs per chunk
  • Auto-splitting: Transparent to queries
  • Parallel loading: Chunks load on demand

Entity Registry

High-performance deduplication system for streaming data:

Registry Structure

// __entity_registry__.json
{
 "mappings": {
 "did:plc:alice123": "550e8400-e29b-41d4-a716-446655440000",
 "handle:alice.bsky.social": "550e8400-e29b-41d4-a716-446655440000"
 },
 "stats": {
 "totalMappings": 10000,
 "lastSync": 1699564234567
 }
}

Performance Characteristics

  • Lookup: O(1) in-memory hash map
  • Persistence: Configurable (memory/storage/hybrid)
  • Cache: LRU with configurable TTL
  • Sync: Periodic or on-demand

Ensures durability and enables recovery:

{
 "timestamp": 1699564234567,
 "operation": "add",
 "data": {
 "id": "550e8400-e29b-41d4-a716-446655440000",
 "content": "...",
 "metadata": {}
 },
 "checksum": "sha256:..."
}

Recovery Process

  1. Replay operations from last checkpoint
  2. Verify checksums for integrity

Storage Optimization

1. Lifecycle Policies (Cloud Storage)

Automatic cost optimization through tier transitions:

// S3: Set lifecycle policy for automatic archival
await storage.setLifecyclePolicy({
 rules: [{
 id: 'archive-old-data',
 prefix: 'entities/',
 status: 'Enabled',
 transitions: [
 { days: 30, storageClass: 'STANDARD_IA' }, // Move to IA after 30 days
 { days: 90, storageClass: 'GLACIER' }, // Archive after 90 days
 { days: 365, storageClass: 'DEEP_ARCHIVE' } // Deep archive after 1 year
 ]
 }]
})

// GCS: Set lifecycle policy
await storage.setLifecyclePolicy({
 rules: [{
 condition: { age: 30 },
 action: { type: 'SetStorageClass', storageClass: 'NEARLINE' }
 }, {
 condition: { age: 90 },
 action: { type: 'SetStorageClass', storageClass: 'COLDLINE' }
 }, {
 condition: { age: 365 },
 action: { type: 'SetStorageClass', storageClass: 'ARCHIVE' }
 }]
})

// Azure: Set lifecycle policy
await storage.setLifecyclePolicy({
 rules: [{
 name: 'archiveOldData',
 enabled: true,
 type: 'Lifecycle',
 definition: {
 filters: { blobTypes: ['blockBlob'] },
 actions: {
 baseBlob: {
 tierToCool: { daysAfterModificationGreaterThan: 30 },
 tierToArchive: { daysAfterModificationGreaterThan: 90 }
 }
 }
 }
 }]
})

Cost Impact (500TB dataset):

Storage Before After Savings
AWS S3 $138,000/yr $5,940/yr 96%
GCS $138,000/yr $8,300/yr 94%
Azure $107,520/yr $5,016/yr 95%

2. Intelligent-Tiering (S3)

Automatic optimization without retrieval fees:

// Enable S3 Intelligent-Tiering
await storage.enableIntelligentTiering('entities/', 'auto-optimize')

// Benefits:
// - Automatic tier transitions based on access patterns
// - No retrieval fees (unlike Glacier)
// - Up to 95% cost savings
// - No performance impact on frequently accessed data

3. Autoclass (GCS)

Google Cloud's intelligent automatic optimization:

// Enable GCS Autoclass
await storage.enableAutoclass({
 terminalStorageClass: 'ARCHIVE' // Optional: Set lowest tier
})

// Benefits:
// - Automatic optimization based on access patterns
// - No data retrieval delays
// - Transparent tier transitions
// - Up to 94% cost savings

4. Compression (FileSystem)

// Enable gzip compression for local storage
const brain = new Brainy({
 storage: {
 type: 'filesystem',
 path: './data',
 compression: true // 60-80% space savings
 }
})

// Performance impact:
// - Write: +10-20ms per file (gzip compression)
// - Read: +5-10ms per file (gzip decompression)
// - Space savings: 60-80% for typical JSON data
// - CPU overhead: Minimal (~5% CPU)

5. Batch Operations

// Efficient batch delete
await storage.batchDelete([
 'entities/nouns/vectors/00/00123456-....json',
 'entities/nouns/metadata/00/00123456-....json',
 // ... up to 1000 objects
])

// Benefits:
// - S3: 1000 objects per request (vs 1 per request)
// - GCS: 100 objects per request
// - Azure: 256 objects per batch
// - Automatic retry logic with exponential backoff
// - Throttling protection

// Batch writes for performance
await brain.addBatch([
 { content: "item1", metadata: {} },
 { content: "item2", metadata: {} },
 { content: "item3", metadata: {} }
])
// Single transaction, optimized I/O

6. Quota Monitoring (OPFS)

// Get quota status for browser storage
const status = await storage.getStorageStatus()

console.log(status)
// {
// type: 'opfs',
// available: true,
// details: {
// usage: 45829120, // 43.7 MB used
// quota: 536870912, // 512 MB available
// usagePercent: 8.5,
// quotaExceeded: false
// }
// }

// Proactive quota management:
// - Monitor usage before writes
// - Warn users when approaching quota
// - Automatically clean up old data

7. Tier Management (Azure)

// Change blob tier for cost optimization
await storage.changeBlobTier(blobPath, 'Cool') // Hot → Cool (50% savings)
await storage.changeBlobTier(blobPath, 'Archive') // Cool → Archive (99% savings)

// Batch tier changes (efficient)
await storage.batchChangeTier([blob1, blob2, blob3], 'Cool')

// Rehydrate from Archive when needed
await storage.rehydrateBlob(blobPath, 'Standard') // Standard or High priority

8. Caching Strategy

// Configure caching per storage type
const brain = new Brainy({
 storage: {
 type: 'filesystem',
 cache: {
 enabled: true,
 maxSize: 1000, // Maximum cached items
 ttl: 300000, // 5 minutes
 strategy: 'lru' // Least recently used
 }
 }
})

Concurrent Access

Locking Mechanism

// Automatic locking for write operations
await brain.storage.withLock('resource-id', async () => {
 // Exclusive access to resource
 await brain.storage.saveNoun(id, data)
})

Read-Write Separation

  • Reads: Non-blocking, parallel
  • Writes: Serialized with locks
  • Hybrid: Read-heavy optimization

Migration and Backup

Export Data

// Export the whole brain to a portable PortableGraph document
const graph = await brain.data().then(d => d.export(undefined, { includeVectors: true }))

Import Data

// Restore a PortableGraph document (dedup-by-id merge by default)
await brain.data().then(d => d.import(graph, { onConflict: 'merge' }))

See the Export & Import guide for partial exports (by id, collection, connected neighbourhood, VFS subtree, or predicate).

Storage Migration

// Migrate between storage types
const oldBrain = new Brainy({ storage: { type: 'filesystem' } })
const newBrain = new Brainy({ storage: { type: 's3' } })

await oldBrain.init()
await newBrain.init()

// Transfer all data via a portable graph
const data = await oldBrain.data().then(d => d.export(undefined, { includeVectors: true }))
await newBrain.data().then(d => d.import(data))

Performance Tuning

Storage-Specific Optimizations

FileSystem

  • Directory sharding: Split files across subdirectories
  • Async I/O: Non-blocking file operations
  • Buffer pooling: Reuse buffers for efficiency

S3

  • Multipart uploads: For large objects
  • Request batching: Combine small operations
  • CDN integration: Edge caching for reads

OPFS

  • Quota management: Monitor and request increases
  • Worker offloading: Heavy operations in workers
  • Transaction batching: Group operations

Monitoring

// Get storage statistics
const stats = await brain.storage.getStatistics()
console.log(stats)
// {
// totalSize: 1048576,
// entityCount: 1000,
// indexSize: 204800,
// walSize: 10240,
// cacheHitRate: 0.85
// }

Best Practices

Choose the Right Adapter

  1. Development: FileSystem with compression (local persistence, small storage footprint)
  2. Production Server: FileSystem with compression or cloud storage with lifecycle policies
  3. Browser Apps: OPFS with quota monitoring
  4. Distributed: S3/GCS/Azure with Intelligent-Tiering/Autoclass

Optimize for Your Use Case

  1. Read-heavy: Enable aggressive caching + cloud CDN
  2. Write-heavy: Batch operations + async writes
  3. Real-time: FileSystem with periodic snapshots
  4. Archival: Cloud storage with lifecycle policies (96% cost savings!)
  5. Large-scale: Metadata/vector separation + UUID sharding + lifecycle policies

Cost Optimization

  1. Enable lifecycle policies for cloud storage (automated cost reduction)
  2. Use Intelligent-Tiering (S3) or Autoclass (GCS) for automatic optimization
  3. Enable compression for FileSystem storage (60-80% space savings)
  4. Monitor quota for OPFS (prevent quota exceeded errors)
  5. Use batch operations for bulk deletions (efficient API usage)
  6. Consider tier management for Azure (Hot/Cool/Archive tiers)

Example Cost Savings (500TB dataset):

  • Without lifecycle policies: $138,000/year
  • With lifecycle policies: $5,940/year
  • Savings: $132,060/year (96%)

Monitor and Maintain

  1. Regular statistics collection
  2. Monitor lifecycle policy effectiveness
  3. Index optimization
  4. Cache tuning based on hit rates
  5. Track storage costs and tier distribution
  6. Review quota usage (OPFS) and storage growth patterns

Production Deployment Checklist

  • Enable lifecycle policies on cloud storage
  • Configure batch delete for cleanup operations
  • Enable compression for FileSystem storage
  • Set up quota monitoring for OPFS
  • Configure appropriate tier transitions
  • Enable Intelligent-Tiering (S3) or Autoclass (GCS)
  • Monitor storage costs and optimize regularly

API Reference

See the Storage API for complete method documentation.


Version: 4.0.0 Last Updated: 2025-10-17 Key Features: Metadata/vector separation, UUID sharding, lifecycle management, tier optimization