Aligned every public doc to the 8.0 contract: filesystem + memory adapters
only, vector index provider terminology (config.vector with recall +
quantization + persistMode knobs), no cloud storage adapters, no closed-
source product names.
Tier 1 — heavier rewrites:
- docs/architecture/storage-architecture.md
- docs/architecture/data-storage-architecture.md
- docs/architecture/distributed-storage.md DELETED — content was 100%
cloud-coordination examples with no 8.0 substance.
- docs/guides/distributed-system.md DELETED — same reason; no inbound refs.
- docs/SCALING.md rewritten for single-node guidance.
- docs/PLUGINS.md, docs/augmentations/{COMPLETE-REFERENCE,README}.md:
HnswProvider→VectorIndexProvider, hnsw→vector key.
- docs/PERFORMANCE.md, docs/BATCHING.md cloud-detection + sharding
sections replaced with single-node vector tuning + filesystem framing.
Tier 2 — surgical renames + cloud-section deletions:
- architecture/{index,initialization-and-rebuild,overview}.md
- transactions.md, DEVELOPER_LEARNING_PATH.md
- vfs/{VFS_API_GUIDE,COMMON_PATTERNS}.md
- api/README.md, guides/{inspection,import-flow}.md
Tier 3 — light edits:
- docs/README.md, architecture/augmentation-system-audit.md
MIGRATION-V3-TO-V4.md untouched (internal migration doc, no stale terms).
547 lines
14 KiB
Markdown
547 lines
14 KiB
Markdown
---
|
|
title: Batch Operations
|
|
slug: guides/batching
|
|
public: true
|
|
category: guides
|
|
template: guide
|
|
order: 5
|
|
description: Eliminate N+1 query patterns with batchGet() and storage-level batch APIs for fast multi-entity reads against filesystem and memory storage.
|
|
next:
|
|
- api/reference
|
|
- guides/find-system
|
|
---
|
|
|
|
# Batch Operations API
|
|
> **Production-Ready** | Zero N+1 Query Patterns
|
|
|
|
## Overview
|
|
|
|
Brainy provides batch operations at the storage layer to eliminate N+1 query patterns for VFS operations, relationship queries, and entity retrieval.
|
|
|
|
### Problem Solved
|
|
|
|
The naive pattern of looping and calling `brain.get(id)` once per item issues sequential reads through the storage layer. Batched APIs collapse that into a single read pass.
|
|
|
|
**IMPORTANT:** The batch optimizations apply **ONLY to `getTreeStructure()`** at the VFS layer and the explicit `batchGet()` / `getNounMetadataBatch()` calls — not to `readFile()` or individual `get()` operations.
|
|
|
|
---
|
|
|
|
## New Public APIs
|
|
|
|
### 1. `brain.batchGet(ids, options?)`
|
|
|
|
Batch retrieval of multiple entities (metadata-only by default).
|
|
|
|
```typescript
|
|
// Fetch multiple entities in a single batched operation
|
|
const ids = ['id1', 'id2', 'id3']
|
|
const results: Map<string, Entity> = await brain.batchGet(ids)
|
|
|
|
// With vectors (falls back to individual gets)
|
|
const resultsWithVectors = await brain.batchGet(ids, { includeVectors: true })
|
|
|
|
// Results map
|
|
results.get('id1') // → Entity or undefined
|
|
results.size // → 3 (number of found entities)
|
|
```
|
|
|
|
**Performance:**
|
|
- Memory storage: Instant (parallel reads)
|
|
- Filesystem storage: Parallel reads, scales with available IOPS
|
|
|
|
**Use Cases:**
|
|
- Loading multiple entities for display
|
|
- Bulk data export operations
|
|
- Relationship traversal (fetch all connected entities)
|
|
|
|
---
|
|
|
|
## Storage-Level APIs
|
|
|
|
### 2. `storage.getNounMetadataBatch(ids)`
|
|
|
|
Batch metadata retrieval with direct O(1) path construction.
|
|
|
|
```typescript
|
|
const storage = brain.storage as BaseStorage
|
|
const ids = ['id1', 'id2', 'id3']
|
|
|
|
const metadataMap: Map<string, NounMetadata> = await storage.getNounMetadataBatch(ids)
|
|
|
|
for (const [id, metadata] of metadataMap) {
|
|
console.log(metadata.noun) // Type: 'document', 'person', etc.
|
|
console.log(metadata.data) // Entity data
|
|
}
|
|
```
|
|
|
|
**Features:**
|
|
- ✅ Direct O(1) path construction from ID (no type lookup needed!)
|
|
- ✅ Sharding preservation (all paths include `{shard}/{id}`)
|
|
- ✅ COW-aware (respects branch paths)
|
|
- ✅ 40x faster than v5.x type-first architecture
|
|
|
|
**Performance:**
|
|
- ~1ms per 100 entities (consistent, no cache misses!)
|
|
- Filesystem: parallel reads bounded by IOPS
|
|
- No type search delays — every ID maps directly to storage path
|
|
|
|
---
|
|
|
|
### 3. `storage.getVerbsBySourceBatch(sourceIds, verbType?)`
|
|
|
|
Batch relationship queries by source entity IDs.
|
|
|
|
```typescript
|
|
const storage = brain.storage as BaseStorage
|
|
|
|
// Get all relationships from multiple sources
|
|
const results: Map<string, GraphVerb[]> = await storage.getVerbsBySourceBatch([
|
|
'person1',
|
|
'person2'
|
|
])
|
|
|
|
// Filter by verb type
|
|
const createsResults = await storage.getVerbsBySourceBatch(
|
|
['person1', 'person2'],
|
|
'creates'
|
|
)
|
|
|
|
// Process results
|
|
for (const [sourceId, verbs] of results) {
|
|
console.log(`${sourceId} has ${verbs.length} relationships`)
|
|
verbs.forEach(verb => {
|
|
console.log(` → ${verb.verb} → ${verb.targetId}`)
|
|
})
|
|
}
|
|
```
|
|
|
|
**Use Cases:**
|
|
- Social graph traversal (fetch all connections for multiple users)
|
|
- Knowledge graph queries (find all relationships of specific type)
|
|
- Bulk export of relationship data
|
|
|
|
**Performance:**
|
|
- Memory storage: <10ms for 1000 relationships
|
|
- Filesystem storage: parallel reads through the metadata index
|
|
|
|
---
|
|
|
|
### 4. `storage.readBatchWithInheritance(paths, targetBranch?)`
|
|
|
|
COW-aware batch path resolution with branch inheritance.
|
|
|
|
```typescript
|
|
const storage = brain.storage as BaseStorage
|
|
|
|
const paths = [
|
|
'entities/nouns/{shard}/id1/metadata.json',
|
|
'entities/nouns/{shard}/id2/metadata.json'
|
|
]
|
|
|
|
// Resolves to: branches/{branch}/entities/nouns/{shard}/{id}/metadata.json
|
|
const results: Map<string, any> = await storage.readBatchWithInheritance(paths, 'my-branch')
|
|
|
|
// Automatically inherits from parent branches for missing entities
|
|
```
|
|
|
|
**Features:**
|
|
- ✅ Branch path resolution (`branches/{branch}/...`)
|
|
- ✅ Write cache integration (read-after-write consistency)
|
|
- ✅ COW inheritance (fallback to parent commits for missing entities)
|
|
- ✅ Adapter-agnostic (works with all storage adapters)
|
|
|
|
---
|
|
|
|
## VFS Integration
|
|
|
|
VFS operations automatically use batch APIs for maximum performance.
|
|
|
|
### Directory Traversal
|
|
|
|
```typescript
|
|
// Tree traversal uses batched reads under the hood
|
|
const tree = await brain.vfs.getTreeStructure('/my-dir')
|
|
// ✅ PathResolver.getChildren() uses brain.batchGet() internally
|
|
// ✅ Parallel traversal of directories at the same tree level
|
|
// ✅ 2-3 batched calls instead of 22 sequential calls
|
|
```
|
|
|
|
**Architecture:**
|
|
|
|
```
|
|
VFS.getTreeStructure()
|
|
↓ PARALLEL (breadth-first traversal)
|
|
→ PathResolver.getChildren() [all dirs at level processed in parallel]
|
|
↓ BATCHED
|
|
→ brain.batchGet(childIds) [1 call instead of N]
|
|
↓ BATCHED
|
|
→ storage.getNounMetadataBatch(ids) [1 call instead of N]
|
|
↓ ADAPTER
|
|
→ Filesystem: Promise.all() parallel reads
|
|
→ Memory: Promise.all() parallel reads
|
|
```
|
|
|
|
---
|
|
|
|
## Advanced Features Compatibility
|
|
|
|
### ✅ ID-First Storage Architecture
|
|
|
|
All batch operations use direct ID-first paths - no type lookup needed!
|
|
|
|
**ID-First Path Structure:**
|
|
```
|
|
entities/nouns/{SHARD}/{ID}/metadata.json
|
|
entities/verbs/{SHARD}/{ID}/metadata.json
|
|
```
|
|
|
|
**Direct O(1) Path Construction:**
|
|
```typescript
|
|
// Every ID maps directly to exactly ONE path - 40x faster!
|
|
const id = 'abc-123'
|
|
const shard = getShardIdFromUuid(id) // → 'ab' (first 2 hex chars)
|
|
const path = `entities/nouns/${shard}/${id}/metadata.json`
|
|
|
|
// No type cache needed!
|
|
// No type search needed!
|
|
// No multi-type fallback needed!
|
|
// Just pure O(1) lookup!
|
|
```
|
|
|
|
**Benefits:**
|
|
- **40x faster** path lookups (eliminates 42-type sequential search)
|
|
- **Simpler code** - removed 500+ lines of type cache complexity
|
|
- **Scalable** - works at billion-scale without type tracking overhead
|
|
|
|
---
|
|
|
|
### ✅ Sharding
|
|
|
|
All batch paths include shard IDs calculated via `getShardIdFromUuid(id)`:
|
|
|
|
```typescript
|
|
const id = 'a3c4e5f7-...'
|
|
const shard = getShardIdFromUuid(id) // → 'a3' (first 2 hex chars)
|
|
const path = `entities/nouns/${shard}/${id}/metadata.json`
|
|
```
|
|
|
|
**Distribution:** 256 shards (00-ff) for optimal load distribution.
|
|
|
|
---
|
|
|
|
### ✅ COW (Copy-on-Write)
|
|
|
|
Batch operations respect branch isolation and time-travel:
|
|
|
|
```typescript
|
|
// Main branch
|
|
const brain = await Brainy.create({ enableCOW: true })
|
|
await brain.add({ type: 'document', data: 'Main' })
|
|
|
|
// Create fork
|
|
const fork = await brain.fork('experiment')
|
|
|
|
// Batch operations are isolated
|
|
await brain.batchGet([id1, id2]) // → Reads from: branches/main/...
|
|
await fork.batchGet([id1, id2]) // → Reads from: branches/experiment/...
|
|
```
|
|
|
|
**Inheritance:**
|
|
- Entities missing from child branch automatically inherit from parent commits
|
|
- `readBatchWithInheritance()` walks commit history for missing items
|
|
- Preserves fork semantics while maintaining performance
|
|
|
|
---
|
|
|
|
### ✅ fork() and checkout()
|
|
|
|
```typescript
|
|
const fork = await brain.fork('my-branch')
|
|
await fork.add({ type: 'document', data: 'Fork entity' })
|
|
|
|
// Batch operations use correct branch
|
|
const results = await fork.batchGet([id1, id2])
|
|
// → Reads from: branches/my-branch/...
|
|
|
|
// Checkout changes active branch
|
|
await fork.checkout('main')
|
|
const mainResults = await fork.batchGet([id1, id2])
|
|
// → Reads from: branches/main/...
|
|
```
|
|
|
|
---
|
|
|
|
### ✅ asOf() Time-Travel
|
|
|
|
```typescript
|
|
// Create historical snapshot
|
|
await brain.commit('v1.0')
|
|
const snapshot = await brain.asOf('v1.0')
|
|
|
|
// Batch operations on historical data
|
|
const results = await snapshot.batchGet([id1, id2])
|
|
// → Reads from historical tree state
|
|
```
|
|
|
|
Historical queries use `HistoricalStorageAdapter` which wraps batch operations to point at specific commits.
|
|
|
|
---
|
|
|
|
## Performance Benchmarks
|
|
|
|
### VFS Operations (12 Files)
|
|
|
|
| Storage | Before optimization | After optimization | Improvement |
|
|
|---------|---------------------|--------------------|-------------|
|
|
| **Memory** | 150ms | 50ms | **67% faster** |
|
|
| **Filesystem** | ~500ms | ~80ms | **84% faster** |
|
|
|
|
### Entity Batch Retrieval (100 Entities)
|
|
|
|
| Storage | Individual Gets | Batch Get | Improvement |
|
|
|---------|-----------------|-----------|-------------|
|
|
| **Memory** | 180ms | 15ms | **92% faster** |
|
|
| **Filesystem** | ~1.2s | ~120ms | **90% faster** |
|
|
|
|
### Throughput (Entities/Second)
|
|
|
|
| Storage | Individual | Batch | Improvement |
|
|
|---------|------------|-------|-------------|
|
|
| **Memory** | 556 ent/s | 6667 ent/s | **12x** |
|
|
| **Filesystem** | ~80 ent/s | ~800 ent/s | **10x** |
|
|
|
|
---
|
|
|
|
## Error Handling
|
|
|
|
### Partial Batch Failures
|
|
|
|
Batch operations gracefully handle missing or invalid entities:
|
|
|
|
```typescript
|
|
const validId = 'abc-123-...'
|
|
const invalidIds = [
|
|
'11111111-1111-1111-1111-111111111111',
|
|
'22222222-2222-2222-2222-222222222222'
|
|
]
|
|
|
|
const results = await brain.batchGet([validId, ...invalidIds])
|
|
|
|
results.size // → 1 (only valid entity)
|
|
results.has(validId) // → true
|
|
results.has(invalidIds[0]) // → false (silently skipped)
|
|
```
|
|
|
|
**Behavior:**
|
|
- Invalid UUIDs: Silently skipped (not included in results)
|
|
- Missing entities: Silently skipped (not included in results)
|
|
- Storage errors: Logged, entity excluded from results
|
|
- No exceptions thrown for partial failures
|
|
|
|
### Empty Batches
|
|
|
|
```typescript
|
|
const results = await brain.batchGet([])
|
|
results.size // → 0 (empty map)
|
|
```
|
|
|
|
### Duplicate IDs
|
|
|
|
```typescript
|
|
const results = await brain.batchGet(['id1', 'id1', 'id1'])
|
|
results.size // → 1 (deduplicated automatically)
|
|
```
|
|
|
|
---
|
|
|
|
## Migration Guide
|
|
|
|
### From Individual Gets
|
|
|
|
**Before:**
|
|
```typescript
|
|
const entities = []
|
|
for (const id of ids) {
|
|
const entity = await brain.get(id)
|
|
if (entity) entities.push(entity)
|
|
}
|
|
```
|
|
|
|
**After:**
|
|
```typescript
|
|
const results = await brain.batchGet(ids)
|
|
const entities = Array.from(results.values())
|
|
```
|
|
|
|
**Performance Gain:** 10-20x faster on filesystem storage.
|
|
|
|
---
|
|
|
|
### From Individual Relationship Queries
|
|
|
|
**Before:**
|
|
```typescript
|
|
const allVerbs = []
|
|
for (const sourceId of sourceIds) {
|
|
const verbs = await brain.getRelations({ from: sourceId })
|
|
allVerbs.push(...verbs)
|
|
}
|
|
```
|
|
|
|
**After:**
|
|
```typescript
|
|
const storage = brain.storage as BaseStorage
|
|
const results = await storage.getVerbsBySourceBatch(sourceIds)
|
|
|
|
const allVerbs = []
|
|
for (const verbs of results.values()) {
|
|
allVerbs.push(...verbs)
|
|
}
|
|
```
|
|
|
|
**Performance Gain:** 5-10x faster due to batched metadata fetches.
|
|
|
|
---
|
|
|
|
## Best Practices
|
|
|
|
### 1. **Use Batching for Multiple Entity Operations**
|
|
|
|
```typescript
|
|
// ✅ GOOD: Batch fetch
|
|
const results = await brain.batchGet(ids)
|
|
|
|
// ❌ BAD: Individual gets in loop
|
|
for (const id of ids) {
|
|
await brain.get(id)
|
|
}
|
|
```
|
|
|
|
### 2. **Batch Size Recommendations**
|
|
|
|
| Storage | Optimal Batch Size | Max Batch Size |
|
|
|---------|--------------------|----------------|
|
|
| **Memory** | Unlimited | Unlimited |
|
|
| **Filesystem** | 100-500 | 1000 |
|
|
|
|
**Guideline:** For batches >1000, split into chunks of 500-1000.
|
|
|
|
### 3. **Metadata-Only by Default**
|
|
|
|
```typescript
|
|
// Default: Metadata-only (fast)
|
|
const results = await brain.batchGet(ids) // No vectors
|
|
|
|
// Only load vectors if needed
|
|
const withVectors = await brain.batchGet(ids, { includeVectors: true })
|
|
```
|
|
|
|
### 4. **Error Handling**
|
|
|
|
```typescript
|
|
// Batch operations never throw for missing entities
|
|
const results = await brain.batchGet(ids)
|
|
|
|
// Check results
|
|
for (const id of ids) {
|
|
if (results.has(id)) {
|
|
// Entity exists
|
|
const entity = results.get(id)
|
|
} else {
|
|
// Entity missing (not an error)
|
|
console.log(`Entity ${id} not found`)
|
|
}
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Testing
|
|
|
|
Comprehensive test coverage in `tests/integration/storage-batch-operations.test.ts`:
|
|
|
|
```bash
|
|
npx vitest run tests/integration/storage-batch-operations.test.ts
|
|
```
|
|
|
|
**Test Coverage:**
|
|
- ✅ brain.batchGet() high-level API
|
|
- ✅ storage.getNounMetadataBatch() with ID-first paths
|
|
- ✅ COW integration (branch isolation, inheritance)
|
|
- ✅ storage.getVerbsBySourceBatch() relationship queries
|
|
- ✅ VFS integration (PathResolver.getChildren())
|
|
- ✅ Performance benchmarks (N+1 elimination)
|
|
- ✅ Error handling (partial failures, empty batches, duplicates)
|
|
- ✅ ID-first storage verification
|
|
- ✅ Sharding preservation
|
|
|
|
**Results:** 23 tests passing ✅
|
|
|
|
---
|
|
|
|
## Implementation Details
|
|
|
|
### Architecture Layers
|
|
|
|
```
|
|
User Code (brain.batchGet)
|
|
↓
|
|
High-Level API (src/brainy.ts)
|
|
↓
|
|
Storage Layer (src/storage/baseStorage.ts)
|
|
↓
|
|
COW Layer (readBatchWithInheritance)
|
|
↓
|
|
Adapter Layer (readBatchFromAdapter)
|
|
↓
|
|
Storage Adapter (FileSystemStorage / MemoryStorage)
|
|
```
|
|
|
|
### Parallel Reads
|
|
|
|
Both shipped adapters fall back to `Promise.all` over individual reads:
|
|
|
|
```typescript
|
|
// BaseStorage.readBatchFromAdapter()
|
|
return await Promise.all(resolvedPaths.map(path => this.read(path)))
|
|
```
|
|
|
|
**Shipped Adapters:**
|
|
- MemoryStorage
|
|
- FileSystemStorage
|
|
- HistoricalStorageAdapter (delegates to underlying)
|
|
|
|
---
|
|
|
|
## API Summary
|
|
|
|
- `brain.batchGet(ids, options?)` - High-level batch entity retrieval
|
|
- `storage.getNounMetadataBatch(ids)` - Storage-level metadata batch
|
|
- `storage.getVerbsBySourceBatch(sourceIds, verbType?)` - Batch relationship queries
|
|
- `storage.readBatchWithInheritance(paths, targetBranch?)` - COW-aware batch reads
|
|
|
|
**Performance Improvements:**
|
|
- VFS operations: 90%+ faster than the naive per-entity loop
|
|
- Entity retrieval: 10-20x throughput improvement
|
|
- Zero N+1 query patterns
|
|
|
|
**Compatibility:**
|
|
- ✅ ID-first storage
|
|
- ✅ Sharding (256 shards)
|
|
- ✅ COW (branch isolation, inheritance)
|
|
- ✅ fork() and checkout()
|
|
- ✅ asOf() time-travel
|
|
- ✅ All indexes respected (vector, type-aware vector, metadata, graph adjacency, version, deleted items)
|
|
|
|
---
|
|
|
|
## Support
|
|
|
|
- **Documentation:** `/docs/BATCHING.md`, `/docs/PERFORMANCE.md`
|
|
- **Tests:** `/tests/integration/storage-batch-operations.test.ts`
|
|
- **Issues:** https://github.com/soulcraft/brainy/issues
|
|
- **Discussions:** https://github.com/soulcraft/brainy/discussions
|
|
|
|
---
|
|
|
|
**Built with ❤️ for enterprise-scale knowledge graphs**
|