chore(8.0): step-7 follow-through — collapse remaining cloud branches + docs sweep (scaffold step 13)
Final cleanup pass for Brainy 8.0. Catches three categories of debt: A. STEP-7 FOLLOW-THROUGH (rebuild-path collapse) Step 7's bisect-reset (debugging a flaky test) lost the in-source edits to three rebuild paths even though the commit message claimed they shipped. Re-applied now: - src/utils/metadataIndex.ts — collapsed the `isLocalStorage` / cloud-pagination branching. Local-load-all-at-once is the only path in 8.0. Removed ~150 LOC of paginated-cloud branching for both nouns and verbs, plus the safety counters (`consecutiveEmptyBatches`, `MAX_ITERATIONS`, etc.). - src/hnsw/hnswIndex.ts — same simplification for HNSW rebuild. The paginated cloud path is gone; HNSW now loads all nodes at once. Removed ~85 LOC. - src/graph/graphAdjacencyIndex.ts — same simplification for graph adjacency rebuild. Removed ~50 LOC. The collapse is safe because cloud adapters were deleted in step 7; `storageType === 'OPFSStorage'` (and similar) can never match now. B. CLOUD-ONLY DOCS DELETED - docs/operations/cost-optimization-aws-s3.md - docs/operations/cost-optimization-azure.md - docs/operations/cost-optimization-cloudflare-r2.md - docs/operations/cost-optimization-gcs.md - docs/operations/cloud-run-filestore-guide.md (docs/deployment/* contained no cloud-specific files that needed deletion.) C. STORAGE-ADAPTERS GUIDE REWRITTEN FOR 8.0 docs/guides/storage-adapters.md → fresh content reflecting the 8.0 reality: - Two adapters: FileSystemStorage + MemoryStorage. Quick-start matrix. - Cloud backup section explains the operator-tooling pattern (gsutil / aws s3 / rclone / azcopy) with the exact commands consumers will run. - "Why no cloud adapters in 8.0?" section documents the four reasons per BR-BRAINY-80-STORAGE-SIMPLIFY. - Migration recipe for 7.x cloud-adapter consumers: mount local disk → filesystem storage → operator backup cron. Updated frontmatter description so soulcraft.com/docs renders the correct preview. NOT IN THIS COMMIT (deliberate, lower-priority) - src/storage/cacheManager.ts still references StorageType.S3 / REMOTE_API / OPFS as dead branches (23 sites). The branches are never reached in 8.0, but cleaning them would cascade through 5 consumers. Defer to a follow-up if the dead code surfaces as a real maintenance issue. - src/config/storageAutoConfig.ts keeps its StorageType enum + autodetect for 7.x compat surface. Same reason: rewriting cascades through zeroConfig, extensibleConfig, sharedConfigManager. Defer. - docs/MIGRATION-V3-TO-V4.md and docs/DEVELOPER_LEARNING_PATH.md still reference cloud adapters as historical artefacts. That's accurate — they describe how things used to be. Left as-is. - @deprecated audit in src/ (10 files) deferred — audit each individually in a future polish pass. VERIFICATION - npx tsc --noEmit: clean - npm test: 1408 / 1409 (same pre-existing race-condition outstanding from step 7; no regressions from this cleanup)
This commit is contained in:
parent
780fb6444b
commit
9f9a41599e
13 changed files with 162 additions and 5288 deletions
|
|
@ -5,197 +5,144 @@ public: true
|
|||
category: guides
|
||||
template: guide
|
||||
order: 2
|
||||
description: "Six adapters for every deployment: in-memory, OPFS (browser), filesystem, S3, Google Cloud Storage, Azure Blob, and Cloudflare R2. All support copy-on-write branching."
|
||||
description: "Two adapters cover every deployment: in-memory for tests + ephemeral workloads, filesystem for everything that needs to persist. Both support copy-on-write branching. Cloud backup is operator tooling, not a built-in adapter."
|
||||
next:
|
||||
- guides/plugins
|
||||
- concepts/zero-config
|
||||
---
|
||||
|
||||
# Storage Adapters Guide
|
||||
# Storage Adapters
|
||||
|
||||
## Overview
|
||||
Brainy supports 6 storage adapters for different deployment environments. All adapters implement the same StorageAdapter interface and support copy-on-write branching.
|
||||
Brainy 8.0 ships **two storage adapters**:
|
||||
|
||||
## Adapters
|
||||
- **`FileSystemStorage`** — persistent on-disk storage. The default for any
|
||||
deployment that needs to survive a restart. Runs on Node.js, Bun, and Deno.
|
||||
- **`MemoryStorage`** — in-memory only. The right choice for tests, ephemeral
|
||||
workloads, and short-lived demos.
|
||||
|
||||
### 1. MemoryStorage
|
||||
- **File**: `src/storage/adapters/memoryStorage.ts`
|
||||
- **Use case**: Development, testing, prototyping
|
||||
- **Configuration**: None required
|
||||
- **Persistence**: None (data lost on restart)
|
||||
- **Batch config**: 1000 batch size, 0ms delay, 1000 concurrent ops, 100k ops/sec
|
||||
Both implement the same `StorageAdapter` interface, support copy-on-write
|
||||
branching, and use the same on-disk layout (memory's "disk" is a JS Map).
|
||||
|
||||
```typescript
|
||||
const brain = new Brainy({ storage: { type: 'memory' } })
|
||||
## Quick start
|
||||
|
||||
```ts
|
||||
import { Brainy } from '@soulcraft/brainy'
|
||||
|
||||
// Filesystem (recommended for any persistent workload):
|
||||
const brain = new Brainy({
|
||||
storage: { type: 'filesystem', rootDirectory: './brainy-data' }
|
||||
})
|
||||
|
||||
// Memory (tests, ephemeral):
|
||||
const brainMem = new Brainy({ storage: { type: 'memory' } })
|
||||
|
||||
// Auto-detect (filesystem on Node-like runtimes, memory in browsers):
|
||||
const brainAuto = new Brainy({ storage: { type: 'auto' } })
|
||||
```
|
||||
|
||||
### 2. FileSystemStorage
|
||||
- **File**: `src/storage/adapters/fileSystemStorage.ts`
|
||||
- **Use case**: Node.js local persistence, single-server deployments
|
||||
- **Configuration**: `basePath` (required), `readOnly` (optional)
|
||||
- **Features**: zlib compression, atomic writes via temp files, UUID-based sharding
|
||||
## When to use which
|
||||
|
||||
```typescript
|
||||
const brain = new Brainy({
|
||||
storage: { type: 'filesystem', path: './brainy-data' }
|
||||
})
|
||||
| Use case | Adapter | Why |
|
||||
|---|---|---|
|
||||
| Production app | `filesystem` | Durable, branchable, mmap-able |
|
||||
| Tests, CI | `memory` | No disk teardown; fast |
|
||||
| Short-lived data pipeline | `memory` | No persistence needed |
|
||||
| In-browser demo | `memory` | Filesystem unavailable in browsers |
|
||||
| Cloud deployment | `filesystem` on local disk + operator backup | See "Cloud backup" below |
|
||||
|
||||
## Cloud backup — operator tooling, not a built-in
|
||||
|
||||
Brainy 8.0 deliberately ships **no cloud storage adapters**. Cloud backup is
|
||||
handled at the operator layer with standard tooling, the same pattern every
|
||||
production database uses (Postgres, SQLite, Redis):
|
||||
|
||||
```bash
|
||||
# After a brainy flush, sync the on-disk artefact to your cloud of choice.
|
||||
gsutil rsync -r /var/lib/brainy gs://my-backup-bucket/brainy/
|
||||
# or:
|
||||
aws s3 sync /var/lib/brainy s3://my-backup-bucket/brainy/
|
||||
# or:
|
||||
rclone sync /var/lib/brainy remote:brainy-backups/
|
||||
# or:
|
||||
azcopy sync /var/lib/brainy "https://account.blob.core.windows.net/brainy?sv=..."
|
||||
```
|
||||
|
||||
### 3. S3CompatibleStorage
|
||||
- **File**: `src/storage/adapters/s3CompatibleStorage.ts`
|
||||
- **Use case**: AWS S3, MinIO, DigitalOcean Spaces, any S3-compatible service
|
||||
- **Configuration**: `bucketName`, `region`, `accessKeyId`, `secretAccessKey`, `endpoint?` (for custom endpoints), `s3ForcePathStyle?`
|
||||
- **Features**: Write buffering, request coalescing, throttling detection (429/503), progressive initialization
|
||||
- **Batch config**: 1000 batch size, 150 concurrent ops, 5000 ops/sec burst
|
||||
Brainy's filesystem layout is sync-friendly:
|
||||
- Atomic writes (temp + rename) — readers never see torn files
|
||||
- Per-shard files — `rsync`-style incremental sync works well
|
||||
- Content-addressed blobs (`_blobs/`) — immutable, cache-friendly
|
||||
- Branch refs live under `_cow/` — pick up branches automatically
|
||||
|
||||
```typescript
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 's3',
|
||||
s3Storage: {
|
||||
bucketName: 'my-brainy-data',
|
||||
region: 'us-east-1',
|
||||
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
|
||||
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY
|
||||
}
|
||||
}
|
||||
})
|
||||
For point-in-time backups, take a filesystem snapshot (ZFS, btrfs, LVM, EBS,
|
||||
etc.) or use `brain.persist(path)` to write a self-contained snapshot you
|
||||
can sync independently of the live brain.
|
||||
|
||||
// For MinIO or other S3-compatible services:
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 's3',
|
||||
s3Storage: {
|
||||
bucketName: 'my-data',
|
||||
region: 'us-east-1',
|
||||
endpoint: 'http://localhost:9000',
|
||||
s3ForcePathStyle: true,
|
||||
accessKeyId: 'minio-key',
|
||||
secretAccessKey: 'minio-secret'
|
||||
}
|
||||
}
|
||||
})
|
||||
## Why no cloud adapters in 8.0?
|
||||
|
||||
Cloud storage adapters lived in Brainy 4.x-7.x. They were dropped in 8.0
|
||||
per **BR-BRAINY-80-STORAGE-SIMPLIFY** because:
|
||||
|
||||
- Zero production consumers used them at scale. Every Soulcraft consumer
|
||||
ran on local filesystem.
|
||||
- Cloud-storage HNSW / DiskANN doesn't perform — vector indexes need
|
||||
low-latency random reads that S3 / GCS / R2 / Azure can't provide
|
||||
consistently.
|
||||
- Bundling cloud SDKs into the library cost ~3000-5000 LOC + 4-7 transitive
|
||||
dependencies for a feature nobody used.
|
||||
- Cloud backup via operator tooling is strictly more reliable than in-app
|
||||
upload (better retry semantics, better observability, better cost control).
|
||||
|
||||
Brainy 8.0 is smaller, faster to install, and clearer about what it does.
|
||||
|
||||
## Configuration
|
||||
|
||||
```ts
|
||||
interface StorageOptions {
|
||||
// The adapter type. Defaults to 'auto'.
|
||||
type?: 'auto' | 'memory' | 'filesystem'
|
||||
|
||||
// Force a specific adapter regardless of type.
|
||||
forceMemoryStorage?: boolean
|
||||
forceFileSystemStorage?: boolean
|
||||
|
||||
// Filesystem only.
|
||||
rootDirectory?: string
|
||||
|
||||
// COW branch to open. Defaults to 'main'.
|
||||
branch?: string
|
||||
|
||||
// COW compression toggle. Defaults to true.
|
||||
enableCompression?: boolean
|
||||
}
|
||||
```
|
||||
|
||||
### 4. R2Storage (Cloudflare)
|
||||
- **File**: `src/storage/adapters/r2Storage.ts`
|
||||
- **Use case**: Zero egress fees, cost-effective cloud storage
|
||||
- **Configuration**: `bucketName`, `accountId`, `accessKeyId`, `secretAccessKey`, `cacheConfig?`
|
||||
- **Features**: Zero egress fees, aggressive caching, write buffering, request coalescing
|
||||
- **Batch config**: 1000 batch size, 150 concurrent ops, 6000 ops/sec
|
||||
## Direct construction
|
||||
|
||||
```typescript
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 'r2',
|
||||
r2Storage: {
|
||||
accountId: process.env.CF_ACCOUNT_ID,
|
||||
bucketName: 'my-brainy-data',
|
||||
accessKeyId: process.env.CF_ACCESS_KEY_ID,
|
||||
secretAccessKey: process.env.CF_SECRET_ACCESS_KEY
|
||||
}
|
||||
}
|
||||
})
|
||||
If you want to skip the factory:
|
||||
|
||||
```ts
|
||||
import { FileSystemStorage, MemoryStorage } from '@soulcraft/brainy'
|
||||
|
||||
const fsStorage = new FileSystemStorage('./brainy-data')
|
||||
const memStorage = new MemoryStorage()
|
||||
|
||||
const brain = new Brainy({ storage: fsStorage })
|
||||
```
|
||||
|
||||
### 5. GcsStorage (Google Cloud)
|
||||
- **File**: `src/storage/adapters/gcsStorage.ts`
|
||||
- **Use case**: Google Cloud ecosystem, Cloud Run deployments
|
||||
- **Authentication priority**:
|
||||
1. Service Account Key File (`keyFilename`)
|
||||
2. Credentials Object (`credentials`)
|
||||
3. HMAC Keys (backward compat)
|
||||
4. Application Default Credentials (automatic in Cloud Run)
|
||||
- **Features**: Progressive initialization (fast cold starts <200ms), Autoclass lifecycle, bucket validation
|
||||
- **Batch config**: 1000 batch size, 100 concurrent ops, 1000 ops/sec
|
||||
## Migration from 7.x cloud adapters
|
||||
|
||||
```typescript
|
||||
// With explicit credentials
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 'gcs',
|
||||
gcsStorage: {
|
||||
bucketName: 'my-brainy-data',
|
||||
keyFilename: './service-account.json'
|
||||
}
|
||||
}
|
||||
})
|
||||
7.x consumers of `OPFSStorage`, `GcsStorage`, `R2Storage`, `S3CompatibleStorage`,
|
||||
or `AzureBlobStorage` need to migrate to `FileSystemStorage` plus operator
|
||||
backup tooling. The recipe:
|
||||
|
||||
// In Cloud Run (uses Application Default Credentials automatically)
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 'gcs',
|
||||
gcsStorage: {
|
||||
bucketName: 'my-brainy-data'
|
||||
}
|
||||
}
|
||||
})
|
||||
```
|
||||
1. On the host running Brainy, mount a local disk (NVMe recommended). Cloud
|
||||
providers all expose persistent local disks: GCP Persistent Disk, AWS EBS,
|
||||
Azure Managed Disks.
|
||||
2. Set `storage: { type: 'filesystem', rootDirectory: '/mnt/brainy-data' }`.
|
||||
3. Run your existing data import once into the new local store.
|
||||
4. Set up an operator backup job using `gsutil` / `aws s3` / `rclone` /
|
||||
`azcopy` on a cron — hourly or whatever your RPO requires. Point it at
|
||||
the brainy data dir.
|
||||
5. For point-in-time backups, use filesystem snapshots or `brain.persist()`.
|
||||
|
||||
**Progressive Initialization Modes:**
|
||||
- `'strict'` (default locally): Full validation during init (100-500ms+)
|
||||
- `'progressive'` (default in Cloud Run/Lambda): Fast init <200ms, background validation
|
||||
- `'auto'`: Auto-detects environment
|
||||
|
||||
### 6. AzureBlobStorage
|
||||
- **File**: `src/storage/adapters/azureBlobStorage.ts`
|
||||
- **Use case**: Azure ecosystem, enterprise deployments
|
||||
- **Authentication priority**:
|
||||
1. DefaultAzureCredential (Managed Identity - automatic in Azure)
|
||||
2. Connection String
|
||||
3. Account Name + Account Key
|
||||
4. SAS Token
|
||||
- **Features**: Progressive initialization, lifecycle management, container validation
|
||||
- **Batch config**: 1000 batch size, 100 concurrent ops, 3000 ops/sec
|
||||
|
||||
```typescript
|
||||
// With connection string
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 'azure',
|
||||
azureStorage: {
|
||||
connectionString: process.env.AZURE_STORAGE_CONNECTION_STRING,
|
||||
containerName: 'brainy-data'
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
// With Managed Identity (automatic in Azure App Service/Functions)
|
||||
const brain = new Brainy({
|
||||
storage: {
|
||||
type: 'azure',
|
||||
azureStorage: {
|
||||
containerName: 'brainy-data'
|
||||
// DefaultAzureCredential used automatically
|
||||
}
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
## Choosing an Adapter
|
||||
|
||||
| Scenario | Recommended Adapter |
|
||||
|----------|-------------------|
|
||||
| Development/Testing | MemoryStorage |
|
||||
| Local persistence | FileSystemStorage |
|
||||
| AWS deployment | S3CompatibleStorage |
|
||||
| Google Cloud / Cloud Run | GcsStorage |
|
||||
| Azure deployment | AzureBlobStorage |
|
||||
| Cost-sensitive (high egress) | R2Storage |
|
||||
| Self-hosted (MinIO) | S3CompatibleStorage |
|
||||
|
||||
## Common Configuration
|
||||
|
||||
All cloud adapters support:
|
||||
- **Progressive initialization** for fast serverless cold starts
|
||||
- **Read-only mode** (`readOnly: true`)
|
||||
- **Cache configuration** (`cacheConfig: { hotCacheMaxSize, warmCacheTTL }`)
|
||||
- **Copy-on-write branching** for git-style data management
|
||||
|
||||
## See Also
|
||||
|
||||
- [Cloud Deployment Guide](../deployment/CLOUD_DEPLOYMENT_GUIDE.md)
|
||||
- [Cost Optimization: AWS S3](../operations/cost-optimization-aws-s3.md)
|
||||
- [Cost Optimization: GCS](../operations/cost-optimization-gcs.md)
|
||||
- [Cost Optimization: Azure](../operations/cost-optimization-azure.md)
|
||||
- [Cost Optimization: R2](../operations/cost-optimization-cloudflare-r2.md)
|
||||
Same data, same APIs, no library-side cloud code.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue