brainy/docs/guides/storage-adapters.md
David Snelling 606445cd61 feat(8.0): API simplification — remove neural()/Db.search, one storage path key, integration→0
8.0 RC cleanup toward "one place per thing, zero-config, no deprecation":

- Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead
  legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})`
  / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate
  entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor,
  NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept.
- Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams).
  Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` →
  now `find({ query, limit })`.
- Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases
  (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the
  exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver
  feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native
  storage provider resolves the identical root (no split-brain).
- Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now
  applied as a post-filter on `result.score` (the documented way to bound semantic results).
- Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()`
  and failed dimension validation; they are metadata-only updates now.
- Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination
  shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an
  in-place path change that preserves the blob and the entity id, for files and directories.
- Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability.
  Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk
  flushes and emits `progress.queryable`, so imported data is queryable during the import.
- Sweep all docs, comments, and JSDoc for the removed/changed APIs.

Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
2026-06-20 13:31:11 -07:00

6.1 KiB

title slug public category template order description next
Storage Adapters guides/storage-adapters true guides guide 2 Two adapters cover every deployment: in-memory for tests + ephemeral workloads, filesystem for everything that needs to persist. Both share one on-disk contract, including generational history and snapshots. Cloud backup is operator tooling, not a built-in adapter.
guides/plugins
concepts/consistency-model

Storage Adapters

Brainy 8.0 ships two storage adapters:

  • FileSystemStorage — persistent on-disk storage. The default for any deployment that needs to survive a restart. Runs on Node.js, Bun, and Deno.
  • MemoryStorage — in-memory only. The right choice for tests, ephemeral workloads, and short-lived demos.

Both implement the same StorageAdapter interface, support the full Db API (generational history, snapshots, restore — see the consistency model), and use the same on-disk layout (memory's "disk" is a JS Map).

Quick start

import { Brainy } from '@soulcraft/brainy'

// Filesystem (recommended for any persistent workload):
const brain = new Brainy({
  storage: { type: 'filesystem', path: './brainy-data' }
})

// Memory (tests, ephemeral):
const brainMem = new Brainy({ storage: { type: 'memory' } })

// Auto-detect (filesystem on Node-like runtimes, memory in browsers):
const brainAuto = new Brainy({ storage: { type: 'auto' } })

When to use which

Use case Adapter Why
Production app filesystem Durable, snapshot-able, mmap-able
Tests, CI memory No disk teardown; fast
Short-lived data pipeline memory No persistence needed
In-browser demo memory Filesystem unavailable in browsers
Cloud deployment filesystem on local disk + operator backup See "Cloud backup" below

Cloud backup — operator tooling, not a built-in

Brainy 8.0 deliberately ships no cloud storage adapters. Cloud backup is handled at the operator layer with standard tooling, the same pattern every production database uses (Postgres, SQLite, Redis):

# After a brainy flush, sync the on-disk artefact to your cloud of choice.
gsutil rsync -r /var/lib/brainy gs://my-backup-bucket/brainy/
# or:
aws s3 sync /var/lib/brainy s3://my-backup-bucket/brainy/
# or:
rclone sync /var/lib/brainy remote:brainy-backups/
# or:
azcopy sync /var/lib/brainy "https://account.blob.core.windows.net/brainy?sv=..."

Brainy's filesystem layout is sync-friendly:

  • Atomic writes (temp + rename) — readers never see torn files
  • Per-shard files — rsync-style incremental sync works well
  • Immutable generation records (_generations/) — append-only, cache-friendly

For point-in-time backups, take a filesystem snapshot (ZFS, btrfs, LVM, EBS, etc.) or use brain.now().persist(path) to write a self-contained snapshot you can sync independently of the live brain — see Snapshots & Time Travel.

Why no cloud adapters in 8.0?

Cloud storage adapters lived in Brainy 4.x-7.x. They were dropped in 8.0 because:

  • Zero production consumers used them at scale — every known production deployment ran on local filesystem.
  • Cloud-storage HNSW / DiskANN doesn't perform — vector indexes need low-latency random reads that S3 / GCS / R2 / Azure can't provide consistently.
  • Bundling cloud SDKs into the library cost ~3000-5000 LOC + 4-7 transitive dependencies for a feature nobody used.
  • Cloud backup via operator tooling is strictly more reliable than in-app upload (better retry semantics, better observability, better cost control).

Brainy 8.0 is smaller, faster to install, and clearer about what it does.

Configuration

// BrainyConfig['storage'] — either a config object or a pre-constructed adapter:
storage?:
  | {
      // The adapter type. Optional — a top-level `path` implies 'filesystem',
      // so { path: '/data' } works without it. Defaults to 'auto'
      // (filesystem on Node-like runtimes, memory otherwise).
      type?: 'auto' | 'memory' | 'filesystem'

      // CANONICAL directory for filesystem storage. The rest of the API already
      // speaks `path` (persist(path), Brainy.load(path), asOf(path),
      // restore(path)). Specifying it implies type: 'filesystem'. Passed through
      // to storage factories, including plugin-provided ones.
      path?: string

      // REMOVED in 8.0 — passing `rootDirectory` THROWS. Use `path`.
      rootDirectory?: string

      // REMOVED in 8.0 — a nested `options.{path,rootDirectory}` THROWS. Use `path`.
      options?: any
    }
  | StorageAdapter   // e.g. storage: new MemoryStorage()

The canonical — and only — key is the top-level path. The pre-8.0 aliases rootDirectory, options.{path,rootDirectory}, and fileSystemStorage.{path,rootDirectory} were removed in 8.0: passing one now throws with the exact rename (never a silent default that would misplace data on upgrade). Because path implies filesystem, { path: '/data' } is a complete config; the type is optional.

Direct construction

If you want to skip the factory:

import { FileSystemStorage, MemoryStorage } from '@soulcraft/brainy'

const fsStorage = new FileSystemStorage('./brainy-data')
const memStorage = new MemoryStorage()

const brain = new Brainy({ storage: fsStorage })

Migration from 7.x cloud adapters

7.x consumers of OPFSStorage, GcsStorage, R2Storage, S3CompatibleStorage, or AzureBlobStorage need to migrate to FileSystemStorage plus operator backup tooling. The recipe:

  1. On the host running Brainy, mount a local disk (NVMe recommended). Cloud providers all expose persistent local disks: GCP Persistent Disk, AWS EBS, Azure Managed Disks.
  2. Set storage: { type: 'filesystem', path: '/mnt/brainy-data' }.
  3. Run your existing data import once into the new local store.
  4. Set up an operator backup job using gsutil / aws s3 / rclone / azcopy on a cron — hourly or whatever your RPO requires. Point it at the brainy data dir.
  5. For point-in-time backups, use filesystem snapshots or brain.now().persist(path).

Same data, same APIs, no library-side cloud code.