brainy/docs/concepts/multi-process.md
David Snelling 606445cd61 feat(8.0): API simplification — remove neural()/Db.search, one storage path key, integration→0
8.0 RC cleanup toward "one place per thing, zero-config, no deprecation":

- Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead
  legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})`
  / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate
  entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor,
  NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept.
- Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams).
  Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` →
  now `find({ query, limit })`.
- Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases
  (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the
  exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver
  feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native
  storage provider resolves the identical root (no split-brain).
- Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now
  applied as a post-filter on `result.score` (the documented way to bound semantic results).
- Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()`
  and failed dimension validation; they are metadata-only updates now.
- Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination
  shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an
  in-place path change that preserves the blob and the entity id, for files and directories.
- Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability.
  Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk
  flushes and emits `progress.queryable`, so imported data is queryable during the import.
- Sweep all docs, comments, and JSDoc for the removed/changed APIs.

Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
2026-06-20 13:31:11 -07:00

5.5 KiB

title slug public category template order description next
Multi-Process Model concepts/multi-process true concepts concept 5 How Brainy coordinates a single writer with any number of readers on a filesystem data directory — and how to safely inspect a live store.
guides/inspection

Multi-Process Model

Brainy is a single-writer, many-reader database when backed by filesystem storage. This page explains the model, the guarantees, and the safe ways to inspect a live store from a second process.

The rule

For one data directory:

  • One writer at a time. The writer acquires an exclusive lock on the directory at init() and releases it on close().
  • Any number of readers, concurrent with each other and with the writer. Readers open via Brainy.openReadOnly() — they never touch the writer lock.

Any attempt to open a second writer on the same directory throws:

BrainyError: Another writer holds this Brainy directory.
  PID:        1774431 on host app-host-1
  Started:    2026-05-15T14:22:11Z
  Heartbeat:  2026-05-15T14:22:34Z
  Version:    7.21.0
  Directory:  /data/brain

This is intentional. Two writers sharing a directory would silently corrupt in-memory indexes and produce wrong query results — the worst possible default for an operations tool.

Why a lock?

Brainy keeps its primary indexes (HNSW, metadata, graph adjacency) in memory. On disk, those indexes are persisted incrementally as writes flush. A second process opening the same directory:

  • Loads the persisted state into a fresh in-memory copy.
  • Has no awareness of writes the first process buffered but hasn't flushed.
  • Will overwrite the persisted state on its own next flush, racing the first process and corrupting whichever wins.

The fix is the lock: refuse to open a second writer. SQLite has done the same since the late 1990s (SQLITE_BUSY).

What about Cortex?

Brainy + Cortex compose cleanly under this model:

  • Cortex stores its column-index segments inside the same rootDir (under indexes/_column_index/{field}/).
  • Segments (*.cidx files) are immutable once written. Cortex mmaps them read-only.
  • The MANIFEST.json per field is updated via atomic rename — readers see either the old or new manifest, never a torn file.

A reader process can safely mmap Cortex segments alongside a live writer without coordination. The single Brainy writer lock at <rootDir>/locks/_writer.lock covers Cortex too, because Cortex segment writes happen on the writer's side.

Stale-lock detection

If a writer crashes or is forcibly killed, its lock file is left behind. To avoid a permanently-jammed directory, Brainy treats a lock as stale when:

  1. The recorded hostname equals the current host (cross-host PID checks are unsafe), AND
  2. The recorded pid is no longer alive (process.kill(pid, 0) returns ESRCH), OR the lastHeartbeat field is older than 60 seconds.

A live writer rewrites lastHeartbeat every 10 seconds, so a hung writer that's missed several heartbeats is treated as dead. Stale locks are overwritten with a warning.

If stale detection cannot prove the existing lock is dead — for example, a crashed writer on a different host writing to a shared filesystem — pass { force: true } to override. A warning is logged either way.

Heartbeat and shutdown

The heartbeat interval rewrites the lock file every 10 seconds. The timer is unref'd, so it does not keep the event loop alive on its own.

On normal shutdown the writer releases the lock in close(). The shutdown hooks Brainy registers for SIGTERM, SIGINT, and beforeExit also release the lock so a container restart doesn't strand the directory.

How to inspect a live writer

Use Brainy.openReadOnly(). It does not acquire the writer lock, so it coexists with whatever the writer is doing:

const reader = await Brainy.openReadOnly({
  storage: { type: 'filesystem', path: '/data/brain' }
})

const stats = await reader.stats()
const bookings = await reader.find({ where: { entityType: 'booking' } })

await reader.close()

What the reader sees reflects the writer's most recent flush to disk. If you need fresher state, ask the writer to flush before opening:

const reader = await Brainy.openReadOnly({
  storage: { type: 'filesystem', path: '/data/brain' }
})

const acked = await reader.requestFlush({ timeoutMs: 5000 })
if (!acked) {
  console.warn('Writer did not respond; results reflect last natural flush.')
}

const fresh = await reader.find({ where: { entityType: 'booking' } })

The CLI brainy inspect subcommands all do this for you by default (--no-fresh to opt out).

What's not enforced (yet)

  • Non-filesystem backends are out of scope in 8.0, which ships only the filesystem and memory adapters. A custom BaseStorage subclass that is not filesystem-backed does not enforce multi-process locking by default: two processes can both succeed at init() in writer mode and clobber each other's writes. A best-effort warning is logged in writer mode against a non-filesystem backend.
  • Long-running readers do not automatically pick up new Cortex segments the writer publishes. One-shot inspector calls re-open the store and see fresh segments; a reader that stays open for hours sees its column store as-of the time it opened.

Reading material

  • Brainy.openReadOnly()API reference
  • brainy inspectinspection guide
  • Cortex columnar storage — see node_modules/@soulcraft/cortex/README.md