brainy/docs/guides/inspection.md
David Snelling 2a03fae0e2 feat: brain.auditGraph() — read-only graph-truth audit
- New public API: walks every canonical relationship record, queries the same
  read path applications use (related() with all visibility tiers included),
  and classifies every discrepancy into its failure family: missingFromReads
  (records the read path omits — stale adjacency), danglingEndpoints
  (endpoint entity gone — the partial-delete scar class), readOnlyVerbIds
  (read-path edges with no canonical record — ghosts). Design-hidden
  internal/system edges are counted separately so intentional hiding is never
  misclassified as loss. Counts exact; example lists capped with an explicit
  truncatedExamples flag; loud narration on incoherence; mutates nothing.
- Classification core lives in src/graph/graphAudit.ts behind injected seams
  (canonical walks + the end-to-end read) so the discrepancy logic is testable
  in isolation; brainy wires the seams to storage walks and related().
- getNounIdsWithPagination now refuses an undecodable resume cursor loudly —
  the third and final pagination walk brought under the 8.5.2 cursor contract.
- Types exported: GraphAuditReport, GraphAuditDiscrepancy. Docs: the
  inspection guide gains the audit -> repair -> audit verification ritual.
2026-07-17 16:34:45 -07:00

7.1 KiB

title slug public category template order description next
Inspecting a Live Brainy guides/inspection true guides guide 30 Operator recipes for diagnosing a running Brainy data directory — counts, queries, query plans, health checks, and snapshots — without stopping the live writer.
concepts/multi-process

Inspecting a Live Brainy

When something is wrong in production, you need to see what's actually in the store. This guide covers the safe ways to query a running Brainy directory.

The cardinal rule

Never open a second writer on the same directory. Filesystem storage will throw, and any other write path will corrupt the live writer's state. Use Brainy.openReadOnly() or the brainy inspect CLI instead.

The CLI is the fastest path

# What's in this brain?
brainy inspect stats /data/brain

# Find specific entities
brainy inspect find /data/brain --type Event --where '{"status":"paid"}' --limit 20

# Single entity by ID
brainy inspect get /data/brain 0b7a9...

# Why is this query returning empty?
brainy inspect explain /data/brain --where '{"entityType":"booking"}'

# Quick invariants
brainy inspect health /data/brain

# Random sample (no query needed)
brainy inspect sample /data/brain --type Event --n 20

# Tail new writes as they happen
brainy inspect watch /data/brain --type Event

# Save a snapshot
brainy inspect backup /data/brain /backups/brain-$(date +%Y%m%d).tar

Every subcommand internally:

  1. Asks the live writer to flush via the cross-process RPC (skip with --no-fresh).
  2. Opens the data directory via Brainy.openReadOnly().
  3. Runs the query.
  4. Closes cleanly.

Results are JSON by default. Add --pretty for indented output.

When a query returns surprising results

If find() returns 0 for a query you expect to match: run inspect explain first. It shows which index path will serve each where clause:

$ brainy inspect explain /data/brain --where '{"entityType":"booking","status":"paid"}'
{
  "query": { "where": { "entityType": "booking", "status": "paid" } },
  "fieldPlan": [
    { "field": "entityType", "path": "none", "notes": "No index entries for field..." },
    { "field": "status", "path": "column-store", "notes": "O(log n) binary search..." }
  ],
  "warnings": [
    "Field \"entityType\" has no index entries. find() will return [] silently."
  ]
}

The "path": "none" is the smoking gun. It means the field has no column store manifest and no sparse chunked index — so find() will return [] regardless of what's actually on disk. Likely causes:

  • The writer registered the field in memory but hasn't flushed. Run brain.requestFlush() from the writer side, or use brainy inspect --fresh (default).
  • The field name has a typo or wrong casing.
  • The field is genuinely absent from every entity.

Health checks

inspect health runs a fixed battery of cheap invariant checks:

$ brainy inspect health /data/brain
{
  "overall": "warn",
  "checks": [
    { "name": "index-parity", "status": "pass", "message": "Vector (1851) and metadata (1851) agree." },
    { "name": "field-registry", "status": "pass", "message": "23 fields registered for 1851 entities." },
    { "name": "seeded-records", "status": "warn", "message": "15 entities tagged _seeded:true." },
    { "name": "writer-heartbeat", "status": "pass", "message": "Writer healthy (PID 1774431...)." }
  ]
}

Each check returns pass, warn, or fail. The exit code is 2 when any check fails — useful for piping into monitoring or CI.

Programmatic inspection

import { Brainy } from '@soulcraft/brainy'

const reader = await Brainy.openReadOnly({
  storage: { type: 'filesystem', path: '/data/brain' }
})

// Force the writer to flush before reading
await reader.requestFlush({ timeoutMs: 5000 })

// What's in there?
const stats = await reader.stats()
console.log(`${stats.entityCount} entities`, stats.entitiesByType)

// Why is this query empty?
const plan = await reader.explain({ where: { entityType: 'booking' } })
for (const f of plan.fieldPlan) {
  console.log(`${f.field} -> ${f.path}`)
}

// Run invariants
const health = await reader.health()
console.log(health.overall)

await reader.close()

Every mutation method (add, update, remove, relate, transact, restore, ...) throws on a read-only instance with a clear message.

Backups

brainy inspect backup asks the writer to flush first, then tars the directory. The snapshot reflects the writer's state at the moment of the flush:

brainy inspect backup /data/brain /backups/brain-2026-05-15.tar

For periodic backups (hourly, daily), schedule this via cron or your container scheduler. For point-in-time recovery, use the Db API's db.persist(path) — a self-contained hard-link snapshot that later writes can never alter, restorable with brain.restore(path, { confirm: true }). See Snapshots & Time Travel.

Comparing two stores

brainy inspect diff returns a JSON summary of counts and a sample of entity IDs present in one but not the other. Useful when debugging replication or migrations:

brainy inspect diff /data/brain-prod /data/brain-staging

Sample-based — for a full diff, dump both with inspect dump and compare the JSONL.

Auditing graph-read truth

brain.auditGraph() (8.6.0+) proves — or disproves — that relationship reads return canonical truth on a given brain, without mutating anything. It walks every stored relationship record, asks the same read path your application uses (related(), VFS readdir) with every visibility tier included, and classifies every discrepancy:

const report = await brain.auditGraph()

report.coherent              // true = related()/readdir can be trusted on this brain
report.missingFromReadsCount // records the read path omits — stale index
report.danglingEndpointsCount // relationships whose endpoint entity is gone
report.readOnlyCount         // read-path edges with NO stored record — ghosts
report.visibilityHiddenCount // internal/system edges hidden by design (not a fault)

Counts are always exact; the example lists (missingFromReads, danglingEndpoints, readOnlyVerbIds) are capped at maxExamples (default 100) and truncatedExamples says so when they are.

Run it after any engine upgrade, restore, or migration. If it reports discrepancies, run brain.repairIndex() and audit again — a coherent report after the repair is the verified statement that the heal worked. Cost: one relationship-record walk plus one indexed read per distinct source entity — safe on a live brain.

Repairing a corrupted store

If invariants fail and you suspect index corruption, inspect repair opens the store in writer mode and rebuilds all indexes from raw storage. Stop the live writer firstrepair will throw if another writer holds the lock. Add --force only if you have personally verified the existing lock is stale.

brainy inspect repair /data/brain

Multi-process safety summary

See concepts/multi-process for the lock semantics, heartbeat behavior, and what's not yet enforced on cloud backends.