Two production-reported spine fixes:
- Canonical delete is a FULL removal: remove()/removeMany() deleted the
metadata leg but never the canonical vectors.json leg or the <id>/
directory, leaving ghost rows that inflated enumerated counts forever,
read as damage scars, and caused duplicate readdir entries for
re-created VFS paths (the stale-unpost path). The delete operation now
routes through storage.deleteNoun — both legs + the entity container —
with a full two-leg before-image rollback; deleteVerb symmetric; the
blind 'no metadata file' catch that masked real faults is gone. The
generation log still holds the delete's before-image, so asOf() history
is unchanged. repairIndex() gains a conservative, loud orphan-prune
sweep (containers with no metadata content leg) + count recompute for
stores damaged by earlier versions.
- Family-scoped migration gate: every read blocked on the whole-brain
migration lock even when the migrating index was irrelevant. The gate
now scopes to the index families a read actually consults — canonical
reads never wait; find() waits only on its query shape's families;
graph traversals wait only on the graph family. Writes keep the
conservative whole-brain wait. A read that needs the migrating family
still blocks (bounded) with the retryable MigrationInProgressError.
Replaces rc.8's no-freeze deference with an automatic, observable, coordinated
migration LOCK (David's reversal: "unknown/halfway states are more dangerous
than blocking"). While a native provider runs its one-time 7.x→8.0
rebuild-from-canonical (`isMigrating()===true`), brainy blocks/queues data-plane
reads AND writes so no operation ever touches a half-built index — closing the
three seam-map gaps (lost mid-rebuild writes, degraded reads, read-time rebuild).
Mechanism (one choke point):
- awaitMigrationLock() at ensureInitialized() covers all ~40 data-plane methods;
a non-migrating brain pays one boolean check. Polls isMigrating() at 250ms;
after migrationWaitTimeoutMs (default 30s) throws a retryable, exported
MigrationInProgressError. The timeout bounds the CALLER'S WAIT, not the
migration — the rebuild is unbounded and never interrupted.
- init() awaits the lock before VFS bootstrap + before serving ("not ready until
upgraded"); a rebuild past the budget surfaces MigrationInProgressError at init
(raise the budget, or run the offline migrator).
Observability (readiness-probe correct — never gated):
- getIndexStatus() gains `migrating` + `migration` (MigrationProgress); a probe
maps migrating→HTTP 503+Retry-After, not 500. health()/checkHealth() lock-exempt
(health() reports a `warn` migration check). stampBrainFormat()/close() are
ungated so cor can clear the lock — no deadlock.
- Optional provider migrationStatus() is relayed verbatim for a live %.
verifyGraphAdjacencyLive() honors the lock (no self-rebuild mid-migration); the
stamp is still withheld while any provider migrates (cor stamps after verify).
Adversarially verified: fixed a critical init/VFS-bootstrap deadlock, a mixed-
provider busy-spin (dropped the event-driven signal → pure poll), a not-yet-
initialized getIndexStatus crash, and once-per-window log/clock resets. 10 lock
tests (incl. the init-during-migration regression). Gates: typecheck 0, build 0,
test:unit 1753/1753.