fix(storage): a clean close is recorded, and the writer lock is always given up
Some checks are pending
CI / Node 22 (push) Waiting to run
CI / Node 24 (push) Waiting to run
CI / Integration + conformance (Node 22) (push) Waiting to run
CI / Bun (latest) (push) Waiting to run

A production restart made this necessary: a service stopped with exit code 0,
having awaited close() on every pooled brain, and its next boot announced
"Overwriting stale writer lock ... appears dead" for every store it owned.
Nothing had crashed. "The recorded pid is gone" is equally true of an orderly
restart and of a crash, so the verdict could not tell an operator which one
they had — and when the OS recycles a pid it fails the other way, refusing to
open a store whose writer died days ago.

Three changes, all at the law:

- close() is two parts, and the second is unconditional. The durable steps
  (flush, markers, component close, plugin deactivate, buffer drain) move to
  closeDurableSteps(); the terminal releases — the flush-request watcher, the
  WRITER LOCK, the VFS timers, the terminal `closed` flag — always run. The
  original failure is narrated with what it costs the next open, then rethrown.
- releaseWriterLock() writes a CLEAN-CLOSE RECORD (`locks/_writer.close`)
  naming the lock generation it released; the next claim consumes it, so a
  record can never vouch for a later crash. An open reads the record instead
  of guessing: recorded → nothing to recover; absent → say so, and name the
  crash recovery this open will now run.
- The signal path stops failing in a batch. It was one try around a loop over
  every open brain, so the first instance whose flush rejected stranded every
  remaining brain's lock and markers — at exit code 0. Now: per-instance
  isolation, the generation store's close (the clean-shutdown marker, without
  which the next open folds the whole log) is part of shutdown, the lock is
  given up in a finally, and the handler no longer calls process.exit() when
  the host application has its own signal handler — that race truncated the
  host's own close() mid-flight.

Pins: tests/integration/writer-lock-clean-close.test.ts — completed close
leaves no lock and a consumed-once record with a silent reopen; a failing
durable step still releases and still rethrows; SIGKILL leaves the lock with
no record and the reopen names the crash; a host SIGTERM handler runs to
completion.

Branch plan (10 lines):
 1. writer lock: clean-close record + always-release close  [this commit]
 2. open narration: an always-on channel; production clamps prodLog to ERROR,
    which is why a three-minute open printed nothing
 3. open narration: per-phase lines as each phase ENDS, with progress cadence
 4. measure both real-store fixtures on the box, before/after
 5. move the generation-log fold out of the foreground where the serving law
    allows; durable resumable progress marker
 6. same for the VFS bootstrap
 7. counts: a legacy container-rule ledger must not keep serving wrong
    denominators; counts.json written atomically
 8. counts pin with scar directories; two copies of one archive agree
 9. docs/canonical-layout-ratification.md — 12 facts confirmed/corrected
10. report: MEASURED before/after, findings, and whether this is 10.4.4
This commit is contained in:
David Snelling 2026-08-28 10:17:20 -07:00
parent 38c3397b60
commit e652162c1f
4 changed files with 641 additions and 108 deletions

View file

@ -125,6 +125,36 @@ export interface WriterLockInfo {
rootDir?: string // Convenience for log lines / error messages
}
/**
* THE CLEAN-CLOSE RECORD. Written by `releaseWriterLock()` at the instant it
* gives up the writer lock, naming the lock identity it released. The next
* `acquireWriterLock()` reads it and can then say from a RECORD, not from a
* guess whether the previous writer left on purpose.
*
* Why a record and not PID liveness: "the recorded PID is no longer alive" is
* true of every orderly restart AND of every crash, so the two were reported
* identically ("appears dead") and neither could be trusted. Worse, the same
* inference fails the other way when the operating system RECYCLES the pid
* a live unrelated process makes a long-dead writer's lock look held, and the
* store refuses to open naming a pid that was never Brainy. A record settles
* both: matched the previous writer closed cleanly, nothing to recover;
* absent say so, and name what recovery the open will now run.
*
* Lifecycle: written at release, consumed (deleted) by the next successful
* lock claim a record must never outlive the lock generation it describes,
* or it would vouch for a later crash.
*/
export interface WriterCloseRecord {
pid: number
hostname: string
/** `startedAt` of the lock this close released — the identity match key. */
startedAt: string
/** ISO timestamp at which the lock was released. */
closedAt: string
/** Brainy version that performed the close. */
version: string
}
/**
* FNV-1a hash returning a 2-char hex bucket (00-ff).
* Distributes system keys across 256 sub-prefixes to avoid