docs(releases): 10.4.4 consumer notes — correctness and observability, with the performance line stated exactly
This commit is contained in:
parent
c8e05189e3
commit
a95cc5849a
1 changed files with 109 additions and 0 deletions
109
RELEASES.md
109
RELEASES.md
|
|
@ -31,6 +31,115 @@ is sometimes cited as a 7.x removal — those methods never existed on 7.x; the
|
|||
|
||||
---
|
||||
|
||||
## v10.4.4 — 2026-08-28
|
||||
|
||||
**A correctness and observability release.** The headline is not speed: it is that a
|
||||
restart now tells you the truth about itself, a store stops lying about how much it
|
||||
holds, and the engine stops doing work nobody asked for. There is a performance
|
||||
improvement and it is modest; it is stated exactly below rather than rounded up.
|
||||
|
||||
### The dark restart — fixed at the root
|
||||
|
||||
A service could stop cleanly, exit 0, having awaited `close()` on every store it held,
|
||||
and its next boot would announce `Overwriting stale writer lock … appears dead` for
|
||||
every one of them. Nothing had crashed. Two deployments hit this; the same defect also
|
||||
made those boots pay a crash-recovery fold they did not owe.
|
||||
|
||||
The cause was not the lock. `close()` released it correctly — when it got there. A
|
||||
failure part-way through close skipped both the release AND the clean-shutdown marker,
|
||||
and "the recorded pid is gone" reads identically for an orderly restart and a crash.
|
||||
|
||||
- `close()` is now two parts and the second is unconditional: the flush-request watcher,
|
||||
the **writer lock**, the VFS timers and the terminal `closed` flag are released whether
|
||||
the durable steps succeeded or not. The original failure is narrated with what it costs
|
||||
the next open, then rethrown.
|
||||
- Releasing the lock writes a **clean-close record** naming the lock generation it gave
|
||||
up. The next open reads that record instead of guessing: recorded → nothing to recover;
|
||||
absent → it says so, and names the recovery it is about to run. This also ends two
|
||||
long-standing false alarms — a recycled pid locking a store out of its own reopen, and
|
||||
`Re-acquiring writer lock … this is a bug` after a perfectly clean close.
|
||||
- The signal path stopped failing in a batch. One store's failing flush used to strand
|
||||
every remaining store's lock and markers — at exit code 0. Now: per-store isolation, the
|
||||
generation store's close (the marker) is part of shutdown, the lock goes in a `finally`,
|
||||
and the handler no longer calls `process.exit()` when the host application has its own
|
||||
signal handler, a race that truncated the host's own shutdown mid-flight.
|
||||
|
||||
### The count ledger stops lying, and `counts.json` is written atomically
|
||||
|
||||
The all-tier scalars are the denominator a coverage check subtracts against. A ledger
|
||||
derived under the old rule — one entity per id DIRECTORY — counted ghost and scar
|
||||
containers as rows, and was only FLAGGED suspect: it went on serving wrong numbers for
|
||||
the life of the store. Two copies of one archive could disagree, and a downstream index
|
||||
heal reported remaining work that did not exist.
|
||||
|
||||
- Such a ledger now derives itself honestly **in the background** after the open, counting
|
||||
identity records, and persists the correction stamped. Nothing waits for it, because no
|
||||
read is served from a denominator.
|
||||
- A derivation that raced a write refuses to stamp its number: one retry on a quiet store,
|
||||
then the ledger stays SUSPECT and names `repairIndex()` as the door that recounts under
|
||||
a barrier.
|
||||
- `counts.json` is written temp+rename. A truncating write left a window in which a
|
||||
concurrent reader saw the file EMPTY — and an unparseable ledger sends the next open
|
||||
down the full-rescan path, so the cheapest file in the store was buying the most
|
||||
expensive recovery.
|
||||
|
||||
### An open and a repair narrate themselves — on a channel a log level cannot silence
|
||||
|
||||
A store could open for three minutes and print nothing at all. The phase timings existed;
|
||||
they were written to a channel that every production-looking environment clamps away.
|
||||
|
||||
- Narration moved to an always-visible channel. An open now heartbeats the phase it is in,
|
||||
names each phase as it ends with what it was paying for, and names the expensive STEP
|
||||
inside a phase. `repairIndex()` does the same and its receipt carries a per-family
|
||||
`durationMs` — a repair that ran for half an hour with no output could only be watched
|
||||
through `top`.
|
||||
- A brain nobody has written to now does nothing: a flush over a clean store is a no-op
|
||||
and says nothing, the graph index's auto-flush asks before it acts, and the
|
||||
cross-process flush-request watch is **event-driven** (`fs.watch`) instead of polling a
|
||||
directory every 500 ms per store forever, with a slow safety sweep behind it and a
|
||||
narrated fall back to polling where a filesystem cannot be watched.
|
||||
- A provider that is REBUILDING ITSELF is no longer confused with a broken one. `init()`
|
||||
does not wait for it, every other family serves, and that family's doors refuse **by
|
||||
name, carrying the provider's own progress**, saying plainly that they open by
|
||||
themselves and no action is needed. Health narration dedupes by content, so an unchanged
|
||||
verdict is silent however a provider's generation counter moves.
|
||||
|
||||
### For operators — one behaviour change
|
||||
|
||||
**Four `where` operators that previously returned an empty page now raise
|
||||
`INVALID_QUERY`:** `startsWith`, `endsWith`, `matches` and `length`. An equality/range
|
||||
posting index cannot evaluate a substring, a pattern or an array length without reading
|
||||
every row, and it now refuses by name instead of answering with an empty result that
|
||||
looks like an answer.
|
||||
|
||||
**Three that previously returned an empty page are now SERVED:** `hasAll`, `noneOf` and
|
||||
`excludes`. All 25 accepted operator tokens now agree between this engine and its
|
||||
accelerated counterpart.
|
||||
|
||||
### Performance — stated exactly
|
||||
|
||||
Measured on a 14,056-noun / 72,679-verb production-shaped store, both builds solo under
|
||||
an exclusive lock:
|
||||
|
||||
- **Warm reopen after a clean close: 85.7 s → 77.0 s (−10.2%).** The whole of that gain is
|
||||
one fix — generation discovery reads directory NAMES instead of recursively walking the
|
||||
entire generation log (−9.2 s, and it scales with history rather than row count). The
|
||||
VFS phase is **unchanged**.
|
||||
- **Cold open: −31.4 s** (518.1 s → 486.7 s), of which the count-ledger derivation moving
|
||||
off the critical path accounts for storage-init dropping 5,941 ms → 25 ms.
|
||||
- **A dominant ~38 s remains, diagnosed and NOT fixed.** It is not the VFS — the VFS's own
|
||||
init is under 2 s of that phase. It is the log-authority adoption and/or the
|
||||
pending-embed log recovery, both now instrumented so the next measurement names the
|
||||
culprit outright.
|
||||
|
||||
Continuing work, named so nobody has to rediscover it: that ~38 s term; making the
|
||||
generation store's committed-range set lazy; the hydration path that substitutes
|
||||
`Date.now()` for an unreadable stored timestamp (inventing data); and a VFS path-prefix
|
||||
filter built with a `$startsWith` spelling no operator set accepts, so
|
||||
`searchFiles({ path })` throws today.
|
||||
|
||||
---
|
||||
|
||||
## v10.4.3 — 2026-08-27 (Open Brainy's first release)
|
||||
|
||||
**`@soulcraftlabs/brainy` 10.4.3 is the same engine as `@soulcraft/brainy` 10.4.2, byte for
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue