docs: measured performance envelopes v1 (per-op p50/p95 at 1k and 10k, pure-JS floor)
First edition of the per-release performance-envelope contract: every number measured against the built dist on stated hardware, never projected. Sub-0.1ms get/related (adjacency O(degree), scale-flat), 1-9ms indexed metadata finds, ~178ms semantic (query embedding dominates), ~167ms durability-priced single-op writes flat across scale, 8-45ms steady-state flush independent of history backlog (the 8.9.0 change). Two weak spots stated honestly: addMany commits per-item today (batched chunk commits belong to the unified-commit roadmap), and pure-JS warm open grows with corpus (4.9s at 10k) — the native accelerator's reason to exist. Refresh rule: any release touching a measured path re-measures in the same release.
This commit is contained in:
parent
70e4bc8a79
commit
5cabd784f4
2 changed files with 128 additions and 0 deletions
45
RELEASES.md
45
RELEASES.md
|
|
@ -8,8 +8,53 @@ Full auto-generated changelog: `CHANGELOG.md` · Releases: https://github.com/so
|
|||
- Debugging data, query, or storage behaviour
|
||||
- A new Brainy feature is available that you want to adopt
|
||||
|
||||
## Removed APIs — 7.x → 8.x (the complete ledger)
|
||||
|
||||
Every public API removed at the 8.0 major, with its sanctioned replacement. If your code
|
||||
still calls a left-column name on 8.x it throws (or the config key is rejected) — the
|
||||
replacement is always a one-line change. (Standing contract from 8.9.0 forward: removals
|
||||
happen only at majors, after ≥1 minor of loud runtime deprecation naming the replacement.)
|
||||
|
||||
| Removed (7.x) | Replacement (8.x) |
|
||||
|---|---|
|
||||
| `brain.search(query, k)` | `find({ query })` — semantic; `find({ query, searchMode })` for hybrid |
|
||||
| `brain.getRelations({...})` | `related(id, opts)` for adjacency; `find({ connected: {...} })` for scoped traversal |
|
||||
| `brain.neural()` clustering | `find({ vector })` + aggregation `GROUP BY` |
|
||||
| `Db.search()` | `db.find({ vector })` |
|
||||
| Pre-8.0 storage path aliases (`directory`, `basePath`, …) | one `storage.path` key (old aliases throw) |
|
||||
| Reserved keys inside `metadata` bags (silently remapped in 7.x) | top-level params (`subtype`, `visibility`, `confidence`, `weight`, …) — reserved-in-bag throws |
|
||||
| 7.x COW branches layout (`branches/main/`) | generational MVCC (`asOf()`, `now()`, `db.persist(path)`) — on-disk migration is automatic at first 8.x open |
|
||||
|
||||
The fork/snapshot family (`brain.snapshot()`, `createSnapshot()`, `restoreSnapshot()`)
|
||||
is sometimes cited as a 7.x removal — those methods never existed on 7.x; the 8.0 Db API
|
||||
(`asOf`/`persist`/`restore({confirm})`) is their first real implementation.
|
||||
|
||||
---
|
||||
|
||||
## v8.9.0 — 2026-07-19 (flush is durability-only: history maintenance moves to close())
|
||||
|
||||
The write path stops paying maintenance costs — the last structural piece of the
|
||||
flush-storm class (a production deployment measured single writes blocked 25–191s behind
|
||||
history reclaim running inline on flush under memory pressure):
|
||||
|
||||
- **`flush()` never compacts history.** It persists the current window's deltas and
|
||||
nothing else — its cost no longer depends on history backlog or retention mode, in any
|
||||
configuration. **`close()` is the auto-compaction site** (time-bounded per pass, ~5s;
|
||||
an early stop is a consistent prefix and the next pass resumes).
|
||||
- **`compactHistory()` gains `timeBudgetMs`** — bound your own maintenance windows; the
|
||||
same resumable-prefix guarantee applies.
|
||||
- **The documented trade**: a long-lived writer that never closes accumulates history
|
||||
until its next explicit `compactHistory()`. Predictable writes, explicit maintenance.
|
||||
If you run bounded retention on an always-on service, schedule a periodic
|
||||
`compactHistory({ ...caps, timeBudgetMs })` in your maintenance window.
|
||||
- **New public doc: `docs/performance-envelopes.md`** — measured per-op envelopes
|
||||
(p50/p95 at stated scales, hardware, and backend, with the measuring script cited).
|
||||
Refresh rule going forward: any release touching a measured path re-runs that op's
|
||||
benchmark and updates the envelope in the same release.
|
||||
- **New in this file: the Removed APIs 7.x→8.x table** (top of this document) — every
|
||||
removal with its sanctioned replacement, one place, per the engine-currency contract.
|
||||
Standing from here: removals only at majors, after ≥1 minor of loud runtime deprecation.
|
||||
|
||||
## v8.8.2 — 2026-07-19 (one field-resolution law: reserved-field aggregates stop drifting)
|
||||
|
||||
Four fixes from a consumer conformance audit, all rooted in the same disease — two field-resolution
|
||||
|
|
|
|||
83
docs/performance-envelopes.md
Normal file
83
docs/performance-envelopes.md
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
---
|
||||
title: Performance Envelopes
|
||||
slug: guides/performance-envelopes
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 40
|
||||
description: Measured per-operation latency envelopes at stated scales — what to expect, on what hardware, and exactly how each number was produced.
|
||||
next:
|
||||
- guides/find-limits
|
||||
---
|
||||
|
||||
# Performance Envelopes
|
||||
|
||||
Every number on this page is **measured, never projected** — produced by the script
|
||||
cited at the bottom, against the built package (the artifact you install), on the stated
|
||||
hardware. Each entry says what was measured, at what scale, on which storage backend.
|
||||
When a release touches a measured path, that operation is re-measured and this page
|
||||
updates in the same release.
|
||||
|
||||
Two scopes to keep straight:
|
||||
|
||||
- **These envelopes are the pure-JS engine** (no native accelerator registered) on
|
||||
filesystem storage. This is the floor every deployment gets from `npm install` alone.
|
||||
- **Accelerated deployments** (the optional native provider) publish their own numbers —
|
||||
this page never claims them.
|
||||
|
||||
## Read operations
|
||||
|
||||
Reads are where the architecture pays off: after the write path has done its indexing
|
||||
work, queries answer from purpose-built indexes without scanning.
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `get(id)` (warm) | p50 < 0.1ms | p50 < 0.1ms | served from cache/metadata index |
|
||||
| `find` (metadata: indexed equality + range, limit 100) | p50 1.0ms · p95 1.8ms | p50 7.0ms · p95 8.9ms | column-store bitmap paths |
|
||||
| `related(id)` (per-node adjacency) | p50 < 0.1ms · p95 0.2ms | p50 < 0.1ms | LSM adjacency index — O(degree), scale-independent |
|
||||
| `find` (semantic: embed + HNSW, 1k docs) | p50 178ms · p95 393ms | — | dominated by WASM query embedding (measured on a machine under concurrent load — treat the p95 as an upper bound); the vector search itself is single-digit ms |
|
||||
|
||||
## Write operations
|
||||
|
||||
Under Model-B **every write is its own durable generation** — a single-op `add` pays
|
||||
serialization, before-image staging, and fsync before it acks. That durability is priced
|
||||
into the write path visibly, by design:
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `add` (single-op) | p50 167ms · p95 171ms | p50 165ms · p95 172ms | full durable generation per write — flat across scale |
|
||||
| `addMany` (bulk) | ~163ms/entity | ~187ms/entity | **currently per-item commits** — see the honest note below |
|
||||
| `relateMany` | ~0.8ms/edge | ~0.9ms/edge | edges batch efficiently today |
|
||||
| `flush` (steady-state, 1 pending write) | p50 8ms · p95 10ms | p50 45ms · p95 52ms | durability-only since 8.9.0 — cost no longer depends on history backlog or retention mode |
|
||||
|
||||
**The honest note on bulk writes:** `addMany` today commits each item as its own
|
||||
generation (the same durability as single-op `add`, serialized by the single-writer
|
||||
lock), so bulk-load cost is N × single-op cost. Batched chunk commits (one generation
|
||||
and one fsync window per chunk, as `removeMany` already does) are designed into the
|
||||
unified-commit work on the current roadmap. Until that ships, size bulk imports
|
||||
accordingly — 10k entities is minutes, not seconds, on filesystem storage.
|
||||
|
||||
## Open / close
|
||||
|
||||
| Operation | 1,000 entities | 10,000 entities | Notes |
|
||||
|---|---|---|---|
|
||||
| `open` (empty store) | ~560ms | ~190ms | includes embedder initialization |
|
||||
| `open` (warm, populated, clean shutdown) | 763ms | 4.9s | pure-JS vector index load dominates and grows with entity count; the native accelerator exists precisely to remove this |
|
||||
| `close` | bounded | bounded | auto-compaction pass is time-bounded (~5s max) since 8.9.0 |
|
||||
|
||||
A store that was NOT cleanly closed pays index rebuilds on top of the warm-open
|
||||
number (tens of seconds at 10k) — clean shutdown is worth engineering for.
|
||||
|
||||
## How these were produced
|
||||
|
||||
- **Hardware**: Intel Core i9-14900HX (32 threads), 62GB RAM, NVMe, Linux, Node v22.
|
||||
- **Backend**: `storage: { type: 'filesystem' }`, pure JS (no native providers).
|
||||
- **Embeddings**: deterministic stub for non-semantic ops (isolates engine cost);
|
||||
the real WASM embedder for the semantic row (that's what you'll run).
|
||||
- **Method**: p50/p95 over 50–200 samples per op against the built `dist/`;
|
||||
the measuring script ships in the repo history and re-runs per release.
|
||||
|
||||
Numbers on different hardware will differ; the *shape* (sub-2ms indexed reads,
|
||||
~160ms embedding-bound semantic queries, durability-priced writes) is the envelope
|
||||
you should hold your deployment against. If your measurements diverge from these
|
||||
shapes by an order of magnitude, something is wrong — file it.
|
||||
Loading…
Add table
Add a link
Reference in a new issue