brainy/docs/performance-envelopes.md
David Snelling 5cabd784f4 docs: measured performance envelopes v1 (per-op p50/p95 at 1k and 10k, pure-JS floor)
First edition of the per-release performance-envelope contract: every
number measured against the built dist on stated hardware, never
projected. Sub-0.1ms get/related (adjacency O(degree), scale-flat),
1-9ms indexed metadata finds, ~178ms semantic (query embedding
dominates), ~167ms durability-priced single-op writes flat across
scale, 8-45ms steady-state flush independent of history backlog
(the 8.9.0 change). Two weak spots stated honestly: addMany commits
per-item today (batched chunk commits belong to the unified-commit
roadmap), and pure-JS warm open grows with corpus (4.9s at 10k) —
the native accelerator's reason to exist. Refresh rule: any release
touching a measured path re-measures in the same release.
2026-07-19 13:35:04 -07:00

4.5 KiB
Raw Blame History

title slug public category template order description next
Performance Envelopes guides/performance-envelopes true guides guide 40 Measured per-operation latency envelopes at stated scales — what to expect, on what hardware, and exactly how each number was produced.
guides/find-limits

Performance Envelopes

Every number on this page is measured, never projected — produced by the script cited at the bottom, against the built package (the artifact you install), on the stated hardware. Each entry says what was measured, at what scale, on which storage backend. When a release touches a measured path, that operation is re-measured and this page updates in the same release.

Two scopes to keep straight:

  • These envelopes are the pure-JS engine (no native accelerator registered) on filesystem storage. This is the floor every deployment gets from npm install alone.
  • Accelerated deployments (the optional native provider) publish their own numbers — this page never claims them.

Read operations

Reads are where the architecture pays off: after the write path has done its indexing work, queries answer from purpose-built indexes without scanning.

Operation 1,000 entities 10,000 entities Notes
get(id) (warm) p50 < 0.1ms p50 < 0.1ms served from cache/metadata index
find (metadata: indexed equality + range, limit 100) p50 1.0ms · p95 1.8ms p50 7.0ms · p95 8.9ms column-store bitmap paths
related(id) (per-node adjacency) p50 < 0.1ms · p95 0.2ms p50 < 0.1ms LSM adjacency index — O(degree), scale-independent
find (semantic: embed + HNSW, 1k docs) p50 178ms · p95 393ms dominated by WASM query embedding (measured on a machine under concurrent load — treat the p95 as an upper bound); the vector search itself is single-digit ms

Write operations

Under Model-B every write is its own durable generation — a single-op add pays serialization, before-image staging, and fsync before it acks. That durability is priced into the write path visibly, by design:

Operation 1,000 entities 10,000 entities Notes
add (single-op) p50 167ms · p95 171ms p50 165ms · p95 172ms full durable generation per write — flat across scale
addMany (bulk) ~163ms/entity ~187ms/entity currently per-item commits — see the honest note below
relateMany ~0.8ms/edge ~0.9ms/edge edges batch efficiently today
flush (steady-state, 1 pending write) p50 8ms · p95 10ms p50 45ms · p95 52ms durability-only since 8.9.0 — cost no longer depends on history backlog or retention mode

The honest note on bulk writes: addMany today commits each item as its own generation (the same durability as single-op add, serialized by the single-writer lock), so bulk-load cost is N × single-op cost. Batched chunk commits (one generation and one fsync window per chunk, as removeMany already does) are designed into the unified-commit work on the current roadmap. Until that ships, size bulk imports accordingly — 10k entities is minutes, not seconds, on filesystem storage.

Open / close

Operation 1,000 entities 10,000 entities Notes
open (empty store) ~560ms ~190ms includes embedder initialization
open (warm, populated, clean shutdown) 763ms 4.9s pure-JS vector index load dominates and grows with entity count; the native accelerator exists precisely to remove this
close bounded bounded auto-compaction pass is time-bounded (~5s max) since 8.9.0

A store that was NOT cleanly closed pays index rebuilds on top of the warm-open number (tens of seconds at 10k) — clean shutdown is worth engineering for.

How these were produced

  • Hardware: Intel Core i9-14900HX (32 threads), 62GB RAM, NVMe, Linux, Node v22.
  • Backend: storage: { type: 'filesystem' }, pure JS (no native providers).
  • Embeddings: deterministic stub for non-semantic ops (isolates engine cost); the real WASM embedder for the semantic row (that's what you'll run).
  • Method: p50/p95 over 50200 samples per op against the built dist/; the measuring script ships in the repo history and re-runs per release.

Numbers on different hardware will differ; the shape (sub-2ms indexed reads, ~160ms embedding-bound semantic queries, durability-priced writes) is the envelope you should hold your deployment against. If your measurements diverge from these shapes by an order of magnitude, something is wrong — file it.