test(budgets): iron-honest wall-clock budgets — 3x the worst honest-iron measurement
Some checks failed
CI / Node 24 (push) Successful in 12m30s
CI / Node 22 (push) Successful in 12m36s
CI / Integration + conformance (Node 22) (push) Failing after 13m45s
CI / Bun (latest) (push) Successful in 12m19s

Seven micro-budget tests were calibrated on one fast desktop and failed on
other honest iron with zero functional failures (bisect-proven pre-existing;
David-waived for 10.1/10.2 with this recalibration filed as the cure). Every
budget is now at least 3x the worst measurement observed across three
machines, each with a comment naming its calibration basis; the find-unified
micro-comparison of two sub-millisecond timings becomes a ratio assertion
(absolute equality of microsecond pairs can never be stable). The
inference-bound trim-history correctness test gets a timeout covering its
slowest observed run (174s) — its assertions are exact and untouched.

These remain order-of-magnitude guards; real perf enforcement lives in the
dedicated perf lanes with iron-specific budgets, per the gate-speed standard.

Known non-test artifact, documented not hidden: on slow-inference machines a
minutes-long awaited-embed loop can trip vitest's worker-RPC 60s tolerance
('Timeout calling onTaskUpdate') — all tests pass, vitest exits 1 on the
unhandled orchestration error. The CI lanes on faster iron exit clean; if a
lane ever trips it, the test moves to deterministic embeddings (its
assertions are size-bookkeeping, not embedding quality).
This commit is contained in:
David Snelling 2026-08-18 09:36:21 -07:00
parent 292e7c0406
commit 314e0e6c29
7 changed files with 46 additions and 17 deletions

View file

@ -343,9 +343,11 @@ describe('NaturalLanguageProcessor', () => {
const duration = Date.now() - startTime
expect(result).toBeDefined()
expect(duration).toBeLessThan(200) // Should be fast
// order-of-magnitude guard: worst honest-iron measurement 4.8s
// (CPU-only inference path, 32-core box); 15s budget covers 3x that
expect(duration).toBeLessThan(15000)
})
it('should handle multiple queries efficiently', async () => {
const queries = Array(10).fill('Find AI research')
@ -356,8 +358,10 @@ describe('NaturalLanguageProcessor', () => {
const duration = Date.now() - startTime
expect(results).toHaveLength(10)
expect(duration).toBeLessThan(2000) // Should handle batch in reasonable time
})
// order-of-magnitude guard: worst honest-iron measurement 48.2s for 10
// concurrent inference-path queries (CPU-only, 32-core box); ~3x headroom
expect(duration).toBeLessThan(150000)
}, 200000)
it('should cache pattern matching for performance', async () => {
const query = 'Find machine learning papers'