fix: transaction timeouts are a typed no-hot-retry contract; engine-side non-retry pinned; dead transaction path removed
A production incident: a native-provider op ground 38-40s inside a transaction, blew the apply-phase budget, rolled back, and a downstream pipeline hot-retried the identical operation into a 6-minute CPU storm. Brainy itself never auto-retried the timeout; the gap was that TransactionTimeoutError only said "retryable" in prose, with nothing machine-readable for a caller to branch on. - TransactionTimeoutError gains two typed, always-true fields: retryable (a later attempt may succeed once the slowness resolves or the budget is raised) and hotRetryUnsafe (an immediate identical retry re-pays the full cost that just timed out and can cascade into a CPU storm -- callers must latch and back off, never loop). context's existing telemetry fields (timeoutMs, operationIndex, elapsedMs, totalOperations, operationName) are now documented as the caller's backoff inputs. - Updated the "retryable" doc-prose sites (transact()'s timeoutMs option, transactionBudgetFloorMs, Transaction.execute()'s contract) to point at the new fields instead of bare prose. - Regression pin (tests/unit/transaction/timeout-never-internally-retried.test.ts): an execution counter proves the engine never re-drives a timed-out operation, through both the single-op engine TransactionManager/Transaction drives for every single-record write, and add()'s upsert-race retry loop (which must exit on the first TransactionTimeoutError, never treat it like the lost-insert-race signal it retries on). - Removed TransactionManager.executeTransactionWithResult -- zero callers anywhere in the codebase.
This commit is contained in:
parent
22702b81c0
commit
003e2a74ea
7 changed files with 199 additions and 76 deletions
|
|
@ -62,9 +62,10 @@ const DEFAULT_BUDGET_FLOOR_MS = 30_000
|
|||
* NEXT operation may start (see {@link Transaction.execute}), never whether
|
||||
* already-completed work is rolled back after the fact. A trip mid-batch
|
||||
* still rolls back every operation applied so far, atomically, and throws a
|
||||
* retryable, fully-labeled TransactionTimeoutError — that zero-loss guarantee
|
||||
* doesn't change; only the point at which the clock stops mattering does (at
|
||||
* the last operation, not one check later).
|
||||
* fully-labeled `TransactionTimeoutError` — retryable-with-latch, never
|
||||
* hot-retry (see its `retryable` and `hotRetryUnsafe` fields) — that
|
||||
* zero-loss guarantee doesn't change; only the point at which the clock
|
||||
* stops mattering does (at the last operation, not one check later).
|
||||
*
|
||||
* @param opCount - Number of operations in the batch.
|
||||
* @param override - A full override for this call; wins over everything else.
|
||||
|
|
|
|||
Reference in a new issue