fix: transaction timeouts are a typed no-hot-retry contract; engine-side non-retry pinned; dead transaction path removed
A production incident: a native-provider op ground 38-40s inside a transaction, blew the apply-phase budget, rolled back, and a downstream pipeline hot-retried the identical operation into a 6-minute CPU storm. Brainy itself never auto-retried the timeout; the gap was that TransactionTimeoutError only said "retryable" in prose, with nothing machine-readable for a caller to branch on. - TransactionTimeoutError gains two typed, always-true fields: retryable (a later attempt may succeed once the slowness resolves or the budget is raised) and hotRetryUnsafe (an immediate identical retry re-pays the full cost that just timed out and can cascade into a CPU storm -- callers must latch and back off, never loop). context's existing telemetry fields (timeoutMs, operationIndex, elapsedMs, totalOperations, operationName) are now documented as the caller's backoff inputs. - Updated the "retryable" doc-prose sites (transact()'s timeoutMs option, transactionBudgetFloorMs, Transaction.execute()'s contract) to point at the new fields instead of bare prose. - Regression pin (tests/unit/transaction/timeout-never-internally-retried.test.ts): an execution counter proves the engine never re-drives a timed-out operation, through both the single-op engine TransactionManager/Transaction drives for every single-record write, and add()'s upsert-race retry loop (which must exit on the first TransactionTimeoutError, never treat it like the lost-insert-race signal it retries on). - Removed TransactionManager.executeTransactionWithResult -- zero callers anywhere in the codebase.
This commit is contained in:
parent
22702b81c0
commit
003e2a74ea
7 changed files with 199 additions and 76 deletions
|
|
@ -121,9 +121,13 @@ export interface TransactOptions {
|
|||
* with the batch: `max(30 000, opCount × 2 000)` — production imports on
|
||||
* network-attached disks measure ~2 s per operation, so a flat 30 s budget
|
||||
* silently capped honest bulk work at ~15 operations. A tripped budget
|
||||
* rolls the whole batch back and throws a retryable
|
||||
* `TransactionTimeoutError` naming the operation it stopped at, the batch
|
||||
* size, and the elapsed/budget times.
|
||||
* rolls the whole batch back and throws a `TransactionTimeoutError` naming
|
||||
* the operation it stopped at, the batch size, and the elapsed/budget
|
||||
* times. That error is retryable-with-latch, never hot-retry: its
|
||||
* `retryable` field says a later attempt may succeed, its
|
||||
* `hotRetryUnsafe` field says an immediate identical retry re-pays the
|
||||
* full cost that just timed out — callers must latch and back off, never
|
||||
* loop.
|
||||
*/
|
||||
timeoutMs?: number
|
||||
}
|
||||
|
|
|
|||
|
|
@ -62,9 +62,10 @@ const DEFAULT_BUDGET_FLOOR_MS = 30_000
|
|||
* NEXT operation may start (see {@link Transaction.execute}), never whether
|
||||
* already-completed work is rolled back after the fact. A trip mid-batch
|
||||
* still rolls back every operation applied so far, atomically, and throws a
|
||||
* retryable, fully-labeled TransactionTimeoutError — that zero-loss guarantee
|
||||
* doesn't change; only the point at which the clock stops mattering does (at
|
||||
* the last operation, not one check later).
|
||||
* fully-labeled `TransactionTimeoutError` — retryable-with-latch, never
|
||||
* hot-retry (see its `retryable` and `hotRetryUnsafe` fields) — that
|
||||
* zero-loss guarantee doesn't change; only the point at which the clock
|
||||
* stops mattering does (at the last operation, not one check later).
|
||||
*
|
||||
* @param opCount - Number of operations in the batch.
|
||||
* @param override - A full override for this call; wins over everything else.
|
||||
|
|
|
|||
|
|
@ -19,7 +19,6 @@
|
|||
import { Transaction } from './Transaction.js'
|
||||
import {
|
||||
TransactionFunction,
|
||||
TransactionResult,
|
||||
TransactionOptions
|
||||
} from './types.js'
|
||||
import { TransactionError } from './errors.js'
|
||||
|
|
@ -105,34 +104,6 @@ export class TransactionManager {
|
|||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Execute a transaction and return detailed result
|
||||
*/
|
||||
async executeTransactionWithResult<T>(
|
||||
fn: TransactionFunction<T>,
|
||||
options?: TransactionOptions
|
||||
): Promise<TransactionResult<T>> {
|
||||
const startTime = Date.now()
|
||||
const transaction = new Transaction(options)
|
||||
|
||||
try {
|
||||
const value = await fn(transaction)
|
||||
await transaction.execute()
|
||||
|
||||
const executionTimeMs = Date.now() - startTime
|
||||
|
||||
return {
|
||||
value,
|
||||
operationCount: transaction.getOperationCount(),
|
||||
executionTimeMs
|
||||
}
|
||||
|
||||
} catch (error) {
|
||||
// Transaction failed
|
||||
throw error
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Get transaction statistics
|
||||
*/
|
||||
|
|
|
|||
|
|
@ -73,14 +73,47 @@ export class InvalidTransactionStateError extends TransactionError {
|
|||
|
||||
/**
|
||||
* Error for transaction timeout
|
||||
*
|
||||
* Machine-readable no-hot-retry contract: {@link retryable} and
|
||||
* {@link hotRetryUnsafe} are both always `true` on this class — they exist
|
||||
* so a caller can branch on the *shape* of the error instead of parsing
|
||||
* message text. Read them together: the operation may eventually succeed,
|
||||
* but never by looping on it immediately.
|
||||
*
|
||||
* `context` (inherited from {@link TransactionError}) carries the caller's
|
||||
* backoff inputs — see the field docs below.
|
||||
*/
|
||||
export class TransactionTimeoutError extends TransactionError {
|
||||
/**
|
||||
* The failed operation MAY succeed on a later attempt — once the
|
||||
* underlying slowness resolves (e.g. a cold page cache warms up) or the
|
||||
* budget is deliberately raised (`transactionBudgetFloorMs`, or a larger
|
||||
* `timeoutMs` override on the batch). This is a statement about eventual
|
||||
* retryability, not a license to retry now — see {@link hotRetryUnsafe}.
|
||||
*/
|
||||
public readonly retryable = true
|
||||
|
||||
/**
|
||||
* An immediate, identical retry re-pays the FULL cost of the work that
|
||||
* just timed out — it does not resume partway. Looping on this error
|
||||
* (hot-retrying) repeats that full cost every attempt and can cascade
|
||||
* into a CPU/resource storm on the caller's side. Callers MUST latch: on
|
||||
* this error, record `{ at: Date.now(), error }`, surface one loud
|
||||
* failure to their own caller, and hold a cooldown window before any
|
||||
* re-attempt (clearing the latch only on success). Never retry this error
|
||||
* in a tight loop.
|
||||
*/
|
||||
public readonly hotRetryUnsafe = true
|
||||
|
||||
constructor(
|
||||
timeoutMs: number,
|
||||
operationIndex: number,
|
||||
telemetry?: {
|
||||
/** Milliseconds elapsed in the transaction when the budget tripped. */
|
||||
elapsedMs?: number
|
||||
/** Total number of operations in the batch that timed out. */
|
||||
totalOperations?: number
|
||||
/** Name of the operation the batch was about to start when it tripped, if named. */
|
||||
operationName?: string
|
||||
}
|
||||
) {
|
||||
|
|
@ -93,8 +126,16 @@ export class TransactionTimeoutError extends TransactionError {
|
|||
telemetry?.elapsedMs !== undefined ? `${telemetry.elapsedMs}ms elapsed, ` : ''
|
||||
super(
|
||||
`Transaction timed out at operation ${progress}${name} — ${elapsed}budget ${timeoutMs}ms. ` +
|
||||
`The batch rolled back atomically; retry with a higher timeoutMs or a smaller batch.`,
|
||||
{ timeoutMs, operationIndex, ...telemetry }
|
||||
`The batch rolled back atomically; retryable after the underlying slowness resolves or ` +
|
||||
`the budget is raised, but hot-retry-unsafe — latch and back off, never loop.`,
|
||||
{
|
||||
// Caller backoff inputs — all present on every instance:
|
||||
/** Configured budget (ms) that was exceeded. */
|
||||
timeoutMs,
|
||||
/** Index of the operation the batch was about to start when it tripped. */
|
||||
operationIndex,
|
||||
...telemetry
|
||||
}
|
||||
)
|
||||
this.name = 'TransactionTimeoutError'
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1658,8 +1658,10 @@ export interface BrainyConfig {
|
|||
* **start** — never whether already-completed work gets rolled back after
|
||||
* the fact (a single-op write can never time out post-hoc: it either runs
|
||||
* or it commits). A trip mid-batch still rolls back every applied operation
|
||||
* atomically and throws a retryable `TransactionTimeoutError`; only the
|
||||
* floor of the formula is configurable here.
|
||||
* atomically and throws a `TransactionTimeoutError` that is
|
||||
* retryable-with-latch, never hot-retry (see its `retryable` and
|
||||
* `hotRetryUnsafe` fields); only the floor of the formula is configurable
|
||||
* here.
|
||||
*
|
||||
* Raise this when a cold store's first writes after a restart legitimately
|
||||
* take longer than 30s per operation (e.g. page-cache-cold canonical writes
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue