fix: prevent circuit breaker activation and data loss during bulk imports

Storage-aware batching system prevents rate limiting issues on cloud storage (GCS, S3, R2, Azure). Replaces entity-by-entity creation with addMany()/relateMany() batch operations in ImportCoordinator. Separate read/write circuit breakers prevent read lockouts during write throttling. Each storage adapter auto-configures optimal batch sizes and delays. Fixes silent data loss and 30+ second lockouts on 1000+ row imports.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
David Snelling 2025-10-30 08:54:04 -07:00
parent 3c5f622d64
commit 14231554e1
12 changed files with 551 additions and 112 deletions

View file

@ -25,6 +25,7 @@ import {
} from '../../coreTypes.js'
import {
BaseStorage,
StorageBatchConfig,
NOUNS_DIR,
VERBS_DIR,
METADATA_DIR,
@ -162,6 +163,32 @@ export class AzureBlobStorage extends BaseStorage {
}
}
/**
* Get Azure Blob-optimized batch configuration
*
* Azure Blob Storage has moderate rate limits between GCS and S3:
* - Medium batch sizes (75 items)
* - Parallel processing supported
* - Moderate delays (75ms)
*
* Azure can handle ~2000 operations/second with good performance
*
* @returns Azure Blob-optimized batch configuration
* @since v4.11.0
*/
public getBatchConfig(): StorageBatchConfig {
return {
maxBatchSize: 75,
batchDelayMs: 75,
maxConcurrent: 75,
supportsParallelWrites: true, // Azure handles parallel reasonably
rateLimit: {
operationsPerSecond: 2000, // Moderate limits
burstCapacity: 500
}
}
}
/**
* Initialize the storage adapter
*/