fix: prevent circuit breaker activation and data loss during bulk imports

Storage-aware batching system prevents rate limiting issues on cloud storage (GCS, S3, R2, Azure). Replaces entity-by-entity creation with addMany()/relateMany() batch operations in ImportCoordinator. Separate read/write circuit breakers prevent read lockouts during write throttling. Each storage adapter auto-configures optimal batch sizes and delays. Fixes silent data loss and 30+ second lockouts on 1000+ row imports.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
David Snelling 2025-10-30 08:54:04 -07:00
parent 3c5f622d64
commit 14231554e1
12 changed files with 551 additions and 112 deletions

View file

@ -18,6 +18,7 @@ import {
} from '../../coreTypes.js'
import {
BaseStorage,
StorageBatchConfig,
NOUNS_DIR,
VERBS_DIR,
METADATA_DIR,
@ -210,6 +211,32 @@ export class S3CompatibleStorage extends BaseStorage {
this.verbCacheManager = new CacheManager<Edge>(options.cacheConfig)
}
/**
* Get S3-optimized batch configuration
*
* S3 has higher throughput than GCS and handles parallel writes efficiently:
* - Larger batch sizes (100 items)
* - Parallel processing supported
* - Shorter delays between batches (50ms)
*
* S3 can handle ~3500 operations/second per bucket with good performance
*
* @returns S3-optimized batch configuration
* @since v4.11.0
*/
public getBatchConfig(): StorageBatchConfig {
return {
maxBatchSize: 100,
batchDelayMs: 50,
maxConcurrent: 100,
supportsParallelWrites: true, // S3 handles parallel writes efficiently
rateLimit: {
operationsPerSecond: 3500, // S3 is more permissive than GCS
burstCapacity: 1000
}
}
}
/**
* Initialize the storage adapter
*/