PRODUCTION BLOCKER: Workshop team reported data corruption at 450+ entities during bulk imports (1400 files) with 1000+ concurrent operations. ## Root Cause FileSystemStorage lacked mutex locks for HNSW operations, causing read-modify-write race conditions at production scale. Memory and OPFS adapters already had mutex locks (v4.9.2), but FileSystemStorage only had atomic rename which prevents torn writes but NOT lost updates. ## The Race Condition Without mutex, concurrent operations on same entity: 1. Thread A reads file (connections: [1,2,3]) 2. Thread B reads file (connections: [1,2,3]) 3. Thread A adds connection 4, writes [1,2,3,4] 4. Thread B adds connection 5, writes [1,2,3,5] ← Connection 4 LOST Result: Corrupted HNSW graph, lost connections, undefined entity IDs ## Why Previous Fixes Failed v4.9.2: Added atomic rename (prevents torn writes, NOT lost updates) v4.10.0: Made problem worse by increasing concurrency without mutex ## The Fix Added mutex locks to FileSystemStorage matching Memory/OPFS (v4.9.2): - fileSystemStorage.ts:90 - Added hnswLocks Map - fileSystemStorage.ts:2609-2626 - Mutex wraps saveHNSWData() - fileSystemStorage.ts:2719-2731 - Mutex wraps saveHNSWSystem() Mutex serializes concurrent operations PER ENTITY while maintaining atomic rename for crash safety. ## Workshop Bug Symptoms (Now Fixed) 1. Entity IDs undefined (300+ errors) - Fixed by preventing data loss 2. JSON truncation at 8KB (position 8192) - Fixed by serializing writes 3. Field index lock contention (100+ indexes) - Fixed by preventing corruption cascade 4. Corruption starts at ~450 entities - Fixed by handling hub node contention ## Test Coverage Added 3 production-scale tests (hnswConcurrency.test.ts:533-688): - 1000 concurrent saveHNSWData() on shared hub node (Workshop scenario) - 500 entity corruption threshold test (crosses 450-entity limit) - 100 concurrent system updates (entry point changes) Test results: 16/16 passing - 1000 concurrent ops: 177ms, 0 errors - 500 entities: 0 undefined IDs, 0 corrupted data, 0 truncation - All production-scale scenarios pass ## Evidence Workshop bug report: brain-cloud/apps/workshop/BRAINY_V4.9.2_HNSW_CONCURRENCY_BUG_REPORT.md Test Gap: Previous tests used 20 concurrent ops vs 1000 in production (50× difference) Fix Pattern: Matches memoryStorage.ts:828 and opfsStorage.ts:2033 mutex implementation ## Breaking Changes None - fully backward compatible ## Migration No migration needed. Existing data compatible. For corrupted v4.9.2 data, recommend clean slate re-import for guaranteed consistency. |
||
|---|---|---|
| .. | ||
| hnswConcurrency.test.ts | ||
| typeAwareStorageAdapter.test.ts | ||