brainy/src/vfs/importers/DirectoryImporter.ts

387 lines
11 KiB
TypeScript
Raw Normal View History

feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
/**
* Directory Importer for VFS
*
* Efficiently imports real directories into VFS with:
* - Batch processing for performance
* - Progress tracking
* - Error recovery
* - Parallel processing
*/
import { promises as fs } from 'fs'
import * as path from 'path'
import { VirtualFileSystem } from '../VirtualFileSystem.js'
import { Brainy } from '../../brainy.js'
import { NounType } from '../../types/graphTypes.js'
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
import { v4 as uuidv4 } from '../../universal/uuid.js'
feat: add ImageHandler with EXIF extraction and comprehensive MIME detection (v5.2.0) Implements Phase 1.5 (Comprehensive MIME Type Detection) and adds built-in image processing support to IntelligentImportAugmentation. **New Features:** - ImageHandler: Extracts image metadata (dimensions, format, color space) using sharp - EXIF extraction: Camera data, GPS, timestamps using exifr library - Support for JPEG, PNG, WebP, GIF, TIFF, BMP, SVG, HEIC, AVIF formats - MimeTypeDetector: Unified MIME type detection with magic byte support - FormatDetector: Enhanced with image format detection via MIME + magic bytes **Architecture Fixes:** - Fixed brain.import() augmentation pipeline integration (src/brainy.ts:3140-3154) - Added parameter spreading for ImportSource objects to enable augmentation access - Fixed metadata propagation through ImportCoordinator to final results - Added augmentation data check in ImportCoordinator.extract() **Integration:** - ImageHandler registered as built-in handler alongside CSV, Excel, PDF - Images import as 'media' entities with 'image' subtype - Full metadata preserved in knowledge graph entities - Configuration options: enableImage, extractEXIF, imageDefaults **Test Coverage:** - 15 integration tests (image-import.test.ts) - 100% passing - 27 unit tests (image-handler.test.ts) - 100% passing - Format detection tests for all supported image types - Error handling and resilience tests **Breaking Changes:** None - backward compatible Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-03 14:06:17 -08:00
import { mimeDetector } from '../MimeTypeDetector.js'
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
export interface ImportOptions {
targetPath?: string // VFS target path (default: '/')
recursive?: boolean // Import subdirectories (default: true)
skipHidden?: boolean // Skip hidden files (default: false)
skipNodeModules?: boolean // Skip node_modules (default: true)
batchSize?: number // Files per batch (default: 100)
generateEmbeddings?: boolean // Generate embeddings (default: true)
extractMetadata?: boolean // Extract metadata (default: true)
showProgress?: boolean // Log progress (default: false)
filter?: (path: string) => boolean // Custom filter function
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
// v4.10.0: Import tracking
importId?: string // Unique import identifier (auto-generated if not provided)
projectId?: string // Project identifier grouping related imports
customMetadata?: Record<string, any> // Custom metadata to attach
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
}
export interface ImportResult {
imported: string[] // Successfully imported paths
failed: Array<{ // Failed imports
path: string
error: Error
}>
skipped: string[] // Skipped paths
totalSize: number // Total bytes imported
duration: number // Time taken in ms
filesProcessed: number // Total files processed
directoriesCreated: number // Total directories created
}
export interface ImportProgress {
type: 'progress' | 'complete' | 'error'
processed: number
total?: number
current?: string
error?: Error
}
export class DirectoryImporter {
constructor(
private vfs: VirtualFileSystem,
private brain: Brainy
) {}
/**
* Import a directory or file into VFS
*/
async import(sourcePath: string, options: ImportOptions = {}): Promise<ImportResult> {
const startTime = Date.now()
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
// v4.10.0: Generate tracking metadata
const importId = options.importId || uuidv4()
const projectId = options.projectId || this.deriveProjectId(options.targetPath || '/')
const trackingMetadata = {
importIds: [importId],
projectId,
importedAt: Date.now(),
importSource: sourcePath,
...(options.customMetadata || {})
}
// Store tracking metadata in options for use in helper methods
const enhancedOptions = { ...options, _trackingMetadata: trackingMetadata }
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
const result: ImportResult = {
imported: [],
failed: [],
skipped: [],
totalSize: 0,
duration: 0,
filesProcessed: 0,
directoriesCreated: 0
}
try {
const stats = await fs.stat(sourcePath)
if (stats.isFile()) {
await this.importFile(sourcePath, options.targetPath || '/', result)
} else if (stats.isDirectory()) {
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
await this.importDirectory(sourcePath, enhancedOptions as any, result)
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
}
} catch (error) {
result.failed.push({
path: sourcePath,
error: error as Error
})
}
result.duration = Date.now() - startTime
return result
}
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
/**
* Derive project ID from target path
*/
private deriveProjectId(targetPath: string): string {
const segments = targetPath.split('/').filter(s => s.length > 0)
return segments.length > 0 ? segments[0] : 'default_project'
}
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
/**
* Import with progress tracking (generator)
*/
async *importStream(sourcePath: string, options: ImportOptions = {}): AsyncGenerator<ImportProgress> {
const files = await this.collectFiles(sourcePath, options)
const total = files.length
const batchSize = options.batchSize || 100
let processed = 0
// Process in batches
for (let i = 0; i < files.length; i += batchSize) {
const batch = files.slice(i, i + batchSize)
try {
await this.processBatch(batch, options)
processed += batch.length
yield {
type: 'progress',
processed,
total,
current: batch[batch.length - 1]
}
} catch (error) {
yield {
type: 'error',
processed,
total,
error: error as Error
}
}
}
yield {
type: 'complete',
processed,
total
}
}
/**
* Import a directory recursively
*/
private async importDirectory(
dirPath: string,
options: ImportOptions,
result: ImportResult
): Promise<void> {
const targetPath = options.targetPath || '/'
// Create VFS directory structure
await this.createDirectoryStructure(dirPath, targetPath, options, result)
// Collect all files
const files = await this.collectFiles(dirPath, options)
// Process files in batches
const batchSize = options.batchSize || 100
for (let i = 0; i < files.length; i += batchSize) {
const batch = files.slice(i, i + batchSize)
await this.processBatch(batch, options, result)
if (options.showProgress && i % (batchSize * 10) === 0) {
console.log(`Imported ${i} / ${files.length} files...`)
}
}
}
/**
* Create directory structure in VFS
*/
private async createDirectoryStructure(
sourcePath: string,
targetPath: string,
options: ImportOptions,
result: ImportResult
): Promise<void> {
// Walk directory tree and create all directories first
const dirsToCreate: string[] = []
const collectDirs = async (dir: string, vfsPath: string) => {
dirsToCreate.push(vfsPath)
const entries = await fs.readdir(dir, { withFileTypes: true })
for (const entry of entries) {
if (entry.isDirectory()) {
if (this.shouldSkip(entry.name, path.join(dir, entry.name), options)) {
continue
}
const childPath = path.join(dir, entry.name)
const childVfsPath = path.posix.join(vfsPath, entry.name)
if (options.recursive !== false) {
await collectDirs(childPath, childVfsPath)
}
}
}
}
await collectDirs(sourcePath, targetPath)
// Create all directories
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
const trackingMetadata = (options as any)._trackingMetadata || {}
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
for (const dirPath of dirsToCreate) {
try {
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
await this.vfs.mkdir(dirPath, {
recursive: true,
metadata: trackingMetadata // v4.10.0: Add tracking metadata
})
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
result.directoriesCreated++
} catch (error: any) {
if (error.code !== 'EEXIST') {
result.failed.push({ path: dirPath, error })
}
}
}
}
/**
* Collect all files to be imported
*/
private async collectFiles(dirPath: string, options: ImportOptions): Promise<string[]> {
const files: string[] = []
const walk = async (dir: string) => {
const entries = await fs.readdir(dir, { withFileTypes: true })
for (const entry of entries) {
const fullPath = path.join(dir, entry.name)
if (this.shouldSkip(entry.name, fullPath, options)) {
continue
}
if (entry.isFile()) {
files.push(fullPath)
} else if (entry.isDirectory() && options.recursive !== false) {
await walk(fullPath)
}
}
}
await walk(dirPath)
return files
}
/**
* Process a batch of files
*/
private async processBatch(
files: string[],
options: ImportOptions,
result?: ImportResult
): Promise<void> {
const imports = await Promise.allSettled(
files.map(filePath => this.importSingleFile(filePath, options))
)
if (result) {
for (let i = 0; i < imports.length; i++) {
const importResult = imports[i]
const filePath = files[i]
if (importResult.status === 'fulfilled') {
result.imported.push(importResult.value.vfsPath)
result.totalSize += importResult.value.size
result.filesProcessed++
} else {
result.failed.push({
path: filePath,
error: importResult.reason
})
}
}
}
}
/**
* Import a single file
*/
private async importSingleFile(
filePath: string,
options: ImportOptions
): Promise<{ vfsPath: string, size: number }> {
const stats = await fs.stat(filePath)
const content = await fs.readFile(filePath)
// Calculate VFS path
const relativePath = path.relative(process.cwd(), filePath)
const vfsPath = path.posix.join(options.targetPath || '/', relativePath)
// Generate embedding if requested
let embedding: number[] | undefined
if (options.generateEmbeddings !== false) {
try {
// Use first 10KB for embedding
const text = content.toString('utf8', 0, Math.min(10240, content.length))
// Generate embedding using brain's embed method
const embedResult = await this.brain.embed({ data: text })
embedding = embedResult
} catch {
// Continue without embedding if generation fails
}
}
// Write to VFS
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
const trackingMetadata = (options as any)._trackingMetadata || {}
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
await this.vfs.writeFile(vfsPath, content, {
generateEmbedding: options.generateEmbeddings,
extractMetadata: options.extractMetadata,
metadata: {
originalPath: filePath,
originalSize: stats.size,
fix: resolve HNSW concurrency race condition across all storage adapters Fixes critical P0 bug causing data corruption during bulk imports with 50+ concurrent operations. The non-atomic read-modify-write pattern in saveHNSWData() combined with fire-and-forget neighbor updates was causing 16-32 concurrent writes per entity, resulting in lost HNSW connections and corrupted graph structure. **Root Cause:** - saveHNSWData() used non-atomic read-modify-write - HNSW neighbor updates fired without await (16-32 concurrent writes/entity) - Popular nodes became hotspots (100 concurrent imports = 3,400 concurrent saveHNSWData calls) - Result: Lost neighbor connections, 0 search results **Atomic Write Strategies by Adapter:** FileSystemStorage: - Atomic rename with temp files - Write to {file}.tmp.{timestamp}.{random} - POSIX-guaranteed atomic rename(temp, final) GCSStorage: - Optimistic locking with generation numbers - preconditionOpts: { ifGenerationMatch } - 5 retries with exponential backoff (50ms→800ms) S3/R2/AzureStorage: - ETag-based optimistic locking - IfMatch/conditions preconditions - 5 retries with exponential backoff MemoryStorage + OPFSStorage: - Mutex locks per entity path - Serializes async operations even in single-threaded environments HNSW Index: - Changed fire-and-forget .catch() to await - Serializes 16-32 neighbor updates per entity - Trade-off: 20-30% slower bulk import vs 100% data integrity **Sharding Compatibility:** - ✅ Works with deterministic UUID sharding (256 shards, always on) - ✅ Works with distributed multi-node sharding (optional) - ✅ All atomic strategies work in both single-node and distributed deployments **Index Impact:** - Only HNSW index modified (saveHNSWData, saveHNSWSystem) - Other 4 indexes unaffected (Metadata, Graph Adjacency, Deleted Items, Entity ID Mapper) - No regression risk - isolated code paths **Testing:** - 8/8 unit tests passing (real concurrent operations, no mocks) - Tests verify data integrity after 20 concurrent updates - Tests verify temp file cleanup and mutex serialization **Files Modified:** - All 8 storage adapters (FileSystem, GCS, S3, R2, Azure, Memory, OPFS) - HNSW Index (neighbor update serialization) - New test: tests/unit/storage/hnswConcurrency.test.ts (8 passing tests) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-29 15:24:20 -07:00
originalModified: stats.mtime.getTime(),
...trackingMetadata // v4.10.0: Add tracking metadata
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
}
})
return { vfsPath, size: stats.size }
}
/**
* Import a single file (for non-directory imports)
*/
private async importFile(
filePath: string,
targetPath: string,
result: ImportResult
): Promise<void> {
try {
const imported = await this.importSingleFile(filePath, { targetPath })
result.imported.push(imported.vfsPath)
result.totalSize += imported.size
result.filesProcessed++
} catch (error) {
result.failed.push({
path: filePath,
error: error as Error
})
}
}
/**
* Check if a path should be skipped
*/
private shouldSkip(name: string, fullPath: string, options: ImportOptions): boolean {
// Skip hidden files if requested
if (options.skipHidden && name.startsWith('.')) {
return true
}
// Skip node_modules by default
if (name === 'node_modules' && options.skipNodeModules !== false) {
return true
}
// Apply custom filter
if (options.filter && !options.filter(fullPath)) {
return true
}
return false
}
feat: add ImageHandler with EXIF extraction and comprehensive MIME detection (v5.2.0) Implements Phase 1.5 (Comprehensive MIME Type Detection) and adds built-in image processing support to IntelligentImportAugmentation. **New Features:** - ImageHandler: Extracts image metadata (dimensions, format, color space) using sharp - EXIF extraction: Camera data, GPS, timestamps using exifr library - Support for JPEG, PNG, WebP, GIF, TIFF, BMP, SVG, HEIC, AVIF formats - MimeTypeDetector: Unified MIME type detection with magic byte support - FormatDetector: Enhanced with image format detection via MIME + magic bytes **Architecture Fixes:** - Fixed brain.import() augmentation pipeline integration (src/brainy.ts:3140-3154) - Added parameter spreading for ImportSource objects to enable augmentation access - Fixed metadata propagation through ImportCoordinator to final results - Added augmentation data check in ImportCoordinator.extract() **Integration:** - ImageHandler registered as built-in handler alongside CSV, Excel, PDF - Images import as 'media' entities with 'image' subtype - Full metadata preserved in knowledge graph entities - Configuration options: enableImage, extractEXIF, imageDefaults **Test Coverage:** - 15 integration tests (image-import.test.ts) - 100% passing - 27 unit tests (image-handler.test.ts) - 100% passing - Format detection tests for all supported image types - Error handling and resilience tests **Breaking Changes:** None - backward compatible Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-03 14:06:17 -08:00
// v5.2.0: MIME detection moved to MimeTypeDetector service
// Removed detectMimeType() - now using mimeDetector singleton
feat: implement complete VFS with Knowledge Layer integration Add production-ready Virtual File System with intelligent Knowledge Layer: Core VFS Features: - Complete file system operations (read, write, mkdir, etc.) - Intelligent PathResolver with 4-layer caching system - Chunked storage for large files with real compression - Embedding generation for semantic operations - File relationships and metadata tracking - Import functionality from local filesystem Knowledge Layer Integration: - EventRecorder for complete file history and temporal coupling - SemanticVersioning with content-based change detection - PersistentEntitySystem for character/entity tracking across files - ConceptSystem for universal concept mapping and graphs - GitBridge for import/export between VFS and Git repositories Architecture: - KnowledgeAugmentation properly integrated into Brainy augmentation system - KnowledgeLayer wrapper provides real-time VFS operation interception - Background processing ensures VFS operations remain fast - All components use real Brainy embed() method for embeddings - Support for creative writing, coding projects, and project management Technical Implementation: - Fixed all stub/mock implementations with real working code - TypeScript compilation passes without errors - Comprehensive test suite demonstrating all features - Documentation covering architecture and usage patterns - Backwards compatible with existing Brainy functionality This enables scenarios like writing books with persistent characters, managing coding projects with concept tracking, and complete project coordination with intelligent file relationships.
2025-09-24 17:31:48 -07:00
}