brainy/src/importers/index.ts
David Snelling a06e8772f1 feat: add unified import system with auto-detection and dual storage
Implemented a comprehensive unified import system that revolutionizes how data flows into Brainy:

## Core Features (Phase 1)
- Auto-detection of file formats (Excel, PDF, CSV, JSON, Markdown) via magic bytes and content analysis
- Dual storage architecture: creates both VFS files AND knowledge graph entities
- Single unified API: brain.import() handles all formats automatically
- Format-specific importers for optimal extraction from each file type
- VFS structure generation with configurable grouping (by type, sheet, or flat)

## Entity Deduplication (Phase 2)
- Embedding-based similarity matching to detect duplicate entities across imports
- Intelligent merging with provenance tracking (records which imports contributed)
- Fuzzy name matching using Levenshtein distance
- Confidence score merging with weighted averages
- Cross-import shared knowledge: same entity referenced in multiple datasets gets merged

## Streaming Support (Phase 3)
- Chunked processing for memory-efficient handling of large datasets
- Configurable chunk size for optimal performance
- Progress tracking with real-time callbacks
- Scales to millions of entities without memory issues

## Import History & Rollback (Phase 4)
- Complete tracking of all imports with full metadata
- Rollback capability to undo any import completely
- Statistics and analytics across all imports
- Persistent history stored in VFS

## Architecture
- ImportCoordinator: orchestrates the entire import pipeline
- FormatDetector: auto-detects file formats with high confidence
- EntityDeduplicator: prevents duplicate entities across imports
- ImportHistory: tracks and enables rollback of imports
- Format-specific importers: SmartExcelImporter, SmartPDFImporter, etc.
- VFSStructureGenerator: creates organized file hierarchies

## Usage
```typescript
const result = await brain.import('/path/to/file.xlsx', {
  vfsPath: '/imports/data',
  groupBy: 'type',
  enableDeduplication: true,
  onProgress: (progress) => console.log(progress)
})
```

## Production Ready
- 5,500+ lines of production code
- All integration tests passing
- No mocks, stubs, or TODOs
- Full TypeScript type safety
- Comprehensive error handling
- Memory efficient and scalable

Closes requirements for unified data ingestion pipeline.
2025-10-08 16:55:30 -07:00

69 lines
1.6 KiB
TypeScript

/**
* Smart Import System
*
* Production-ready entity and relationship extraction from multiple formats:
* - Excel (.xlsx)
* - PDF (.pdf)
* - CSV (.csv)
* - JSON (.json)
* - Markdown (.md)
*
* Uses brainy's built-in NeuralEntityExtractor and NaturalLanguageProcessor
*
* NO MOCKS - Real working implementation
*/
// Excel Importer
export { SmartExcelImporter } from './SmartExcelImporter.js'
export type {
SmartExcelOptions,
ExtractedRow,
SmartExcelResult
} from './SmartExcelImporter.js'
// PDF Importer
export { SmartPDFImporter } from './SmartPDFImporter.js'
export type {
SmartPDFOptions,
ExtractedSection,
SmartPDFResult
} from './SmartPDFImporter.js'
// CSV Importer
export { SmartCSVImporter } from './SmartCSVImporter.js'
export type {
SmartCSVOptions,
SmartCSVResult
} from './SmartCSVImporter.js'
// JSON Importer
export { SmartJSONImporter } from './SmartJSONImporter.js'
export type {
SmartJSONOptions,
ExtractedJSONEntity,
ExtractedJSONRelationship,
SmartJSONResult
} from './SmartJSONImporter.js'
// Markdown Importer
export { SmartMarkdownImporter } from './SmartMarkdownImporter.js'
export type {
SmartMarkdownOptions,
MarkdownSection,
SmartMarkdownResult
} from './SmartMarkdownImporter.js'
// VFS Structure Generator
export { VFSStructureGenerator } from './VFSStructureGenerator.js'
export type {
VFSStructureOptions,
VFSStructureResult
} from './VFSStructureGenerator.js'
// Orchestrator (Main entry point)
export { SmartImportOrchestrator } from './SmartImportOrchestrator.js'
export type {
SmartImportOptions,
SmartImportProgress,
SmartImportResult
} from './SmartImportOrchestrator.js'