brainy/examples/unified-import-example.ts

157 lines
5 KiB
TypeScript
Raw Permalink Normal View History

feat: add unified import system with auto-detection and dual storage Implemented a comprehensive unified import system that revolutionizes how data flows into Brainy: ## Core Features (Phase 1) - Auto-detection of file formats (Excel, PDF, CSV, JSON, Markdown) via magic bytes and content analysis - Dual storage architecture: creates both VFS files AND knowledge graph entities - Single unified API: brain.import() handles all formats automatically - Format-specific importers for optimal extraction from each file type - VFS structure generation with configurable grouping (by type, sheet, or flat) ## Entity Deduplication (Phase 2) - Embedding-based similarity matching to detect duplicate entities across imports - Intelligent merging with provenance tracking (records which imports contributed) - Fuzzy name matching using Levenshtein distance - Confidence score merging with weighted averages - Cross-import shared knowledge: same entity referenced in multiple datasets gets merged ## Streaming Support (Phase 3) - Chunked processing for memory-efficient handling of large datasets - Configurable chunk size for optimal performance - Progress tracking with real-time callbacks - Scales to millions of entities without memory issues ## Import History & Rollback (Phase 4) - Complete tracking of all imports with full metadata - Rollback capability to undo any import completely - Statistics and analytics across all imports - Persistent history stored in VFS ## Architecture - ImportCoordinator: orchestrates the entire import pipeline - FormatDetector: auto-detects file formats with high confidence - EntityDeduplicator: prevents duplicate entities across imports - ImportHistory: tracks and enables rollback of imports - Format-specific importers: SmartExcelImporter, SmartPDFImporter, etc. - VFSStructureGenerator: creates organized file hierarchies ## Usage ```typescript const result = await brain.import('/path/to/file.xlsx', { vfsPath: '/imports/data', groupBy: 'type', enableDeduplication: true, onProgress: (progress) => console.log(progress) }) ``` ## Production Ready - 5,500+ lines of production code - All integration tests passing - No mocks, stubs, or TODOs - Full TypeScript type safety - Comprehensive error handling - Memory efficient and scalable Closes requirements for unified data ingestion pipeline.
2025-10-08 16:55:30 -07:00
/**
* Unified Import System Example
*
* Demonstrates the new brain.import() method that:
* - Auto-detects file formats
* - Creates both VFS structure and Knowledge Graph
* - Links files to entities
* - Works with all formats (Excel, PDF, CSV, JSON, Markdown)
*/
import { Brainy } from '../src/brainy.js'
import * as fs from 'fs'
import * as path from 'path'
async function main() {
console.log('🧠 Brainy Unified Import System Demo\n')
// Initialize Brainy with in-memory storage for demo
const brain = new Brainy({
storage: { type: 'memory' as const }
})
await brain.init()
// Example 1: Import JSON object (no file needed!)
console.log('📥 Example 1: Import JSON object')
const jsonData = {
entities: [
{
name: 'John Smith',
type: 'person',
description: 'Software engineer interested in AI and machine learning'
},
{
name: 'San Francisco',
type: 'location',
description: 'City in California known for tech companies'
}
]
}
const jsonResult = await brain.import(jsonData, {
vfsPath: '/imports/demo-json',
onProgress: (progress) => {
console.log(` ${progress.stage}: ${progress.message}`)
}
})
console.log(`✅ Imported ${jsonResult.stats.entitiesExtracted} entities`)
console.log(` Created ${jsonResult.stats.graphNodesCreated} graph nodes`)
console.log(` Created ${jsonResult.stats.vfsFilesCreated} VFS files`)
console.log()
// Example 2: Import Markdown content
console.log('📥 Example 2: Import Markdown content')
const markdown = `
# AI Technologies
## Machine Learning
Machine learning is a subset of artificial intelligence that enables systems to learn from data.
## Neural Networks
Neural networks are computational models inspired by the human brain, used in deep learning.
## Natural Language Processing
NLP is a branch of AI that helps computers understand human language.
`
const mdResult = await brain.import(markdown, {
format: 'markdown', // Optional - will auto-detect anyway
vfsPath: '/imports/demo-markdown',
onProgress: (progress) => {
if (progress.stage === 'complete') {
console.log(`${progress.message}`)
}
}
})
console.log(`✅ Imported ${mdResult.stats.entitiesExtracted} entities`)
console.log(` Format detected: ${mdResult.format} (confidence: ${mdResult.formatConfidence})`)
console.log()
// Example 3: Import from file (optional - requires local file)
// Set TEST_EXCEL_FILE environment variable to test with your own Excel file
const testFile = process.env.TEST_EXCEL_FILE
if (testFile && fs.existsSync(testFile)) {
console.log('📥 Example 3: Import Excel file (auto-detection)')
const fileResult = await brain.import(testFile, {
vfsPath: '/imports/excel-data',
groupBy: 'type', // Group by entity type (Places/, Characters/, etc.)
onProgress: (progress) => {
if (progress.stage === 'extracting' && progress.processed && progress.total) {
process.stdout.write(`\r Extracting: ${progress.processed}/${progress.total}`)
} else if (progress.stage === 'complete') {
console.log(`\n ✅ ${progress.message}`)
}
}
})
console.log(`✅ Format: ${fileResult.format}`)
console.log(` Entities: ${fileResult.stats.entitiesExtracted}`)
console.log(` Relationships: ${fileResult.stats.graphEdgesCreated}`)
console.log(` VFS directories: ${fileResult.vfs.directories.length}`)
console.log()
}
// Example 4: Query the imported data
console.log('🔍 Querying imported entities...')
// Find entities in the graph
const machineEntity = await brain.find({
query: 'machine learning',
limit: 1
})
if (machineEntity.length > 0) {
console.log(` Found: "${machineEntity[0].metadata.name}"`)
console.log(` VFS Path: ${machineEntity[0].metadata.vfsPath}`)
console.log(` Type: ${machineEntity[0].metadata.type}`)
}
console.log()
// Example 5: Browse VFS structure
console.log('📂 VFS Structure:')
try {
const vfs = brain.vfs()
feat: add confidence/weight to Entity and flatten Result fields for convenient access Add confidence and weight properties to Entity interface and flatten Result fields to top level for improved developer experience and API consistency. Breaking Changes: None (all changes are backward compatible) Phase 2 - Entity Confidence & Weight: - Add confidence (type classification certainty) and weight (entity importance) to Entity interface - Add confidence/weight parameters to AddParams and UpdateParams - Update convertNounToEntity() to extract confidence/weight from storage - Update add() and update() methods to preserve confidence/weight in metadata - Enable developers to specify and access entity confidence/weight scores Phase 3 - Result Field Flattening: - Flatten commonly-used entity fields (type, metadata, data, confidence, weight) to Result top level - Add createResult() helper for consistent Result construction - Update all find() code paths to use createResult() - Enable direct access: result.metadata instead of result.entity.metadata - Preserve full entity in result.entity for backward compatibility VFS Fix (from previous work): - Fix VFSStructureGenerator to use brain.vfs() cached instance instead of creating separate instance - Improve VFS error messages with step-by-step guidance - Update examples to show correct vfs.init() usage - Add comprehensive VFS import verification tests Documentation Updates: - Update API_REFERENCE.md with confidence/weight examples and flattened Result documentation - Enhance JSDoc for add(), get(), find(), similar() with v4.3.0 examples - Document Result structure changes and backward compatibility - Add migration examples showing both old and new access patterns Tests: - Add 16 comprehensive tests for Entity confidence/weight exposure - Add tests for Result field flattening - Add tests for backward compatibility - All tests passing (16/16) API Consistency: - Entity: direct access to confidence/weight - Result: flattened fields + nested entity (both work) - Relation: already had confidence/weight (consistent) - VFS: inherits from Entity (automatic) Files Changed: - src/types/brainy.types.ts - Updated Entity, AddParams, UpdateParams, Result interfaces - src/brainy.ts - Updated implementation and JSDoc for all affected methods - tests/integration/entity-confidence-weight.test.ts - 16 comprehensive tests - docs/API_REFERENCE.md - Updated with v4.3.0 examples - src/importers/VFSStructureGenerator.ts - VFS fix - src/vfs/VirtualFileSystem.ts - Improved error messages - examples/unified-import-example.ts - Added vfs.init() example - tests/integration/vfs-*-verification.test.ts - VFS verification tests 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-23 12:19:50 -07:00
// IMPORTANT: Initialize VFS before querying!
// This is required even after import (idempotent - safe to call multiple times)
await vfs.init()
feat: add unified import system with auto-detection and dual storage Implemented a comprehensive unified import system that revolutionizes how data flows into Brainy: ## Core Features (Phase 1) - Auto-detection of file formats (Excel, PDF, CSV, JSON, Markdown) via magic bytes and content analysis - Dual storage architecture: creates both VFS files AND knowledge graph entities - Single unified API: brain.import() handles all formats automatically - Format-specific importers for optimal extraction from each file type - VFS structure generation with configurable grouping (by type, sheet, or flat) ## Entity Deduplication (Phase 2) - Embedding-based similarity matching to detect duplicate entities across imports - Intelligent merging with provenance tracking (records which imports contributed) - Fuzzy name matching using Levenshtein distance - Confidence score merging with weighted averages - Cross-import shared knowledge: same entity referenced in multiple datasets gets merged ## Streaming Support (Phase 3) - Chunked processing for memory-efficient handling of large datasets - Configurable chunk size for optimal performance - Progress tracking with real-time callbacks - Scales to millions of entities without memory issues ## Import History & Rollback (Phase 4) - Complete tracking of all imports with full metadata - Rollback capability to undo any import completely - Statistics and analytics across all imports - Persistent history stored in VFS ## Architecture - ImportCoordinator: orchestrates the entire import pipeline - FormatDetector: auto-detects file formats with high confidence - EntityDeduplicator: prevents duplicate entities across imports - ImportHistory: tracks and enables rollback of imports - Format-specific importers: SmartExcelImporter, SmartPDFImporter, etc. - VFSStructureGenerator: creates organized file hierarchies ## Usage ```typescript const result = await brain.import('/path/to/file.xlsx', { vfsPath: '/imports/data', groupBy: 'type', enableDeduplication: true, onProgress: (progress) => console.log(progress) }) ``` ## Production Ready - 5,500+ lines of production code - All integration tests passing - No mocks, stubs, or TODOs - Full TypeScript type safety - Comprehensive error handling - Memory efficient and scalable Closes requirements for unified data ingestion pipeline.
2025-10-08 16:55:30 -07:00
const rootContents = await vfs.readdir('/')
console.log(' Root directories:', rootContents.filter(f => !f.includes('.')))
if (rootContents.includes('imports')) {
const imports = await vfs.readdir('/imports')
console.log(' Import directories:', imports)
}
feat: add confidence/weight to Entity and flatten Result fields for convenient access Add confidence and weight properties to Entity interface and flatten Result fields to top level for improved developer experience and API consistency. Breaking Changes: None (all changes are backward compatible) Phase 2 - Entity Confidence & Weight: - Add confidence (type classification certainty) and weight (entity importance) to Entity interface - Add confidence/weight parameters to AddParams and UpdateParams - Update convertNounToEntity() to extract confidence/weight from storage - Update add() and update() methods to preserve confidence/weight in metadata - Enable developers to specify and access entity confidence/weight scores Phase 3 - Result Field Flattening: - Flatten commonly-used entity fields (type, metadata, data, confidence, weight) to Result top level - Add createResult() helper for consistent Result construction - Update all find() code paths to use createResult() - Enable direct access: result.metadata instead of result.entity.metadata - Preserve full entity in result.entity for backward compatibility VFS Fix (from previous work): - Fix VFSStructureGenerator to use brain.vfs() cached instance instead of creating separate instance - Improve VFS error messages with step-by-step guidance - Update examples to show correct vfs.init() usage - Add comprehensive VFS import verification tests Documentation Updates: - Update API_REFERENCE.md with confidence/weight examples and flattened Result documentation - Enhance JSDoc for add(), get(), find(), similar() with v4.3.0 examples - Document Result structure changes and backward compatibility - Add migration examples showing both old and new access patterns Tests: - Add 16 comprehensive tests for Entity confidence/weight exposure - Add tests for Result field flattening - Add tests for backward compatibility - All tests passing (16/16) API Consistency: - Entity: direct access to confidence/weight - Result: flattened fields + nested entity (both work) - Relation: already had confidence/weight (consistent) - VFS: inherits from Entity (automatic) Files Changed: - src/types/brainy.types.ts - Updated Entity, AddParams, UpdateParams, Result interfaces - src/brainy.ts - Updated implementation and JSDoc for all affected methods - tests/integration/entity-confidence-weight.test.ts - 16 comprehensive tests - docs/API_REFERENCE.md - Updated with v4.3.0 examples - src/importers/VFSStructureGenerator.ts - VFS fix - src/vfs/VirtualFileSystem.ts - Improved error messages - examples/unified-import-example.ts - Added vfs.init() example - tests/integration/vfs-*-verification.test.ts - VFS verification tests 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-23 12:19:50 -07:00
} catch (error: any) {
console.log(` Error: ${error.message}`)
feat: add unified import system with auto-detection and dual storage Implemented a comprehensive unified import system that revolutionizes how data flows into Brainy: ## Core Features (Phase 1) - Auto-detection of file formats (Excel, PDF, CSV, JSON, Markdown) via magic bytes and content analysis - Dual storage architecture: creates both VFS files AND knowledge graph entities - Single unified API: brain.import() handles all formats automatically - Format-specific importers for optimal extraction from each file type - VFS structure generation with configurable grouping (by type, sheet, or flat) ## Entity Deduplication (Phase 2) - Embedding-based similarity matching to detect duplicate entities across imports - Intelligent merging with provenance tracking (records which imports contributed) - Fuzzy name matching using Levenshtein distance - Confidence score merging with weighted averages - Cross-import shared knowledge: same entity referenced in multiple datasets gets merged ## Streaming Support (Phase 3) - Chunked processing for memory-efficient handling of large datasets - Configurable chunk size for optimal performance - Progress tracking with real-time callbacks - Scales to millions of entities without memory issues ## Import History & Rollback (Phase 4) - Complete tracking of all imports with full metadata - Rollback capability to undo any import completely - Statistics and analytics across all imports - Persistent history stored in VFS ## Architecture - ImportCoordinator: orchestrates the entire import pipeline - FormatDetector: auto-detects file formats with high confidence - EntityDeduplicator: prevents duplicate entities across imports - ImportHistory: tracks and enables rollback of imports - Format-specific importers: SmartExcelImporter, SmartPDFImporter, etc. - VFSStructureGenerator: creates organized file hierarchies ## Usage ```typescript const result = await brain.import('/path/to/file.xlsx', { vfsPath: '/imports/data', groupBy: 'type', enableDeduplication: true, onProgress: (progress) => console.log(progress) }) ``` ## Production Ready - 5,500+ lines of production code - All integration tests passing - No mocks, stubs, or TODOs - Full TypeScript type safety - Comprehensive error handling - Memory efficient and scalable Closes requirements for unified data ingestion pipeline.
2025-10-08 16:55:30 -07:00
}
console.log()
console.log('✨ Demo complete!')
console.log()
console.log('Key features demonstrated:')
console.log(' ✅ Auto-detection of formats (JSON, Markdown, Excel)')
console.log(' ✅ Dual storage (VFS + Knowledge Graph)')
console.log(' ✅ Entity extraction and relationship inference')
console.log(' ✅ VFS files linked to graph entities')
console.log(' ✅ Simple unified API: brain.import()')
}
main().catch(console.error)