feat: comprehensive import progress tracking for all 7 formats

Add real-time progress reporting throughout the entire import pipeline
with a standardized API that works across all supported formats.

Workshop Team Feature Request:
- Eliminates "0% complete" hangs during AI extraction
- Shows continuous progress with entities/sec, throughput, ETA
- Reports contextual messages ("Processing page 5 of 23")
- Standardized progress API for CSV, PDF, Excel, JSON, Markdown, YAML, DOCX

Core Changes:
- Add FormatHandlerProgressHooks interface for extensible progress
- Wire up all 3 binary format handlers (CSV, PDF, Excel) with 7+ progress points
- Wire up all 4 text format importers (JSON, Markdown, YAML, DOCX)
- Add ImportProgress interface with stage, message, counts, throughput, ETA
- ImportCoordinator normalizes all format progress to standard interface

CLI Improvements:
- Import command now uses brain.import() directly with full progress
- Add --include-vfs flag to find command (v4.4.0 compatibility)
- Add --confidence and --weight options to add command

Documentation:
- docs/guides/standard-import-progress.md - Universal API guide
- docs/guides/import-progress-implementation.md - Developer guide
- docs/guides/import-progress-examples.md - Practical examples
- JSDoc on brain.import() with universal handler examples

Result: ONE progress handler works for ALL 7 formats with zero format-specific code!

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
David Snelling 2025-10-24 14:45:46 -07:00
parent e7ea9c4e4b
commit d5576ffb56
20 changed files with 3967 additions and 52 deletions

View file

@ -181,7 +181,43 @@ export class SmartPDFImporter {
}
// Parse PDF using existing handler
const processedData = await this.pdfHandler.process(buffer, options)
// v4.5.0: Pass progress hooks to handler for file parsing progress
const processedData = await this.pdfHandler.process(buffer, {
...options,
totalBytes: buffer.length,
progressHooks: {
onBytesProcessed: (bytes) => {
// Handler reports bytes processed during parsing
opts.onProgress?.({
processed: 0,
total: 0,
entities: 0,
relationships: 0,
phase: `Parsing PDF (${Math.round((bytes / buffer.length) * 100)}%)`
})
},
onCurrentItem: (message) => {
// Handler reports current processing step (e.g., "Processing page 5 of 23")
opts.onProgress?.({
processed: 0,
total: 0,
entities: 0,
relationships: 0,
phase: message
})
},
onDataExtracted: (count, total) => {
// Handler reports items extracted (paragraphs + tables)
opts.onProgress?.({
processed: 0,
total: total || count,
entities: 0,
relationships: 0,
phase: `Extracted ${count} items from PDF`
})
}
}
})
const data = processedData.data
const pdfMetadata = processedData.metadata.additionalInfo?.pdfMetadata || {}