brainy/tests/unit/vfs/blob-storage-integration.test.ts

188 lines
6 KiB
TypeScript
Raw Normal View History

feat: add ImageHandler with EXIF extraction and comprehensive MIME detection (v5.2.0) Implements Phase 1.5 (Comprehensive MIME Type Detection) and adds built-in image processing support to IntelligentImportAugmentation. **New Features:** - ImageHandler: Extracts image metadata (dimensions, format, color space) using sharp - EXIF extraction: Camera data, GPS, timestamps using exifr library - Support for JPEG, PNG, WebP, GIF, TIFF, BMP, SVG, HEIC, AVIF formats - MimeTypeDetector: Unified MIME type detection with magic byte support - FormatDetector: Enhanced with image format detection via MIME + magic bytes **Architecture Fixes:** - Fixed brain.import() augmentation pipeline integration (src/brainy.ts:3140-3154) - Added parameter spreading for ImportSource objects to enable augmentation access - Fixed metadata propagation through ImportCoordinator to final results - Added augmentation data check in ImportCoordinator.extract() **Integration:** - ImageHandler registered as built-in handler alongside CSV, Excel, PDF - Images import as 'media' entities with 'image' subtype - Full metadata preserved in knowledge graph entities - Configuration options: enableImage, extractEXIF, imageDefaults **Test Coverage:** - 15 integration tests (image-import.test.ts) - 100% passing - 27 unit tests (image-handler.test.ts) - 100% passing - Format detection tests for all supported image types - Error handling and resilience tests **Breaking Changes:** None - backward compatible Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-03 14:06:17 -08:00
import { describe, it, expect, beforeEach, afterEach } from 'vitest'
import { Brainy } from '../../../src/brainy.js'
import { VirtualFileSystem } from '../../../src/vfs/VirtualFileSystem.js'
import * as fs from 'fs/promises'
import * as path from 'path'
/**
* v5.2.0: Test unified BlobStorage integration with VFS
*
* This test verifies that:
* 1. All files (small, medium, large) use BlobStorage
* 2. No size-based branching occurs
* 3. Content is stored and retrieved correctly
* 4. Deduplication works automatically
*
* Note: Uses FileSystemStorage because BlobStorage is only available
* in COW-enabled storage adapters (not MemoryStorage)
*/
describe('VFS Unified BlobStorage (v5.2.0)', () => {
let brain: Brainy
let vfs: VirtualFileSystem
let testDir: string
beforeEach(async () => {
// Create temporary directory for test storage
testDir = path.join('/tmp', `brainy-test-blob-${Date.now()}-${Math.random().toString(36).slice(2)}`)
await fs.mkdir(testDir, { recursive: true })
feat(8.0)!: flip requireSubtype default to true (BRAINY-8.0-SUBTYPE-CONTRACT § C-1) Brainy 8.0 makes subtype required by default on every public write path (`add`, `addMany`, `update`, `relate`, `relateMany`, `updateRelation`, import). Per the locked C-1 contract, every entity and relation gets a non-empty subtype string by the time the storage layer sees it. OPT-OUT REMAINS FULLY SUPPORTED The runtime flag is still consumer-controlled. Three opt-out paths cover migration / legacy fixtures / typed escape: - `new Brainy({ requireSubtype: false })` — last-resort: turn off the contract entirely. Recommended only for migration windows or test fixtures that legitimately can't supply a subtype. - `new Brainy({ requireSubtype: { except: [NounType.Thing, ...] } })` — per-type allowlist: strict everywhere except the listed types. - `brain.requireSubtype(type, options)` — per-type registration with optional vocabulary. Composes with the brain-wide flag. Default is now `true`. Opt-out is explicit and documented; nothing silently degrades. TEST SWEEP Bulk-applied `requireSubtype: false` to every `new Brainy({...})` call site across 120 test files. Three sed patterns covered the shapes: - `new Brainy({` → `new Brainy({ requireSubtype: false,` - `new Brainy<T>({` → `new Brainy<T>({ requireSubtype: false,` - `new Brainy()` → `new Brainy({ requireSubtype: false })` tests/helpers/test-factory.ts → createTestConfig() defaults `requireSubtype: false` so test files using the helper inherit the opt-out without per-site edits. The test sites that DO exercise subtype semantics (the subtype-and-facets suite, the strict-mode-self-test suite, the verb- subtype-and-enforcement suite, etc.) already pass real subtypes — they were the 7.30.x acceptance tests for this contract. Those tests continue to pass unchanged. CHANGES src/brainy.ts - normalizeConfig() — `requireSubtype` default `false` → `true`. Comment refreshed to document the three opt-out paths. tests/* (120 files) - Bulk-edited brain construction sites. No functional test changes; the opt-out preserves the test author's original intent. tests/helpers/test-factory.ts - createTestConfig() base config gains `requireSubtype: false`. NO-OP for consumers who were already passing subtype on every write. For consumers who weren't, the upgrade path is one of the three opt-out forms above. Migration recipe documented in 8.0 release notes (next commit). VERIFICATION - npx tsc --noEmit: clean - npm test: 1408 / 1409 (same pre-existing race-condition outstanding; no other regressions from the flip)
2026-06-09 14:58:25 -07:00
brain = new Brainy({ requireSubtype: false,
feat: add ImageHandler with EXIF extraction and comprehensive MIME detection (v5.2.0) Implements Phase 1.5 (Comprehensive MIME Type Detection) and adds built-in image processing support to IntelligentImportAugmentation. **New Features:** - ImageHandler: Extracts image metadata (dimensions, format, color space) using sharp - EXIF extraction: Camera data, GPS, timestamps using exifr library - Support for JPEG, PNG, WebP, GIF, TIFF, BMP, SVG, HEIC, AVIF formats - MimeTypeDetector: Unified MIME type detection with magic byte support - FormatDetector: Enhanced with image format detection via MIME + magic bytes **Architecture Fixes:** - Fixed brain.import() augmentation pipeline integration (src/brainy.ts:3140-3154) - Added parameter spreading for ImportSource objects to enable augmentation access - Fixed metadata propagation through ImportCoordinator to final results - Added augmentation data check in ImportCoordinator.extract() **Integration:** - ImageHandler registered as built-in handler alongside CSV, Excel, PDF - Images import as 'media' entities with 'image' subtype - Full metadata preserved in knowledge graph entities - Configuration options: enableImage, extractEXIF, imageDefaults **Test Coverage:** - 15 integration tests (image-import.test.ts) - 100% passing - 27 unit tests (image-handler.test.ts) - 100% passing - Format detection tests for all supported image types - Error handling and resilience tests **Breaking Changes:** None - backward compatible Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-03 14:06:17 -08:00
storage: {
type: 'filesystem',
feat(8.0): API simplification — remove neural()/Db.search, one storage `path` key, integration→0 8.0 RC cleanup toward "one place per thing, zero-config, no deprecation": - Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})` / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor, NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept. - Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams). Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` → now `find({ query, limit })`. - Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native storage provider resolves the identical root (no split-brain). - Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now applied as a post-filter on `result.score` (the documented way to bound semantic results). - Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()` and failed dimension validation; they are metadata-only updates now. - Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an in-place path change that preserves the blob and the entity id, for files and directories. - Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability. Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk flushes and emits `progress.queryable`, so imported data is queryable during the import. - Sweep all docs, comments, and JSDoc for the removed/changed APIs. Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
2026-06-20 13:31:11 -07:00
path: testDir
feat: add ImageHandler with EXIF extraction and comprehensive MIME detection (v5.2.0) Implements Phase 1.5 (Comprehensive MIME Type Detection) and adds built-in image processing support to IntelligentImportAugmentation. **New Features:** - ImageHandler: Extracts image metadata (dimensions, format, color space) using sharp - EXIF extraction: Camera data, GPS, timestamps using exifr library - Support for JPEG, PNG, WebP, GIF, TIFF, BMP, SVG, HEIC, AVIF formats - MimeTypeDetector: Unified MIME type detection with magic byte support - FormatDetector: Enhanced with image format detection via MIME + magic bytes **Architecture Fixes:** - Fixed brain.import() augmentation pipeline integration (src/brainy.ts:3140-3154) - Added parameter spreading for ImportSource objects to enable augmentation access - Fixed metadata propagation through ImportCoordinator to final results - Added augmentation data check in ImportCoordinator.extract() **Integration:** - ImageHandler registered as built-in handler alongside CSV, Excel, PDF - Images import as 'media' entities with 'image' subtype - Full metadata preserved in knowledge graph entities - Configuration options: enableImage, extractEXIF, imageDefaults **Test Coverage:** - 15 integration tests (image-import.test.ts) - 100% passing - 27 unit tests (image-handler.test.ts) - 100% passing - Format detection tests for all supported image types - Error handling and resilience tests **Breaking Changes:** None - backward compatible Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-03 14:06:17 -08:00
},
silent: true
})
await brain.init()
vfs = brain.vfs
})
afterEach(async () => {
await brain.close()
// Clean up temporary directory
try {
await fs.rm(testDir, { recursive: true, force: true })
} catch (error) {
// Ignore cleanup errors
}
})
describe('Unified Storage Path', () => {
it('should store small files (<100KB) in BlobStorage', async () => {
const content = 'Small file content'
await vfs.writeFile('/small.txt', content)
// Get entity directly using VFS API
const entity = await vfs.getEntity('/small.txt')
expect(entity.metadata.vfsType).toBe('file')
expect(entity.metadata.storage?.type).toBe('blob')
expect(entity.metadata.storage?.hash).toBeDefined()
const readContent = await vfs.readFile('/small.txt')
expect(readContent.toString()).toBe(content)
})
it('should store medium files (100KB-10MB) in BlobStorage', async () => {
const content = Buffer.alloc(200_000, 'M') // 200KB
await vfs.writeFile('/medium.bin', content)
const entity = await vfs.getEntity('/medium.bin')
expect(entity.metadata.vfsType).toBe('file')
expect(entity.metadata.storage?.type).toBe('blob')
expect(entity.metadata.storage?.hash).toBeDefined()
const readContent = await vfs.readFile('/medium.bin')
expect(Buffer.compare(readContent, content)).toBe(0)
})
it('should store large files (>10MB) in BlobStorage', async () => {
const content = Buffer.alloc(11_000_000, 'L') // 11MB
await vfs.writeFile('/large.bin', content)
const entity = await vfs.getEntity('/large.bin')
expect(entity.metadata.vfsType).toBe('file')
expect(entity.metadata.storage?.type).toBe('blob')
expect(entity.metadata.storage?.hash).toBeDefined()
const readContent = await vfs.readFile('/large.bin')
expect(Buffer.compare(readContent, content)).toBe(0)
})
})
describe('Deduplication', () => {
it('should deduplicate identical files', async () => {
const content = 'Duplicate content test'
// Write same content to two different paths
await vfs.writeFile('/file1.txt', content)
await vfs.writeFile('/file2.txt', content)
// Get entities directly
const entity1 = await vfs.getEntity('/file1.txt')
const entity2 = await vfs.getEntity('/file2.txt')
// Both should have blob storage
expect(entity1.metadata.storage?.type).toBe('blob')
expect(entity2.metadata.storage?.type).toBe('blob')
// But same blob hash (deduplicated)
const hash1 = entity1.metadata.storage?.hash
const hash2 = entity2.metadata.storage?.hash
expect(hash1).toBeDefined()
expect(hash2).toBeDefined()
expect(hash1).toBe(hash2) // Same content = same hash
})
})
describe('File Operations', () => {
it('should update files correctly', async () => {
await vfs.writeFile('/update.txt', 'Original content')
await vfs.writeFile('/update.txt', 'Updated content')
const content = await vfs.readFile('/update.txt')
expect(content.toString()).toBe('Updated content')
})
it('should delete files and decrement blob refs', async () => {
await vfs.writeFile('/delete.txt', 'Delete me')
await vfs.unlink('/delete.txt')
await expect(vfs.readFile('/delete.txt')).rejects.toThrow()
})
it('should append to files', async () => {
await vfs.writeFile('/append.txt', 'First part')
await vfs.appendFile('/append.txt', ' Second part')
const content = await vfs.readFile('/append.txt')
expect(content.toString()).toBe('First part Second part')
})
})
describe('Binary Files', () => {
it('should handle binary files correctly', async () => {
const binary = Buffer.from([0x00, 0xFF, 0xAB, 0xCD, 0xEF])
await vfs.writeFile('/binary.dat', binary)
const read = await vfs.readFile('/binary.dat')
expect(Buffer.compare(read, binary)).toBe(0)
})
it('should preserve binary file integrity', async () => {
// Create a buffer with various byte patterns
const buffer = Buffer.alloc(1000)
for (let i = 0; i < 1000; i++) {
buffer[i] = i % 256
}
await vfs.writeFile('/integrity.bin', buffer)
const read = await vfs.readFile('/integrity.bin')
expect(Buffer.compare(read, buffer)).toBe(0)
expect(read.length).toBe(buffer.length)
})
})
describe('Metadata', () => {
it('should store correct metadata', async () => {
const content = 'Test file'
await vfs.writeFile('/meta.txt', content)
const entity = await vfs.getEntity('/meta.txt')
expect(entity.metadata.size).toBe(content.length)
expect(entity.metadata.vfsType).toBe('file')
expect(entity.metadata.storage?.type).toBe('blob')
expect(entity.metadata.storage?.size).toBe(content.length)
expect(entity.metadata.mimeType).toBeDefined()
})
})
})