- **Core**: - Added `check-database.js` to verify database status and validate search functionality. - Created `fix-dimension-mismatch.js` to handle re-embedding of existing data to resolve dimension mismatch from 3 to 512. - Improved test cases by updating vector operations to support 512 dimensions, replacing previously hardcoded dimensions. - **Migration**: - Developed `DIMENSION_MISMATCH_SUMMARY.md`, detailing the root cause, solution, and preventive strategies for dimension mismatch issues. - Added `production-migration-guide.md` for structured production migration with detailed steps on re-embedding strategies, batching, and error handling. - **Tests**: - Enhanced test coverage with 512-dimensional vector validation. - Introduced helper functions for consistent vector testing behavior and streamlined search test cases. - **Documentation**: - Updated project documentation to highlight the resolution process for dimension mismatches, emphasizing preventive mechanisms such as auto-migration and version tracking. **Purpose**: Address critical dimension mismatch issues caused by embedding changes, restore functionality, and provide a roadmap for robust prevention strategies and migration processes.
4.6 KiB
Dimension Mismatch Issue: Summary and Recommendations
What Happened
The search functionality in Brainy stopped working because of a dimension mismatch between stored vectors and the expected dimensions in the current version of the codebase:
- Previous State: The system was using vectors with 3 dimensions.
- Current State: The system now expects 512-dimensional vectors from the Universal Sentence Encoder.
- Code Change: Recent updates (around July 16, 2025) introduced dimension validation during initialization, which skips vectors with mismatched dimensions.
- Result: During initialization, vectors with 3 dimensions were skipped, resulting in an empty search index and no search results.
Root Cause Analysis
The root cause was identified by examining the codebase:
-
In
brainyData.ts, theinit()method checks if vector dimensions match the expected dimensions (line 400):if (noun.vector.length !== this._dimensions) { console.warn( `Skipping noun ${noun.id} due to dimension mismatch: expected ${this._dimensions}, got ${noun.vector.length}` ) // Skip this noun and continue with the next one return; } -
The default dimension is set to 512 in the constructor (line 200):
this._dimensions = config.dimensions || 512 -
The
UniversalSentenceEncoderclass inembedding.tsproduces 512-dimensional vectors (lines 358-359):// Return a zero vector of appropriate dimension (512 is the default for USE) return new Array(512).fill(0) -
Git history shows that on July 16, 2025, a commit was made that added dimension validation:
Added a `dimensions` property to `BrainyDataConfig` for specifying vector dimensions. Introduced validation for vector dimensions during database creation and insertion to ensure consistency. Enhanced error handling and logging for dimension mismatches.
This indicates that the system previously used 3-dimensional vectors, but after the update, it expects 512-dimensional vectors. The existing data was not migrated, causing the search functionality to break.
Solution Implemented
We created and tested a fix script (fix-dimension-mismatch.js) that:
- Creates a backup of the existing data
- Reads all noun files directly from the filesystem
- For each noun:
- Extracts text from metadata
- Deletes the existing noun
- Re-adds the noun with the same ID but using the current embedding function
- Recreates all verb relationships between the re-embedded nouns
- Verifies that search works by performing a test search
The script successfully fixed the issue by re-embedding all data with the correct dimensions, and search functionality was restored.
Production Recommendations
For production environments, we recommend:
1. Use the Enhanced Migration Script
We've created a comprehensive production migration guide (production-migration-guide.md) that includes:
- Enhanced backup strategies with metadata
- Batching for large datasets
- Robust error handling and recovery
- Progress monitoring and reporting
- A parallel database approach for mission-critical systems
2. Implement Preventive Measures
To prevent similar issues in the future:
- Version Tracking: Add version information to stored vectors
- Auto-Migration: Enhance initialization to automatically re-embed mismatched vectors
- Regular Validation: Implement a database validation process
- Documentation: Document embedding changes in release notes
3. Scheduling and Communication
- Schedule the migration during a maintenance window
- Communicate the change to all stakeholders
- Have a rollback plan in case of issues
- Monitor the system after the migration
Conclusion
The dimension mismatch issue was caused by a change in the embedding function that increased vector dimensions from 3 to 512. The solution is to re-embed all existing data using the current embedding function, which can be done using the provided fix-dimension-mismatch.js script with the enhancements suggested for production environments.
By implementing the preventive measures outlined in the production migration guide, you can avoid similar issues in the future and ensure smoother transitions when embedding functions or vector dimensions change.
Files Created
check-database.js- Script to verify database status and search functionalityfix-dimension-mismatch.js- Script to fix the dimension mismatch issueproduction-migration-guide.md- Comprehensive guide for production migrationDIMENSION_MISMATCH_SUMMARY.md- This summary document