- Added support for field-specific and prioritized searches in `brainyData`: - Introduced `searchField` option to enable targeted field-level searches. - Implemented `priorityFields` option for weighted vectorization and query relevance. - Developed utilities in `jsonProcessing.ts` and `fieldNameTracking.ts`: - `extractTextFromJson` for text extraction with customizable depth and field prioritization. - `extractFieldFromJson` to target specific fields in JSON documents. - `prepareJsonForVectorization` for optimized JSON vectorization. - Enhanced management of field names and mappings: - Integrated `trackFieldNames` to associate fields with their services. - Supported cross-service consistency through `standardFieldMappings`. - Updated documentation: - Added detailed guides for JSON search enhancements and HNSW limitations. - Extended usage examples in `README.md` and `json-search-test.js`. - Verified improvements with comprehensive tests: - Created unit and integration tests demonstrating search behavior improvements. - Addressed previous TypeScript errors related to search parameters. **Purpose**: Improve search accuracy and usability when working with complex JSON documents by enabling field-specific searches and enhancing contextual relevance.
2.7 KiB
2.7 KiB
No, HNSW (Hierarchical Navigable Small World) does not natively support field-level search. Here's why:
How HNSW Actually Works
HNSW is a vector similarity search algorithm that operates purely on high-dimensional vectors. It:
- Stores only vectors: Each node in the HNSW graph contains just an ID and a vector (as shown in the
HNSWNouninterface) - Performs proximity search: Finds vectors that are closest in vector space using distance functions like cosine similarity
- Has no concept of fields: The algorithm doesn't understand document structure, field names, or metadata
How Brainy Implements "Field-Level Search"
The field-level search functionality in Brainy is implemented above the HNSW layer, not within HNSW itself:
1. Pre-Processing Approach
- JSON documents are processed by
prepareJsonForVectorization()before being converted to vectors - Field names and values are combined into a text representation
- Priority fields get more weight in the text representation
- The entire processed text is then vectorized into a single 512-dimensional vector
2. Query-Time Processing
- When you search for a specific field like
searchField: "company", the system:- Extracts text from that field using
extractFieldFromJson() - Creates a vector from just that field's content
- Searches the HNSW index using standard vector similarity
- Extracts text from that field using
3. Storage Layer Enhancement
- Field names and mappings are tracked in the storage layer, not in HNSW
- The storage system maintains metadata about available fields
- Standard field mappings are handled outside of the vector index
The Fundamental Limitation
This approach has inherent limitations because:
- Single Vector Per Document: HNSW stores one vector per document, which is a "flattened" representation of all the document's content
- No Structural Awareness: The vector space doesn't preserve field boundaries or hierarchical structure
- Approximation: Field-specific searches are approximations based on how well the original vectorization captured field-specific information
Alternative Approaches for True Field-Level Search
For genuine field-level search, you would typically use:
- Hybrid search systems that combine vector search with traditional indexing
- Multi-vector approaches where each field gets its own vector
- Specialized vector databases that support structured data natively
- Traditional search engines like Elasticsearch for structured queries combined with vector search
The current implementation is a clever workaround that provides field-aware functionality on top of a pure vector similarity engine, but it's not true field-level search in the traditional database sense.