brainy/docs/guides/hnsw-field-search.md
David Snelling f86295eab8 **feat(search): enhance JSON document search with field-level filtering and prioritization**
- Added support for field-specific and prioritized searches in `brainyData`:
  - Introduced `searchField` option to enable targeted field-level searches.
  - Implemented `priorityFields` option for weighted vectorization and query relevance.
- Developed utilities in `jsonProcessing.ts` and `fieldNameTracking.ts`:
  - `extractTextFromJson` for text extraction with customizable depth and field prioritization.
  - `extractFieldFromJson` to target specific fields in JSON documents.
  - `prepareJsonForVectorization` for optimized JSON vectorization.
- Enhanced management of field names and mappings:
  - Integrated `trackFieldNames` to associate fields with their services.
  - Supported cross-service consistency through `standardFieldMappings`.
- Updated documentation:
  - Added detailed guides for JSON search enhancements and HNSW limitations.
  - Extended usage examples in `README.md` and `json-search-test.js`.
- Verified improvements with comprehensive tests:
  - Created unit and integration tests demonstrating search behavior improvements.
  - Addressed previous TypeScript errors related to search parameters.

**Purpose**: Improve search accuracy and usability when working with complex JSON documents by enabling field-specific searches and enhancing contextual relevance.
2025-08-01 08:27:39 -07:00

57 lines
2.7 KiB
Markdown

No, HNSW (Hierarchical Navigable Small World) does **not** natively support field-level search. Here's why:
## How HNSW Actually Works
HNSW is a **vector similarity search algorithm** that operates purely on high-dimensional vectors. It:
1. **Stores only vectors**: Each node in the HNSW graph contains just an ID and a vector (as shown in the `HNSWNoun`
interface)
2. **Performs proximity search**: Finds vectors that are closest in vector space using distance functions like cosine
similarity
3. **Has no concept of fields**: The algorithm doesn't understand document structure, field names, or metadata
## How Brainy Implements "Field-Level Search"
The field-level search functionality in Brainy is implemented **above** the HNSW layer, not within HNSW itself:
### 1. **Pre-Processing Approach**
- JSON documents are processed by `prepareJsonForVectorization()` before being converted to vectors
- Field names and values are combined into a text representation
- Priority fields get more weight in the text representation
- The entire processed text is then vectorized into a single 512-dimensional vector
### 2. **Query-Time Processing**
- When you search for a specific field like `searchField: "company"`, the system:
- Extracts text from that field using `extractFieldFromJson()`
- Creates a vector from just that field's content
- Searches the HNSW index using standard vector similarity
### 3. **Storage Layer Enhancement**
- Field names and mappings are tracked in the **storage layer**, not in HNSW
- The storage system maintains metadata about available fields
- Standard field mappings are handled outside of the vector index
## The Fundamental Limitation
This approach has inherent limitations because:
1. **Single Vector Per Document**: HNSW stores one vector per document, which is a "flattened" representation of all the
document's content
2. **No Structural Awareness**: The vector space doesn't preserve field boundaries or hierarchical structure
3. **Approximation**: Field-specific searches are approximations based on how well the original vectorization captured
field-specific information
## Alternative Approaches for True Field-Level Search
For genuine field-level search, you would typically use:
- **Hybrid search systems** that combine vector search with traditional indexing
- **Multi-vector approaches** where each field gets its own vector
- **Specialized vector databases** that support structured data natively
- **Traditional search engines** like Elasticsearch for structured queries combined with vector search
The current implementation is a clever workaround that provides field-aware functionality on top of a pure vector
similarity engine, but it's not true field-level search in the traditional database sense.