fix: exclude __words__ keyword index from corruption detection and getStats()
The __words__ keyword index stores 50-5000 entries per entity (one per word), which inflated avg entries/entity well above the corruption threshold of 100. This caused: 1. validateConsistency() to falsely detect corruption on every startup, triggering unnecessary clearAllIndexData() + rebuild() cycles 2. getStats() to log false "Metadata index may be corrupted" warnings and report inflated totalEntries/totalIds stats Both methods now skip __words__ when counting, so stats and health checks reflect metadata fields only (noun, type, createdAt, etc.). Keyword search is unaffected since the __words__ field index itself is not modified.
This commit is contained in:
parent
32dbdcec61
commit
364360d447
128 changed files with 5637 additions and 5682 deletions
|
|
@ -1,5 +1,5 @@
|
|||
/**
|
||||
* Unified Index Interface (v3.35.0+)
|
||||
* Unified Index Interface
|
||||
*
|
||||
* Standardizes index lifecycle across all index types in Brainy.
|
||||
* All indexes (HNSW Vector, Graph Adjacency, Metadata Field) implement this interface
|
||||
|
|
@ -57,7 +57,7 @@ export interface RebuildOptions {
|
|||
* - On-demand: Large datasets loaded adaptively via UnifiedCache
|
||||
*
|
||||
* This option is kept for backwards compatibility but is ignored.
|
||||
* The system always uses adaptive caching (v3.36.0+).
|
||||
* The system always uses adaptive caching.
|
||||
*/
|
||||
lazy?: boolean
|
||||
|
||||
|
|
@ -104,7 +104,7 @@ export interface IIndex {
|
|||
* - Provide progress reporting for large datasets
|
||||
* - Recover gracefully from partial failures
|
||||
*
|
||||
* Adaptive Caching (v3.36.0+):
|
||||
* Adaptive Caching:
|
||||
* System automatically chooses optimal strategy:
|
||||
* - Small datasets: Preload all data at init for zero-latency access
|
||||
* - Large datasets: Load on-demand via UnifiedCache for memory efficiency
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue