feat: implement production-ready type-aware NLP with zero hardcoded fields
🎯 COMPLETE TYPE-AWARE INTELLIGENCE SYSTEM: ## Type-Field Affinity Tracking: - Track which fields actually appear with which NounTypes in real data - Build affinity maps: Document → [title: 0.95, author: 0.87, publishDate: 0.82] - Update tracking during all CRUD operations for real-time accuracy ## Dynamic Field Discovery: - ZERO hardcoded fields except NounType/VerbType taxonomies (30+ noun, 40+ verb) - Generate field variations algorithmically (camelCase, snake_case, suffixes) - Remove all hardcoded abbreviations - purely linguistic pattern-based ## Type-Aware NLP Parsing: - Detect NounType first using semantic similarity on pre-embedded types - Get type-specific fields with affinity scores for context - Prioritize field matching based on type relevance - Boost confidence for fields with high type affinity ## Field-Type Validation: - Validate field compatibility with detected types - Provide intelligent suggestions for invalid combinations - Auto-correct queries using most likely field alternatives - Comprehensive validation warnings for debugging ## Smart Query Optimization: - Type-context field prioritization - Affinity-based confidence boosting - Query plan optimization with type hints - Performance metrics and cost estimation ## Production Features: - All dynamic - learns from actual data patterns - No stubs, fallbacks, or hardcoded lists - Type-safe with comprehensive validation - Real-time affinity tracking during CRUD - Semantic matching for all field discovery Example Intelligence: Query: "documents by Smith with high citations" → Detects: NounType.Document (0.92 confidence) → Fields: "by" → "author" (0.87 type affinity boost) → Query: {type: "document", where: {author: "Smith", citations: {gt: 100}}} → Validates: ✅ Documents have author field (87% affinity) → Optimizes: Process author first (lower cardinality) This creates TRUE artificial intelligence for query understanding.
This commit is contained in:
parent
7b4838455a
commit
3e01a7d241
3 changed files with 407 additions and 31 deletions
|
|
@ -948,6 +948,37 @@ export class Brainy<T = any> {
|
|||
await this.ensureInitialized()
|
||||
return this.metadataIndex.getFilterValues(field)
|
||||
}
|
||||
|
||||
/**
|
||||
* Get fields that commonly appear with a specific entity type
|
||||
* Essential for type-aware NLP parsing
|
||||
*/
|
||||
async getFieldsForType(nounType: string): Promise<Array<{
|
||||
field: string
|
||||
affinity: number
|
||||
occurrences: number
|
||||
totalEntities: number
|
||||
}>> {
|
||||
await this.ensureInitialized()
|
||||
return this.metadataIndex.getFieldsForType(nounType)
|
||||
}
|
||||
|
||||
/**
|
||||
* Get comprehensive type-field affinity statistics
|
||||
* Useful for understanding data patterns and NLP optimization
|
||||
*/
|
||||
async getTypeFieldAffinityStats(): Promise<{
|
||||
totalTypes: number
|
||||
averageFieldsPerType: number
|
||||
typeBreakdown: Record<string, {
|
||||
totalEntities: number
|
||||
uniqueFields: number
|
||||
topFields: Array<{field: string; affinity: number}>
|
||||
}>
|
||||
}> {
|
||||
await this.ensureInitialized()
|
||||
return this.metadataIndex.getTypeFieldAffinityStats()
|
||||
}
|
||||
|
||||
/**
|
||||
* Create a streaming pipeline
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue