# Neural Entity Extraction Guide **Version:** 5.7.6+ **Status:** Production-Ready **Performance:** ~15-20ms per extraction --- ## Overview Brainy's neural extraction system uses a **4-signal ensemble architecture** to classify entities and relationships with high accuracy. The system is production-tested and handles 7 different document formats with format-specific intelligence. ### Key Components 1. **`brain.extractEntities()`** - Simplest API, use for 95% of cases 2. **`SmartExtractor`** - Direct entity type classifier (advanced) 3. **`SmartRelationshipExtractor`** - Relationship type classifier 4. **`NeuralEntityExtractor`** - Full extraction orchestrator --- ## Quick Start ### Method 1: Brain Instance (Recommended) ```typescript import { Brainy, NounType } from '@soulcraft/brainy' const brain = new Brainy() await brain.init() // Extract all entities const entities = await brain.extractEntities('Sarah Chen founded Acme Corp') // Returns (fast pattern + embedding ensemble; confidences are approximate): // [ // { text: 'Sarah Chen', type: NounType.Person, confidence: 0.68 }, // { text: 'Acme Corp', type: NounType.Organization, confidence: 0.85 } // ] // Extract with filters const people = await brain.extractEntities('...', { types: [NounType.Person], confidence: 0.8, neuralMatching: true }) ``` > **What this is (and isn't).** `extractEntities` is a fast, dependency-free **heuristic > ensemble** (regex/pattern + type-embedding similarity + context), not a trained NER model. > Each candidate is typed by its own span — a type indicator in a neighbour's text (e.g. "Corp" > in "Sarah Chen founded Acme Corp") will **not** bleed onto another candidate. > > **Known limitation:** a bare proper-noun place name with no structural cue (e.g. "New York", > no comma, state code, or "in"/"at" preposition handling) can be typed as `Person` by the > generic full-name pattern. For high-precision typing, pass `types` to constrain results, supply > richer context, or post-validate. For state-of-the-art NER, drive extraction from an LLM at the > application layer and store the results in Brainy. ### Method 2: Direct Import (Advanced) ```typescript import { SmartExtractor, SmartRelationshipExtractor } from '@soulcraft/brainy' // Or use subpath imports: import { SmartExtractor } from '@soulcraft/brainy/neural/SmartExtractor' const brain = new Brainy() await brain.init() const extractor = new SmartExtractor(brain, { minConfidence: 0.7 }) const result = await extractor.extract('CEO') // { type: NounType.Role, confidence: 0.89, source: 'ensemble', evidence: '...' } ``` --- ## Architecture ### 4-Signal Ensemble Both `SmartExtractor` and `SmartRelationshipExtractor` use parallel signal execution: | Signal | Weight | Description | Speed | |--------|--------|-------------|-------| | **ExactMatch** | 40% | Dictionary lookups, aliases | ~1ms | | **Embedding** | 35% | Semantic similarity (384-dim vectors) | ~8ms | | **Pattern** | 20% | Regex patterns, format-aware | ~2ms | | **Context** | 5% | Surrounding text hints | ~4ms | **Total Execution Time:** ~15-20ms (parallel) ### Format Intelligence The system adapts to 7 document formats: ```typescript await extractor.extract('CEO', { formatContext: { format: 'excel', columnHeader: 'Job Title', // Boosts Role type columnIndex: 3, adjacentHeaders: ['Name', 'Department'] } }) ``` **Supported Formats:** - Excel - Column headers, position intelligence - CSV - Same as Excel - PDF - Page structure, section detection - YAML - Key paths, nesting levels - DOCX - Styles, headings, lists - JSON - Key paths, schema hints - Markdown - Headers, lists, links --- ## API Reference ### brain.extractEntities() **Primary extraction method.** Handles candidate detection + classification automatically. ```typescript async brain.extractEntities( text: string, options?: { types?: NounType[] // Filter by types confidence?: number // Min confidence (0-1) includeVectors?: boolean // Add embeddings to results neuralMatching?: boolean // Enable ensemble scoring } ): Promise ``` **Returns:** ```typescript interface ExtractedEntity { text: string // Original text type: NounType // Classified type confidence: number // Score (0-1) start?: number // Character offset end?: number // Character offset vector?: number[] // 384-dim embedding (if requested) } ``` **Examples:** ```typescript // Extract all entities const all = await brain.extractEntities(markdown) // Extract only people const people = await brain.extractEntities(text, { types: [NounType.Person] }) // High-confidence only const highConf = await brain.extractEntities(text, { confidence: 0.9 }) // Include vectors for similarity const withVectors = await brain.extractEntities(text, { includeVectors: true }) ``` --- ### SmartExtractor **Direct entity type classifier.** Use when you have pre-detected candidates or need custom configuration. ```typescript import { SmartExtractor, FormatContext } from '@soulcraft/brainy' const extractor = new SmartExtractor(brain, { minConfidence: 0.7, // Threshold enableEnsemble: true, // Use all signals enableExactMatch: true, // Dictionary lookups enableEmbedding: true, // Semantic similarity enablePattern: true, // Regex patterns enableContext: true, // Context hints weights: { // Custom signal weights exactMatch: 0.5, // 50% embedding: 0.3, // 30% pattern: 0.15, // 15% context: 0.05 // 5% } }) // Extract entity type const result = await extractor.extract( 'CEO', // Candidate text { formatContext: { // Optional format hints format: 'excel', columnHeader: 'Title' }, contextWindow: 'John is the CEO' // Optional context } ) ``` **Returns:** ```typescript interface ExtractionResult { type: NounType // Classified type confidence: number // Score (0-1) source: string // 'exact-match' | 'embedding' | 'ensemble' evidence: string // Human-readable explanation signalScores?: { // Individual signal scores exactMatch?: number embedding?: number pattern?: number context?: number } } ``` --- ### SmartRelationshipExtractor **Relationship type classifier.** Determines verb/relationship types between entities. ```typescript import { SmartRelationshipExtractor } from '@soulcraft/brainy' const relExtractor = new SmartRelationshipExtractor(brain, { minConfidence: 0.6, enableEnsemble: true }) // Infer relationship type const relationship = await relExtractor.infer( 'Alice', // Subject 'UCSF', // Object 'Alice works as a researcher at UCSF', // Context { subjectType: NounType.Person, // Optional type hints objectType: NounType.Organization } ) ``` **Returns:** ```typescript interface RelationshipExtractionResult { type: VerbType // Classified relationship confidence: number // Score (0-1) source: string // Signal source evidence: string // Explanation signalScores?: { // Individual scores exactMatch?: number embedding?: number pattern?: number context?: number } } ``` **Example:** ```typescript const rel = await relExtractor.infer( 'John', 'Acme Corp', 'John Smith is the CEO of Acme Corp' ) // { // type: VerbType.WorksFor, // confidence: 0.87, // source: 'ensemble', // evidence: 'Ensemble: exact-match (CEO pattern) + embedding (0.89 similarity)' // } ``` --- ### NeuralEntityExtractor **Full extraction orchestrator.** Handles candidate detection, classification, and deduplication. ```typescript import { NeuralEntityExtractor } from '@soulcraft/brainy' const extractor = new NeuralEntityExtractor(brain) // Full pipeline const entities = await extractor.extract(text, { types: [NounType.Person, NounType.Organization], confidence: 0.7, neuralMatching: true }) ``` **When to use:** - Need automatic candidate detection - Want deduplication ("John Smith" and "Smith" → same entity) - Building custom extraction pipelines - Advanced configuration requirements **Typically accessed via `brain.extractEntities()` instead.** --- ## Import Preview Mode Extract entities **without persisting** them to the database: ```typescript // Method 1: Using import() with preview mode const result = await brain.import(markdownContent, { format: 'markdown', enableNeuralExtraction: true, // Enable extraction enableConceptExtraction: true, // Enable concepts createEntities: false, // DON'T persist to database vfsPath: null, // DON'T create VFS structure returnExtracted: true // Return extracted data }) // Access extracted entities const entities = result.extractedEntities // [ // { text: 'John Smith', type: NounType.Person, confidence: 0.95 }, // ... // ] // Method 2: Direct extraction (simpler) const entities = await brain.extractEntities(markdownContent, { confidence: 0.7 }) ``` **Preview Mode Options:** | Option | Effect | |--------|--------| | `createEntities: false` | Don't add to database | | `vfsPath: null` | Don't create VFS files/folders | | `returnExtracted: true` | Include extraction results | | `enableNeuralExtraction: true` | Run entity extraction | | `enableConceptExtraction: true` | Extract concepts/tags | --- ## Confidence Scoring ### How Confidence is Calculated **Ensemble Mode (default):** ``` confidence = ( exactMatch × 0.40 + embedding × 0.35 + pattern × 0.20 + context × 0.05 ) ``` **Signal Scores:** - **ExactMatch:** 0.0 (no match) or 1.0 (exact match) - **Embedding:** Cosine similarity (0.0-1.0) - **Pattern:** Pattern match confidence (0.5-1.0) - **Context:** Context relevance (0.0-1.0) **Example Calculation:** ``` Text: "CEO" ExactMatch: 1.0 (in dictionary) Embedding: 0.89 (similar to "Role") Pattern: 0.8 (job title pattern) Context: 0.3 (mentioned near name) Final: 1.0×0.40 + 0.89×0.35 + 0.8×0.20 + 0.3×0.05 = 0.40 + 0.3115 + 0.16 + 0.015 = 0.8865 (87% confidence) ``` ### Recommended Thresholds | Use Case | Threshold | Precision | Recall | |----------|-----------|-----------|--------| | High precision | 0.9 | 95% | 70% | | Balanced | 0.7 | 85% | 85% | | High recall | 0.5 | 75% | 95% | --- ## NounType Detection ### 42 Universal Types Brainy supports 42 noun types covering 95% of all domains: **Core 7:** - `Person` - Human individuals - `Organization` - Companies, institutions - `Location` - Places, addresses - `Thing` - Physical objects - `Concept` - Abstract ideas - `Event` - Occurrences, meetings - `Agent` - Software, bots, AI **Extended 35:** - `Document`, `Media`, `File` - Content types - `Message`, `Collection`, `Dataset` - Data structures - `Product`, `Service` - Commercial - `Task`, `Project`, `Process` - Work - `State`, `Role`, `Language` - Properties - `Currency`, `Measurement` - Quantitative - `Hypothesis`, `Experiment` - Scientific - `Contract`, `Regulation` - Legal - ... [see types/graphTypes.ts for complete list] ### Type Detection Methods **1. ExactMatch Signal** (40% weight) - Dictionary: 10,000+ aliases per type - Examples: "CEO" → Role, "USD" → Currency - Speed: ~1ms **2. Embedding Signal** (35% weight) - Semantic similarity to type embeddings - 384-dimensional vectors - Examples: "Chief Executive" → Role (cosine: 0.92) - Speed: ~8ms **3. Pattern Signal** (20% weight) - Regex patterns for each type - Format-aware (email → Message, URL → Document) - Examples: `\d{4}-\d{2}-\d{2}` → Event - Speed: ~2ms **4. Context Signal** (5% weight) - Surrounding word patterns - Examples: "works at [X]" → X is Organization - Speed: ~4ms --- ## Performance Optimization ### Caching All extractors use LRU caching: ```typescript const extractor = new SmartExtractor(brain, { cache: { maxSize: 10000, // Max cached items ttl: 3600000 // 1 hour TTL } }) // Cache stats const stats = extractor.getCacheStats() // { hits: 8432, misses: 1568, hitRate: 0.843 } ``` ### Batch Processing ```typescript // Process multiple candidates in parallel const candidates = ['CEO', 'Alice', 'Acme Corp', 'New York'] const results = await Promise.all( candidates.map(text => extractor.extract(text)) ) ``` ### Format Context Reuse ```typescript const formatContext = { format: 'excel' as const, columnHeader: 'Title' } // Reuse context for entire column for (const cell of column) { await extractor.extract(cell, { formatContext }) } ``` --- ## Advanced: Custom Signal Weights Adjust weights for domain-specific extraction: ```typescript // Medical domain: Boost pattern matching const medicalExtractor = new SmartExtractor(brain, { weights: { exactMatch: 0.3, embedding: 0.2, pattern: 0.45, // High for medical codes context: 0.05 } }) // Legal domain: Boost exact matching const legalExtractor = new SmartExtractor(brain, { weights: { exactMatch: 0.6, // High for legal terms embedding: 0.25, pattern: 0.10, context: 0.05 } }) ``` --- ## Troubleshooting ### Low Confidence Scores **Problem:** Entities extracted with confidence <0.5 **Solutions:** 1. Add format context hints 2. Provide more surrounding context 3. Lower confidence threshold 4. Add domain-specific aliases ```typescript // Before: Generic extraction const result = await extractor.extract('PM') // { confidence: 0.45 } // After: With context const result = await extractor.extract('PM', { formatContext: { format: 'excel', columnHeader: 'Job Title' }, contextWindow: 'The PM leads the project team' }) // { confidence: 0.87 } ``` ### Type Misclassification **Problem:** "John Smith" classified as Organization instead of Person **Solutions:** 1. Provide type hints 2. Add more context 3. Check for name patterns ```typescript // Force type filtering const people = await brain.extractEntities(text, { types: [NounType.Person] // Only consider Person type }) ``` ### Slow Extraction **Problem:** Extraction taking >100ms **Solutions:** 1. Enable caching 2. Reduce context window 3. Disable unused signals 4. Use batch processing ```typescript const fastExtractor = new SmartExtractor(brain, { enableContext: false, // Disable slowest signal cache: { maxSize: 50000 } // Large cache }) ``` --- ## Examples ### Example 1: PDF Resume Extraction ```typescript const resume = ` John Smith Senior Software Engineer Acme Corp (2020-2024) Skills: Python, TypeScript, React Location: San Francisco, CA ` const entities = await brain.extractEntities(resume, { types: [NounType.Person, NounType.Organization, NounType.Location, NounType.Role], confidence: 0.7 }) // Filter by type const person = entities.find(e => e.type === NounType.Person) const companies = entities.filter(e => e.type === NounType.Organization) const locations = entities.filter(e => e.type === NounType.Location) ``` ### Example 2: Excel Data Classification ```typescript import { SmartExtractor } from '@soulcraft/brainy' const extractor = new SmartExtractor(brain) // Process Excel column const results = [] for (let i = 0; i < cells.length; i++) { const result = await extractor.extract(cells[i], { formatContext: { format: 'excel', columnHeader: headers[columnIndex], columnIndex, rowIndex: i } }) results.push(result) } ``` ### Example 3: Relationship Extraction ```typescript import { SmartRelationshipExtractor } from '@soulcraft/brainy' const relExtractor = new SmartRelationshipExtractor(brain) const text = 'Alice works as a researcher at UCSF' // Extract relationship const rel = await relExtractor.infer('Alice', 'UCSF', text, { subjectType: NounType.Person, objectType: NounType.Organization }) // Create relationship in brain if (rel.confidence > 0.7) { await brain.relate({ from: aliceId, to: ucsfId, type: rel.type, metadata: { confidence: rel.confidence } }) } ``` --- ## Best Practices 1. **Use `brain.extractEntities()` for 95% of cases** - Handles everything automatically - Optimal for general use 2. **Use direct extractors for:** - Custom signal weights - Format-specific extraction - Batch processing optimization 3. **Always provide context when possible** - Improves confidence by 10-20% - Especially important for ambiguous terms 4. **Enable caching for production** - 80-90% cache hit rate typical - 10x speedup for repeated extractions 5. **Filter by types when you know the domain** - Reduces false positives - Improves performance 6. **Monitor confidence distributions** - Adjust thresholds per use case - Balance precision vs recall --- ## See Also - [API Reference](./api/README.md) - [Type System](./types/README.md) - [Import System](./import/README.md) - [VFS System](./vfs/README.md) --- **Questions or Issues?** https://github.com/soulcraftlabs/brainy/issues