feat: expose neural entity extraction APIs (v5.7.6 - Workshop request)
Addresses Workshop team's request for direct access to neural extraction classes.
**Changes:**
1. **New Exports** (src/index.ts):
- `NeuralEntityExtractor` - Full extraction orchestrator
- `SmartExtractor` - Entity type classifier (4-signal ensemble)
- `SmartRelationshipExtractor` - Relationship type classifier
- Types: `ExtractedEntity`, `ExtractionResult`, `RelationshipExtractionResult`, etc.
2. **Package.json Subpath Exports**:
```typescript
// Enable direct imports:
import { NeuralEntityExtractor } from '@soulcraft/brainy/neural/entityExtractor'
import { SmartExtractor } from '@soulcraft/brainy/neural/SmartExtractor'
import { SmartRelationshipExtractor } from '@soulcraft/brainy/neural/SmartRelationshipExtractor'
```
3. **New brain.extractEntities() Method** (brainy.ts:3254):
- Alias for `brain.extract()` with clearer naming
- Documented with examples and architecture details
- 4-signal ensemble: ExactMatch (40%) + Embedding (35%) + Pattern (20%) + Context (5%)
4. **Comprehensive Documentation** (docs/neural-extraction.md):
- Complete neural extraction guide (200+ lines)
- API reference for all extraction classes
- Performance optimization tips
- Import preview mode documentation
- Confidence scoring explanation
- 42 NounType detection methods
- Troubleshooting guide
- Real-world examples
5. **README Updates**:
- Added "Entity Extraction" section with examples
- Links to neural extraction guide
- Import preview mode link
**Features:**
- ⚡ Fast extraction: ~15-20ms per entity
- 🎯 4-signal ensemble architecture
- 📊 Format intelligence (Excel, CSV, PDF, YAML, DOCX, JSON, Markdown)
- 🌍 42 universal noun types + 127 verb types
- 💾 LRU caching built-in
- 🧪 Production-tested in import pipeline
**Usage:**
```typescript
// Simple API (recommended)
const entities = await brain.extractEntities('John Smith founded Acme Corp', {
types: [NounType.Person, NounType.Organization],
confidence: 0.7
})
// Advanced API (custom configuration)
import { SmartExtractor } from '@soulcraft/brainy'
const extractor = new SmartExtractor(brain, { minConfidence: 0.8 })
const result = await extractor.extract('CEO', {
formatContext: { format: 'excel', columnHeader: 'Title' }
})
```
**Backward Compatible:** All existing APIs unchanged. New exports are pure additions.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
parent
201fbed78c
commit
8cca096d7e
5 changed files with 794 additions and 0 deletions
44
README.md
44
README.md
|
|
@ -135,6 +135,50 @@ const results = await brain.find({
|
|||
|
||||
---
|
||||
|
||||
## Entity Extraction (NEW in v5.7.6)
|
||||
|
||||
**Extract entities from text with AI-powered classification:**
|
||||
|
||||
```javascript
|
||||
import { Brainy, NounType } from '@soulcraft/brainy'
|
||||
|
||||
const brain = new Brainy()
|
||||
await brain.init()
|
||||
|
||||
// Extract all entities
|
||||
const entities = await brain.extractEntities('John Smith founded Acme Corp in New York')
|
||||
// Returns:
|
||||
// [
|
||||
// { text: 'John Smith', type: NounType.Person, confidence: 0.95 },
|
||||
// { text: 'Acme Corp', type: NounType.Organization, confidence: 0.92 },
|
||||
// { text: 'New York', type: NounType.Location, confidence: 0.88 }
|
||||
// ]
|
||||
|
||||
// Extract with filters
|
||||
const people = await brain.extractEntities(resume, {
|
||||
types: [NounType.Person],
|
||||
confidence: 0.8
|
||||
})
|
||||
|
||||
// Advanced: Direct access to extractors
|
||||
import { SmartExtractor } from '@soulcraft/brainy'
|
||||
|
||||
const extractor = new SmartExtractor(brain, { minConfidence: 0.7 })
|
||||
const result = await extractor.extract('CEO', {
|
||||
formatContext: { format: 'excel', columnHeader: 'Title' }
|
||||
})
|
||||
```
|
||||
|
||||
**Features:**
|
||||
- 🎯 **4-Signal Ensemble** - ExactMatch (40%) + Embedding (35%) + Pattern (20%) + Context (5%)
|
||||
- 📊 **Format Intelligence** - Adapts to Excel, CSV, PDF, YAML, DOCX, JSON, Markdown
|
||||
- ⚡ **Fast** - ~15-20ms per extraction with LRU caching
|
||||
- 🌍 **42 Types** - Person, Organization, Location, Document, and 38 more
|
||||
|
||||
**→ [Neural Extraction Guide](docs/neural-extraction.md)** | **[Import Preview Mode](docs/neural-extraction.md#import-preview-mode)**
|
||||
|
||||
---
|
||||
|
||||
## From Prototype to Planet Scale
|
||||
|
||||
**The same API. Zero rewrites. Any scale.**
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue