feat(8.0): API simplification — remove neural()/Db.search, one storage path key, integration→0

8.0 RC cleanup toward "one place per thing, zero-config, no deprecation":

- Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead
  legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})`
  / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate
  entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor,
  NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept.
- Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams).
  Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` →
  now `find({ query, limit })`.
- Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases
  (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the
  exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver
  feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native
  storage provider resolves the identical root (no split-brain).
- Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now
  applied as a post-filter on `result.score` (the documented way to bound semantic results).
- Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()`
  and failed dimension validation; they are metadata-only updates now.
- Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination
  shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an
  in-place path change that preserves the blob and the entity id, for files and directories.
- Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability.
  Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk
  flushes and emits `progress.queryable`, so imported data is queryable during the import.
- Sweep all docs, comments, and JSDoc for the removed/changed APIs.

Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
This commit is contained in:
David Snelling 2026-06-20 13:31:11 -07:00
parent 0c4a51c24e
commit 606445cd61
74 changed files with 712 additions and 7470 deletions

View file

@ -1,306 +0,0 @@
# 🧠 **Brainy Clustering Algorithms - Complete Analysis**
## 🎯 **Current State & Capabilities**
### **✅ Existing Infrastructure (Excellent Foundation)**
#### **1. HNSW Hierarchical Clustering**
```typescript
// ALREADY IMPLEMENTED & OPTIMIZED
brain.neural.clusters({ algorithm: 'hierarchical', level: 2 })
```
**How it works:**
- **Leverages HNSW natural hierarchy**: Uses existing index levels as natural cluster boundaries
- **O(n) performance**: Much faster than O(n²) traditional clustering
- **Multi-level granularity**: Higher levels = fewer, broader clusters; Lower levels = more, specific clusters
- **Representative sampling**: Uses HNSW level nodes as natural cluster centers
**Performance characteristics:**
- **Excellent for large datasets** (millions of items)
- **Preserves semantic relationships** from vector space
- **Automatic granularity control** via level parameter
#### **2. Distance-Based Algorithms**
```typescript
// COMPREHENSIVE DISTANCE FUNCTIONS AVAILABLE
euclideanDistance, cosineDistance, manhattanDistance, dotProductDistance
```
**Optimized implementations:**
- **Batch processing**: `calculateDistancesBatch()` with parallelization
- **Multiple metrics**: Choose optimal distance function per use case
- **Performance optimized**: Faster than GPU for small vectors due to no transfer overhead
#### **3. Rich Semantic Taxonomy**
```typescript
// 25+ NOUN TYPES & 35+ VERB TYPES
NounType: Person, Organization, Document, Concept, Event, Media, etc.
VerbType: RelatedTo, Contains, PartOf, Causes, CreatedBy, etc.
```
**Semantic clustering capabilities:**
- **Type-based clustering**: Group by semantic categories
- **Cross-type relationships**: Use verb types to find semantic bridges
- **Hierarchical taxonomies**: Natural clustering within and across types
#### **4. Graph Structure**
```typescript
// VERB RELATIONSHIPS CREATE RICH GRAPH
await brain.relate(sourceId, targetId, VerbType.Causes, { strength: 0.8 })
```
**Graph-based clustering potential:**
- **Connected components**: Find strongly connected groups
- **Community detection**: Use relationship strength for clustering
- **Multi-modal clustering**: Combine graph + vector + taxonomy
## 🚀 **Advanced Clustering Algorithms We Can Implement**
### **1. ✅ Already Implemented: HNSW Hierarchical**
```typescript
// PRODUCTION READY - Uses existing HNSW levels
const clusters = await brain.neural.clusters({
algorithm: 'hierarchical',
level: 2, // Control granularity
maxClusters: 15
})
```
**Performance:** **A+** - O(n) leveraging existing index structure
### **2. 🔥 Semantic Taxonomy Clustering**
```typescript
// IMPLEMENT: Fast type-based clustering with cross-type bridges
const clusters = await brain.neural.clusterByDomain('nounType', {
preserveTypeBoundaries: false, // Allow cross-type clusters
bridgeStrength: 0.8, // Minimum relationship strength for bridges
hybridWeighting: {
taxonomy: 0.4, // 40% weight to type similarity
vector: 0.4, // 40% weight to vector similarity
graph: 0.2 // 20% weight to relationship strength
}
})
```
**Algorithm approach:**
1. **Primary clustering by taxonomy**: Group by NounType/VerbType first
2. **Vector refinement**: Sub-cluster within types using vector similarity
3. **Cross-type bridging**: Find relationships that bridge type boundaries
4. **Weighted fusion**: Combine taxonomy + vector + graph signals
**Performance:** **A+** - O(n log n) - taxonomy grouping is O(n), refinement is HNSW-accelerated
### **3. 🔥 Graph Community Detection**
```typescript
// IMPLEMENT: Relationship-based clustering
const clusters = await brain.neural.clusterByConnections({
algorithm: 'modularity', // or 'louvain', 'leiden'
minCommunitySize: 3,
relationshipWeights: {
[VerbType.Creates]: 1.0,
[VerbType.PartOf]: 0.8,
[VerbType.RelatedTo]: 0.5
}
})
```
**Algorithm approach:**
1. **Build weighted graph**: Use verbs as edges, weights from relationship types + metadata
2. **Community detection**: Apply Louvain or Leiden algorithm for modularity optimization
3. **Semantic enhancement**: Use vector similarity to refine community boundaries
**Performance:** **A** - O(n log n) for sparse graphs, handles millions of relationships efficiently
### **4. 🔥 Multi-Modal Fusion Clustering**
```typescript
// IMPLEMENT: Best of all worlds
const clusters = await brain.neural.clusters({
algorithm: 'multimodal',
signals: {
vector: { weight: 0.5, metric: 'cosine' },
graph: { weight: 0.3, algorithm: 'modularity' },
taxonomy: { weight: 0.2, crossTypeThreshold: 0.8 }
},
fusion: 'weighted_ensemble' // or 'consensus', 'hierarchical'
})
```
**Algorithm approach:**
1. **Independent clustering**: Run HNSW, graph, and taxonomy clustering separately
2. **Consensus building**: Find agreement between different clustering results
3. **Conflict resolution**: Use weighted voting or hierarchical merging for disagreements
4. **Quality optimization**: Iteratively refine based on silhouette scores
**Performance:** **A** - O(n log n) - parallel execution of component algorithms
### **5. 💎 Temporal Pattern Clustering**
```typescript
// IMPLEMENT: Time-aware clustering using existing infrastructure
const clusters = await brain.neural.clusterByTime('createdAt', [
{ start: new Date('2024-01-01'), end: new Date('2024-06-30'), label: 'H1 2024' },
{ start: new Date('2024-07-01'), end: new Date('2024-12-31'), label: 'H2 2024' }
], {
evolution: 'track', // Track how clusters evolve over time
stability: 0.7, // Minimum stability threshold
trendAnalysis: true // Include trend detection
})
```
**Algorithm approach:**
1. **Time window clustering**: Apply HNSW clustering within each time window
2. **Cluster evolution tracking**: Match clusters across time windows using vector similarity
3. **Trend analysis**: Detect growing, shrinking, merging, splitting patterns
4. **Stability scoring**: Measure cluster consistency over time
**Performance:** **A+** - O(k*n log n) where k = number of time windows
### **6. 💎 DBSCAN with Adaptive Parameters**
```typescript
// IMPLEMENT: Density-based clustering with smart parameter selection
const clusters = await brain.neural.clusters({
algorithm: 'dbscan',
autoParams: true, // Automatically select eps and minPts
distanceMetric: 'cosine',
outlierHandling: 'soft' // Soft assignment instead of hard outliers
})
```
**Algorithm approach:**
1. **Adaptive parameter selection**: Use HNSW k-NN distances to estimate optimal eps
2. **Multi-scale analysis**: Run DBSCAN at multiple scales and merge results
3. **Soft outlier assignment**: Assign outliers to nearest clusters with confidence scores
**Performance:** **A** - O(n log n) using HNSW for neighbor queries
## 📊 **Performance Comparison Matrix**
| Algorithm | Time Complexity | Space | Large Scale | Semantic Quality | Graph Aware |
|-----------|----------------|-------|-------------|------------------|-------------|
| **HNSW Hierarchical** | O(n) | O(n) | ✅ Excellent | ✅ Very Good | ❌ No |
| **Taxonomy Fusion** | O(n log n) | O(n) | ✅ Excellent | 🔥 Exceptional | ⚡ Partial |
| **Graph Communities** | O(n log n) | O(e) | ✅ Very Good | ✅ Very Good | 🔥 Exceptional |
| **Multi-Modal** | O(n log n) | O(n) | ✅ Very Good | 🔥 Exceptional | 🔥 Exceptional |
| **Temporal Patterns** | O(k*n log n) | O(n) | ⚡ Good | ✅ Very Good | ⚡ Partial |
| **Adaptive DBSCAN** | O(n log n) | O(n) | ✅ Very Good | ✅ Very Good | ❌ No |
## 🎯 **Specific Improvements Using Existing Capabilities**
### **1. Enhanced HNSW Clustering (Easy Win)**
```typescript
// IMPROVE EXISTING: Add semantic post-processing
private async enhanceHNSWClusters(clusters: SemanticCluster[]): Promise<SemanticCluster[]> {
return Promise.all(clusters.map(async cluster => {
// Get actual metadata for cluster members
const members = await this.brain.getNouns(cluster.members.map(id => ({ id })))
// Analyze semantic characteristics
const semanticProfile = this.analyzeSemanticProfile(members)
// Generate meaningful cluster labels
const label = await this.generateClusterLabel(members, semanticProfile)
// Calculate cluster coherence using multiple signals
const coherence = this.calculateMultiModalCoherence(members)
return {
...cluster,
label,
semanticProfile,
coherence,
quality: coherence.overall
}
}))
}
```
### **2. Intelligent Algorithm Selection**
```typescript
// IMPLEMENT: Smart routing based on data characteristics
private selectOptimalAlgorithm(dataCharacteristics: {
size: number,
dimensionality: number,
graphDensity: number,
typeDistribution: Record<string, number>
}): string {
if (dataCharacteristics.size > 100000) {
return 'hierarchical' // HNSW scales best
}
if (dataCharacteristics.graphDensity > 0.1) {
return 'multimodal' // Rich graph structure
}
if (Object.keys(dataCharacteristics.typeDistribution).length > 10) {
return 'taxonomy' // Diverse semantic types
}
return 'hierarchical' // Safe default
}
```
### **3. Streaming Cluster Updates**
```typescript
// IMPLEMENT: Incremental clustering using existing infrastructure
public async updateClusters(newItems: string[]): Promise<SemanticCluster[]> {
// Use HNSW nearest neighbor for fast cluster assignment
const assignments = await Promise.all(
newItems.map(async itemId => {
const neighbors = await this.brain.neural.neighbors(itemId, { limit: 5 })
return this.assignToNearestCluster(itemId, neighbors, this.existingClusters)
})
)
// Incrementally update cluster centroids and boundaries
return this.updateClusterBoundaries(assignments)
}
```
## 🏆 **Recommended Implementation Priority**
### **🔥 Phase 1: High Impact, Easy Implementation**
1. **Enhanced HNSW Clustering**: Add semantic post-processing to existing algorithm
2. **Taxonomy-Aware Clustering**: Leverage existing NounType/VerbType enums
3. **Intelligent Algorithm Selection**: Route based on data characteristics
### **⚡ Phase 2: Advanced Features**
4. **Graph Community Detection**: Use existing verb relationships
5. **Multi-Modal Fusion**: Combine all signals intelligently
6. **Streaming Updates**: Incremental cluster maintenance
### **💎 Phase 3: Cutting Edge**
7. **Temporal Pattern Analysis**: Track cluster evolution over time
8. **Adaptive DBSCAN**: Dynamic parameter selection
9. **Explainable Clustering**: Generate cluster explanations and reasoning
## 🎯 **Key Advantages of Our Approach**
### **✅ Leverages Existing Infrastructure**
- **HNSW index**: Already optimized for large-scale vector operations
- **Distance functions**: Battle-tested and performance-optimized
- **Semantic taxonomy**: Rich type system with 60+ semantic categories
- **Graph structure**: Relationship network from verb connections
### **✅ Multiple Clustering Paradigms**
- **Vector similarity**: Traditional embedding-based clustering
- **Graph structure**: Relationship-based community detection
- **Semantic taxonomy**: Type-aware intelligent grouping
- **Temporal patterns**: Time-aware cluster evolution
- **Multi-modal fusion**: Best of all worlds
### **✅ Scalability & Performance**
- **O(n) hierarchical clustering**: Leveraging HNSW levels
- **Parallel processing**: Batch distance calculations optimized
- **Streaming support**: Real-time cluster updates
- **Memory efficient**: Existing index structures reused
**Our clustering algorithms are not just competitive - they're architecturally superior by leveraging Brainy's unique multi-modal semantic infrastructure.**

View file

@ -19,7 +19,7 @@ storage keys its internal map by the identical path strings).
## 1. Directory Tree
A real 8.0 filesystem store (`rootDirectory`, default `./brainy-data`):
A real 8.0 filesystem store (`path`, default `./brainy-data`):
```
brainy-data/

View file

@ -71,7 +71,7 @@ Brainy 8.0 ships two adapters, both implementing the same `StorageAdapter` inter
const brain = new Brainy({
storage: {
type: 'filesystem',
rootDirectory: './data'
path: './data'
}
})
```
@ -99,15 +99,15 @@ const brain = new Brainy({
const brain = new Brainy({
storage: {
type: 'auto',
rootDirectory: './data'
path: './data'
}
})
```
`'auto'` picks `'filesystem'` when running on Node.js with a writable `rootDirectory`, and falls back to `'memory'` otherwise.
`'auto'` picks `'filesystem'` when running on Node.js with a writable `path`, and falls back to `'memory'` otherwise.
## Backup and Off-Site Replication
Brainy 8.0 does not embed cloud SDKs. The on-disk artifact at `rootDirectory` is a plain directory tree of JSON files, so backup is an operator-layer concern. Typical patterns:
Brainy 8.0 does not embed cloud SDKs. The on-disk artifact at `path` is a plain directory tree of JSON files, so backup is an operator-layer concern. Typical patterns:
- `gsutil rsync -r ./data gs://my-bucket/brainy-data`
- `aws s3 sync ./data s3://my-bucket/brainy-data`
@ -180,7 +180,7 @@ High-performance deduplication system for streaming data:
## Durability
Brainy persists writes to disk through the filesystem adapter. Each save is a rename-based atomic write of a JSON file under the appropriate shard. Operators that need point-in-time recovery should snapshot `rootDirectory` (see [Backup and Off-Site Replication](#backup-and-off-site-replication)).
Brainy persists writes to disk through the filesystem adapter. Each save is a rename-based atomic write of a JSON file under the appropriate shard. Operators that need point-in-time recovery should snapshot `path` (see [Backup and Off-Site Replication](#backup-and-off-site-replication)).
## Storage Optimization
@ -210,7 +210,7 @@ await brain.addBatch([
const brain = new Brainy({
storage: {
type: 'filesystem',
rootDirectory: './data',
path: './data',
cache: {
enabled: true,
maxSize: 1000, // Maximum cached items
@ -260,7 +260,7 @@ await brain.restore('/backups/2026-06-11', { confirm: true })
### Move to a new directory
```typescript
// A snapshot directory is a complete store: restore it into a fresh brain
const brain = new Brainy({ storage: { type: 'filesystem', rootDirectory: './new' } })
const brain = new Brainy({ storage: { type: 'filesystem', path: './new' } })
await brain.init()
await brain.restore('/backups/2026-06-11', { confirm: true })
```
@ -298,7 +298,7 @@ console.log(stats)
1. **Read-heavy**: Enable caching and let the OS page cache do its job
2. **Write-heavy**: Batch operations and tune the cache `maxSize`
3. **Real-time**: FileSystem with periodic snapshots
4. **Archival**: Snapshot `rootDirectory` to cold object storage on a schedule
4. **Archival**: Snapshot `path` to cold object storage on a schedule
5. **Large-scale**: Rely on metadata/vector separation + UUID sharding
### Monitor and Maintain

View file

@ -68,7 +68,7 @@ const articles = await brain.find("verified articles by John Smith about machine
#### Simple Vector Search
```typescript
const results = await brain.search("machine learning concepts")
const results = await brain.find("machine learning concepts")
```
#### Combined Intelligence Query
@ -161,7 +161,7 @@ Brainy includes 220+ embedded patterns for natural language understanding:
```typescript
// Natural language automatically parsed
const results = await brain.search(
const results = await brain.find(
"show me recent AI papers from Stanford published this year"
)
// Automatically converts to:
@ -189,10 +189,10 @@ The NLP processor identifies query intent:
Successful execution plans are cached:
```typescript
// First query: 50ms (plan generation + execution)
await brain.search("machine learning papers")
await brain.find("machine learning papers")
// Subsequent similar queries: 10ms (cached plan)
await brain.search("deep learning papers")
await brain.find("deep learning papers")
```
### Self-Optimization

View file

@ -65,7 +65,7 @@ With no `storage` option, Brainy uses `type: 'auto'`:
```typescript
// Explicit override when you want a specific root
const brain = new Brainy({
storage: { type: 'filesystem', rootDirectory: './brainy-data' }
storage: { type: 'filesystem', path: './brainy-data' }
})
```
@ -135,7 +135,7 @@ explicit constructor option:
```typescript
const brain = new Brainy({
storage: { type: 'filesystem', rootDirectory: '/var/lib/brainy' },
storage: { type: 'filesystem', path: '/var/lib/brainy' },
vector: {
recall: 'accurate',
persistMode: 'immediate'