feat(8.0): API simplification — remove neural()/Db.search, one storage path key, integration→0

8.0 RC cleanup toward "one place per thing, zero-config, no deprecation":

- Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead
  legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})`
  / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate
  entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor,
  NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept.
- Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams).
  Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` →
  now `find({ query, limit })`.
- Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases
  (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the
  exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver
  feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native
  storage provider resolves the identical root (no split-brain).
- Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now
  applied as a post-filter on `result.score` (the documented way to bound semantic results).
- Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()`
  and failed dimension validation; they are metadata-only updates now.
- Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination
  shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an
  in-place path change that preserves the blob and the entity id, for files and directories.
- Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability.
  Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk
  flushes and emits `progress.queryable`, so imported data is queryable during the import.
- Sweep all docs, comments, and JSDoc for the removed/changed APIs.

Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
This commit is contained in:
David Snelling 2026-06-20 13:31:11 -07:00
parent 0c4a51c24e
commit 606445cd61
74 changed files with 712 additions and 7470 deletions

View file

@ -429,23 +429,7 @@ fusionResults.forEach(r => {
}
})
// 5. NEURAL API: Automatic clustering
console.log('\n\n🤖 NEURAL API: Automatic Clustering')
const neural = brain.neural()
const clusters = await neural.clusters({
maxClusters: 3,
minClusterSize: 1
})
console.log(`Found ${clusters.length} semantic clusters:`)
clusters.forEach((cluster, i) => {
console.log(`\n Cluster ${i + 1}: ${cluster.label || cluster.id}`)
console.log(` Members: ${cluster.members.length}`)
console.log(` Centroid topics: ${cluster.metadata?.topics?.join(', ') || 'N/A'}`)
})
// 6. SIMILARITY: Find similar documents
// 5. SIMILARITY: Find similar documents
console.log('\n\n🔍 SIMILARITY: Find Similar Documents')
const similarTo = await brain.similar({
to: paper1, // Entity ID of first AI paper
@ -459,19 +443,6 @@ similarTo.forEach(r => {
console.log(` [${r.score.toFixed(3)}] ${r.entity.data?.substring(0, 50)}...`)
})
// 7. OUTLIER DETECTION
console.log('\n\n🚨 OUTLIER DETECTION')
const outliers = await neural.outliers({
method: 'statistical',
threshold: 2.0 // 2 standard deviations
})
console.log(`Found ${outliers.length} outlier documents:`)
outliers.forEach(o => {
const entity = await brain.get(o.id)
console.log(` [Anomaly score: ${o.score.toFixed(3)}] ${entity?.data?.substring(0, 50)}...`)
})
await brain.close()
```
@ -869,7 +840,7 @@ console.log('Initializing production storage...\n')
const brain = new Brainy({
storage: {
type: 'filesystem',
rootDirectory: '/var/lib/brainy'
path: '/var/lib/brainy'
},
// Performance tuning
@ -963,26 +934,6 @@ console.log(` Failed: ${batchResult.failed.length}`)
console.log(` Duration: ${duration}ms`)
console.log(` Throughput: ${(batchResult.successful.length / (duration / 1000)).toFixed(0)} entities/sec`)
// 5. ADVANCED CLUSTERING (Large Dataset)
console.log('\n\n🤖 Clustering 1000 entities...\n')
const neural = brain.neural()
// Use fast clustering for large datasets
const clusters = await neural.clusterFast({
maxClusters: 5
})
console.log(`Found ${clusters.length} clusters:`)
clusters.forEach((cluster, i) => {
console.log(`\n Cluster ${i + 1}: ${cluster.label || cluster.id}`)
console.log(` Size: ${cluster.members.length} members`)
console.log(` Density: ${(cluster.density || 0).toFixed(3)}`)
if (cluster.metadata?.keywords) {
console.log(` Keywords: ${cluster.metadata.keywords.slice(0, 5).join(', ')}`)
}
})
// 6. PRODUCTION STATISTICS
console.log('\n\n📊 Production Statistics:\n')
@ -1040,7 +991,7 @@ console.log('✅ Brain closed cleanly')
console.log('\n\n🎓 Production Deployment Complete!')
console.log('\n📚 Key Production Learnings:')
console.log(' 1. Use filesystem storage and snapshot rootDirectory off-site from your scheduler')
console.log(' 1. Use filesystem storage and snapshot the data directory off-site from your scheduler')
console.log(' 2. Batch operations = 100x faster than individual ops')
console.log(' 3. Metadata query optimization for complex filters')
console.log(' 4. Monitor query performance with explain: true')
@ -1067,7 +1018,7 @@ console.log(' 7. Stream large imports with progress callbacks')
*/15 * * * * gsutil rsync -r /var/lib/brainy gs://my-bucket/brainy-backup
```
Brainy itself never reaches out to an object store. Snapshot `rootDirectory` from your scheduler.
Brainy itself never reaches out to an object store. Snapshot `path` from your scheduler.
#### 3. **Performance Optimization**
@ -1141,7 +1092,7 @@ const brain = new Brainy({ verbose: true })
#### Before Deployment
- [ ] Provision a writable `rootDirectory` on the host
- [ ] Provision a writable `path` on the host
- [ ] Set up an off-site snapshot job (`rclone` / `aws s3 sync` / `gsutil rsync`)
- [ ] Configure caching
- [ ] Test batch operations