feat(8.0): API simplification — remove neural()/Db.search, one storage path key, integration→0

8.0 RC cleanup toward "one place per thing, zero-config, no deprecation":

- Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead
  legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})`
  / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate
  entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor,
  NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept.
- Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams).
  Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` →
  now `find({ query, limit })`.
- Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases
  (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the
  exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver
  feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native
  storage provider resolves the identical root (no split-brain).
- Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now
  applied as a post-filter on `result.score` (the documented way to bound semantic results).
- Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()`
  and failed dimension validation; they are metadata-only updates now.
- Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination
  shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an
  in-place path change that preserves the blob and the entity id, for files and directories.
- Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability.
  Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk
  flushes and emits `progress.queryable`, so imported data is queryable during the import.
- Sweep all docs, comments, and JSDoc for the removed/changed APIs.

Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
This commit is contained in:
David Snelling 2026-06-20 13:31:11 -07:00
parent 0c4a51c24e
commit 606445cd61
74 changed files with 712 additions and 7470 deletions

View file

@ -920,7 +920,7 @@ export class ImportCoordinator {
// Starts at 100, increases to 1000 at 1K entities, then 5000 at 10K
// This works for both known totals (files) and unknown totals (streaming APIs)
let currentFlushInterval = 100 // Start with frequent updates for better UX
let entitiesSinceFlush = 0
let entitiesSinceFlush = 0 // used by the dedup slow path below
let totalFlushes = 0
console.log(
@ -1026,42 +1026,60 @@ export class ImportCoordinator {
}
})
// Batch create all entities (storage-aware batching handles rate limits automatically)
const addResult = await this.brain.addMany({
items: entityParams,
continueOnError: true,
onProgress: (done, total) => {
options.onProgress?.({
stage: 'storing-graph',
message: `Creating entities: ${done}/${total}`,
processed: done,
total,
entities: done
})
}
})
// Map results to entities array and update rows with new IDs
for (let i = 0; i < addResult.successful.length; i++) {
const entityId = addResult.successful[i]
const row = rows[i]
const entity = row.entity || row
const vfsFile = vfsResult.files.find((f: any) => f.entityId === entity.id)
entity.id = entityId
entities.push({
id: entityId,
name: entity.name,
type: entity.type,
vfsPath: vfsFile?.path,
metadata: entity.metadata // Include metadata in return (for ImageHandler, etc)
// Chunked batch creation with PROGRESSIVE FLUSH so imported data becomes
// queryable DURING the import (the always-on streaming contract): after
// each chunk we flush the indexes and emit a `queryable: true` progress
// event. The interval widens with volume (100 → 1000 → 5000) to keep large
// imports fast while small ones stay responsive. (storage-aware batching
// inside addMany still handles rate limits within each chunk.)
let failedCount = 0
for (let offset = 0; offset < entityParams.length; offset += currentFlushInterval) {
const chunk = entityParams.slice(offset, offset + currentFlushInterval)
const addResult = await this.brain.addMany({
items: chunk,
continueOnError: true
})
newCount++
// Map this chunk's results back to their source rows.
for (let j = 0; j < addResult.successful.length; j++) {
const entityId = addResult.successful[j]
const row = rows[offset + j]
const entity = row.entity || row
const vfsFile = vfsResult.files.find((f: any) => f.entityId === entity.id)
entity.id = entityId
entities.push({
id: entityId,
name: entity.name,
type: entity.type,
vfsPath: vfsFile?.path,
metadata: entity.metadata // Include metadata in return (for ImageHandler, etc)
})
newCount++
}
failedCount += addResult.failed.length
// Flush so the just-added chunk is immediately queryable, then signal it.
await this.brain.flush()
totalFlushes++
const processed = Math.min(offset + currentFlushInterval, entityParams.length)
options.onProgress?.({
stage: 'storing-graph',
message: `Creating entities: ${processed}/${entityParams.length}`,
processed,
total: entityParams.length,
entities: entities.length,
queryable: true
})
// Progressive interval: widen as the import grows.
if (entities.length >= 10000) currentFlushInterval = 5000
else if (entities.length >= 1000) currentFlushInterval = 1000
}
// Handle failed entities
if (addResult.failed.length > 0) {
console.warn(`⚠️ ${addResult.failed.length} entities failed to create`)
if (failedCount > 0) {
console.warn(`⚠️ ${failedCount} entities failed to create`)
}
// Create provenance links in batch