With plugins unset (the default), init() probes for the first-party accelerator: not installed means plain brainy with zero noise; installed and healthy means it loads and announces itself; installed but broken (import failure, invalid shape, failed activation, version mismatch) makes init() THROW. An installed accelerator never silently vanishes behind the JS engines - the anti-drift posture of the explicit list, applied to detection. plugins: []/false stays a true opt-out (no probe); an explicit list keeps its required-and-loud semantics. The import runs through an importPluginPackage seam (variable specifier - bundlers cannot static-resolve the optional package; tests simulate all outcomes without it installed). Also: when a plugin activates but registers zero native providers (e.g. a licensing gate declining to engage), the provider summary now warns loudly instead of leaving every query silently on the JS engines. Supersedes 8.0.8's explicit-opt-in wording; README/PLUGINS/types now document the guarded contract. 7 new tests (tests/unit/plugin-autodetect.test.ts).
217 lines
11 KiB
Markdown
217 lines
11 KiB
Markdown
<p align="center">
|
||
<img src="https://raw.githubusercontent.com/soulcraftlabs/brainy/main/brainy.png" alt="Brainy" width="180">
|
||
</p>
|
||
|
||
<h1 align="center">Brainy</h1>
|
||
|
||
<p align="center">
|
||
<b>Three database paradigms. One API. Zero configuration.</b><br>
|
||
The in-process knowledge database for TypeScript — vector search, graph traversal,<br>
|
||
and metadata filtering unified in a single query.
|
||
</p>
|
||
|
||
<p align="center">
|
||
<a href="https://www.npmjs.com/package/@soulcraft/brainy"><img src="https://img.shields.io/npm/v/@soulcraft/brainy.svg" alt="npm version"></a>
|
||
<a href="https://www.npmjs.com/package/@soulcraft/brainy"><img src="https://img.shields.io/npm/dm/@soulcraft/brainy.svg" alt="npm downloads"></a>
|
||
<a href="https://github.com/soulcraftlabs/brainy/actions/workflows/ci.yml"><img src="https://github.com/soulcraftlabs/brainy/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
|
||
<a href="https://soulcraft.com/docs"><img src="https://img.shields.io/badge/docs-soulcraft.com-blue.svg" alt="Documentation"></a>
|
||
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="MIT License"></a>
|
||
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/%3C%2F%3E-TypeScript-%230074c1.svg" alt="TypeScript"></a>
|
||
</p>
|
||
|
||
<p align="center">
|
||
<a href="#quick-start">Quick start</a> ·
|
||
<a href="#one-query-three-engines">One query</a> ·
|
||
<a href="#feature-tour">Features</a> ·
|
||
<a href="#from-laptop-to-hundreds-of-millions">Scale with Cor</a> ·
|
||
<a href="#documentation">Docs</a>
|
||
</p>
|
||
|
||
---
|
||
|
||
Built because we were tired of stitching a vector store to a graph database to a document store — and spending weeks on plumbing before writing a line of business logic. Brainy indexes every fact **three ways at once** and lets one call query them together:
|
||
|
||
| You write | Brainy indexes it as | You query it with |
|
||
|---|---|---|
|
||
| `data: 'Ada wrote the first program'` | a **384-dim vector** (local embedding — no API key) | `find({ query: 'computing pioneers' })` |
|
||
| `metadata: { field: 'CS', year: 1843 }` | **structured fields** (O(1) exact, O(log n) range) | `find({ where: { year: { lessThan: 1900 } } })` |
|
||
| `relate({ from: ada, to: babbage })` | a **typed, directed graph edge** | `find({ connected: { to: babbage, depth: 2 } })` |
|
||
|
||
It runs **inside your process** — no server, no Docker, nothing to operate — and persists to plain files you can snapshot with a hard link.
|
||
|
||
**New here?** → **[What is Brainy? — plain-language overview, no jargon](docs/eli5.md)**
|
||
|
||
## Quick start
|
||
|
||
```bash
|
||
bun add @soulcraft/brainy # Bun ≥ 1.1 — recommended
|
||
npm install @soulcraft/brainy # Node.js ≥ 22
|
||
```
|
||
|
||
```javascript
|
||
import { Brainy, NounType, VerbType } from '@soulcraft/brainy'
|
||
|
||
const brain = new Brainy() // in-memory; one line swaps to disk
|
||
await brain.init()
|
||
|
||
// Text auto-embeds locally; metadata auto-indexes
|
||
const react = await brain.add({
|
||
data: 'React is a JavaScript library for building user interfaces',
|
||
type: NounType.Concept,
|
||
subtype: 'library',
|
||
metadata: { category: 'frontend', year: 2013 }
|
||
})
|
||
|
||
const next = await brain.add({
|
||
data: 'Next.js is a React framework with server-side rendering',
|
||
type: NounType.Concept,
|
||
subtype: 'framework',
|
||
metadata: { category: 'frontend', year: 2016 }
|
||
})
|
||
|
||
await brain.relate({ from: next, to: react, type: VerbType.DependsOn, subtype: 'runtime' })
|
||
```
|
||
|
||
## One query, three engines
|
||
|
||
```javascript
|
||
const results = await brain.find({
|
||
query: 'modern frontend frameworks', // vector — what it means
|
||
where: { year: { greaterThan: 2015 } }, // metadata — what it is
|
||
connected: { to: react, depth: 2 } // graph — what it touches
|
||
})
|
||
```
|
||
|
||
Every clause is optional; any combination composes. Under the hood Brainy plans the query across an HNSW vector index, a roaring-bitmap field index, and an adjacency graph index — and re-validates every result against your predicate before returning it, so a corrupt index can never hand you a wrong answer.
|
||
|
||
## Feature tour
|
||
|
||
### The database is a value
|
||
|
||
Pin it, rewind it, fork it. Snapshot isolation without a server.
|
||
|
||
```javascript
|
||
const db = brain.now() // pin current state — O(1)
|
||
|
||
await brain.transact([ // atomic all-or-nothing, CAS-guarded
|
||
{ op: 'update', id: order, metadata: { status: 'paid' } },
|
||
{ op: 'relate', from: invoice, to: order, type: VerbType.References, subtype: 'billing' }
|
||
], { ifAtGeneration: db.generation })
|
||
|
||
await db.get(order) // still 'pending' — pinned forever
|
||
await brain.get(order) // 'paid' — live
|
||
|
||
const lastWeek = await brain.asOf(Date.now() - 7 * 86_400_000) // full query surface, past state
|
||
const whatIf = await db.with([{ op: 'remove', id: order }]) // speculative — never touches disk
|
||
await brain.now().persist('/backups/today') // instant hard-link snapshot
|
||
```
|
||
|
||
**[Consistency model](docs/concepts/consistency-model.md)** · **[Snapshots & time travel](docs/guides/snapshots-and-time-travel.md)**
|
||
|
||
### Local embeddings — no API keys
|
||
|
||
Strings embed on-device with a bundled MiniLM model (WASM). Semantic search works offline, in CI, and on air-gapped machines, at zero cost per call. Hybrid keyword + semantic ranking is the default:
|
||
|
||
```javascript
|
||
await brain.find({ query: 'David Smith' }) // auto: text + semantic
|
||
await brain.find({ query: 'AI concepts', searchMode: 'semantic' }) // semantic only
|
||
```
|
||
|
||
### A typed graph, not a bag of edges
|
||
|
||
42 entity types × 127 relationship types form a shared vocabulary for any domain — healthcare (`Patient → diagnoses → Condition`), finance (`Account → transfers → Transaction`), yours. Your own taxonomy layers on with `subtype`, enforced at write time:
|
||
|
||
```javascript
|
||
await brain.add({ data: 'Avery Brooks', type: NounType.Person, subtype: 'employee' })
|
||
|
||
brain.counts.bySubtype(NounType.Person) // O(1) — { employee: 12, customer: 847 }
|
||
brain.requireSubtype(NounType.Person, { values: ['employee', 'customer'], required: true })
|
||
```
|
||
|
||
**[Type system](docs/architecture/noun-verb-taxonomy.md)** · **[Subtypes & facets](docs/guides/subtypes-and-facets.md)**
|
||
|
||
### Graph analytics built in
|
||
|
||
```javascript
|
||
await brain.graph.rank() // which entities matter most (centrality)
|
||
await brain.graph.communities() // natural clusters
|
||
await brain.graph.path(a, b) // how two things connect
|
||
await brain.graph.subgraph([seed], { depth: 2 }) // bounded neighborhood → { nodes, edges }
|
||
await brain.graph.export() // whole graph, one O(N+E) streaming pass
|
||
```
|
||
|
||
### Write-time aggregations
|
||
|
||
`SUM` / `COUNT` / `AVG` / `MIN` / `MAX` with `GROUP BY` and time windows, maintained incrementally on every write — reads are O(1) lookups, not scans. **[Aggregation guide](docs/guides/aggregation.md)**
|
||
|
||
### Import anything
|
||
|
||
```javascript
|
||
await brain.import('customers.csv')
|
||
await brain.import('sales.xlsx') // every sheet
|
||
await brain.import('research-paper.pdf') // tables extracted
|
||
await brain.import('https://api.example.com/data.json')
|
||
```
|
||
|
||
Entities auto-classify on the way in; `brain.extractEntities(text)` exposes the same NER ensemble directly. **[Import guide](docs/guides/import-anything.md)**
|
||
|
||
### A filesystem that understands content
|
||
|
||
```javascript
|
||
await brain.vfs.writeFile('/docs/readme.md', 'Project documentation')
|
||
await brain.vfs.search('React components with hooks') // semantic file search
|
||
```
|
||
|
||
**[VFS quick start](docs/vfs/QUICK_START.md)**
|
||
|
||
### Operations-grade by default
|
||
|
||
- **Single-writer, many-reader** — an exclusive lock protects the data directory; `Brainy.openReadOnly()` and the `brainy inspect` CLI examine a live brain from another process, safely.
|
||
- **Self-upgrading data files** — a 7.x brain opens under 8.x and migrates itself behind an observable lock (`getIndexStatus().migration`), with an automatic pre-upgrade backup. No migration scripts.
|
||
- **No silent wrong answers** — cold-open guards self-heal or throw typed errors (`MetadataIndexNotReadyError`, `GraphIndexNotReadyError`); they never return `[]` for data that exists.
|
||
|
||
**[Multi-process model](docs/concepts/multi-process.md)** · **[Inspection guide](docs/guides/inspection.md)**
|
||
|
||
## From laptop to hundreds of millions
|
||
|
||
Brainy's TypeScript engines take you a long way. When you outgrow them, add the native engine — **the API doesn't change**:
|
||
|
||
```bash
|
||
npm install @soulcraft/cor
|
||
```
|
||
|
||
```javascript
|
||
const brain = new Brainy({ storage: { type: 'filesystem', path: './data' } })
|
||
await brain.init() // @soulcraft/cor detected — same code, native engines underneath
|
||
```
|
||
|
||
Installing the package is the opt-in: if `@soulcraft/cor` is present, it loads and announces itself in the init log; if it's present but broken, `init()` **throws** — an installed accelerator never silently vanishes behind the JS engines. Opt out with `plugins: []`, or pin exactly what loads with `plugins: ['@soulcraft/cor']`. [`@soulcraft/cor`](https://www.npmjs.com/package/@soulcraft/cor) (Brainy 8.x ↔ Cor 3.x, version-matched) registers Rust implementations behind every provider seam: SIMD distance kernels, memory-mapped storage, a disk-native vector index that doesn't need your dataset in RAM, durable LSM field/graph indexes that serve cold opens instantly, and native aggregation. Recall@10 measured **0.99 / 0.96 / 0.96 at 1M / 10M / 100M vectors** in Cor's release gate.
|
||
|
||
Open core, commercial accelerator: Brainy is MIT and complete on its own; Cor is licensed and funds both.
|
||
|
||
## Performance
|
||
|
||
- JS distance kernels: **~6× faster cosine, ~1.4× euclidean** than 7.x (measured: [`tests/benchmarks/distance-microbench.mjs`](tests/benchmarks/distance-microbench.mjs), 384-dim, median of 41).
|
||
- Whole-graph reads are single **O(N + E)** cursor walks — a consumer-measured 19k-edge export dropped from ~27 s of per-node calls to one scan.
|
||
- Full numbers and capacity planning: **[docs/PERFORMANCE.md](docs/PERFORMANCE.md)** · **[docs/SCALING.md](docs/SCALING.md)**
|
||
|
||
## Use cases
|
||
|
||
**AI agent memory** — persistent semantic recall with relationship tracking · **Knowledge bases** — auto-linking and meaning-aware navigation · **Semantic search** over codebases, documents, media · **Enterprise data** — CRM, catalogs, institutional memory · **Games & simulations** — worlds and characters that remember.
|
||
|
||
## Documentation
|
||
|
||
| Start | Core | Going deeper |
|
||
|---|---|---|
|
||
| [Brainy explained simply](docs/eli5.md) | [API reference](docs/api/README.md) | [Architecture overview](docs/architecture/overview.md) |
|
||
| [Installation](docs/guides/installation.md) | [Data model](docs/DATA_MODEL.md) | [Consistency model](docs/concepts/consistency-model.md) |
|
||
| [Natural-language queries](docs/guides/natural-language.md) | [Query operators](docs/QUERY_OPERATORS.md) | [Multi-process model](docs/concepts/multi-process.md) |
|
||
| | [Find system](docs/FIND_SYSTEM.md) | [Scaling](docs/SCALING.md) |
|
||
|
||
## Requirements
|
||
|
||
**Bun ≥ 1.1** (recommended) or **Node.js ≥ 22**. Brainy 8.x is server-only; the 7.x line remains on npm for browser use.
|
||
|
||
## Contributing & license
|
||
|
||
Contributions welcome — see **[CONTRIBUTING.md](CONTRIBUTING.md)**. MIT © Brainy Contributors.
|