brainy/README.md
David Snelling 99bc4990bd Set package.json private flag to false for npm publication
Changed `"private"` to `false` in package.json to allow publishing the package to npm. The `"access": "restricted"` setting ensures that access remains limited to @soulcraft organization members. Updated README to reflect the new configuration.
2025-05-23 11:30:19 -07:00

511 lines
17 KiB
Markdown

# Soulcraft Brainy
A vector database that runs in a browser or Node.js and utilizes Origin Private File System (OPFS) for storage, with HNSW (Hierarchical Navigable Small World) for efficient vector indexing.
## Features
- **Cross-platform**: Works in both browsers and Node.js
- **Persistent storage**: Uses Origin Private File System (OPFS) in browsers, with fallback to in-memory storage
- **Efficient vector search**: Implements HNSW (Hierarchical Navigable Small World) algorithm for fast approximate nearest neighbor search
- **Automatic embedding**: Converts text and other data to vectors using embedding models
- **TensorFlow.js integration**: Uses Universal Sentence Encoder for high-quality text embeddings
- **Metadata support**: Store and retrieve metadata alongside vectors
- **TypeScript support**: Fully typed API with generics for metadata types
- **Multiple distance functions**: Supports cosine, Euclidean, Manhattan, and dot product distance metrics
- **Augmentation system**: Extensible architecture for adding specialized capabilities
- **Graph data model**: Structured representation of entities and relationships
## Installation
```bash
npm install @soulcraft/brainy
```
## Usage
### Basic Example
```typescript
import {BrainyData} from '@soulcraft/brainy';
// Create a new vector database
const db = new BrainyData();
await db.init();
// Add vectors with metadata
const catId = await db.add([0.2, 0.3, 0.4, 0.1], {type: 'mammal', name: 'cat'});
const dogId = await db.add([0.3, 0.2, 0.4, 0.2], {type: 'mammal', name: 'dog'});
const fishId = await db.add([0.1, 0.1, 0.8, 0.2], {type: 'fish', name: 'fish'});
// Add text directly - it will be automatically embedded
const lionDescId = await db.add("Lions are large cats with a golden mane", {type: 'mammal', name: 'lion'});
const tigerDescId = await db.add("Tigers are large cats with striped fur", {type: 'mammal', name: 'tiger'});
// Search for similar vectors
const results = await db.search([0.2, 0.3, 0.4, 0.1], 2);
console.log(results);
// [
// { id: 'cat-id', score: 0, vector: [0.2, 0.3, 0.4, 0.1], metadata: { type: 'mammal', name: 'cat' } },
// { id: 'dog-id', score: 0.1, vector: [0.3, 0.2, 0.4, 0.2], metadata: { type: 'mammal', name: 'dog' } }
// ]
// Search with text directly - it will be automatically embedded
const catResults = await db.search("cat", 2);
console.log(catResults);
// Results will include vectors similar to the embedding of "cat"
// Get a vector by ID
const cat = await db.get(catId);
console.log(cat);
// { id: 'cat-id', vector: [0.2, 0.3, 0.4, 0.1], metadata: { type: 'mammal', name: 'cat' } }
// Update metadata
await db.updateMetadata(catId, {type: 'mammal', name: 'cat', color: 'orange'});
// Delete a vector
await db.delete(fishId);
// Clear the database
await db.clear();
```
### Configuration Options
```typescript
import {
BrainyData,
euclideanDistance,
UniversalSentenceEncoder,
createEmbeddingFunction
} from '@soulcraft/brainy';
// Configure the vector database
const db = new BrainyData({
// HNSW index configuration
hnsw: {
M: 16, // Max number of connections per node
efConstruction: 200, // Size of dynamic candidate list during construction
efSearch: 50, // Size of dynamic candidate list during search
ml: 16 // Max level
},
// Distance function to use (default is cosineDistance)
distanceFunction: euclideanDistance,
// Custom embedding function (optional)
// By default, it uses the Universal Sentence Encoder for high-quality text embeddings
// You can use the SimpleEmbedding for a basic character-based embedding:
// embeddingFunction: createEmbeddingFunction(new SimpleEmbedding()),
// Or create your own custom embedding function:
// embeddingFunction: async (data) => {
// // Convert data to a vector
// return [0.1, 0.2, 0.3, 0.4]; // Return a vector
// },
// Custom storage adapter (optional)
// By default, it uses OPFS in browsers, FileSystemStorage in Node.js,
// or falls back to in-memory storage if neither is available
// storageAdapter: myCustomStorageAdapter
// You can also explicitly use the FileSystemStorage with a custom directory:
// import { FileSystemStorage } from '@soulcraft/brainy/storage/fileSystemStorage';
// storageAdapter: new FileSystemStorage('/custom/path')
});
```
## Publishing and Using as a Private NPM Package
Soulcraft Brainy is configured as a private NPM package with restricted access. This section provides information on how to publish and use it within your organization.
### Publishing the Package
To publish updates to the package:
1. Ensure you have the appropriate npm credentials and access to the @soulcraft organization
2. Update the version in package.json
3. Build the package:
```bash
npm run build
```
4. Publish the package:
```bash
npm publish
```
Note that the package has the following configuration in package.json:
```json
"private": false,
"publishConfig": {
"access": "restricted"
}
```
This ensures that the package is only accessible to users with appropriate permissions within the @soulcraft organization. The `"access": "restricted"` setting limits access to the package to members of the @soulcraft organization, while `"private": false` allows the package to be published to npm.
### Installing the Private Package
To install the package in another project:
1. Ensure you have access to the @soulcraft organization on npm
2. Add the package to your project:
```bash
npm install @soulcraft/brainy
```
3. If you're using a private npm registry, you may need to configure npm to use your organization's registry:
```bash
npm config set @soulcraft:registry https://your-private-registry.com/
```
### Requirements
- Node.js >= 18.0.0
## Augmentation System
Brainy includes a powerful augmentation system that allows extending its capabilities through specialized modules. Each augmentation implements a specific interface and provides additional functionality.
### Base Augmentation Interface
All augmentations implement the `IAugmentation` interface:
```typescript
interface IAugmentation {
readonly name: string; // Unique identifier for the augmentation
readonly description: string; // Human-readable description
initialize(): Promise<void>; // Called when Brainy starts up
shutDown(): Promise<void>; // Called when shutting down
getStatus(): Promise<'active' | 'inactive' | 'error'>; // Current status
}
```
### WebSocket Support
Augmentations can optionally implement WebSocket support:
```typescript
interface IWebSocketSupport {
connectWebSocket(url: string, protocols?: string | string[]): Promise<WebSocketConnection>;
sendWebSocketMessage(connectionId: string, data: unknown): Promise<void>;
onWebSocketMessage(connectionId: string, callback: DataCallback<unknown>): Promise<void>;
closeWebSocket(connectionId: string, code?: number, reason?: string): Promise<void>;
}
```
### Specialized Augmentation Types
Brainy supports several specialized augmentation types:
#### Cognition Augmentations
For reasoning, inference, and logical operations:
```typescript
interface ICognitionAugmentation extends IAugmentation {
reason(query: string, context?: Record<string, unknown>): AugmentationResponse<{
inference: string;
confidence: number;
}>;
infer(dataSubset: Record<string, unknown>): AugmentationResponse<Record<string, unknown>>;
executeLogic(ruleId: string, input: Record<string, unknown>): AugmentationResponse<boolean>;
}
```
#### Sense Augmentations
For processing raw, unstructured data:
```typescript
interface ISenseAugmentation extends IAugmentation {
processRawData(rawData: Buffer | string, dataType: string): AugmentationResponse<{
nouns: string[];
verbs: string[];
}>;
listenToFeed(
feedUrl: string,
callback: DataCallback<{ nouns: string[]; verbs: string[] }>
): Promise<void>;
}
```
#### Perception Augmentations
For interpreting and contextualizing data:
```typescript
interface IPerceptionAugmentation extends IAugmentation {
interpret(
nouns: string[],
verbs: string[],
context?: Record<string, unknown>
): AugmentationResponse<Record<string, unknown>>;
organize(
data: Record<string, unknown>,
criteria?: Record<string, unknown>
): AugmentationResponse<Record<string, unknown>>;
generateVisualization(
data: Record<string, unknown>,
visualizationType: string
): AugmentationResponse<string | Buffer | Record<string, unknown>>;
}
```
#### Activation Augmentations
For triggering actions and generating outputs:
```typescript
interface IActivationAugmentation extends IAugmentation {
triggerAction(
actionName: string,
parameters?: Record<string, unknown>
): AugmentationResponse<unknown>;
generateOutput(knowledgeId: string, format: string): AugmentationResponse<string | Record<string, unknown>>;
interactExternal(systemId: string, payload: Record<string, unknown>): AugmentationResponse<unknown>;
}
```
#### Dialog Augmentations
For natural language understanding and generation:
```typescript
interface IDialogAugmentation extends IAugmentation {
processUserInput(naturalLanguageQuery: string, sessionId?: string): AugmentationResponse<{
intent: string;
nouns: string[];
verbs: string[];
context: Record<string, unknown>;
}>;
generateResponse(
interpretedInput: Record<string, unknown>,
knowledgeContext: Record<string, unknown>,
sessionId?: string
): AugmentationResponse<string>;
manageContext(sessionId: string, contextUpdate: Record<string, unknown>): Promise<void>;
}
```
#### Conduit Augmentations
For establishing data exchange channels:
```typescript
interface IConduitAugmentation extends IAugmentation {
establishConnection(
targetSystemId: string,
config: Record<string, unknown>
): AugmentationResponse<WebSocketConnection>;
readData(
query: Record<string, unknown>,
options?: Record<string, unknown>
): AugmentationResponse<unknown>;
writeData(
data: Record<string, unknown>,
options?: Record<string, unknown>
): AugmentationResponse<unknown>;
monitorStream(streamId: string, callback: DataCallback<unknown>): Promise<void>;
}
```
## Graph Data Model
Brainy uses a graph-based data model to represent entities and relationships. This model consists of nouns (nodes) and verbs (edges).
### Common Types
#### Timestamp
Used for tracking creation and update times:
```typescript
interface Timestamp {
seconds: number;
nanoseconds: number;
}
```
#### CreatorMetadata
Tracks which augmentation and model created an element:
```typescript
interface CreatorMetadata {
augmentation: string; // Name of the augmentation that created this element
version: string; // Version of the augmentation
model: string; // Model identifier used in creation
modelVersion: string; // Version of the model
}
```
### Graph Elements
#### GraphNoun
Base interface for nodes (entities) in the graph:
```typescript
interface GraphNoun {
id: string; // Unique identifier for the noun
createdBy: CreatorMetadata; // Information about what created this noun
noun: NounType; // Type classification of the noun
createdAt: Timestamp; // When the noun was created
updatedAt: Timestamp; // When the noun was last updated
data?: Record<string, unknown>; // Additional flexible data storage
embedding?: number[]; // Vector representation of the noun
}
```
#### GraphVerb
Base interface for edges (relationships) in the graph:
```typescript
interface GraphVerb {
id: string; // Unique identifier for the verb
source: string; // ID of the source noun
target: string; // ID of the target noun
label?: string; // Optional descriptive label
verb: VerbType; // Type of relationship
createdAt: Timestamp; // When the verb was created
updatedAt: Timestamp; // When the verb was last updated
data?: Record<string, unknown>; // Additional flexible data storage
embedding?: number[]; // Vector representation of the relationship
confidence?: number; // Confidence score (0-1)
weight?: number; // Strength/importance of the relationship
}
```
### Noun Types
Brainy supports the following noun types:
- **Person**: Represents a person entity
- **Place**: Represents a physical location
- **Thing**: Represents a physical or virtual object
- **Event**: Represents an event or occurrence
- **Concept**: Represents an abstract concept or idea
- **Content**: Represents content (text, media, etc.)
### Verb Types
Brainy supports the following verb types:
- **AttributedTo**: Indicates attribution or authorship
- **Controls**: Indicates control or ownership
- **Created**: Indicates creation or authorship
- **Earned**: Indicates achievement or acquisition
- **Owns**: Indicates ownership
## Examples
The repository includes several examples to help you get started:
### Modern UI Demo
A complete web application that demonstrates all the features of Soulcraft Brainy with a modern user interface:
- Initialize the database with different distance functions
- Configure HNSW parameters
- Add sample vectors and custom vectors with metadata
- Search for similar vectors
- Get, update, and delete vectors
- View database size and clear the database
To run the Modern UI Demo:
1. Clone the repository
2. Build the project with `npm run build`
3. Open `examples/demo.html` in a browser
### Node.js Examples
The repository also includes TypeScript examples for Node.js:
- `src/examples/basicUsage.ts`: Demonstrates basic vector operations
- `src/examples/customStorage.ts`: Shows how to use a custom storage adapter
## How It Works
### HNSW Indexing
The Hierarchical Navigable Small World (HNSW) algorithm is used for efficient approximate nearest neighbor search. It creates a multi-layered graph structure that allows for logarithmic-time search complexity.
Key features of the HNSW implementation:
- Hierarchical graph structure for efficient navigation
- Configurable parameters for tuning performance vs. accuracy
- Support for different distance metrics
### Origin Private File System (OPFS) Storage
In browser environments, the database uses the Origin Private File System (OPFS) API for persistent storage. This provides:
- Fast, local storage that persists between sessions
- Isolation from other origins for security
- Efficient file operations
In Node.js environments, the database uses a file system-based storage adapter that stores data in JSON files. This provides:
- Persistent storage between application restarts
- Efficient file operations using Node.js fs module
- Configurable storage location
In environments where neither OPFS nor Node.js file system is available, the database automatically falls back to in-memory storage.
## API Reference
### BrainyData
The main class for interacting with the vector database.
#### Constructor
```typescript
constructor(config?: BrainyDataConfig)
```
#### Methods
- `init(): Promise<void>` - Initialize the database
- `add(vectorOrData: Vector | any, metadata?: T, options?: { forceEmbed?: boolean }): Promise<string>` - Add a vector or data to the database
- `addBatch(items: Array<{ vectorOrData: Vector | any, metadata?: T }>, options?: { forceEmbed?: boolean }): Promise<string[]>` - Add multiple vectors or data items
- `search(queryVectorOrData: Vector | any, k?: number, options?: { forceEmbed?: boolean }): Promise<SearchResult<T>[]>` - Search for similar vectors
- `get(id: string): Promise<VectorDocument<T> | null>` - Get a vector by ID
- `delete(id: string): Promise<boolean>` - Delete a vector
- `updateMetadata(id: string, metadata: T): Promise<boolean>` - Update metadata
- `clear(): Promise<void>` - Clear the database
- `size(): number` - Get the number of vectors in the database
### Distance Functions
- `euclideanDistance(a: Vector, b: Vector): number` - Euclidean (L2) distance
- `cosineDistance(a: Vector, b: Vector): number` - Cosine distance
- `manhattanDistance(a: Vector, b: Vector): number` - Manhattan (L1) distance
- `dotProductDistance(a: Vector, b: Vector): number` - Dot product distance
### Embedding Models
- `SimpleEmbedding` - A simple character-based embedding model for text
- `UniversalSentenceEncoder` - TensorFlow Universal Sentence Encoder for high-quality text embeddings
### Embedding Functions
- `createEmbeddingFunction(model: EmbeddingModel): EmbeddingFunction` - Create an embedding function from an embedding model
- `defaultEmbeddingFunction` - Default embedding function using UniversalSentenceEncoder
## Browser Compatibility
The Soulcraft Brainy database works in all modern browsers that support the Origin Private File System API:
- Chrome 86+
- Edge 86+
- Opera 72+
- Chrome for Android 86+
For browsers without OPFS support, the database will automatically fall back to in-memory storage.
## License
MIT