brainy/docs/transactions.md

568 lines
16 KiB
Markdown
Raw Normal View History

---
title: Transactions & Atomicity
slug: guides/transactions
public: true
category: guides
template: guide
order: 10
description: How Brainy keeps every write atomic — automatic per-operation transactions with rollback, and brain.transact() for atomic multi-write batches with compare-and-swap.
next:
- concepts/consistency-model
- guides/optimistic-concurrency
---
# Transaction System
**Status:** ✅ Production Ready
## Overview
Brainy's transaction system provides **atomic operations** with automatic rollback on failure. All operations within a transaction either succeed completely or fail completely - there are no partial failures.
### Key Benefits
- **Atomicity**: All operations succeed or all rollback
- **Consistency**: Indexes and storage remain consistent
- **Automatic**: Transparently used by all `brain.add()`, `brain.update()`, `brain.remove()`, and `brain.relate()` operations
- **Composable**: `brain.transact()` runs a multi-write batch through the same machinery as exactly one atomic commit
## Architecture
```
User Code (brain.add(), brain.update(), brain.transact(), etc.)
Transaction Manager (orchestration)
Operations (SaveNounMetadataOperation, SaveNounOperation, etc.)
Storage Adapter (sharding, ID-first routing)
```
### How It Works
Every write operation in Brainy automatically uses transactions:
```typescript
// Internally, this uses a transaction
const id = await brain.add({
data: { name: 'Alice', role: 'Engineer' },
type: NounType.Person
})
```
**Transaction Flow:**
1. **Begin Transaction**: TransactionManager creates new transaction
2. **Add Operations**: Operations added to transaction (SaveNounMetadataOperation, SaveNounOperation)
3. **Execute**: Each operation executes in sequence
4. **Commit**: All operations succeeded → changes persist
5. **Rollback**: Any operation failed → all changes reverted
### Rollback Mechanism
Each operation implements both **execute** and **undo**:
```typescript
class SaveNounMetadataOperation {
async execute(): Promise<void> {
// Save new metadata
await this.storage.saveNounMetadata(this.id, this.metadata)
}
async undo(): Promise<void> {
// Restore previous metadata (or delete if new entity)
if (this.previousMetadata) {
await this.storage.saveNounMetadata(this.id, this.previousMetadata)
} else {
await this.storage.deleteNounMetadata(this.id)
}
}
}
```
**On failure:**
- Operations rolled back in **reverse order**
- Previous state fully restored
- Indexes updated to reflect rollback
## Compatibility with Advanced Features
### Multi-Write Batches: `brain.transact()`
**The 8.0 path for atomic multi-entity writes**
Single-operation methods each commit their own transaction. When several
writes must succeed or fail **together**, use `brain.transact()` — a
declarative batch that commits as exactly one generation, with optional
whole-store compare-and-swap and durable transaction metadata:
```typescript
const db = await brain.transact([
{ op: 'add', id: orderId, type: NounType.Document, subtype: 'order', data: 'Order #1042' },
{ op: 'update', id: customerId, metadata: { lastOrderAt: Date.now() }, ifRev: customer._rev },
{ op: 'relate', from: customerId, to: orderId, type: VerbType.Creates, subtype: 'purchase' }
], { meta: { author: 'order-service' } })
db.receipt.ids // resolved id per operation, in input order
```
**How It Works:**
- The batch executes through the same TransactionManager as single
operations, wrapped in the generational commit protocol: before-images are
staged and fsynced first, and the atomic manifest rename is the commit
point — a crash anywhere before it rolls back to the exact
pre-transaction bytes.
- Per-entity `ifRev` and whole-store `ifAtGeneration` provide
compare-and-swap at two granularities; any conflict rejects the entire
batch before anything is staged.
- The returned `Db` is a pinned, snapshot-isolated view of the committed
state.
See the **[consistency model](concepts/consistency-model.md)** for the
full guarantees (snapshot isolation, time travel, snapshots) and
**[Snapshots & Time Travel](guides/snapshots-and-time-travel.md)** for
recipes.
### Sharding
**Fully Compatible**
Transactions work across multiple shards:
```typescript
// Entities with different UUID prefixes go to different shards
const id1 = 'aaa00000-1111-4111-8111-111111111111' // Shard: aaa
const id2 = 'bbb00000-2222-4222-8222-222222222222' // Shard: bbb
await brain.add({ id: id1, data: { name: 'Entity A' }, type: NounType.Thing })
await brain.relate({ from: id1, to: id2, type: VerbType.RelatesTo })
// Transaction handles cross-shard atomicity automatically
```
**How It Works:**
- Sharding is transparent to transactions
- `analyzeKey()` method routes to correct shard based on UUID
- Transaction operations don't need to know about shards
- Rollback works across all shards involved
### ID-First Storage
**Fully Compatible**
feat: ID-first storage architecture + remove memory-unsafe APIs (v6.0.0) BREAKING CHANGES: **ID-First Storage Paths** - Direct O(1) entity access without type lookups - Before: entities/nouns/{TYPE}/metadata/{SHARD}/{ID}.json - After: entities/nouns/{SHARD}/{ID}/metadata.json - Migration handled automatically on first init() **Removed Memory-Unsafe APIs** - Removed brain.merge() - loaded all entities into memory - Removed brain.diff() - loaded all entities into memory - Removed brain.data().backup() - loaded all entities into memory - Removed brain.data().restore() - depended on backup() - Removed CLI commands: backup, restore, cow merge **Migration Paths** - merge() → Use checkout() or manually copy entities with pagination - diff() → Use asOf() with manual paginated comparison - backup() → Use fork() for instant COW snapshots - restore() → Use checkout() to switch to snapshot branch Core Improvements: - ✅ All 8 storage adapters properly call super.init() - ✅ GraphAdjacencyIndex integration in BaseStorage.init() - ✅ Fixed ID-first path bugs (vector.json → vectors.json) - ✅ Fixed MemoryStorage.initializeCounts() for ID-first paths - ✅ New VFS APIs: du(), access(), find() - ✅ Comprehensive documentation with migration guides Storage Adapters Fixed: - MemoryStorage, FileSystemStorage, AzureBlobStorage - GCSStorage, R2Storage, S3CompatibleStorage - OPFSStorage, HistoricalStorageAdapter Files Changed: 28 files, +1,075/-1,933 lines (net -858) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-19 16:46:11 -08:00
Transactions work with direct ID-first paths - no type routing needed!
```typescript
feat: ID-first storage architecture + remove memory-unsafe APIs (v6.0.0) BREAKING CHANGES: **ID-First Storage Paths** - Direct O(1) entity access without type lookups - Before: entities/nouns/{TYPE}/metadata/{SHARD}/{ID}.json - After: entities/nouns/{SHARD}/{ID}/metadata.json - Migration handled automatically on first init() **Removed Memory-Unsafe APIs** - Removed brain.merge() - loaded all entities into memory - Removed brain.diff() - loaded all entities into memory - Removed brain.data().backup() - loaded all entities into memory - Removed brain.data().restore() - depended on backup() - Removed CLI commands: backup, restore, cow merge **Migration Paths** - merge() → Use checkout() or manually copy entities with pagination - diff() → Use asOf() with manual paginated comparison - backup() → Use fork() for instant COW snapshots - restore() → Use checkout() to switch to snapshot branch Core Improvements: - ✅ All 8 storage adapters properly call super.init() - ✅ GraphAdjacencyIndex integration in BaseStorage.init() - ✅ Fixed ID-first path bugs (vector.json → vectors.json) - ✅ Fixed MemoryStorage.initializeCounts() for ID-first paths - ✅ New VFS APIs: du(), access(), find() - ✅ Comprehensive documentation with migration guides Storage Adapters Fixed: - MemoryStorage, FileSystemStorage, AzureBlobStorage - GCSStorage, R2Storage, S3CompatibleStorage - OPFSStorage, HistoricalStorageAdapter Files Changed: 28 files, +1,075/-1,933 lines (net -858) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-19 16:46:11 -08:00
// Entities stored with direct ID-first paths
const personId = await brain.add({
data: { name: 'John Doe' },
type: NounType.Person // → entities/nouns/{shard}/{id}/metadata.json
})
const orgId = await brain.add({
data: { name: 'Acme Corp' },
type: NounType.Organization // → entities/nouns/{shard}/{id}/metadata.json
})
feat: ID-first storage architecture + remove memory-unsafe APIs (v6.0.0) BREAKING CHANGES: **ID-First Storage Paths** - Direct O(1) entity access without type lookups - Before: entities/nouns/{TYPE}/metadata/{SHARD}/{ID}.json - After: entities/nouns/{SHARD}/{ID}/metadata.json - Migration handled automatically on first init() **Removed Memory-Unsafe APIs** - Removed brain.merge() - loaded all entities into memory - Removed brain.diff() - loaded all entities into memory - Removed brain.data().backup() - loaded all entities into memory - Removed brain.data().restore() - depended on backup() - Removed CLI commands: backup, restore, cow merge **Migration Paths** - merge() → Use checkout() or manually copy entities with pagination - diff() → Use asOf() with manual paginated comparison - backup() → Use fork() for instant COW snapshots - restore() → Use checkout() to switch to snapshot branch Core Improvements: - ✅ All 8 storage adapters properly call super.init() - ✅ GraphAdjacencyIndex integration in BaseStorage.init() - ✅ Fixed ID-first path bugs (vector.json → vectors.json) - ✅ Fixed MemoryStorage.initializeCounts() for ID-first paths - ✅ New VFS APIs: du(), access(), find() - ✅ Comprehensive documentation with migration guides Storage Adapters Fixed: - MemoryStorage, FileSystemStorage, AzureBlobStorage - GCSStorage, R2Storage, S3CompatibleStorage - OPFSStorage, HistoricalStorageAdapter Files Changed: 28 files, +1,075/-1,933 lines (net -858) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-19 16:46:11 -08:00
// Type changes handled atomically (type is just metadata)
await brain.update({
id: personId,
type: NounType.Organization, // Type change
data: { name: 'Doe Corp' }
})
```
**How It Works:**
feat: ID-first storage architecture + remove memory-unsafe APIs (v6.0.0) BREAKING CHANGES: **ID-First Storage Paths** - Direct O(1) entity access without type lookups - Before: entities/nouns/{TYPE}/metadata/{SHARD}/{ID}.json - After: entities/nouns/{SHARD}/{ID}/metadata.json - Migration handled automatically on first init() **Removed Memory-Unsafe APIs** - Removed brain.merge() - loaded all entities into memory - Removed brain.diff() - loaded all entities into memory - Removed brain.data().backup() - loaded all entities into memory - Removed brain.data().restore() - depended on backup() - Removed CLI commands: backup, restore, cow merge **Migration Paths** - merge() → Use checkout() or manually copy entities with pagination - diff() → Use asOf() with manual paginated comparison - backup() → Use fork() for instant COW snapshots - restore() → Use checkout() to switch to snapshot branch Core Improvements: - ✅ All 8 storage adapters properly call super.init() - ✅ GraphAdjacencyIndex integration in BaseStorage.init() - ✅ Fixed ID-first path bugs (vector.json → vectors.json) - ✅ Fixed MemoryStorage.initializeCounts() for ID-first paths - ✅ New VFS APIs: du(), access(), find() - ✅ Comprehensive documentation with migration guides Storage Adapters Fixed: - MemoryStorage, FileSystemStorage, AzureBlobStorage - GCSStorage, R2Storage, S3CompatibleStorage - OPFSStorage, HistoricalStorageAdapter Files Changed: 28 files, +1,075/-1,933 lines (net -858) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-19 16:46:11 -08:00
- Type information stored in metadata.noun field
- Storage layer uses O(1) ID-first path construction
- No type cache needed (removed in a previous version)
- Type counters adjusted on commit/rollback
- 40x faster path lookups (eliminates 42-type search)
### Storage Adapter Interface
**Fully Compatible**
Transactions go through the `StorageAdapter` interface, so both shipped adapters (filesystem, memory) and any custom plugin adapter inherit the same atomicity guarantees:
```typescript
const brain = new Brainy({
feat(8.0): API simplification — remove neural()/Db.search, one storage `path` key, integration→0 8.0 RC cleanup toward "one place per thing, zero-config, no deprecation": - Remove the `brain.neural()` clustering namespace (ImprovedNeuralAPI + the dead legacy NeuralAPI + the neural CLI + neural-only types). Similarity is `find({vector})` / `similar({to})`; attribute grouping is the aggregation `GROUP BY` engine. The separate entity-extraction / smart-import feature (NeuralImport, NeuralEntityExtractor, SmartExtractor, NaturalLanguageProcessor, `brain.extract()`/`brain.nlp()`) is kept. - Remove `Db.search()`; `find()` is the one query verb (accepts a bare string or FindParams). Fix the bundled MCP client, which called a non-existent `brain.search(query, limit)` → now `find({ query, limit })`. - Storage config: collapse to one canonical top-level `path` key. The pre-8.0 aliases (`rootDirectory`, `options.*`, `fileSystemStorage.*`) are removed and now THROW with the exact rename instead of silently defaulting to `./brainy-data` on upgrade. A single resolver feeds createStorage, the 7.x→8.0 migration probe, and the plugin-factory handoff, so a native storage provider resolves the identical root (no split-brain). - Fix `similar({ threshold })`: the min-similarity filter was silently dropped; it is now applied as a post-filter on `result.score` (the documented way to bound semantic results). - Fix `vfs.rename()` on a directory: child path updates spread the entity vector into `update()` and failed dimension validation; they are metadata-only updates now. - Fix `vfs.move()`: copy+delete orphaned the content-addressed content blob (the destination shared the source hash, then unlink removed it). `move()` now delegates to `rename()` — an in-place path change that preserves the blob and the entity id, for files and directories. - Fix streaming import: the bulk fast path never flushed mid-import nor signalled queryability. Entity writes are now chunked by a progressive flush interval (100 → 1000 → 5000); each chunk flushes and emits `progress.queryable`, so imported data is queryable during the import. - Sweep all docs, comments, and JSDoc for the removed/changed APIs. Integration suite: 49 files / 588 passed / 0 failed. Unit: 80 files / 1456 passed, no type errors.
2026-06-20 13:31:11 -07:00
storage: { type: 'filesystem', path: './data' }
})
await brain.add({ data: { name: 'Entity' }, type: NounType.Thing })
```
**How It Works:**
- Transactions operate through `StorageAdapter` interface
- Custom adapters registered via the plugin system implement the same interface
- Atomicity guaranteed at the write-coordinator level
- Read-after-write consistency maintained inside a single Brainy process
## Examples
### Basic Add Operation
```typescript
import { Brainy } from '@soulcraft/brainy'
import { NounType } from '@soulcraft/brainy/types'
const brain = new Brainy()
await brain.init()
// Automatically uses transaction
const id = await brain.add({
data: { name: 'Alice', role: 'Engineer' },
type: NounType.Person
})
// If add fails, all changes rolled back automatically
```
### Update with Type Change
```typescript
// Original entity
const id = await brain.add({
data: { name: 'John Smith', category: 'individual' },
type: NounType.Person
})
// Update with type change (atomic)
await brain.update({
id,
type: NounType.Organization, // Type change
data: { name: 'Smith Corp', category: 'business' }
})
// If update fails, original type and data restored
```
### Creating Relationships
```typescript
const personId = await brain.add({
data: { name: 'Alice' },
type: NounType.Person
})
const projectId = await brain.add({
data: { name: 'Project X' },
type: NounType.Thing
})
// Create relationship (atomic)
await brain.relate({
from: personId,
to: projectId,
type: VerbType.WorksOn
})
// If relate fails, no partial relationship created
```
### Batch Operations
```typescript
// Multiple operations, all atomic
for (let i = 0; i < 100; i++) {
await brain.add({
data: { name: `Entity ${i}`, index: i },
type: NounType.Thing
})
}
// Each add() is a separate transaction
// If any add fails, only that specific add is rolled back
```
### Delete with Cascade
```typescript
const personId = await brain.add({
data: { name: 'Bob' },
type: NounType.Person
})
const projectId = await brain.add({
data: { name: 'Project Y' },
type: NounType.Thing
})
await brain.relate({
from: personId,
to: projectId,
type: VerbType.WorksOn
})
// Delete person (atomic - deletes entity + relationships)
await brain.remove(personId)
// If delete fails, both entity and relationships remain
```
## Error Handling
Transactions automatically handle errors and rollback:
```typescript
try {
await brain.add({
data: { name: 'Test Entity' },
type: NounType.Thing,
vector: [1, 2, 3] // Wrong dimension → error
})
} catch (error) {
// Transaction automatically rolled back
// No partial data in storage or indexes
console.error('Add failed:', error.message)
}
```
**Common Error Scenarios:**
- **Invalid vector dimension**: Automatic rollback
- **Type validation failure**: Automatic rollback
- **Storage write failure**: Automatic rollback
- **Index update failure**: Automatic rollback
## Performance Considerations
### Transaction Overhead
**What a transaction costs:**
- A typical single-operation write wraps 2-8 operations (metadata + data + indexes) in one transaction
- The overhead is bookkeeping (operation objects + undo state), not extra I/O on the success path
- Rollback cost is proportional to the operations already applied (each is undone in reverse order)
**Optimization:**
- Operations executed sequentially (not parallel) for consistency
- Rollback only happens on failure (success path is fast)
- Index updates batched within transaction
### Auditing Committed Batches
Every committed `brain.transact()` batch is recorded in the transaction
log, newest first:
```typescript
await brain.transact(ops, { meta: { author: 'import-job' } })
const entries = await brain.transactionLog({ limit: 10 })
// [{ generation: 1042, timestamp: 1765432100000, meta: { author: 'import-job' } }]
```
Single-operation writes advance the generation counter but do not append
log entries — see the [consistency model](concepts/consistency-model.md)
for the history-granularity contract.
## Best Practices
### 1. Let Brainy Handle Transactions
```typescript
// ✅ Recommended: Use Brainy's API (transactions automatic)
await brain.add({ data, type })
await brain.update({ id, data })
await brain.remove(id)
// ❌ Avoid: Direct storage access bypasses transactions
await brain.storage.saveNoun(noun) // No transaction protection
```
### 2. Handle Errors Gracefully
```typescript
// ✅ Recommended: Catch errors, transaction rolls back automatically
try {
const id = await brain.add({ data, type })
return id
} catch (error) {
console.error('Add failed, rolled back:', error)
// Decide how to handle (retry, log, alert user)
}
```
### 3. Validate Before Operations
```typescript
// ✅ Recommended: Validate early to avoid unnecessary rollbacks
if (!isValidVector(vector, brain.dimension)) {
throw new Error(`Vector must have ${brain.dimension} dimensions`)
}
await brain.add({ data, type, vector })
```
### 4. Batch Related Writes with `transact()`
```typescript
// ✅ Recommended: writes that must land together go in one batch
await brain.transact([
{ op: 'add', id: orderId, type: NounType.Document, subtype: 'order', data: 'Order #1042' },
{ op: 'relate', from: customerId, to: orderId, type: VerbType.Creates, subtype: 'purchase' }
])
// ❌ Avoid: sequential single operations when partial application is unacceptable
const id = await brain.add({ ... }) // commits alone
await brain.relate({ ... }) // a crash here leaves the entity unlinked
```
### 5. Understand Atomicity Guarantees
**What Transactions GUARANTEE:**
- ✅ Atomicity within a single Brainy process
- ✅ Consistent state across all indexes
- ✅ Automatic rollback on failure
- ✅ Works with all storage adapters (filesystem, memory, custom plugin adapters)
**What Transactions DON'T Provide:**
- ❌ Two-phase commit across multiple Brainy instances
- ❌ Distributed locking across processes
- ❌ Cross-datacenter ACID guarantees
**Design:** Transactions ensure atomicity at the **write-coordinator level** inside one process. Cross-instance coordination, if you need it, lives in your service layer.
## Testing Transactions
### Unit Tests
```typescript
import { describe, it, expect } from 'vitest'
import { Brainy } from '@soulcraft/brainy'
describe('Transaction Tests', () => {
it('should rollback on failure', async () => {
const brain = new Brainy()
await brain.init()
const id1 = await brain.add({ data: { name: 'Entity 1' }, type: NounType.Thing })
try {
await brain.add({
data: null as any, // Invalid - will fail
type: NounType.Thing
})
} catch (e) {
// Expected failure
}
// First entity should still exist (rollback didn't affect it)
const entity1 = await brain.get(id1)
expect(entity1).toBeTruthy()
})
})
```
### Integration Tests
See `tests/transaction/integration/` for comprehensive integration tests covering:
- Sharding integration (`sharding-transactions.test.ts`)
- Type-aware integration (`typeaware-transactions.test.ts`)
The atomicity guarantees of `brain.transact()` — including crash recovery
through the real recovery path — are proven in
`tests/integration/db-mvcc.test.ts`.
## Troubleshooting
### High Rollback Rate
**Symptom:** a high share of writes throw and roll back
**Possible Causes:**
1. Invalid vector dimensions
2. Type validation errors
3. Storage write failures (disk full, network issues)
4. Index corruption
**Solutions:**
- Validate data before operations
- Check storage adapter health
- Monitor disk space and network connectivity
- Review error logs for patterns
### Slow Transaction Performance
**Symptom:** Operations take > 100ms per transaction
**Possible Causes:**
1. Large metadata objects
2. Remote storage latency
3. Many indexes enabled
4. Disk I/O bottleneck
**Solutions:**
- Optimize metadata size
- Use local caching for remote storage
- Disable unused indexes
- Use SSD storage
## Architecture Details
### Transaction Lifecycle
```
1. BEGIN
2. ADD OPERATIONS
- SaveNounMetadataOperation
- SaveNounOperation
- UpdateGraphIndexOperation
3. EXECUTE (sequential)
- Execute operation 1 → Success
- Execute operation 2 → Success
- Execute operation 3 → FAILURE
4. ROLLBACK (reverse order)
- Undo operation 2
- Undo operation 1
5. THROW ERROR
```
### Operation Types
| Operation | Description | Undo Behavior |
|-----------|-------------|---------------|
| `SaveNounMetadataOperation` | Save entity metadata | Restore previous metadata or delete if new |
| `SaveNounOperation` | Save entity data | Restore previous data or delete if new |
| `UpdateGraphIndexOperation` | Update graph index | Restore previous index state |
| `SaveVerbMetadataOperation` | Save relationship metadata | Restore previous metadata or delete if new |
| `SaveVerbOperation` | Save relationship data | Restore previous data or delete if new |
### Storage Adapter Integration
Transactions use the `StorageAdapter` interface:
```typescript
interface StorageAdapter {
saveNounMetadata(id: string, metadata: NounMetadata): Promise<void>
saveNoun(noun: Noun): Promise<void>
deleteNounMetadata(id: string): Promise<void>
deleteNoun(id: string): Promise<void>
// ... other methods
}
```
**Key Insight:** Both shipped storage adapters (filesystem, memory) — and any custom plugin adapter — implement this interface. Transactions work with **any** storage adapter automatically.
## Additional Resources
- **Unit Tests:** `tests/transaction/Transaction.test.ts`, `tests/transaction/TransactionManager.test.ts`
- **Integration Tests:** `tests/transaction/integration/`
- **MVCC Proofs:** `tests/integration/db-mvcc.test.ts` (atomicity, CAS, crash recovery for `brain.transact()`)
- **Consistency Model:** [docs/concepts/consistency-model.md](concepts/consistency-model.md)
## Summary
Brainy's transaction system provides **production-ready atomic operations** with automatic rollback. Every single-operation write is transactional out of the box, and `brain.transact()` extends the same guarantee to multi-write batches — one atomic commit, with compare-and-swap and durable transaction metadata.
**Key Takeaways:**
-**Automatic**: No manual transaction management needed for single operations
-**Atomic**: All operations succeed or all rollback — per operation and per `transact()` batch
-**Compatible**: Works with all storage adapters and features
-**Coordinated**: Per-entity `ifRev` and whole-store `ifAtGeneration` CAS reject conflicting batches before anything is staged
Start using transactions today - they're already built into `brain.add()`, `brain.update()`, `brain.remove()`, and `brain.relate()` — and reach for `brain.transact()` whenever several writes must land together.