brainy/src/utils/entityIdMapper.ts

281 lines
8 KiB
TypeScript
Raw Normal View History

/**
feat: stable EntityIdMapper — rebuild() no longer renumbers UUID→int Previously metadataIndex.rebuild() called idMapper.clear() which reset nextId to 1 and renumbered every UUID by re-insertion order. Any consumer that had persisted int-keyed data against the old map was silently invalidated — and 2.4.0's vector mmap store (#20), graph link compression (#21), and column-store JS↔native interchange all need persisted int indices that survive a rebuild. Remove the unconditional clear() in rebuild(). The rebuild already re-iterates every entity via idMapper.getOrAssign(uuid), which returns the existing int unchanged for known UUIDs. Stale UUID→int entries for entities no longer in storage persist as harmless memory overhead; a dedicated prune step can be added if it ever matters. clearAllIndexData() — the explicit nuclear recovery path — keeps its existing idMapper.clear() call (renumbering is intentional there), and now logs a prodLog.warn making it explicit that any persisted int-keyed data is invalidated and must be rebuilt from canonical sources. Strengthened the EntityIdMapper class JSDoc to document the stability guarantee as a contract — append-only getOrAssign, monotonic nextId, remove() leaves permanent holes, rebuild() never renumbers, only clear() does. Added tests/regression/entity-id-mapper-stability.test.ts pinning down the five-point contract: (1) single-rebuild stability; (2) many-rebuild stability; (3) post-rebuild adds get fresh monotonic ints; (4) removes leave permanent holes — new entities never recycle; (5) clearAllIndexData() explicitly renumbers (the documented destructive path). Foundation for 2.4.0 #2-#4. Full test suite (62 files, 1417 tests) green.
2026-05-28 09:45:22 -07:00
* EntityIdMapper - Bidirectional mapping between UUID strings and integer IDs.
*
feat: stable EntityIdMapper — rebuild() no longer renumbers UUID→int Previously metadataIndex.rebuild() called idMapper.clear() which reset nextId to 1 and renumbered every UUID by re-insertion order. Any consumer that had persisted int-keyed data against the old map was silently invalidated — and 2.4.0's vector mmap store (#20), graph link compression (#21), and column-store JS↔native interchange all need persisted int indices that survive a rebuild. Remove the unconditional clear() in rebuild(). The rebuild already re-iterates every entity via idMapper.getOrAssign(uuid), which returns the existing int unchanged for known UUIDs. Stale UUID→int entries for entities no longer in storage persist as harmless memory overhead; a dedicated prune step can be added if it ever matters. clearAllIndexData() — the explicit nuclear recovery path — keeps its existing idMapper.clear() call (renumbering is intentional there), and now logs a prodLog.warn making it explicit that any persisted int-keyed data is invalidated and must be rebuilt from canonical sources. Strengthened the EntityIdMapper class JSDoc to document the stability guarantee as a contract — append-only getOrAssign, monotonic nextId, remove() leaves permanent holes, rebuild() never renumbers, only clear() does. Added tests/regression/entity-id-mapper-stability.test.ts pinning down the five-point contract: (1) single-rebuild stability; (2) many-rebuild stability; (3) post-rebuild adds get fresh monotonic ints; (4) removes leave permanent holes — new entities never recycle; (5) clearAllIndexData() explicitly renumbers (the documented destructive path). Foundation for 2.4.0 #2-#4. Full test suite (62 files, 1417 tests) green.
2026-05-28 09:45:22 -07:00
* Roaring bitmaps require 32-bit unsigned integers, but Brainy uses UUID strings
* as canonical entity IDs. This class provides efficient O(1) bidirectional
* mapping with persistence and, importantly, a stability guarantee that any
* persisted int-keyed data can rely on.
*
* **Stability guarantee (the foundation 2.4.0 vector-mmap, graph-link-compression,
* and column-store interchange all key off):**
*
* - `getOrAssign(uuid)` is **append-only**: once a UUID is assigned an int, the
* mapping never changes. Subsequent `getOrAssign` calls for the same UUID
* return the same int.
* - `nextId` is **monotonically increasing**. New UUIDs always get an int greater
* than any previously assigned, so a removed-then-re-added UUID is treated as
* a fresh entity (and gets a fresh int there is no automatic "revive").
* - `remove(uuid)` removes the mapping but does **not** decrement `nextId` or
* recycle the int. The removed int becomes a permanent hole in `intToUuid`
* downstream consumers seeing `getUuid(int) === undefined` know the entity
* was deleted.
* - A metadata-index `rebuild()` does **not** clear the mapper (the rebuild path
* re-iterates entities via `getOrAssign`, which returns existing ints unchanged).
* Only the explicit `clear()` method renumbers used by `clearAllIndexData()`
* as the nuclear recovery path with a documented warning.
*
* Features:
feat: stable EntityIdMapper — rebuild() no longer renumbers UUID→int Previously metadataIndex.rebuild() called idMapper.clear() which reset nextId to 1 and renumbered every UUID by re-insertion order. Any consumer that had persisted int-keyed data against the old map was silently invalidated — and 2.4.0's vector mmap store (#20), graph link compression (#21), and column-store JS↔native interchange all need persisted int indices that survive a rebuild. Remove the unconditional clear() in rebuild(). The rebuild already re-iterates every entity via idMapper.getOrAssign(uuid), which returns the existing int unchanged for known UUIDs. Stale UUID→int entries for entities no longer in storage persist as harmless memory overhead; a dedicated prune step can be added if it ever matters. clearAllIndexData() — the explicit nuclear recovery path — keeps its existing idMapper.clear() call (renumbering is intentional there), and now logs a prodLog.warn making it explicit that any persisted int-keyed data is invalidated and must be rebuilt from canonical sources. Strengthened the EntityIdMapper class JSDoc to document the stability guarantee as a contract — append-only getOrAssign, monotonic nextId, remove() leaves permanent holes, rebuild() never renumbers, only clear() does. Added tests/regression/entity-id-mapper-stability.test.ts pinning down the five-point contract: (1) single-rebuild stability; (2) many-rebuild stability; (3) post-rebuild adds get fresh monotonic ints; (4) removes leave permanent holes — new entities never recycle; (5) clearAllIndexData() explicitly renumbers (the documented destructive path). Foundation for 2.4.0 #2-#4. Full test suite (62 files, 1417 tests) green.
2026-05-28 09:45:22 -07:00
* - O(1) lookup in both directions.
* - Persistent storage via storage adapter.
* - Atomic, monotonic, append-only int counter.
* - Serialization/deserialization support.
*
* @module utils/entityIdMapper
*/
import type { StorageAdapter } from '../coreTypes.js'
import type { EntityIdMapperProvider } from '../plugin.js'
export interface EntityIdMapperOptions {
storage: StorageAdapter
storageKey?: string
}
export interface EntityIdMapperData {
nextId: number
uuidToInt: Record<string, number>
intToUuid: Record<number, string>
}
/**
* Maps entity UUIDs to integer IDs for use with Roaring Bitmaps.
*
* Implements {@link EntityIdMapperProvider}: the surface a registered
* `'entityIdMapper'` provider (e.g. Cortex's native mapper) must also satisfy.
*/
export class EntityIdMapper implements EntityIdMapperProvider {
private storage: StorageAdapter
private storageKey: string
// Bidirectional maps
private uuidToInt = new Map<string, number>()
private intToUuid = new Map<number, string>()
// Atomic counter for next ID
private nextId = 1
// Dirty flag for persistence
private dirty = false
constructor(options: EntityIdMapperOptions) {
this.storage = options.storage
this.storageKey = options.storageKey || 'brainy:entityIdMapper'
}
/**
* Initialize the mapper by loading from storage
*/
async init(): Promise<void> {
try {
feat(v4.0.0): Complete metadata/vector separation architecture with Azure support This commit completes the core v4.0.0 architecture changes for billion-scale performance with metadata/vector separation. NO RELEASE YET - remaining optimizations and testing required before production release. ## Core v4.0.0 Architecture Changes ### Type System Updates - Fixed all TypeScript compilation errors (zero errors achieved) - Updated HNSWNoun/HNSWVerb to separate core fields from metadata - Implemented HNSWNounWithMetadata/HNSWVerbWithMetadata for API boundaries - Added required 'noun' field to NounMetadata for semantic structure - Renamed verb.type to verb.verb for consistency ### Storage Adapter Updates **All adapters updated for v4.0.0 two-file storage pattern:** - memoryStorage: Proper metadata/vector separation - fileSystemStorage: Two-file pattern with sharding - opfsStorage: Browser persistent storage updated - s3CompatibleStorage: AWS/MinIO/DigitalOcean support - r2Storage: Cloudflare R2 optimization - gcsStorage: Google Cloud with ADC support - **azureBlobStorage: NEW - Full Azure Blob Storage support** ### Storage Features - BaseStorage: Internal vs public method separation (_getNoun vs getNoun) - Two-file storage: Vectors in one file, metadata in another - Change tracking: getChangesSince return type updated - Pagination: getNounsWithPagination returns WithMetadata types ### Azure Blob Storage Integration (NEW) - Native @azure/storage-blob SDK integration - Four authentication methods: * DefaultAzureCredential (Managed Identity) - recommended * Connection String - simplest setup * Account Name + Key - traditional auth * SAS Token - delegated access - High-volume mode with write buffering - Adaptive backpressure for throttling - UUID-based sharding for billion-scale - Full HNSW support with graph persistence ### Utility Updates - EmbeddingManager: Updated to accept Record<string, unknown> - LSMTree: Wrapped data in NounMetadata structure with 'noun' field - EntityIdMapper: Fixed nested metadata.data structure access - MetadataIndex: Fixed field type inference integration - PeriodicCleanup: Updated for new metadata structure ### Core API Updates - Brainy: Updated verb property access from v.type to v.verb - ConfigAPI: Fixed NounMetadata access patterns - DataAPI: Updated metadata handling ### Documentation Updates - CREATING-AUGMENTATIONS.md: v4.0.0 breaking changes guide - DEVELOPER-GUIDE.md: Migration checklist and examples - COMPLETE-REFERENCE.md: v4.0.0 architecture improvements - **finite-type-system.md: NEW - Revolutionary type system benefits** ### Build & Dependencies - Zero TypeScript compilation errors - Added @azure/storage-blob and @azure/identity - 591 tests passing (23 timeout in long-running neural tests) ## What's NOT in This Release This is a work-in-progress commit. Before v4.0.0 release we need: - Storage adapter optimizations (batch operations, compression) - Azure blob tier management (Hot/Cool/Archive) - Cost optimization implementations - Additional performance testing at billion-scale - Migration guides for v3.x users ## Testing - Clean build: ✅ - Type checking: ✅ (zero errors) - Test suite: ✅ (591/614 passing, timeouts in neural tests only) 🔐 Generated with Claude Code https://claude.com/claude-code Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-17 12:29:27 -07:00
const metadata = await this.storage.getMetadata(this.storageKey)
// metadata IS the data (no nested 'data' property)
if (metadata && (metadata as any).nextId !== undefined) {
const data = metadata as any as EntityIdMapperData
this.nextId = data.nextId
// Rebuild maps from serialized data
this.uuidToInt = new Map(Object.entries(data.uuidToInt).map(([k, v]) => [k, Number(v)]))
this.intToUuid = new Map(Object.entries(data.intToUuid).map(([k, v]) => [Number(k), v]))
} else {
// Guard: mapper file missing but entities may exist on disk.
// If we start from nextId=1 with existing entities, roaring bitmap
// queries will use wrong integer IDs → silent data corruption.
// Probe storage to detect this case and log a warning.
try {
const probe = await this.storage.getNouns({ pagination: { limit: 1, offset: 0 } })
if ((probe.totalCount ?? 0) > 0 || probe.items.length > 0) {
console.warn(
`[EntityIdMapper] Mapper file missing but entities exist on disk. ` +
`IDs will be rebuilt during metadata index reconstruction.`
)
}
} catch {
// Storage not ready
}
}
} catch (error) {
// First time initialization - maps are empty, nextId = 1
}
}
/**
* Get integer ID for UUID, assigning a new ID if not exists
*/
getOrAssign(uuid: string): number {
const existing = this.uuidToInt.get(uuid)
if (existing !== undefined) {
return existing
}
// Assign new ID
const newId = this.nextId++
this.uuidToInt.set(uuid, newId)
this.intToUuid.set(newId, uuid)
this.dirty = true
return newId
}
/**
* Get integer ID for UUID with immediate persistence guarantee
* Unlike getOrAssign(), this method flushes to storage immediately after assigning
* a new ID. This prevents UUIDint mapping divergence if the process crashes
* before a normal flush() occurs.
*
* Use this for critical operations where data integrity is paramount.
* Normal operations can use getOrAssign() with batched flushing for better performance.
*/
async getOrAssignSync(uuid: string): Promise<number> {
const id = this.getOrAssign(uuid)
// If a new ID was assigned, immediately persist to storage
if (this.dirty) {
await this.flush()
}
return id
}
/**
* Get UUID for integer ID
*/
getUuid(intId: number): string | undefined {
return this.intToUuid.get(intId)
}
/**
* Get integer ID for UUID (without assigning if not exists)
*/
getInt(uuid: string): number | undefined {
return this.uuidToInt.get(uuid)
}
/**
* Check if UUID has been assigned an integer ID
*/
has(uuid: string): boolean {
return this.uuidToInt.has(uuid)
}
/**
* Remove mapping for UUID
*/
remove(uuid: string): boolean {
const intId = this.uuidToInt.get(uuid)
if (intId === undefined) {
return false
}
this.uuidToInt.delete(uuid)
this.intToUuid.delete(intId)
this.dirty = true
return true
}
/**
* Get total number of mappings
*/
get size(): number {
return this.uuidToInt.size
}
/**
* Get all mapped integer IDs without storage reads.
* Used for bitmap negation operations (ne, exists:false, missing:true)
* to avoid full-table getAllIds() scans.
*/
getAllIntIds(): number[] {
return Array.from(this.intToUuid.keys())
}
/**
* Convert array of UUIDs to array of integers
*/
uuidsToInts(uuids: string[]): number[] {
return uuids.map(uuid => this.getOrAssign(uuid))
}
/**
* Convert array of integers to array of UUIDs
*/
intsToUuids(ints: number[]): string[] {
const result: string[] = []
for (const intId of ints) {
const uuid = this.intToUuid.get(intId)
if (uuid) {
result.push(uuid)
}
}
return result
}
/**
* Convert iterable of integers to array of UUIDs (for roaring bitmap iteration)
*/
intsIterableToUuids(ints: Iterable<number>): string[] {
const result: string[] = []
for (const intId of ints) {
const uuid = this.intToUuid.get(intId)
if (uuid) {
result.push(uuid)
}
}
return result
}
/**
* Flush mappings to storage
*/
async flush(): Promise<void> {
if (!this.dirty) {
return
}
// Convert maps to plain objects for serialization
// Add required 'noun' property for NounMetadata
feat(v4.0.0): Complete metadata/vector separation architecture with Azure support This commit completes the core v4.0.0 architecture changes for billion-scale performance with metadata/vector separation. NO RELEASE YET - remaining optimizations and testing required before production release. ## Core v4.0.0 Architecture Changes ### Type System Updates - Fixed all TypeScript compilation errors (zero errors achieved) - Updated HNSWNoun/HNSWVerb to separate core fields from metadata - Implemented HNSWNounWithMetadata/HNSWVerbWithMetadata for API boundaries - Added required 'noun' field to NounMetadata for semantic structure - Renamed verb.type to verb.verb for consistency ### Storage Adapter Updates **All adapters updated for v4.0.0 two-file storage pattern:** - memoryStorage: Proper metadata/vector separation - fileSystemStorage: Two-file pattern with sharding - opfsStorage: Browser persistent storage updated - s3CompatibleStorage: AWS/MinIO/DigitalOcean support - r2Storage: Cloudflare R2 optimization - gcsStorage: Google Cloud with ADC support - **azureBlobStorage: NEW - Full Azure Blob Storage support** ### Storage Features - BaseStorage: Internal vs public method separation (_getNoun vs getNoun) - Two-file storage: Vectors in one file, metadata in another - Change tracking: getChangesSince return type updated - Pagination: getNounsWithPagination returns WithMetadata types ### Azure Blob Storage Integration (NEW) - Native @azure/storage-blob SDK integration - Four authentication methods: * DefaultAzureCredential (Managed Identity) - recommended * Connection String - simplest setup * Account Name + Key - traditional auth * SAS Token - delegated access - High-volume mode with write buffering - Adaptive backpressure for throttling - UUID-based sharding for billion-scale - Full HNSW support with graph persistence ### Utility Updates - EmbeddingManager: Updated to accept Record<string, unknown> - LSMTree: Wrapped data in NounMetadata structure with 'noun' field - EntityIdMapper: Fixed nested metadata.data structure access - MetadataIndex: Fixed field type inference integration - PeriodicCleanup: Updated for new metadata structure ### Core API Updates - Brainy: Updated verb property access from v.type to v.verb - ConfigAPI: Fixed NounMetadata access patterns - DataAPI: Updated metadata handling ### Documentation Updates - CREATING-AUGMENTATIONS.md: v4.0.0 breaking changes guide - DEVELOPER-GUIDE.md: Migration checklist and examples - COMPLETE-REFERENCE.md: v4.0.0 architecture improvements - **finite-type-system.md: NEW - Revolutionary type system benefits** ### Build & Dependencies - Zero TypeScript compilation errors - Added @azure/storage-blob and @azure/identity - 591 tests passing (23 timeout in long-running neural tests) ## What's NOT in This Release This is a work-in-progress commit. Before v4.0.0 release we need: - Storage adapter optimizations (batch operations, compression) - Azure blob tier management (Hot/Cool/Archive) - Cost optimization implementations - Additional performance testing at billion-scale - Migration guides for v3.x users ## Testing - Clean build: ✅ - Type checking: ✅ (zero errors) - Test suite: ✅ (591/614 passing, timeouts in neural tests only) 🔐 Generated with Claude Code https://claude.com/claude-code Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-17 12:29:27 -07:00
const data = {
noun: 'EntityIdMapper',
nextId: this.nextId,
uuidToInt: Object.fromEntries(this.uuidToInt),
intToUuid: Object.fromEntries(this.intToUuid)
}
feat(v4.0.0): Complete metadata/vector separation architecture with Azure support This commit completes the core v4.0.0 architecture changes for billion-scale performance with metadata/vector separation. NO RELEASE YET - remaining optimizations and testing required before production release. ## Core v4.0.0 Architecture Changes ### Type System Updates - Fixed all TypeScript compilation errors (zero errors achieved) - Updated HNSWNoun/HNSWVerb to separate core fields from metadata - Implemented HNSWNounWithMetadata/HNSWVerbWithMetadata for API boundaries - Added required 'noun' field to NounMetadata for semantic structure - Renamed verb.type to verb.verb for consistency ### Storage Adapter Updates **All adapters updated for v4.0.0 two-file storage pattern:** - memoryStorage: Proper metadata/vector separation - fileSystemStorage: Two-file pattern with sharding - opfsStorage: Browser persistent storage updated - s3CompatibleStorage: AWS/MinIO/DigitalOcean support - r2Storage: Cloudflare R2 optimization - gcsStorage: Google Cloud with ADC support - **azureBlobStorage: NEW - Full Azure Blob Storage support** ### Storage Features - BaseStorage: Internal vs public method separation (_getNoun vs getNoun) - Two-file storage: Vectors in one file, metadata in another - Change tracking: getChangesSince return type updated - Pagination: getNounsWithPagination returns WithMetadata types ### Azure Blob Storage Integration (NEW) - Native @azure/storage-blob SDK integration - Four authentication methods: * DefaultAzureCredential (Managed Identity) - recommended * Connection String - simplest setup * Account Name + Key - traditional auth * SAS Token - delegated access - High-volume mode with write buffering - Adaptive backpressure for throttling - UUID-based sharding for billion-scale - Full HNSW support with graph persistence ### Utility Updates - EmbeddingManager: Updated to accept Record<string, unknown> - LSMTree: Wrapped data in NounMetadata structure with 'noun' field - EntityIdMapper: Fixed nested metadata.data structure access - MetadataIndex: Fixed field type inference integration - PeriodicCleanup: Updated for new metadata structure ### Core API Updates - Brainy: Updated verb property access from v.type to v.verb - ConfigAPI: Fixed NounMetadata access patterns - DataAPI: Updated metadata handling ### Documentation Updates - CREATING-AUGMENTATIONS.md: v4.0.0 breaking changes guide - DEVELOPER-GUIDE.md: Migration checklist and examples - COMPLETE-REFERENCE.md: v4.0.0 architecture improvements - **finite-type-system.md: NEW - Revolutionary type system benefits** ### Build & Dependencies - Zero TypeScript compilation errors - Added @azure/storage-blob and @azure/identity - 591 tests passing (23 timeout in long-running neural tests) ## What's NOT in This Release This is a work-in-progress commit. Before v4.0.0 release we need: - Storage adapter optimizations (batch operations, compression) - Azure blob tier management (Hot/Cool/Archive) - Cost optimization implementations - Additional performance testing at billion-scale - Migration guides for v3.x users ## Testing - Clean build: ✅ - Type checking: ✅ (zero errors) - Test suite: ✅ (591/614 passing, timeouts in neural tests only) 🔐 Generated with Claude Code https://claude.com/claude-code Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-17 12:29:27 -07:00
await this.storage.saveMetadata(this.storageKey, data as any)
this.dirty = false
}
/**
* Clear all mappings
*/
async clear(): Promise<void> {
this.uuidToInt.clear()
this.intToUuid.clear()
this.nextId = 1
this.dirty = true
await this.flush()
}
/**
* Get statistics about the mapper
*/
getStats() {
return {
mappings: this.uuidToInt.size,
nextId: this.nextId,
dirty: this.dirty,
memoryEstimate: this.uuidToInt.size * (36 + 8 + 4 + 8) // uuid string + map overhead + int + map overhead
}
}
}