brainy/src/utils/paramValidation.ts
David Snelling c0d326b36d feat: verb subtype + updateRelation + requireSubtype enforcement
Brings verbs to first-class parity with nouns. The 7.29.0 subtype primitive
shipped for entities only; this release ships the symmetric verb mirror plus
the enforcement layer for ensuring every entity AND every relationship has
both type AND subtype.

Layer V1 — verb subtype mirror
- HNSWVerbWithMetadata.subtype + STANDARD_VERB_FIELDS set + resolveVerbField()
- Relation<T>.subtype, RelateParams<T>.subtype, UpdateRelationParams<T> extended,
  GetRelationsParams.subtype, GraphConstraints.subtype (for find connected)
- relate() persists subtype on verbMetadata + GraphVerb + transaction ops
- getRelations({ type, subtype }) fast-path filter with set membership
- find({ connected: { via, subtype, depth } }) traversal filter (depth-1 on
  the JS path; explicit error on depth > 1 pointing at Cortex native)
- verbsToRelations + storage destructure sites surface subtype to top-level
- All three graph-index fast-path queries (getVerbsBySource/ByTarget) enrich
  with subtype from metadata

Layer V2 — updateRelation() closes a pre-7.30 gap
- New first-class verb update method (parallel to update() for nouns)
- Changes subtype/type/weight/confidence/data/metadata in place
- Re-indexes in graph adjacency when verb type changes; id preserved
- validateUpdateRelationParams enforces id + at-least-one-field-to-update

Layer V3 — verb subtype storage rollup
- verbSubtypeCountsByType: Map<number, Map<string, number>> on BaseStorage
- verbSubtypeByIdCache for self-heal during update/delete
- incrementVerbSubtypeCount + decrementVerbSubtypeCount maintain state
- loadVerbSubtypeStatistics + saveVerbSubtypeStatistics persist to
  _system/verb-subtype-statistics.json (mirrors noun-side shape)
- rebuildVerbSubtypeCounts for poison recovery / explicit repair
- getVerbSubtypeCountsByType accessor for the public counts API
- Wired into init() / flushCounts() / saveVerbMetadata / deleteVerbMetadata

Layer V4 — verb counts API + relationshipSubtypesOf
- brain.counts.byRelationshipSubtype(verb, subtype?) — O(1) breakdown or point
- brain.counts.topRelationshipSubtypes(verb, n) — top N by count
- brain.relationshipSubtypesOf(verb) — sorted distinct subtypes

Layer V5 — migrateField extended to verbs
- New entityKind?: 'noun' | 'verb' | 'both' option (default 'noun')
- Mirror verb iteration via storage.getVerbs() with same path semantics
- verbToRelationLike + buildRelationMigrationUpdate helpers project the
  storage verb shape onto the Entity<T>-shaped surface readPath understands
- Routes through new updateRelation() for the verb-side rewrite

Enforcement (opt-in in 7.30, default in 8.0)
- brain.requireSubtype(type, options) — unified API for NounType OR VerbType.
  Registers per-type rules with optional values whitelist; composes with the
  brain-wide flag.
- new Brainy({ requireSubtype: true }) — brain-wide strict mode. Every public
  write path validates the pairing guarantee.
- { except: [NounType.Thing, ...] } form for catch-all type exemptions
- Atomic-fail semantics on addMany / relateMany — pre-validate every item
  before any storage write, throw on first failure with item index
- Per-type rules + brain-wide flag both throw with descriptive messages
- VFS infrastructure bypass via metadata.isVFSEntity / isVFS markers so
  brain's own VFS writes don't get rejected when strict mode is on

VFS labeling — concrete subtypes for infrastructure entities
- VFS root: NounType.Collection + subtype: 'vfs-root' (was bare Collection)
- VFS directories: subtype: 'vfs-directory'
- VFS files: subtype: 'vfs-file' (NounType still mime-based)
- VFS containment edges: VerbType.Contains + subtype: 'vfs-contains'
- Lets consumers cleanly enumerate VFS state via find({ subtype: 'vfs-file' })
  and distinguish Brainy's VFS Collections from user-created Collections

Docs
- docs/guides/subtypes-and-facets.md extended with Layer V (Verbs) section +
  Enforcement section. New full reference at the bottom split into Layer 1
  (nouns), Layer V (verbs), Layer 2 (facets), Layer 3 (migration), Enforcement.
- docs/api/README.md adds updateRelation(), getRelations({ subtype }), the
  three verb-side counts methods, requireSubtype(), and the brain-wide
  constructor option. relate() params include subtype.
- docs/DATA_MODEL.md adds a Subtype-for-VerbType section + STANDARD_VERB_FIELDS
- docs/architecture/finite-type-system.md extends Principle 1a to verbs
- docs/QUERY_OPERATORS.md adds a verb-subtype filter section covering
  getRelations and find({connected, subtype}) traversal
- README.md "Subtypes" section now shows both noun + verb in one example +
  the enforcement APIs
- RELEASES.md v7.30.0 entry with the full noun/verb capability parity matrix

Tests
- tests/integration/verb-subtype-and-enforcement.test.ts — 30 new tests
  covering V1 round-trips, V1 set membership, updateRelation in place,
  updateRelation preservation, V2 counts breakdown + point + topN + distinct,
  V2 decrements on unrelate, V2 re-routes on updateRelation, V3 depth-1
  traversal filter, V3 depth>1 explicit error, V4 verb migration, V4 both
  entity kinds, V4 readBoth preservation, V5 per-type required rejection,
  V5 vocabulary rejection, V5 on-vocab acceptance, V5 verb-side enforcement,
  V5 addMany atomic-fail, V5 relateMany atomic-fail, V5 update enforcement,
  V5 updateRelation enforcement, V5 brain-wide strict mode, V5 except clause.

Verification
- Unit suite: 1468/1468 passing
- Noun subtype integration (7.29 carryover): 26/26 passing
- Verb subtype + enforcement integration: 30/30 passing
- Type-check: clean
- Build: clean
- Public closed-source reference audit: clean

Internal 8.0 spec
- .strategy/BRAINY-8.0-SUBTYPE-CONTRACT.md (gitignored, not in npm artifact)
  documents the contract upgrade Cortex 3.0 implements against: required-by-
  default subtype, SubtypeRegistry typing hook, native simplification,
  multi-hop traversal native fast path, brain.fillSubtypes() migration helper.
  Coordinated via PLATFORM-HANDOFF rows CTX-SUBTYPE-PARITY-V2 (7.30 parallel
  work) and CTX-SUBTYPE-8.0-CONTRACT (8.0 spec).
2026-06-05 11:15:52 -07:00

487 lines
No EOL
14 KiB
TypeScript

/**
* Zero-Config Parameter Validation
*
* Self-configuring validation that adapts to system capabilities
* Only enforces universal truths, learns everything else
*/
import { FindParams, AddParams, UpdateParams, RelateParams, UpdateRelationParams } from '../types/brainy.types.js'
import { NounType, VerbType } from '../types/graphTypes.js'
// Dynamic import for Node.js os and fs modules
let os: any = null
let fs: any = null
if (typeof window === 'undefined') {
try {
os = await import('node:os')
fs = await import('node:fs')
} catch (e) {
// OS/FS modules not available
}
}
// Browser-safe memory detection
const getSystemMemory = (): number => {
if (os) {
return os.totalmem()
}
// Browser fallback: assume 4GB
return 4 * 1024 * 1024 * 1024
}
const getAvailableMemory = (): number => {
if (os) {
return os.freemem()
}
// Browser fallback: assume 2GB available
return 2 * 1024 * 1024 * 1024
}
/**
* Detect container memory limit (Docker/Kubernetes/Cloud Run)
*
* Production-grade detection for containerized environments.
* Supports:
* - cgroup v1 (legacy Docker/K8s)
* - cgroup v2 (modern systems)
* - Environment variables (Cloud Run, GCP, AWS, Azure)
*
* @returns Container memory limit in bytes, or null if not containerized
*/
const getContainerMemoryLimit = (): number | null => {
// Not in Node.js environment
if (!fs) {
return null
}
try {
// 1. Check environment variables first (fastest, most reliable for Cloud Run)
// Google Cloud Run
if (process.env.CLOUD_RUN_MEMORY) {
// Format: "512Mi", "1Gi", "2Gi", "4Gi"
const match = process.env.CLOUD_RUN_MEMORY.match(/^(\d+)(Mi|Gi)$/)
if (match) {
const value = parseInt(match[1])
const unit = match[2]
return unit === 'Gi' ? value * 1024 * 1024 * 1024 : value * 1024 * 1024
}
}
// Generic MEMORY_LIMIT env var (bytes)
if (process.env.MEMORY_LIMIT) {
const limit = parseInt(process.env.MEMORY_LIMIT)
if (!isNaN(limit) && limit > 0) {
return limit
}
}
// 2. Check cgroup v2 (modern Docker/K8s)
try {
const cgroupV2Path = '/sys/fs/cgroup/memory.max'
const cgroupV2Content = fs.readFileSync(cgroupV2Path, 'utf8').trim()
// "max" means no limit, otherwise it's bytes
if (cgroupV2Content !== 'max') {
const limit = parseInt(cgroupV2Content)
if (!isNaN(limit) && limit > 0) {
return limit
}
}
} catch (e) {
// cgroup v2 not available, try v1
}
// 3. Check cgroup v1 (legacy Docker/K8s)
try {
const cgroupV1Path = '/sys/fs/cgroup/memory/memory.limit_in_bytes'
const cgroupV1Content = fs.readFileSync(cgroupV1Path, 'utf8').trim()
const limit = parseInt(cgroupV1Content)
// Very large values (> 1 PB) indicate no limit
const ONE_PETABYTE = 1024 * 1024 * 1024 * 1024 * 1024
if (!isNaN(limit) && limit > 0 && limit < ONE_PETABYTE) {
return limit
}
} catch (e) {
// cgroup v1 not available
}
// Not containerized or no limit set
return null
} catch (e) {
// Error reading cgroup files
return null
}
}
/**
* Configuration options for ValidationConfig
*/
export interface ValidationConfigOptions {
/**
* Explicit maximum query limit override
* Bypasses all auto-detection
*/
maxQueryLimit?: number
/**
* Memory reserved for query operations (in bytes)
* Bypasses auto-detection but still applies safety limits
*/
reservedQueryMemory?: number
}
/**
* Auto-configured limits based on system resources
* These adapt to available memory and observed performance
*/
export class ValidationConfig {
private static instance: ValidationConfig
// Dynamic limits based on system
public maxLimit: number
public maxQueryLength: number
public maxVectorDimensions: number
// Tracking for diagnostics
public limitBasis: 'override' | 'reservedMemory' | 'containerMemory' | 'freeMemory'
public detectedContainerLimit: number | null
// Performance observations
private avgQueryTime: number = 0
private queryCount: number = 0
private constructor(options?: ValidationConfigOptions) {
// Vector dimensions (standard for all-MiniLM-L6-v2)
this.maxVectorDimensions = 384
// Detect container memory limit
this.detectedContainerLimit = getContainerMemoryLimit()
// Priority 1: Explicit override (highest priority)
if (options?.maxQueryLimit !== undefined) {
this.maxLimit = Math.min(options.maxQueryLimit, 100000) // Still cap at 100k for safety
this.limitBasis = 'override'
// Scale query length with limit
this.maxQueryLength = Math.min(50000, this.maxLimit * 5)
return
}
// Priority 2: Reserved memory specified
if (options?.reservedQueryMemory !== undefined) {
this.maxLimit = Math.min(
100000,
Math.floor(options.reservedQueryMemory / (1024 * 1024 * 100)) * 1000
)
this.limitBasis = 'reservedMemory'
this.maxQueryLength = Math.min(
50000,
Math.floor(options.reservedQueryMemory / (1024 * 1024 * 10)) * 1000
)
return
}
// Priority 3: Container detected (smart containerized behavior)
if (this.detectedContainerLimit) {
// In containers, assume 75% used by graph data (EXPECTED)
// Reserve 25% for query operations
const queryMemory = this.detectedContainerLimit * 0.25
this.maxLimit = Math.min(
100000,
Math.floor(queryMemory / (1024 * 1024 * 100)) * 1000
)
this.limitBasis = 'containerMemory'
this.maxQueryLength = Math.min(
50000,
Math.floor(queryMemory / (1024 * 1024 * 10)) * 1000
)
return
}
// Priority 4: Free memory (fallback, current behavior)
const availableMemory = getAvailableMemory()
this.maxLimit = Math.min(
100000,
Math.floor(availableMemory / (1024 * 1024 * 100)) * 1000
)
this.limitBasis = 'freeMemory'
this.maxQueryLength = Math.min(
50000,
Math.floor(availableMemory / (1024 * 1024 * 10)) * 1000
)
}
static getInstance(options?: ValidationConfigOptions): ValidationConfig {
if (!ValidationConfig.instance) {
ValidationConfig.instance = new ValidationConfig(options)
}
return ValidationConfig.instance
}
/**
* Reset singleton (for testing or reconfiguration)
*/
static reset(): void {
ValidationConfig.instance = null as any
}
/**
* Reconfigure with new options
*/
static reconfigure(options: ValidationConfigOptions): ValidationConfig {
ValidationConfig.instance = new ValidationConfig(options)
return ValidationConfig.instance
}
/**
* Learn from actual usage to adjust limits
*/
recordQuery(duration: number, resultCount: number) {
this.queryCount++
this.avgQueryTime = (this.avgQueryTime * (this.queryCount - 1) + duration) / this.queryCount
// Only auto-adjust if not using explicit overrides
if (this.limitBasis !== 'override') {
// If queries are consistently fast with large results, increase limits
if (this.avgQueryTime < 100 && resultCount > this.maxLimit * 0.8) {
this.maxLimit = Math.min(this.maxLimit * 1.5, 100000)
}
// If queries are slow, reduce limits
if (this.avgQueryTime > 1000) {
this.maxLimit = Math.max(this.maxLimit * 0.8, 1000)
}
}
}
}
/**
* Universal validations - things that are always invalid
* These are mathematical/logical truths, not configuration
*/
export function validateFindParams(params: FindParams): void {
const config = ValidationConfig.getInstance()
// Universal truth: negative pagination never makes sense
if (params.limit !== undefined) {
if (params.limit < 0) {
throw new Error('limit must be non-negative')
}
if (params.limit > config.maxLimit) {
throw new Error(`limit exceeds auto-configured maximum of ${config.maxLimit} (based on available memory)`)
}
}
if (params.offset !== undefined && params.offset < 0) {
throw new Error('offset must be non-negative')
}
// Universal truth: probability/similarity must be 0-1
if (params.near?.threshold !== undefined) {
const t = params.near.threshold
if (t < 0 || t > 1) {
throw new Error('threshold must be between 0 and 1')
}
}
// Universal truth: can't specify both query and vector (they're alternatives)
if (params.query !== undefined && params.vector !== undefined) {
throw new Error('cannot specify both query and vector - they are mutually exclusive')
}
// Universal truth: can't use both cursor and offset pagination
if (params.cursor !== undefined && params.offset !== undefined) {
throw new Error('cannot use both cursor and offset pagination simultaneously')
}
// Auto-limit query length based on memory
if (params.query && params.query.length > config.maxQueryLength) {
throw new Error(`query exceeds auto-configured maximum length of ${config.maxQueryLength} characters`)
}
// Validate vector dimensions if provided
if (params.vector && params.vector.length !== config.maxVectorDimensions) {
throw new Error(`vector must have exactly ${config.maxVectorDimensions} dimensions`)
}
// Validate enum types if specified
if (params.type) {
const types = Array.isArray(params.type) ? params.type : [params.type]
for (const type of types) {
if (!Object.values(NounType).includes(type)) {
throw new Error(`invalid NounType: ${type}`)
}
}
}
}
/**
* Validate add parameters
*/
export function validateAddParams(params: AddParams): void {
// Universal truth: must have data or vector
if (!params.data && !params.vector) {
throw new Error(
`Invalid add() parameters: Missing required field 'data'\n` +
`\nReceived: ${JSON.stringify({
type: params.type,
hasMetadata: !!params.metadata,
hasId: !!params.id
}, null, 2)}\n` +
`\nExpected one of:\n` +
` { data: 'text to store', type?: 'note', metadata?: {...} }\n` +
` { vector: [0.1, 0.2, ...], type?: 'embedding', metadata?: {...} }\n` +
`\nExamples:\n` +
` await brain.add({ data: 'Machine learning is AI', type: 'concept' })\n` +
` await brain.add({ data: { title: 'Doc', content: '...' }, type: 'document' })`
)
}
// Validate noun type
if (!Object.values(NounType).includes(params.type)) {
throw new Error(
`Invalid NounType: '${params.type}'\n` +
`\nValid types: ${Object.values(NounType).join(', ')}\n` +
`\nExample: await brain.add({ data: 'text', type: NounType.Document })`
)
}
// Validate vector dimensions if provided
if (params.vector) {
const config = ValidationConfig.getInstance()
if (params.vector.length !== config.maxVectorDimensions) {
throw new Error(`vector must have exactly ${config.maxVectorDimensions} dimensions`)
}
}
}
/**
* Validate update parameters
*/
export function validateUpdateParams(params: UpdateParams): void {
// Universal truth: must have an ID
if (!params.id) {
throw new Error('id is required for update')
}
// Universal truth: must update something
if (
!params.data &&
!params.metadata &&
!params.type &&
!params.vector &&
params.subtype === undefined &&
params.confidence === undefined &&
params.weight === undefined
) {
throw new Error('must specify at least one field to update')
}
// Validate type if changing
if (params.type && !Object.values(NounType).includes(params.type)) {
throw new Error(`invalid NounType: ${params.type}`)
}
// Validate vector dimensions if provided
if (params.vector) {
const config = ValidationConfig.getInstance()
if (params.vector.length !== config.maxVectorDimensions) {
throw new Error(`vector must have exactly ${config.maxVectorDimensions} dimensions`)
}
}
}
/**
* Validate relate parameters
*/
export function validateRelateParams(params: RelateParams): void {
// Universal truths
if (!params.from) {
throw new Error('from entity ID is required')
}
if (!params.to) {
throw new Error('to entity ID is required')
}
// Allow self-referential relationships - they're valid in graph systems
// (e.g., a person can be related to themselves, a file can reference itself, etc.)
// Validate verb type - default to RelatedTo if not specified
if (params.type === undefined) {
params.type = VerbType.RelatedTo
} else if (!Object.values(VerbType).includes(params.type)) {
throw new Error(`invalid VerbType: ${params.type}`)
}
// Universal truth: weight must be 0-1
if (params.weight !== undefined) {
if (params.weight < 0 || params.weight > 1) {
throw new Error('weight must be between 0 and 1')
}
}
}
/**
* Validate UpdateRelationParams. Mirror of validateUpdateParams for verbs —
* requires id + at least one field to change; bounds-checks weight/confidence;
* accepts type/subtype/weight/confidence/data/metadata changes.
*/
export function validateUpdateRelationParams(params: UpdateRelationParams): void {
if (!params.id) {
throw new Error('id is required for updateRelation')
}
if (
!params.data &&
!params.metadata &&
!params.type &&
params.subtype === undefined &&
params.weight === undefined &&
params.confidence === undefined
) {
throw new Error('updateRelation: must specify at least one field to update')
}
if (params.type !== undefined && !Object.values(VerbType).includes(params.type)) {
throw new Error(`invalid VerbType: ${params.type}`)
}
if (params.weight !== undefined && (params.weight < 0 || params.weight > 1)) {
throw new Error('weight must be between 0 and 1')
}
if (params.confidence !== undefined && (params.confidence < 0 || params.confidence > 1)) {
throw new Error('confidence must be between 0 and 1')
}
}
/**
* Get current validation configuration
* Useful for debugging and monitoring
*/
export function getValidationConfig() {
const config = ValidationConfig.getInstance()
return {
maxLimit: config.maxLimit,
maxQueryLength: config.maxQueryLength,
maxVectorDimensions: config.maxVectorDimensions,
systemMemory: getSystemMemory(),
availableMemory: getAvailableMemory()
}
}
/**
* Record query performance for auto-tuning
*/
export function recordQueryPerformance(duration: number, resultCount: number) {
ValidationConfig.getInstance().recordQuery(duration, resultCount)
}