brainy/OFFLINE_MODELS.md
David Snelling f898f0ce7b feat\!: migrate from TensorFlow.js to Transformers.js with ONNX Runtime
BREAKING CHANGE: Complete migration from TensorFlow.js to Transformers.js for embedding generation

This is a major architectural change that replaces TensorFlow.js (USE model) with Transformers.js (all-MiniLM-L6-v2) for significantly improved performance and reduced complexity.

Key Changes:
- Replace TensorFlow.js Universal Sentence Encoder with Transformers.js all-MiniLM-L6-v2
- Reduce model size from 525MB to 87MB (83% reduction)
- Reduce embedding dimensions from 512 to 384 (faster distance calculations)
- Remove TensorFlow.js Float32Array patching (caused ONNX conflicts)
- Implement smart bundled model detection for offline operation
- Add explicit model download script for Docker deployments
- Remove complex environment variables in favor of simple configuration
- Update all distance functions to use optimized pure JavaScript
- Remove TensorFlow-specific utilities and type definitions

Performance Improvements:
- Model loading: 5x faster (87MB vs 525MB)
- Memory usage: 75% reduction (~200-400MB vs ~1.5GB)
- Distance calculations: Faster pure JS vs GPU overhead for small vectors
- Cold start performance: Significantly improved

Files Changed:
- Updated package.json: New dependencies, simplified scripts
- Rewrote src/utils/embedding.ts: Complete Transformers.js implementation
- Updated src/utils/distance.ts: Optimized JavaScript distance functions
- Simplified src/setup.ts: Removed TensorFlow-specific patching
- Simplified src/utils/textEncoding.ts: Only Node.js TextEncoder/Decoder patches
- Deleted src/utils/robustModelLoader.ts: TensorFlow-specific loader
- Deleted src/types/tensorflowTypes.ts: TensorFlow type definitions
- Added scripts/download-models.cjs: Docker-compatible model downloader
- Added comprehensive documentation: README.md, OFFLINE_MODELS.md, analysis docs

Testing:
- All 19 tests passing
- Removed test mocking in favor of real implementation testing
- Updated test environment for Transformers.js compatibility
- Performance tests validate improved efficiency

This migration resolves production issues with Docker egress limitations and provides a more robust, performant foundation for vector operations.
2025-08-05 19:29:59 -07:00

1.7 KiB

Offline Models

Brainy uses Transformers.js with ONNX Runtime for true offline operation - no more TensorFlow.js dependency hell!

How it works

Brainy automatically figures out the best approach:

  1. First use: Downloads models once (~87 MB) to local cache
  2. Subsequent use: Loads from cache (completely offline, zero network calls)
  3. Smart detection: Automatically finds models in cache, bundled, or downloads as needed

Standard usage

npm install @soulcraft/brainy
# Use immediately - models download automatically on first use

Docker with production egress restrictions

For environments where production has no internet but build does:

FROM node:24-slim
WORKDIR /app
COPY package*.json ./
RUN npm install @soulcraft/brainy
RUN npm run download-models  # Download during build (when internet available)
COPY . .
# Production container now works completely offline

Development with immediate offline

If you want models available immediately for development:

npm install @soulcraft/brainy
npm run download-models  # Optional: download now instead of on first use

Key benefits vs TensorFlow.js

  • 95% smaller package - 643 kB vs 12.5 MB
  • 84% smaller models - 87 MB vs 525 MB
  • True offline - Zero network calls after initial download
  • No dependency issues - 5 deps vs 47+, no more --legacy-peer-deps
  • Better performance - ONNX Runtime beats TensorFlow.js
  • Same API - Drop-in replacement

Philosophy

Install and use. Brainy handles the rest.

No configuration files, no environment variables, no complex setup. Brainy detects your environment and does the right thing automatically.