brainy/docs/MODEL_LOADING_QUICK_REFERENCE.md
David Snelling 196690863d fix: update all imports and references from BrainyData to Brainy
- Fixed imports in examples/tests/ to use correct Brainy import
- Fixed imports in tests/benchmarks/ to use correct paths
- Updated bin/brainy-interactive.js to use Brainy instead of BrainyData
- Corrected documentation references throughout codebase
- Removed duplicate imports in benchmark files
- All files now consistently use 'Brainy' class from dist/index.js
2025-09-30 17:09:15 -07:00

2.8 KiB

🤖 Model Loading Quick Reference

🚀 Common Scenarios

Development (Zero Config)

const brain = new Brainy()
await brain.init() // Downloads automatically (FP32 default)

Development (Optimized - v2.8.0+)

// 75% smaller models, 99% accuracy
const brain = new Brainy({
  embeddingOptions: { dtype: 'q8' }
})
await brain.init()

🐳 Docker Production

# Both models (recommended)
RUN npm run download-models

# Or FP32 only (compatibility)
RUN npm run download-models:fp32

# Or Q8 only (space-constrained)
RUN npm run download-models:q8

ENV BRAINY_ALLOW_REMOTE_MODELS=false

☁️ Serverless/Lambda

# Build step
npm run download-models

# Runtime
export BRAINY_ALLOW_REMOTE_MODELS=false

🔒 Air-Gapped/Offline

# Connected machine
npm run download-models
tar -czf brainy-models.tar.gz ./models

# Offline machine  
tar -xzf brainy-models.tar.gz
export BRAINY_ALLOW_REMOTE_MODELS=false

🌐 Browser/CDN

<!-- Automatic - no setup needed -->
<script type="module">
  import { Brainy } from 'brainy'
  const brain = new Brainy()
  await brain.init() // Works in browser
</script>

🚨 Troubleshooting

Error Solution
"Failed to load embedding model" npm run download-models
"ENOENT: no such file" Check BRAINY_MODELS_PATH
"Network timeout" Set BRAINY_ALLOW_REMOTE_MODELS=false
"Permission denied" chmod 755 ./models
"Out of memory" Increase container memory limit

🎯 Environment Variables

Variable Values Purpose
BRAINY_ALLOW_REMOTE_MODELS true/false Allow/block downloads
BRAINY_MODELS_PATH ./models Model storage path
BRAINY_Q8_CONFIRMED true/false Silence Q8 compatibility warnings
NODE_ENV production Environment detection

📦 Model Info

FP32 (Default)

  • Model: All-MiniLM-L6-v2
  • Dimensions: 384 (fixed)
  • Size: 90MB
  • Accuracy: 100% (baseline)
  • Location: ./models/Xenova/all-MiniLM-L6-v2/onnx/model.onnx

Q8 (Optional - v2.8.0+)

  • Model: All-MiniLM-L6-v2 (quantized)
  • Dimensions: 384 (same)
  • Size: 23MB (75% smaller!)
  • Accuracy: ~99% (minimal loss)
  • Location: ./models/Xenova/all-MiniLM-L6-v2/onnx/model_quantized.onnx

⚠️ Important: FP32 and Q8 create different embeddings and are incompatible!

Verification Commands

# Check FP32 model exists
ls ./models/Xenova/all-MiniLM-L6-v2/onnx/model.onnx

# Check Q8 model exists  
ls ./models/Xenova/all-MiniLM-L6-v2/onnx/model_quantized.onnx

# Test offline mode
BRAINY_ALLOW_REMOTE_MODELS=false npm test

# Download fresh models (both)
rm -rf ./models && npm run download-models

# Download specific model variant
rm -rf ./models && npm run download-models:q8