docs: external-backups/sparse-storage guide + generation fact log concept
- New public guide (guides/external-backups): the sparse-mmap reality for external backup tooling — why a brain directory can show 100+ GB apparent size on a small disk, which files are sparse, tar czSf / rsync --sparse / cp --sparse=always, live-store caveats, and what persist()/restore() already handle natively. - New public concept (concepts/generation-fact-log): what facts are (after-image commit records, body-less tombstones), the crash-safety model (checksummed frames, reconcile-to-committed, absent = never committed), the scanFacts()/factSegmentPaths() surfaces with the telemetry shape, family stamps + open-time coherence, and the storage capability seam for plugin authors. - Cross-link from the snapshots guide; sparse notes added to the public persist()/restore() JSDoc (the operator-facing sites).
This commit is contained in:
parent
d60619c83b
commit
593bb8b0f9
5 changed files with 236 additions and 0 deletions
99
docs/guides/external-backups-and-sparse-storage.md
Normal file
99
docs/guides/external-backups-and-sparse-storage.md
Normal file
|
|
@ -0,0 +1,99 @@
|
|||
---
|
||||
title: External Backups & Sparse Storage
|
||||
slug: guides/external-backups
|
||||
public: true
|
||||
category: guides
|
||||
template: guide
|
||||
order: 10
|
||||
description: How to back up a brain directory with external tools (tar, rsync, cp) without exploding sparse files — why a store can show 100+ GB "apparent" size on a small disk, which files are sparse, and how persist()/restore() handle it for you.
|
||||
next:
|
||||
- guides/snapshots-and-time-travel
|
||||
- concepts/storage-adapters
|
||||
---
|
||||
|
||||
# External Backups & Sparse Storage
|
||||
|
||||
The built-in snapshot path — [`db.persist()` and `brain.restore()`](/docs/guides/snapshots-and-time-travel) —
|
||||
already handles everything on this page for you. Read this when you back up a brain directory with
|
||||
**external tools**: `tar`, `rsync`, `cp`, `scp`, or a filesystem-level backup agent.
|
||||
|
||||
## The one-sentence rule
|
||||
|
||||
> **Always use the sparse-aware flag**: `tar czSf` (capital `S`), `rsync --sparse`,
|
||||
> `cp --sparse=always`. A naive copy can turn a 2 GB store into a 100+ GB one — or fail
|
||||
> the disk entirely.
|
||||
|
||||
## Why: some files are sparse
|
||||
|
||||
When a native accelerator plugin is active, parts of the index live in **memory-mapped files**
|
||||
created at a large fixed virtual size — the file's *apparent* size — while the filesystem only
|
||||
allocates blocks that were actually written. A brand-new id-mapper file can report tens of
|
||||
gigabytes in `ls -l` while occupying a few megabytes on disk.
|
||||
|
||||
Check the difference yourself:
|
||||
|
||||
```bash
|
||||
ls -lh brain-data/_id_mapper/ # APPARENT size (can be huge)
|
||||
du -sh brain-data/ # ALLOCATED size (the real footprint)
|
||||
```
|
||||
|
||||
The sparse candidates in a brain directory:
|
||||
|
||||
| Path | What it is |
|
||||
|---|---|
|
||||
| `_id_mapper/` | The native id-mapper's mmap files (large fixed virtual size) |
|
||||
| `_blobs/` | Native index files (vector base, segments) — may be mmap-backed |
|
||||
|
||||
Everything else (entities, `_system`, `_generations`, `_cas` content blobs) is ordinary dense data.
|
||||
|
||||
## Doing it right
|
||||
|
||||
**tar** — the `S` flag detects holes and stores only real data:
|
||||
|
||||
```bash
|
||||
tar czSf brain-backup.tgz /data/brain
|
||||
# restore preserves the holes:
|
||||
tar xzSf brain-backup.tgz -C /data/
|
||||
```
|
||||
|
||||
**rsync**:
|
||||
|
||||
```bash
|
||||
rsync -a --sparse /data/brain/ backup-host:/backups/brain/
|
||||
```
|
||||
|
||||
**cp**:
|
||||
|
||||
```bash
|
||||
cp -a --sparse=always /data/brain /backups/brain
|
||||
```
|
||||
|
||||
**What goes wrong without the flag:** the copy *materializes* every hole as real zero bytes.
|
||||
A store whose apparent size exceeds the target disk fails with `ENOSPC` partway through — and a
|
||||
copy that *does* fit silently costs the full apparent size in storage and transfer time.
|
||||
|
||||
## What the built-in paths do (so you don't have to)
|
||||
|
||||
- **`db.persist(path)`** snapshots via **hard links** — instant and space-shared, since every data
|
||||
file is immutable-by-rename. The handful of append-in-place files (the transaction log, the
|
||||
commit fact log's tail segment) and mmap-mutated directories (`_id_mapper/`) are **byte-copied**
|
||||
instead, so a post-snapshot write can never reach through a shared inode into your backup.
|
||||
- **`brain.restore(path, { confirm: true })`** is **non-destructive and sparse-aware**: the snapshot
|
||||
is copied into a staging area *before* any live data is touched (all-zero blocks stay holes), and
|
||||
only after the copy fully succeeds does an atomic swap move it into place. A failed copy —
|
||||
including `ENOSPC` — leaves the live store exactly as it was. A crash mid-swap completes forward
|
||||
on the next open.
|
||||
|
||||
## Live-store caveats for external tools
|
||||
|
||||
1. **Prefer snapshotting a `persist()` output, not the live directory.** `persist()` produces a
|
||||
crash-consistent, immutable snapshot; running `tar` against a live, actively-written directory
|
||||
can capture a torn mid-write state. If you must archive live, stop writes first (or accept that
|
||||
the archive is only as consistent as the moment's flush state).
|
||||
2. **Never prune or "clean up" files inside a brain directory.** Index files that look stale or
|
||||
redundant are load-bearing; the store protects its declared index families from in-process
|
||||
deletion, but an external `rm` bypasses that fence. If space is the concern, `du -sh` first —
|
||||
the allocated size is usually far smaller than it looks.
|
||||
3. **Verify restores by opening them.** `Brainy.load(path)` opens any snapshot or restored
|
||||
directory read-only — the store verifies its own coherence at open and reports loudly if
|
||||
anything is missing or torn.
|
||||
|
|
@ -9,6 +9,7 @@ description: Recipes for the Db API — instant backups with persist(), restore,
|
|||
next:
|
||||
- concepts/consistency-model
|
||||
- guides/optimistic-concurrency
|
||||
- guides/external-backups
|
||||
---
|
||||
|
||||
# Snapshots & Time Travel
|
||||
|
|
@ -41,6 +42,10 @@ bytes. Cross-device targets fall back to per-file byte copies, and
|
|||
persisting an in-memory brain serializes it to the same directory layout —
|
||||
a real, durable store.
|
||||
|
||||
> Archiving a brain directory with **external tools** (`tar`, `rsync`, `cp`)?
|
||||
> Some index files are sparse and can explode to their apparent size under a
|
||||
> naive copy — see [External Backups & Sparse Storage](/docs/guides/external-backups).
|
||||
|
||||
Two things to know:
|
||||
|
||||
- `persist()` requires the view to still be the store's **latest**
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue