The cutover: new tail segments write format v2 (per-record [type, version, cipherFlag, keyId] envelope; noun/verb after-images carry dense ints MINTED AT APPEND from the id mapper — a rebuilt mapper reproduces assignments exactly; log.genesis opens every new log with the id-space width + a minted brain id; sync() seals to the header-declared sector boundary with reader-invisible pad frames). Existing v1 segments are never rewritten — per-segment decoder dispatch reads both formats and v2 facts map to the exact CommitFact shape all consumers already read. Cutover on a live v1 log: an empty v1 tail re-heads in place; a non-empty one is sealed by rotation, byte-identical. Records reserve the encryption fields (cipherFlag 0 / keyId nil are the only legal values; anything else refuses typed naming the needed newer reader) — crypto-ready with no future bump on the compat surface. Empty-records facts are legal (an all-deduped batch is a real generation — v1 semantics preserved; the refusal there tore a column-store flush mid-commit in the full suite, the consistency guard caught it loudly, and the root is fixed). Golden byte vectors pinned for the second (native) reader implementation. Pins: cutover 5/5 · codec 54 · kill-matrix stays 11/11.
1305 lines
49 KiB
TypeScript
1305 lines
49 KiB
TypeScript
/**
|
|
* @module db/factLogFormat
|
|
* @description Fact-log format v2 (record envelope + sector seals) — the pure
|
|
* encode/decode functions for the versioned on-disk fact-log byte format.
|
|
* No I/O and no storage dependencies live here: this module is the REFERENCE
|
|
* IMPLEMENTATION of the format, and a second (native) reader parses these
|
|
* exact bytes. Byte-level behavior is a two-implementation contract — bytes
|
|
* change only behind a format-version bump, never in place.
|
|
*
|
|
* ## Segment header (32 bytes, both versions)
|
|
*
|
|
* magic "BFACTS\0\0" (8B) | formatVersion:u32 LE | firstGeneration:u64 LE |
|
|
* v1: reserved 12B (ZEROED, verified)
|
|
* v2: sealSize:u16 LE at offset +20 | reserved 10B (ZEROED, verified)
|
|
*
|
|
* V1 segments remain readable forever via the v1 decode path — never rewritten.
|
|
*
|
|
* ## Frame (unchanged from v1)
|
|
*
|
|
* payloadLength:u32 LE | crc32c:u32 LE (of payload) | msgpack payload
|
|
*
|
|
* A bad length (overruns the buffer) or CRC mismatch is a TORN TAIL: it
|
|
* terminates the scan; everything before it is intact.
|
|
*
|
|
* ## V2 fact payload (msgpack, positional — same 5 positions as v1, but
|
|
* position 2 is `records`, not v1's `ops`)
|
|
*
|
|
* fact := [ generation:u64, timestamp:u64, records, meta|nil, blobHashes|nil ]
|
|
* record := [ recordType:u8, recordVersion:u8, cipherFlag:u8, keyId:bin16|nil,
|
|
* ...type-specific fields ]
|
|
*
|
|
* `cipherFlag`/`keyId` are RESERVED crypto envelope fields: `0`/`nil` (a
|
|
* plaintext record) is the ONLY legal combination this release writes or
|
|
* reads. Any nonzero cipherFlag or non-nil keyId refuses with the typed
|
|
* {@link UnknownLogRecordError} ("encrypted records need a newer reader") —
|
|
* so record-level encryption can land later without a format-version bump on
|
|
* the one compat surface. No crypto logic exists here; the bytes are reserved
|
|
* only. Pad records (type 0) are exempt: they are skipped WHOLESALE as
|
|
* length-only filler, so their fields beyond [type, version] are never
|
|
* inspected (this keeps pad frames byte-stable across the envelope change).
|
|
*
|
|
* Record type registry (all recordVersion = 1; type-specific fields listed —
|
|
* every record carries the 4-field envelope above first):
|
|
*
|
|
* 0 pad [] — length-only filler; readers SKIP; crc-covered
|
|
* 1 noun.afterImage [id bin16, entityInt u64, metadata, vectorLeg]
|
|
* 2 noun.tombstone [id bin16]
|
|
* 3 verb.afterImage [id bin16, verbInt u64, metadata, vectorLeg,
|
|
* verb str, sourceId bin16, sourceInt u64,
|
|
* targetId bin16, targetInt u64]
|
|
* 4 verb.tombstone [id bin16]
|
|
* 5 batch.meta [metaMap] — at most ONE per fact
|
|
* 6 embed.pending [id bin16, enqueuedAt u64]
|
|
* 7 embed.landed [id bin16, vector — INLINE float[] only]
|
|
* 8 blob.manifest [hash bin32, size u64, mimeType str, refOp u8 (0=add,1=release)]
|
|
* 9 projection.note [noteMap] — opaque map, reserved consumer
|
|
* 10 bootstrap.baseline [id bin16, kind u8 (0=noun,1=verb), metadata, vectorLeg]
|
|
* 11 log.genesis [idSpaceWidth u8 (32|64), brainId bin16, createdAt u64]
|
|
* — MUST be the first record of the first fact in a
|
|
* v2 log (first-record-of-fact is enforced here; the
|
|
* first-fact-of-log half belongs to the log layer)
|
|
*
|
|
* vectorLeg := float[] | ['ref', sameAsGeneration u64] | nil
|
|
*
|
|
* Integer wire discipline (reference encoder): every field declared u64 above
|
|
* rides as msgpack uint64 (0xcf, fixed 8 bytes); u8 fields ride as minimal
|
|
* msgpack uints (positive fixint). The decoder is liberal and accepts any
|
|
* msgpack unsigned-integer width for these fields. `entityInt`/`verbInt`/
|
|
* `sourceInt`/`targetInt` surface as `bigint` (full u64 range); scalar
|
|
* counters and timestamps surface as `number` and refuse values beyond
|
|
* `Number.MAX_SAFE_INTEGER` loudly.
|
|
*
|
|
* ## Decoder law
|
|
*
|
|
* An unknown recordType, or a recordVersion newer than this reader knows,
|
|
* throws {@link UnknownLogRecordError} — NEVER skip-and-continue (type 0 pad
|
|
* is the sole exception: skipped by definition). A log.genesis whose
|
|
* idSpaceWidth disagrees with the caller's expected width throws
|
|
* {@link GenesisWidthMismatchError} naming both widths.
|
|
*
|
|
* ## Sector seals
|
|
*
|
|
* A "sealed group" is one or more frames padded to the next `sealSize`
|
|
* boundary with ONE pad frame — a frame whose fact is
|
|
* `[0, 0, [[0, 1, filler?]], nil, nil]` (generation 0 marks filler; real
|
|
* facts start at 1). Pad frames are invisible to readers. When the gap to the
|
|
* boundary is smaller than the smallest constructible pad frame, the group is
|
|
* padded through to the boundary AFTER next (one extra sealSize) — chosen as
|
|
* the simpler correct approach over rewriting the previous frame's payload:
|
|
* input frames stay byte-immutable, alignment still holds, and the cost is at
|
|
* most one sector on a rare (<1%) size coincidence.
|
|
*/
|
|
import { encode as msgpackEncode, decode as msgpackDecode } from '@msgpack/msgpack'
|
|
import { crc32c } from '../utils/crc32c.js'
|
|
import type { CommitFact } from './factLog.js'
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Constants
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Segment magic: ASCII "BFACTS" + two NULs (shared by v1 and v2 headers). */
|
|
export const FACT_SEGMENT_MAGIC: Uint8Array = new Uint8Array([
|
|
0x42, 0x46, 0x41, 0x43, 0x54, 0x53, 0x00, 0x00
|
|
])
|
|
|
|
/** Segment format version 1 (ops-shaped facts, 12 zeroed reserved bytes). */
|
|
export const FACT_LOG_FORMAT_V1 = 1
|
|
|
|
/** Segment format version 2 (record envelope + sector seals). */
|
|
export const FACT_LOG_FORMAT_V2 = 2
|
|
|
|
/** Segment header size in bytes (identical for v1 and v2). */
|
|
export const SEGMENT_HEADER_BYTES = 32
|
|
|
|
/** Frame prefix size: payloadLength(4) + crc32c(4). */
|
|
export const FRAME_PREFIX_BYTES = 8
|
|
|
|
/** Default sector-seal size (bytes) when the caller does not probe a device. */
|
|
export const DEFAULT_SEAL_SIZE = 4096
|
|
|
|
/** The record version this reader knows (all registry types are version 1). */
|
|
export const LOG_RECORD_VERSION = 1
|
|
|
|
/**
|
|
* The only legal `cipherFlag` value this release: plaintext. The encoder
|
|
* always writes it (with a nil keyId); the decoder refuses anything else
|
|
* with {@link UnknownLogRecordError} — encrypted records need a newer reader.
|
|
*/
|
|
export const LOG_RECORD_CIPHER_PLAINTEXT = 0
|
|
|
|
/** The v2 record-type registry — wire codes for every record type. */
|
|
export const LOG_RECORD_TYPES = {
|
|
PAD: 0,
|
|
NOUN_AFTER_IMAGE: 1,
|
|
NOUN_TOMBSTONE: 2,
|
|
VERB_AFTER_IMAGE: 3,
|
|
VERB_TOMBSTONE: 4,
|
|
BATCH_META: 5,
|
|
EMBED_PENDING: 6,
|
|
EMBED_LANDED: 7,
|
|
BLOB_MANIFEST: 8,
|
|
PROJECTION_NOTE: 9,
|
|
BOOTSTRAP_BASELINE: 10,
|
|
LOG_GENESIS: 11
|
|
} as const
|
|
|
|
/** A wire code from the v2 record-type registry. */
|
|
export type LogRecordTypeCode = (typeof LOG_RECORD_TYPES)[keyof typeof LOG_RECORD_TYPES]
|
|
|
|
const U64_MAX = (1n << 64n) - 1n
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Errors
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* A record whose type or version this reader does not know. Thrown — never
|
|
* skipped — so an old reader can NEVER silently drop data written by a newer
|
|
* writer. Carries the offending type/version for programmatic handling.
|
|
*/
|
|
export class UnknownLogRecordError extends Error {
|
|
/** The wire recordType that was not understood. */
|
|
public readonly recordType: number
|
|
/** The wire recordVersion that was not understood. */
|
|
public readonly recordVersion: number
|
|
|
|
constructor(recordType: number, recordVersion: number, message: string) {
|
|
super(message)
|
|
this.name = 'UnknownLogRecordError'
|
|
this.recordType = recordType
|
|
this.recordVersion = recordVersion
|
|
}
|
|
}
|
|
|
|
/**
|
|
* A log.genesis record whose id-space width disagrees with the width the
|
|
* caller expects. Decoding across id-space widths is refused loudly — the
|
|
* error names both widths.
|
|
*/
|
|
export class GenesisWidthMismatchError extends Error {
|
|
/** The width the caller expected (32 or 64). */
|
|
public readonly expectedWidth: number
|
|
/** The width the genesis record declares (32 or 64). */
|
|
public readonly actualWidth: number
|
|
|
|
constructor(expectedWidth: number, actualWidth: number) {
|
|
super(
|
|
`fact log v2: log.genesis declares a ${actualWidth}-bit id space but this reader ` +
|
|
`expected ${expectedWidth}-bit — refusing to decode across id-space widths`
|
|
)
|
|
this.name = 'GenesisWidthMismatchError'
|
|
this.expectedWidth = expectedWidth
|
|
this.actualWidth = actualWidth
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Record + fact types (the TS surface of the wire registry)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** A vector reference: "same vector as the one generation N carried inline". */
|
|
export interface VectorRef {
|
|
/** The generation whose record carried the INLINE vector (single-hop only). */
|
|
sameAsGeneration: number
|
|
}
|
|
|
|
/** A record's vector leg: inline floats, a single-hop ref, or none. */
|
|
export type VectorLeg = number[] | VectorRef | null
|
|
|
|
/** Type 1 — the after-image of a noun: what the entity BECAME. */
|
|
export interface NounAfterImageRecord {
|
|
type: 'noun.afterImage'
|
|
id: string
|
|
/** The entity's u64 integer handle (full range — hence bigint). */
|
|
entityInt: bigint
|
|
metadata: unknown
|
|
vectorLeg: VectorLeg
|
|
}
|
|
|
|
/** Type 2 — a body-less noun tombstone: the entity was removed. */
|
|
export interface NounTombstoneRecord {
|
|
type: 'noun.tombstone'
|
|
id: string
|
|
}
|
|
|
|
/** Type 3 — the after-image of a verb (relationship), endpoints included. */
|
|
export interface VerbAfterImageRecord {
|
|
type: 'verb.afterImage'
|
|
id: string
|
|
/** The verb's u64 integer handle (full range — hence bigint). */
|
|
verbInt: bigint
|
|
metadata: unknown
|
|
vectorLeg: VectorLeg
|
|
/** The verb name (relationship type). */
|
|
verb: string
|
|
sourceId: string
|
|
sourceInt: bigint
|
|
targetId: string
|
|
targetInt: bigint
|
|
}
|
|
|
|
/** Type 4 — a body-less verb tombstone: the relationship was removed. */
|
|
export interface VerbTombstoneRecord {
|
|
type: 'verb.tombstone'
|
|
id: string
|
|
}
|
|
|
|
/** Type 5 — batch-level metadata; at most ONE per fact. */
|
|
export interface BatchMetaRecord {
|
|
type: 'batch.meta'
|
|
meta: Record<string, unknown>
|
|
}
|
|
|
|
/** Type 6 — an embedding was enqueued for the id (vector not yet available). */
|
|
export interface EmbedPendingRecord {
|
|
type: 'embed.pending'
|
|
id: string
|
|
/** Enqueue time (epoch ms). */
|
|
enqueuedAt: number
|
|
}
|
|
|
|
/** Type 7 — a deferred embedding landed; carries the INLINE vector only. */
|
|
export interface EmbedLandedRecord {
|
|
type: 'embed.landed'
|
|
id: string
|
|
/** The landed vector — inline floats only; refs are not allowed here. */
|
|
vector: number[]
|
|
}
|
|
|
|
/** Type 8 — a blob reference-count event (content-addressed by hash). */
|
|
export interface BlobManifestRecord {
|
|
type: 'blob.manifest'
|
|
/** The blob's content hash — 64 lowercase hex chars (bin32 on the wire). */
|
|
hash: string
|
|
size: number
|
|
mimeType: string
|
|
refOp: 'add' | 'release'
|
|
}
|
|
|
|
/** Type 9 — an opaque note for a reserved projection consumer. */
|
|
export interface ProjectionNoteRecord {
|
|
type: 'projection.note'
|
|
note: Record<string, unknown>
|
|
}
|
|
|
|
/** Type 10 — a bootstrap baseline row (initial-load after-image). */
|
|
export interface BootstrapBaselineRecord {
|
|
type: 'bootstrap.baseline'
|
|
id: string
|
|
kind: 'noun' | 'verb'
|
|
metadata: unknown
|
|
vectorLeg: VectorLeg
|
|
}
|
|
|
|
/** Type 11 — the log's birth certificate; first record of the first fact. */
|
|
export interface LogGenesisRecord {
|
|
type: 'log.genesis'
|
|
/** The integer-handle width this log's records use. */
|
|
idSpaceWidth: 32 | 64
|
|
brainId: string
|
|
/** Creation time (epoch ms). */
|
|
createdAt: number
|
|
}
|
|
|
|
/** Any decodable v2 record (pads are skipped, never surfaced). */
|
|
export type LogRecord =
|
|
| NounAfterImageRecord
|
|
| NounTombstoneRecord
|
|
| VerbAfterImageRecord
|
|
| VerbTombstoneRecord
|
|
| BatchMetaRecord
|
|
| EmbedPendingRecord
|
|
| EmbedLandedRecord
|
|
| BlobManifestRecord
|
|
| ProjectionNoteRecord
|
|
| BootstrapBaselineRecord
|
|
| LogGenesisRecord
|
|
|
|
/** One committed generation in v2 shape: a record envelope, not v1 ops. */
|
|
export interface CommitFactV2 {
|
|
generation: number
|
|
timestamp: number
|
|
records: LogRecord[]
|
|
meta?: Record<string, unknown>
|
|
blobHashes?: string[]
|
|
}
|
|
|
|
/** A parsed segment header (v1 has no sealSize; v2 always carries one). */
|
|
export interface SegmentHeader {
|
|
formatVersion: number
|
|
firstGeneration: number
|
|
/** Sector-seal size (v2 only) — `undefined` on v1 headers. */
|
|
sealSize?: number
|
|
}
|
|
|
|
/** Options for {@link encodeFactV2}. */
|
|
export interface EncodeFactV2Options {
|
|
/**
|
|
* Single-hop validator for vector refs: the set (or predicate) of
|
|
* generations whose records carried an INLINE vector. REQUIRED whenever any
|
|
* record carries a `VectorRef` — encoding an unverifiable ref is refused.
|
|
*/
|
|
inlineVectorGenerations?: Set<number> | ((generation: number) => boolean)
|
|
}
|
|
|
|
/** Options for the v2 decode path of {@link decodeFact}. */
|
|
export interface DecodeFactV2Options {
|
|
/**
|
|
* The id-space width the caller expects. When set and the fact carries a
|
|
* log.genesis record, a disagreeing width throws
|
|
* {@link GenesisWidthMismatchError}.
|
|
*/
|
|
expectedIdSpaceWidth?: 32 | 64
|
|
}
|
|
|
|
/** The result of decoding a frame group: intact facts + valid byte length. */
|
|
export interface DecodedFrameGroup {
|
|
facts: CommitFactV2[]
|
|
/** Byte length of the intact prefix (whole frames that decoded cleanly). */
|
|
validBytes: number
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// msgpack wire helpers
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* The v2 codec: `useBigInt64` makes bigints ride as fixed 8-byte uint64/int64
|
|
* (the u64 wire discipline) while JS numbers keep exact-value round-trips
|
|
* (integers ≤ 32-bit ride minimal; larger numbers ride float64, which holds
|
|
* every safe integer exactly).
|
|
*/
|
|
const enc = (value: unknown): Uint8Array => msgpackEncode(value, { useBigInt64: true })
|
|
const dec = (bytes: Uint8Array): unknown => msgpackDecode(bytes, { useBigInt64: true })
|
|
|
|
/** Coerce an encode-side u64 field to bigint, refusing out-of-range values. */
|
|
function toWireU64(value: number | bigint, field: string): bigint {
|
|
let big: bigint
|
|
if (typeof value === 'bigint') {
|
|
big = value
|
|
} else if (Number.isSafeInteger(value)) {
|
|
big = BigInt(value)
|
|
} else {
|
|
throw new Error(`fact log v2: ${field} must be a safe integer or bigint; got ${value}`)
|
|
}
|
|
if (big < 0n || big > U64_MAX) {
|
|
throw new Error(`fact log v2: ${field} is out of u64 range: ${big}`)
|
|
}
|
|
return big
|
|
}
|
|
|
|
/** Decode-side u64 → bigint (liberal: accepts any msgpack uint width). */
|
|
function wireToBigint(value: unknown, field: string): bigint {
|
|
if (typeof value === 'bigint') {
|
|
if (value < 0n || value > U64_MAX) {
|
|
throw new Error(`fact log v2: ${field} is out of u64 range: ${value}`)
|
|
}
|
|
return value
|
|
}
|
|
if (typeof value === 'number' && Number.isSafeInteger(value) && value >= 0) {
|
|
return BigInt(value)
|
|
}
|
|
throw new Error(`fact log v2: ${field} is not an unsigned integer`)
|
|
}
|
|
|
|
/** Decode-side u64 → number, refusing values beyond safe-integer range. */
|
|
function wireToNumber(value: unknown, field: string): number {
|
|
const big = wireToBigint(value, field)
|
|
if (big > BigInt(Number.MAX_SAFE_INTEGER)) {
|
|
throw new Error(`fact log v2: ${field} ${big} exceeds Number.MAX_SAFE_INTEGER`)
|
|
}
|
|
return Number(big)
|
|
}
|
|
|
|
/** Decode-side u8 (record types, kinds, flags). */
|
|
function wireToU8(value: unknown, field: string): number {
|
|
const n = typeof value === 'bigint' ? Number(value) : value
|
|
if (typeof n !== 'number' || !Number.isInteger(n) || n < 0 || n > 255) {
|
|
throw new Error(`fact log v2: ${field} is not a u8`)
|
|
}
|
|
return n
|
|
}
|
|
|
|
/** uuid string → 16 raw bytes (bin16 on the wire). */
|
|
function uuidToBytes(id: string): Uint8Array {
|
|
const hex = id.replace(/-/g, '')
|
|
if (hex.length !== 32 || /[^0-9a-fA-F]/.test(hex)) {
|
|
throw new Error(`fact log v2: id is not a uuid: ${id}`)
|
|
}
|
|
const bytes = new Uint8Array(16)
|
|
for (let i = 0; i < 16; i++) {
|
|
bytes[i] = parseInt(hex.slice(i * 2, i * 2 + 2), 16)
|
|
}
|
|
return bytes
|
|
}
|
|
|
|
/** 16 raw bytes → canonical lowercase uuid string. */
|
|
function bytesToUuid(bytes: unknown, field: string): string {
|
|
if (!(bytes instanceof Uint8Array) || bytes.length !== 16) {
|
|
throw new Error(`fact log v2: ${field} is not a bin16 id`)
|
|
}
|
|
let hex = ''
|
|
for (let i = 0; i < 16; i++) hex += bytes[i].toString(16).padStart(2, '0')
|
|
return `${hex.slice(0, 8)}-${hex.slice(8, 12)}-${hex.slice(12, 16)}-${hex.slice(16, 20)}-${hex.slice(20)}`
|
|
}
|
|
|
|
/** 64-hex-char content hash → 32 raw bytes (bin32 on the wire). */
|
|
function hashToBytes(hash: string): Uint8Array {
|
|
if (typeof hash !== 'string' || !/^[0-9a-fA-F]{64}$/.test(hash)) {
|
|
throw new Error(`fact log v2: blob hash must be 64 hex chars; got ${String(hash).slice(0, 80)}`)
|
|
}
|
|
const bytes = new Uint8Array(32)
|
|
for (let i = 0; i < 32; i++) {
|
|
bytes[i] = parseInt(hash.slice(i * 2, i * 2 + 2), 16)
|
|
}
|
|
return bytes
|
|
}
|
|
|
|
/** 32 raw bytes → 64-char lowercase hex content hash. */
|
|
function bytesToHash(bytes: unknown): string {
|
|
if (!(bytes instanceof Uint8Array) || bytes.length !== 32) {
|
|
throw new Error('fact log v2: blob hash is not bin32')
|
|
}
|
|
let hex = ''
|
|
for (let i = 0; i < 32; i++) hex += bytes[i].toString(16).padStart(2, '0')
|
|
return hex
|
|
}
|
|
|
|
/** True for a plain map object (not null/array/binary). */
|
|
function isPlainMap(value: unknown): value is Record<string, unknown> {
|
|
return (
|
|
typeof value === 'object' &&
|
|
value !== null &&
|
|
!Array.isArray(value) &&
|
|
!(value instanceof Uint8Array)
|
|
)
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Segment header (v1 read + v2 read/write)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Build a v2 segment header: magic + formatVersion 2 + firstGeneration u64 LE
|
|
* + sealSize u16 LE at offset +20. The remaining 10 reserved bytes stay zero
|
|
* and are verified by every reader.
|
|
*
|
|
* @param firstGeneration - The first generation this segment will hold.
|
|
* @param sealSize - The sector-seal size groups in this segment align to
|
|
* (device atomic-write probing is the caller's business; default 4096).
|
|
*/
|
|
export function encodeSegmentHeaderV2(
|
|
firstGeneration: number,
|
|
sealSize: number = DEFAULT_SEAL_SIZE
|
|
): Uint8Array {
|
|
if (!Number.isSafeInteger(firstGeneration) || firstGeneration < 0) {
|
|
throw new Error(`fact log v2: firstGeneration must be a non-negative integer; got ${firstGeneration}`)
|
|
}
|
|
assertValidSealSize(sealSize)
|
|
const header = new Uint8Array(SEGMENT_HEADER_BYTES)
|
|
header.set(FACT_SEGMENT_MAGIC, 0)
|
|
const view = new DataView(header.buffer)
|
|
view.setUint32(8, FACT_LOG_FORMAT_V2, true)
|
|
view.setBigUint64(12, BigInt(firstGeneration), true)
|
|
view.setUint16(20, sealSize, true)
|
|
// bytes 22..31 stay zero (reserved, verified)
|
|
return header
|
|
}
|
|
|
|
/**
|
|
* Parse a segment header — reads BOTH v1 (version 1, twelve zeroed reserved
|
|
* bytes, no sealSize) and v2 (version 2, sealSize u16 LE at +20, ten zeroed
|
|
* reserved bytes). Bad magic, non-zero reserved bytes, or an unknown version
|
|
* throw loudly; nothing is guessed.
|
|
*
|
|
* @param bytes - At least the first {@link SEGMENT_HEADER_BYTES} of a segment.
|
|
* @returns The parsed header; `sealSize` is `undefined` for v1 headers.
|
|
*/
|
|
export function parseSegmentHeader(bytes: Uint8Array): SegmentHeader {
|
|
if (bytes.length < SEGMENT_HEADER_BYTES) {
|
|
throw new Error(
|
|
`fact log: segment header needs ${SEGMENT_HEADER_BYTES} bytes; got ${bytes.length}`
|
|
)
|
|
}
|
|
for (let i = 0; i < FACT_SEGMENT_MAGIC.length; i++) {
|
|
if (bytes[i] !== FACT_SEGMENT_MAGIC[i]) {
|
|
throw new Error('fact log: bad magic — not a fact segment')
|
|
}
|
|
}
|
|
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength)
|
|
const formatVersion = view.getUint32(8, true)
|
|
const firstGenerationBig = view.getBigUint64(12, true)
|
|
if (firstGenerationBig > BigInt(Number.MAX_SAFE_INTEGER)) {
|
|
throw new Error(`fact log: firstGeneration ${firstGenerationBig} exceeds Number.MAX_SAFE_INTEGER`)
|
|
}
|
|
const firstGeneration = Number(firstGenerationBig)
|
|
|
|
if (formatVersion === FACT_LOG_FORMAT_V1) {
|
|
assertReservedZero(bytes, 20)
|
|
return { formatVersion, firstGeneration }
|
|
}
|
|
if (formatVersion === FACT_LOG_FORMAT_V2) {
|
|
const sealSize = view.getUint16(20, true)
|
|
assertReservedZero(bytes, 22)
|
|
return { formatVersion, firstGeneration, sealSize }
|
|
}
|
|
throw new Error(
|
|
`fact log: segment formatVersion ${formatVersion}; this build reads 1 and 2 — ` +
|
|
`a newer reader is required`
|
|
)
|
|
}
|
|
|
|
/** Verify header bytes [from, 32) are zero — anything else is unverifiable. */
|
|
function assertReservedZero(bytes: Uint8Array, from: number): void {
|
|
for (let i = from; i < SEGMENT_HEADER_BYTES; i++) {
|
|
if (bytes[i] !== 0) {
|
|
throw new Error('fact log: non-zero reserved header bytes — unverifiable')
|
|
}
|
|
}
|
|
}
|
|
|
|
/** Refuse seal sizes the header cannot carry or a pad frame cannot fill. */
|
|
function assertValidSealSize(sealSize: number): void {
|
|
if (!Number.isInteger(sealSize) || sealSize < 64 || sealSize > 0xffff) {
|
|
throw new Error(
|
|
`fact log v2: sealSize must be an integer in [64, 65535]; got ${sealSize}`
|
|
)
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Frames
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Wrap a msgpack payload in the frame envelope (length + crc32c + payload). */
|
|
function buildFrame(payload: Uint8Array): Uint8Array {
|
|
const frame = new Uint8Array(FRAME_PREFIX_BYTES + payload.length)
|
|
const view = new DataView(frame.buffer)
|
|
view.setUint32(0, payload.length, true)
|
|
view.setUint32(4, crc32c(payload), true)
|
|
frame.set(payload, FRAME_PREFIX_BYTES)
|
|
return frame
|
|
}
|
|
|
|
/**
|
|
* Verify a complete frame (exact length, CRC) and return its msgpack payload
|
|
* (a view into the frame — copy if you outlive the frame). The bridge between
|
|
* frame-level producers ({@link encodeFactV2}, {@link sealGroup}) and the
|
|
* payload-level {@link decodeFact}.
|
|
*/
|
|
export function framePayload(frame: Uint8Array): Uint8Array {
|
|
if (frame.length < FRAME_PREFIX_BYTES) {
|
|
throw new Error(`fact log: frame shorter than its ${FRAME_PREFIX_BYTES}-byte prefix`)
|
|
}
|
|
const view = new DataView(frame.buffer, frame.byteOffset, frame.byteLength)
|
|
const length = view.getUint32(0, true)
|
|
if (FRAME_PREFIX_BYTES + length !== frame.length) {
|
|
throw new Error(
|
|
`fact log: frame declares ${length} payload bytes but carries ${frame.length - FRAME_PREFIX_BYTES}`
|
|
)
|
|
}
|
|
const payload = frame.subarray(FRAME_PREFIX_BYTES)
|
|
const expectedCrc = view.getUint32(4, true)
|
|
if (crc32c(payload) !== expectedCrc) {
|
|
throw new Error('fact log: frame payload fails its crc32c')
|
|
}
|
|
return payload
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// vectorLeg encode/decode
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Encode a vector leg; refs must pass the single-hop validator. */
|
|
function encodeVectorLeg(
|
|
leg: VectorLeg | undefined,
|
|
options: EncodeFactV2Options | undefined,
|
|
context: string
|
|
): unknown {
|
|
if (leg === null || leg === undefined) return null
|
|
if (Array.isArray(leg)) {
|
|
for (const value of leg) {
|
|
if (typeof value !== 'number') {
|
|
throw new Error(`fact log v2: ${context} inline vector has a non-number element`)
|
|
}
|
|
}
|
|
return leg
|
|
}
|
|
if (isPlainMap(leg) && typeof (leg as VectorRef).sameAsGeneration === 'number') {
|
|
const target = (leg as VectorRef).sameAsGeneration
|
|
const validator = options?.inlineVectorGenerations
|
|
if (!validator) {
|
|
throw new Error(
|
|
`fact log v2: ${context} carries a vector ref to generation ${target} but no ` +
|
|
`single-hop validator was provided — refusing to encode an unverifiable ref`
|
|
)
|
|
}
|
|
const targetIsInline = typeof validator === 'function' ? validator(target) : validator.has(target)
|
|
if (!targetIsInline) {
|
|
throw new Error(
|
|
`fact log v2: ${context} vector ref targets generation ${target}, which did not ` +
|
|
`carry an inline vector — refs must be single-hop`
|
|
)
|
|
}
|
|
return ['ref', toWireU64(target, `${context} sameAsGeneration`)]
|
|
}
|
|
throw new Error(`fact log v2: ${context} has a malformed vector leg`)
|
|
}
|
|
|
|
/** Decode a vector leg: floats, a single-hop ref, or null. */
|
|
function decodeVectorLeg(wire: unknown, context: string): VectorLeg {
|
|
if (wire === null || wire === undefined) return null
|
|
if (Array.isArray(wire)) {
|
|
if (wire.length === 2 && wire[0] === 'ref') {
|
|
return { sameAsGeneration: wireToNumber(wire[1], `${context} sameAsGeneration`) }
|
|
}
|
|
return wire.map((value, i) => {
|
|
if (typeof value === 'number') return value
|
|
if (typeof value === 'bigint') return Number(value)
|
|
throw new Error(`fact log v2: ${context} vector element ${i} is not a number`)
|
|
})
|
|
}
|
|
throw new Error(`fact log v2: ${context} has a malformed vector leg`)
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Record encode/decode
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Encode one record into its positional wire array. Every record leads with
|
|
* the 4-field envelope [type, version, cipherFlag, keyId]; this release
|
|
* writes cipherFlag {@link LOG_RECORD_CIPHER_PLAINTEXT} and a nil keyId
|
|
* always (the fields are crypto-RESERVED, carrying no logic yet).
|
|
*/
|
|
function encodeRecord(record: LogRecord, options: EncodeFactV2Options | undefined): unknown[] {
|
|
const T = LOG_RECORD_TYPES
|
|
const V = LOG_RECORD_VERSION
|
|
const C = LOG_RECORD_CIPHER_PLAINTEXT
|
|
const K = null // keyId: nil until record-level encryption exists
|
|
switch (record.type) {
|
|
case 'noun.afterImage':
|
|
return [
|
|
T.NOUN_AFTER_IMAGE,
|
|
V,
|
|
C,
|
|
K,
|
|
uuidToBytes(record.id),
|
|
toWireU64(record.entityInt, 'entityInt'),
|
|
record.metadata ?? null,
|
|
encodeVectorLeg(record.vectorLeg, options, `noun.afterImage ${record.id}`)
|
|
]
|
|
case 'noun.tombstone':
|
|
return [T.NOUN_TOMBSTONE, V, C, K, uuidToBytes(record.id)]
|
|
case 'verb.afterImage': {
|
|
if (typeof record.verb !== 'string' || record.verb.length === 0) {
|
|
throw new Error(`fact log v2: verb.afterImage ${record.id} needs a non-empty verb name`)
|
|
}
|
|
return [
|
|
T.VERB_AFTER_IMAGE,
|
|
V,
|
|
C,
|
|
K,
|
|
uuidToBytes(record.id),
|
|
toWireU64(record.verbInt, 'verbInt'),
|
|
record.metadata ?? null,
|
|
encodeVectorLeg(record.vectorLeg, options, `verb.afterImage ${record.id}`),
|
|
record.verb,
|
|
uuidToBytes(record.sourceId),
|
|
toWireU64(record.sourceInt, 'sourceInt'),
|
|
uuidToBytes(record.targetId),
|
|
toWireU64(record.targetInt, 'targetInt')
|
|
]
|
|
}
|
|
case 'verb.tombstone':
|
|
return [T.VERB_TOMBSTONE, V, C, K, uuidToBytes(record.id)]
|
|
case 'batch.meta':
|
|
if (!isPlainMap(record.meta)) {
|
|
throw new Error('fact log v2: batch.meta requires a map')
|
|
}
|
|
return [T.BATCH_META, V, C, K, record.meta]
|
|
case 'embed.pending':
|
|
return [
|
|
T.EMBED_PENDING,
|
|
V,
|
|
C,
|
|
K,
|
|
uuidToBytes(record.id),
|
|
toWireU64(record.enqueuedAt, 'enqueuedAt')
|
|
]
|
|
case 'embed.landed': {
|
|
if (!Array.isArray(record.vector) || record.vector.some((v) => typeof v !== 'number')) {
|
|
throw new Error(
|
|
`fact log v2: embed.landed ${record.id} carries an INLINE float vector only — ` +
|
|
`refs and nil are not allowed here`
|
|
)
|
|
}
|
|
return [T.EMBED_LANDED, V, C, K, uuidToBytes(record.id), record.vector]
|
|
}
|
|
case 'blob.manifest': {
|
|
if (typeof record.mimeType !== 'string') {
|
|
throw new Error('fact log v2: blob.manifest mimeType must be a string')
|
|
}
|
|
if (record.refOp !== 'add' && record.refOp !== 'release') {
|
|
throw new Error(`fact log v2: blob.manifest refOp must be 'add' or 'release'`)
|
|
}
|
|
return [
|
|
T.BLOB_MANIFEST,
|
|
V,
|
|
C,
|
|
K,
|
|
hashToBytes(record.hash),
|
|
toWireU64(record.size, 'blob size'),
|
|
record.mimeType,
|
|
record.refOp === 'add' ? 0 : 1
|
|
]
|
|
}
|
|
case 'projection.note':
|
|
if (!isPlainMap(record.note)) {
|
|
throw new Error('fact log v2: projection.note requires a map')
|
|
}
|
|
return [T.PROJECTION_NOTE, V, C, K, record.note]
|
|
case 'bootstrap.baseline': {
|
|
if (record.kind !== 'noun' && record.kind !== 'verb') {
|
|
throw new Error(`fact log v2: bootstrap.baseline kind must be 'noun' or 'verb'`)
|
|
}
|
|
return [
|
|
T.BOOTSTRAP_BASELINE,
|
|
V,
|
|
C,
|
|
K,
|
|
uuidToBytes(record.id),
|
|
record.kind === 'noun' ? 0 : 1,
|
|
record.metadata ?? null,
|
|
encodeVectorLeg(record.vectorLeg, options, `bootstrap.baseline ${record.id}`)
|
|
]
|
|
}
|
|
case 'log.genesis': {
|
|
if (record.idSpaceWidth !== 32 && record.idSpaceWidth !== 64) {
|
|
throw new Error(
|
|
`fact log v2: log.genesis idSpaceWidth must be 32 or 64; got ${record.idSpaceWidth}`
|
|
)
|
|
}
|
|
return [
|
|
T.LOG_GENESIS,
|
|
V,
|
|
C,
|
|
K,
|
|
record.idSpaceWidth,
|
|
uuidToBytes(record.brainId),
|
|
toWireU64(record.createdAt, 'createdAt')
|
|
]
|
|
}
|
|
default: {
|
|
// Pads are the sealer's business ({@link sealGroup}); anything else
|
|
// here is an unencodable record — refuse instead of writing bytes a
|
|
// reader would have to guess about.
|
|
const unknown = record as { type?: unknown }
|
|
throw new Error(`fact log v2: cannot encode record type ${String(unknown.type)}`)
|
|
}
|
|
}
|
|
}
|
|
|
|
/** Exact wire arity per record type (envelope of 4 + type-specific fields). */
|
|
const RECORD_ARITY: Record<number, number> = {
|
|
[LOG_RECORD_TYPES.NOUN_AFTER_IMAGE]: 8,
|
|
[LOG_RECORD_TYPES.NOUN_TOMBSTONE]: 5,
|
|
[LOG_RECORD_TYPES.VERB_AFTER_IMAGE]: 13,
|
|
[LOG_RECORD_TYPES.VERB_TOMBSTONE]: 5,
|
|
[LOG_RECORD_TYPES.BATCH_META]: 5,
|
|
[LOG_RECORD_TYPES.EMBED_PENDING]: 6,
|
|
[LOG_RECORD_TYPES.EMBED_LANDED]: 6,
|
|
[LOG_RECORD_TYPES.BLOB_MANIFEST]: 8,
|
|
[LOG_RECORD_TYPES.PROJECTION_NOTE]: 5,
|
|
[LOG_RECORD_TYPES.BOOTSTRAP_BASELINE]: 8,
|
|
[LOG_RECORD_TYPES.LOG_GENESIS]: 7
|
|
}
|
|
|
|
/**
|
|
* Decode one wire record. Returns `null` for pads (skipped by definition).
|
|
* Unknown type / newer version throw {@link UnknownLogRecordError} — never
|
|
* skip-and-continue. The reserved crypto envelope is verified BEFORE the
|
|
* arity check (an encrypted record's field layout is a newer reader's
|
|
* business, not a malformed-record error): any nonzero cipherFlag or non-nil
|
|
* keyId refuses with the same typed error class.
|
|
*/
|
|
function decodeRecord(raw: unknown): LogRecord | null {
|
|
if (!Array.isArray(raw) || raw.length < 2) {
|
|
throw new Error('fact log v2: malformed record envelope (need [type, version, cipherFlag, keyId, ...])')
|
|
}
|
|
const recordType = wireToU8(raw[0], 'recordType')
|
|
const recordVersion = wireToU8(raw[1], 'recordVersion')
|
|
|
|
if (recordType === LOG_RECORD_TYPES.PAD) {
|
|
// Length-only filler: skipped wholesale, filler fields never inspected
|
|
// (pads therefore carry no crypto envelope — by definition, not omission).
|
|
return null
|
|
}
|
|
const arity = RECORD_ARITY[recordType]
|
|
if (arity === undefined) {
|
|
throw new UnknownLogRecordError(
|
|
recordType,
|
|
recordVersion,
|
|
`fact log v2: unknown record type ${recordType} (record version ${recordVersion}) — ` +
|
|
`a newer reader is required to decode this log`
|
|
)
|
|
}
|
|
if (recordVersion > LOG_RECORD_VERSION) {
|
|
throw new UnknownLogRecordError(
|
|
recordType,
|
|
recordVersion,
|
|
`fact log v2: record type ${recordType} carries record version ${recordVersion}; ` +
|
|
`this reader knows version ${LOG_RECORD_VERSION} — a newer reader is required to decode this log`
|
|
)
|
|
}
|
|
if (recordVersion !== LOG_RECORD_VERSION) {
|
|
throw new Error(`fact log v2: record type ${recordType} has invalid record version ${recordVersion}`)
|
|
}
|
|
if (raw.length < 4) {
|
|
throw new Error('fact log v2: malformed record envelope (need [type, version, cipherFlag, keyId, ...])')
|
|
}
|
|
const cipherFlag = wireToU8(raw[2], 'cipherFlag')
|
|
const keyId = raw[3]
|
|
if (cipherFlag !== LOG_RECORD_CIPHER_PLAINTEXT || (keyId !== null && keyId !== undefined)) {
|
|
throw new UnknownLogRecordError(
|
|
recordType,
|
|
recordVersion,
|
|
`fact log v2: record type ${recordType} carries cipherFlag ${cipherFlag}` +
|
|
`${keyId !== null && keyId !== undefined ? ' and a keyId' : ''} — ` +
|
|
`encrypted records need a newer reader`
|
|
)
|
|
}
|
|
if (raw.length !== arity) {
|
|
throw new Error(
|
|
`fact log v2: record type ${recordType} expects ${arity} wire fields; got ${raw.length}`
|
|
)
|
|
}
|
|
|
|
switch (recordType) {
|
|
case LOG_RECORD_TYPES.NOUN_AFTER_IMAGE:
|
|
return {
|
|
type: 'noun.afterImage',
|
|
id: bytesToUuid(raw[4], 'noun.afterImage id'),
|
|
entityInt: wireToBigint(raw[5], 'entityInt'),
|
|
metadata: raw[6] ?? null,
|
|
vectorLeg: decodeVectorLeg(raw[7], 'noun.afterImage')
|
|
}
|
|
case LOG_RECORD_TYPES.NOUN_TOMBSTONE:
|
|
return { type: 'noun.tombstone', id: bytesToUuid(raw[4], 'noun.tombstone id') }
|
|
case LOG_RECORD_TYPES.VERB_AFTER_IMAGE: {
|
|
if (typeof raw[8] !== 'string') {
|
|
throw new Error('fact log v2: verb.afterImage verb name is not a string')
|
|
}
|
|
return {
|
|
type: 'verb.afterImage',
|
|
id: bytesToUuid(raw[4], 'verb.afterImage id'),
|
|
verbInt: wireToBigint(raw[5], 'verbInt'),
|
|
metadata: raw[6] ?? null,
|
|
vectorLeg: decodeVectorLeg(raw[7], 'verb.afterImage'),
|
|
verb: raw[8],
|
|
sourceId: bytesToUuid(raw[9], 'verb.afterImage sourceId'),
|
|
sourceInt: wireToBigint(raw[10], 'sourceInt'),
|
|
targetId: bytesToUuid(raw[11], 'verb.afterImage targetId'),
|
|
targetInt: wireToBigint(raw[12], 'targetInt')
|
|
}
|
|
}
|
|
case LOG_RECORD_TYPES.VERB_TOMBSTONE:
|
|
return { type: 'verb.tombstone', id: bytesToUuid(raw[4], 'verb.tombstone id') }
|
|
case LOG_RECORD_TYPES.BATCH_META: {
|
|
if (!isPlainMap(raw[4])) throw new Error('fact log v2: batch.meta payload is not a map')
|
|
return { type: 'batch.meta', meta: raw[4] }
|
|
}
|
|
case LOG_RECORD_TYPES.EMBED_PENDING:
|
|
return {
|
|
type: 'embed.pending',
|
|
id: bytesToUuid(raw[4], 'embed.pending id'),
|
|
enqueuedAt: wireToNumber(raw[5], 'enqueuedAt')
|
|
}
|
|
case LOG_RECORD_TYPES.EMBED_LANDED: {
|
|
const leg = decodeVectorLeg(raw[5], 'embed.landed')
|
|
if (!Array.isArray(leg)) {
|
|
throw new Error(
|
|
'fact log v2: embed.landed must carry an INLINE float vector — refs and nil are not allowed here'
|
|
)
|
|
}
|
|
return { type: 'embed.landed', id: bytesToUuid(raw[4], 'embed.landed id'), vector: leg }
|
|
}
|
|
case LOG_RECORD_TYPES.BLOB_MANIFEST: {
|
|
if (typeof raw[6] !== 'string') {
|
|
throw new Error('fact log v2: blob.manifest mimeType is not a string')
|
|
}
|
|
const refOp = wireToU8(raw[7], 'refOp')
|
|
if (refOp !== 0 && refOp !== 1) {
|
|
throw new Error(`fact log v2: blob.manifest refOp must be 0 (add) or 1 (release); got ${refOp}`)
|
|
}
|
|
return {
|
|
type: 'blob.manifest',
|
|
hash: bytesToHash(raw[4]),
|
|
size: wireToNumber(raw[5], 'blob size'),
|
|
mimeType: raw[6],
|
|
refOp: refOp === 0 ? 'add' : 'release'
|
|
}
|
|
}
|
|
case LOG_RECORD_TYPES.PROJECTION_NOTE: {
|
|
if (!isPlainMap(raw[4])) throw new Error('fact log v2: projection.note payload is not a map')
|
|
return { type: 'projection.note', note: raw[4] }
|
|
}
|
|
case LOG_RECORD_TYPES.BOOTSTRAP_BASELINE: {
|
|
const kind = wireToU8(raw[5], 'bootstrap.baseline kind')
|
|
if (kind !== 0 && kind !== 1) {
|
|
throw new Error(`fact log v2: bootstrap.baseline kind must be 0 (noun) or 1 (verb); got ${kind}`)
|
|
}
|
|
return {
|
|
type: 'bootstrap.baseline',
|
|
id: bytesToUuid(raw[4], 'bootstrap.baseline id'),
|
|
kind: kind === 0 ? 'noun' : 'verb',
|
|
metadata: raw[6] ?? null,
|
|
vectorLeg: decodeVectorLeg(raw[7], 'bootstrap.baseline')
|
|
}
|
|
}
|
|
case LOG_RECORD_TYPES.LOG_GENESIS: {
|
|
const width = wireToU8(raw[4], 'idSpaceWidth')
|
|
if (width !== 32 && width !== 64) {
|
|
throw new Error(`fact log v2: log.genesis idSpaceWidth must be 32 or 64; got ${width}`)
|
|
}
|
|
return {
|
|
type: 'log.genesis',
|
|
idSpaceWidth: width,
|
|
brainId: bytesToUuid(raw[5], 'log.genesis brainId'),
|
|
createdAt: wireToNumber(raw[6], 'createdAt')
|
|
}
|
|
}
|
|
default:
|
|
// Unreachable: every arity-table type is handled above.
|
|
throw new Error(`fact log v2: unhandled record type ${recordType}`)
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Fact encode/decode
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Encode one committed generation as a complete v2 FRAME (length + crc32c +
|
|
* msgpack payload) ready for appending or sealing.
|
|
*
|
|
* Writer-enforced invariants (refusals, never silent fixes): at least one
|
|
* record; no pad records (pads belong to {@link sealGroup}); at most one
|
|
* batch.meta; log.genesis only as the first record; vector refs only with a
|
|
* passing single-hop validator; embed.landed vectors inline only.
|
|
*
|
|
* @param fact - The fact to encode (generation ≥ 1; generation 0 marks filler).
|
|
* @param options - Single-hop validation for vector refs.
|
|
* @returns The complete frame bytes.
|
|
*/
|
|
export function encodeFactV2(fact: CommitFactV2, options?: EncodeFactV2Options): Uint8Array {
|
|
if (!Number.isSafeInteger(fact.generation) || fact.generation < 1) {
|
|
throw new Error(`fact log v2: generation must be a positive integer; got ${fact.generation}`)
|
|
}
|
|
if (!Number.isSafeInteger(fact.timestamp) || fact.timestamp < 0) {
|
|
throw new Error(`fact log v2: timestamp must be a non-negative integer; got ${fact.timestamp}`)
|
|
}
|
|
// records MAY be empty: a committed generation whose ops all collapsed
|
|
// (e.g. a batch whose relates deduped to no-ops) is still a real
|
|
// generation — v1 encoded empty ops the same way; refusing here would
|
|
// fork the two formats' commit semantics.
|
|
if (!Array.isArray(fact.records)) {
|
|
throw new Error('fact log v2: records must be an array')
|
|
}
|
|
if (fact.meta !== undefined && !isPlainMap(fact.meta)) {
|
|
throw new Error('fact log v2: fact meta must be a map when present')
|
|
}
|
|
if (
|
|
fact.blobHashes !== undefined &&
|
|
(!Array.isArray(fact.blobHashes) || fact.blobHashes.some((h) => typeof h !== 'string'))
|
|
) {
|
|
throw new Error('fact log v2: blobHashes must be an array of strings when present')
|
|
}
|
|
|
|
let batchMetaCount = 0
|
|
const wireRecords = fact.records.map((record, index) => {
|
|
if (record.type === 'batch.meta' && ++batchMetaCount > 1) {
|
|
throw new Error('fact log v2: at most one batch.meta record per fact')
|
|
}
|
|
if (record.type === 'log.genesis' && index !== 0) {
|
|
throw new Error('fact log v2: log.genesis must be the first record of its fact')
|
|
}
|
|
return encodeRecord(record, options)
|
|
})
|
|
|
|
const payload = enc([
|
|
toWireU64(fact.generation, 'generation'),
|
|
toWireU64(fact.timestamp, 'timestamp'),
|
|
wireRecords,
|
|
fact.meta ?? null,
|
|
fact.blobHashes && fact.blobHashes.length > 0 ? fact.blobHashes : null
|
|
])
|
|
return buildFrame(payload)
|
|
}
|
|
|
|
/**
|
|
* Decode one fact PAYLOAD (the msgpack bytes inside a frame — see
|
|
* {@link framePayload}). The segment's formatVersion, read from its header,
|
|
* selects the schema: version 1 decodes the v1 ops shape into a
|
|
* {@link CommitFact}; version 2 decodes the record envelope into a
|
|
* {@link CommitFactV2}. Any other version is refused.
|
|
*/
|
|
export function decodeFact(payload: Uint8Array, segmentFormatVersion: 1): CommitFact
|
|
export function decodeFact(
|
|
payload: Uint8Array,
|
|
segmentFormatVersion: 2,
|
|
options?: DecodeFactV2Options
|
|
): CommitFactV2
|
|
export function decodeFact(
|
|
payload: Uint8Array,
|
|
segmentFormatVersion: number,
|
|
options?: DecodeFactV2Options
|
|
): CommitFact | CommitFactV2
|
|
export function decodeFact(
|
|
payload: Uint8Array,
|
|
segmentFormatVersion: number,
|
|
options?: DecodeFactV2Options
|
|
): CommitFact | CommitFactV2 {
|
|
if (segmentFormatVersion === FACT_LOG_FORMAT_V1) return decodeFactV1(payload)
|
|
if (segmentFormatVersion === FACT_LOG_FORMAT_V2) return decodeFactV2(payload, options)
|
|
throw new Error(
|
|
`fact log: no decoder for segment formatVersion ${segmentFormatVersion} — this build reads 1 and 2`
|
|
)
|
|
}
|
|
|
|
/**
|
|
* The v1 decode path — byte-identical in behavior to the v1 log's own
|
|
* decoder (positional ops, bin16 ids, body-less tombstones). Kept here so v1
|
|
* segments stay readable through the same entry point forever.
|
|
*/
|
|
function decodeFactV1(payload: Uint8Array): CommitFact {
|
|
const raw = msgpackDecode(payload) as unknown[]
|
|
const [generation, timestamp, ops, meta, blobHashes] = raw as [
|
|
number,
|
|
number,
|
|
Array<[number, Uint8Array, [unknown, unknown] | null]>,
|
|
Record<string, unknown> | null,
|
|
string[] | null
|
|
]
|
|
return {
|
|
generation: Number(generation),
|
|
timestamp: Number(timestamp),
|
|
ops: ops.map(([kind, idBytes, record]) => ({
|
|
kind: kind === 0 ? ('noun' as const) : ('verb' as const),
|
|
id: bytesToUuid(idBytes, 'op id'),
|
|
record: record === null ? null : { metadata: record[0] ?? null, vector: record[1] ?? null }
|
|
})),
|
|
...(meta ? { meta } : {}),
|
|
...(blobHashes && blobHashes.length > 0 ? { blobHashes } : {})
|
|
}
|
|
}
|
|
|
|
/** The v2 decode path: record envelope, decoder-law enforcement, pad skip. */
|
|
function decodeFactV2(payload: Uint8Array, options?: DecodeFactV2Options): CommitFactV2 {
|
|
const raw = dec(payload)
|
|
if (!Array.isArray(raw) || raw.length !== 5) {
|
|
throw new Error('fact log v2: fact payload must be a positional array of 5')
|
|
}
|
|
const [genWire, tsWire, recordsWire, metaWire, blobsWire] = raw
|
|
if (!Array.isArray(recordsWire)) {
|
|
throw new Error('fact log v2: fact records position is not an array')
|
|
}
|
|
|
|
const records: LogRecord[] = []
|
|
let batchMetaCount = 0
|
|
recordsWire.forEach((rawRecord, index) => {
|
|
const record = decodeRecord(rawRecord)
|
|
if (record === null) return // pad: length-only filler, skipped by definition
|
|
if (record.type === 'log.genesis') {
|
|
if (index !== 0) {
|
|
throw new Error('fact log v2: log.genesis must be the first record of its fact')
|
|
}
|
|
const expected = options?.expectedIdSpaceWidth
|
|
if (expected !== undefined && record.idSpaceWidth !== expected) {
|
|
throw new GenesisWidthMismatchError(expected, record.idSpaceWidth)
|
|
}
|
|
}
|
|
if (record.type === 'batch.meta' && ++batchMetaCount > 1) {
|
|
throw new Error('fact log v2: at most one batch.meta record per fact')
|
|
}
|
|
records.push(record)
|
|
})
|
|
|
|
let meta: Record<string, unknown> | undefined
|
|
if (metaWire !== null && metaWire !== undefined) {
|
|
if (!isPlainMap(metaWire)) throw new Error('fact log v2: fact meta position is not a map')
|
|
meta = metaWire
|
|
}
|
|
let blobHashes: string[] | undefined
|
|
if (blobsWire !== null && blobsWire !== undefined) {
|
|
if (!Array.isArray(blobsWire) || blobsWire.some((h) => typeof h !== 'string')) {
|
|
throw new Error('fact log v2: fact blobHashes position is not a string array')
|
|
}
|
|
blobHashes = blobsWire
|
|
}
|
|
|
|
return {
|
|
generation: wireToNumber(genWire, 'generation'),
|
|
timestamp: wireToNumber(tsWire, 'timestamp'),
|
|
records,
|
|
...(meta ? { meta } : {}),
|
|
...(blobHashes && blobHashes.length > 0 ? { blobHashes } : {})
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Sector seals
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Smallest constructible pad frame in bytes (frame prefix + the bare pad
|
|
* record fact), memoized. Exported for streaming writers that pad an
|
|
* append-only tail to a seal boundary: a gap smaller than this cannot hold
|
|
* any frame, so the writer pads through one extra sector (the same rule
|
|
* {@link sealGroup} applies).
|
|
*/
|
|
let minPadFrameBytesMemo: number | null = null
|
|
export function minPadFrameBytes(): number {
|
|
if (minPadFrameBytesMemo === null) {
|
|
minPadFrameBytesMemo =
|
|
FRAME_PREFIX_BYTES +
|
|
enc([0n, 0n, [[LOG_RECORD_TYPES.PAD, LOG_RECORD_VERSION]], null, null]).length
|
|
}
|
|
return minPadFrameBytesMemo
|
|
}
|
|
|
|
/**
|
|
* Build a pad frame of EXACTLY `totalBytes`: a filler fact
|
|
* `[0, 0, [[0, 1, filler?]], nil, nil]` sized via a binary filler field.
|
|
* Readers skip pad records by definition, so filler fields are never
|
|
* inspected — only their length matters.
|
|
*/
|
|
function buildPadFrame(totalBytes: number): Uint8Array {
|
|
const targetPayload = totalBytes - FRAME_PREFIX_BYTES
|
|
const attempt = (record: unknown[]): Uint8Array => enc([0n, 0n, [record], null, null])
|
|
|
|
let payload = attempt([LOG_RECORD_TYPES.PAD, LOG_RECORD_VERSION])
|
|
if (payload.length !== targetPayload) {
|
|
// One byte short: a fixint filler adds exactly one byte.
|
|
payload = attempt([LOG_RECORD_TYPES.PAD, LOG_RECORD_VERSION, 0])
|
|
}
|
|
if (payload.length !== targetPayload) {
|
|
// Binary filler: msgpack bin grows byte-for-byte within a size class;
|
|
// iterate to absorb the class-header steps (bin8 → bin16 → bin32).
|
|
let fillerLength = Math.max(0, targetPayload - payload.length - 1)
|
|
let converged = false
|
|
for (let i = 0; i < 8; i++) {
|
|
const candidate = attempt([
|
|
LOG_RECORD_TYPES.PAD,
|
|
LOG_RECORD_VERSION,
|
|
new Uint8Array(fillerLength)
|
|
])
|
|
const diff = targetPayload - candidate.length
|
|
if (diff === 0) {
|
|
payload = candidate
|
|
converged = true
|
|
break
|
|
}
|
|
fillerLength += diff
|
|
if (fillerLength < 0) break
|
|
}
|
|
if (!converged) {
|
|
throw new Error(`fact log v2: a pad frame of ${totalBytes} bytes is not constructible`)
|
|
}
|
|
}
|
|
return buildFrame(payload)
|
|
}
|
|
|
|
/**
|
|
* Build a pad frame of EXACTLY `totalBytes` — the streaming-append counterpart
|
|
* of {@link sealGroup} for writers that append pads directly to a live tail
|
|
* instead of sealing an in-memory group. Refuses sizes smaller than the
|
|
* smallest constructible pad frame ({@link minPadFrameBytes}); readers skip
|
|
* the result by definition (a type-0 record is length-only filler).
|
|
*
|
|
* @param totalBytes - The exact frame size to construct (prefix included).
|
|
* @returns The complete pad frame bytes.
|
|
*/
|
|
export function encodePadFrame(totalBytes: number): Uint8Array {
|
|
if (!Number.isInteger(totalBytes) || totalBytes < minPadFrameBytes()) {
|
|
throw new Error(
|
|
`fact log v2: a pad frame must be at least ${minPadFrameBytes()} bytes; got ${totalBytes}`
|
|
)
|
|
}
|
|
return buildPadFrame(totalBytes)
|
|
}
|
|
|
|
/**
|
|
* Seal a group of frames to a sector boundary: concatenate the frames and pad
|
|
* to the next `sealSize` multiple with ONE pad frame. An already-aligned
|
|
* group gets no pad. When the gap is smaller than the smallest constructible
|
|
* pad frame, the group is padded through to the boundary AFTER next (one
|
|
* extra sealSize) — input frames are never rewritten.
|
|
*
|
|
* @param frames - Complete, well-formed frames (verified; garbage is refused).
|
|
* @param sealSize - The sector-seal size (device probing is the caller's
|
|
* business; default {@link DEFAULT_SEAL_SIZE}).
|
|
* @returns The sector-aligned group (`length % sealSize === 0`).
|
|
*/
|
|
export function sealGroup(frames: Uint8Array[], sealSize: number = DEFAULT_SEAL_SIZE): Uint8Array {
|
|
assertValidSealSize(sealSize)
|
|
if (!Array.isArray(frames) || frames.length === 0) {
|
|
throw new Error('fact log v2: sealGroup needs at least one frame')
|
|
}
|
|
frames.forEach((frame, i) => {
|
|
try {
|
|
framePayload(frame)
|
|
} catch (error) {
|
|
throw new Error(
|
|
`fact log v2: sealGroup frame ${i} is not a well-formed frame: ${(error as Error).message}`
|
|
)
|
|
}
|
|
})
|
|
|
|
const total = frames.reduce((n, f) => n + f.length, 0)
|
|
const remainder = total % sealSize
|
|
let padBytes = remainder === 0 ? 0 : sealSize - remainder
|
|
if (padBytes !== 0 && padBytes < minPadFrameBytes()) {
|
|
padBytes += sealSize // gap too small for any frame — pad through one more sector
|
|
}
|
|
|
|
const sealed = new Uint8Array(total + padBytes)
|
|
let offset = 0
|
|
for (const frame of frames) {
|
|
sealed.set(frame, offset)
|
|
offset += frame.length
|
|
}
|
|
if (padBytes > 0) {
|
|
sealed.set(buildPadFrame(padBytes), offset)
|
|
}
|
|
return sealed
|
|
}
|
|
|
|
/**
|
|
* Decode a sequence of v2 frames (a sealed group, or a segment body after its
|
|
* 32-byte header) with the torn-tail discipline: a frame whose length overruns
|
|
* the buffer or whose CRC fails TERMINATES the walk — everything before it is
|
|
* intact and returned; nothing after it is guessed at. Pad frames are dropped
|
|
* (invisible). CRC-valid frames with unknown record types still throw
|
|
* {@link UnknownLogRecordError} — physical damage truncates, format novelty
|
|
* refuses.
|
|
*/
|
|
export function decodeGroupV2(bytes: Uint8Array, options?: DecodeFactV2Options): DecodedFrameGroup {
|
|
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength)
|
|
const facts: CommitFactV2[] = []
|
|
let offset = 0
|
|
while (offset + FRAME_PREFIX_BYTES <= bytes.length) {
|
|
const length = view.getUint32(offset, true)
|
|
const expectedCrc = view.getUint32(offset + 4, true)
|
|
const start = offset + FRAME_PREFIX_BYTES
|
|
const end = start + length
|
|
if (end > bytes.length) break // torn tail: frame length overruns the buffer
|
|
const payload = bytes.subarray(start, end)
|
|
if (crc32c(payload) !== expectedCrc) break // torn tail: payload CRC mismatch
|
|
const fact = decodeFactV2(payload, options)
|
|
if (fact.records.length > 0) facts.push(fact) // zero-record fact = pad filler
|
|
offset = end
|
|
}
|
|
return { facts, validBytes: offset }
|
|
}
|