feat: add stddev/variance aggregation ops with Welford's online algorithm
- New aggregation ops: stddev and variance (incremental, O(1) per update) - Welford's online algorithm for numerically stable running variance - MetricState extended with m2 field for variance tracking - AggregationProvider interface: defineAggregate, removeAggregate, restoreState, serializeState - Export AggregateGroupState and MetricState types - Plugin docs: aggregation provider, analytics providers (HyperLogLog, t-digest, etc.) - Aggregation architecture and usage guide docs Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
parent
4602960e86
commit
b7c6752388
9 changed files with 1435 additions and 792 deletions
|
|
@ -192,6 +192,23 @@ context.registerProvider('graphIndex', (storage) => {
|
|||
})
|
||||
```
|
||||
|
||||
#### `aggregation`
|
||||
**Type:** `(storage: StorageAdapter) => AggregationProvider-compatible`
|
||||
|
||||
Factory function that creates an aggregation engine for write-time incremental SUM/COUNT/AVG/MIN/MAX with GROUP BY and time windows. The returned object must implement the `AggregationProvider` interface.
|
||||
|
||||
```typescript
|
||||
context.registerProvider('aggregation', (storage) => {
|
||||
return new MyNativeAggregationEngine(storage)
|
||||
})
|
||||
```
|
||||
|
||||
When provided by a native plugin like `@soulcraft/cortex`, this enables:
|
||||
- Compiled source filters (vs per-entity JS object traversal)
|
||||
- Precise MIN/MAX via sorted data structures (vs lazy recompute)
|
||||
- Parallel aggregate rebuild across CPU cores
|
||||
- SIMD-accelerated timestamp bucketing
|
||||
|
||||
### Utility Providers
|
||||
|
||||
#### `cache`
|
||||
|
|
@ -220,6 +237,29 @@ Replacement for the roaring bitmap implementation. Used internally by the metada
|
|||
|
||||
Native msgpack encode/decode for SSTable serialization.
|
||||
|
||||
### Analytics Providers (Native-Only)
|
||||
|
||||
These provider keys have **no JavaScript fallback** — they represent capabilities that require native code (SIMD, mmap, sub-microsecond latency). They are available when a native plugin like `@soulcraft/cortex` is installed.
|
||||
|
||||
Use `brain.getProvider('analytics:hyperloglog')` to check availability. Returns `undefined` if no plugin provides it.
|
||||
|
||||
#### `analytics:hyperloglog`
|
||||
Approximate distinct counts. Count unique values (e.g., unique merchants) across millions of records using ~16KB of memory with ~1% error. Each update is O(1).
|
||||
|
||||
#### `analytics:tdigest`
|
||||
Streaming percentiles. Compute P50/P90/P95/P99 from streaming data without storing all values. Uses ~4KB per digest with ~1% accuracy at the tails.
|
||||
|
||||
#### `analytics:countmin`
|
||||
Frequency estimation. Find the most common values (e.g., top-K merchants) using ~40KB with 0.1% error. O(1) per update.
|
||||
|
||||
#### `analytics:anomaly`
|
||||
Real-time anomaly detection. Flag statistically unusual values at write-time using exponentially weighted moving averages. 64 bytes per group, sub-microsecond decisions.
|
||||
|
||||
#### `aggregation:mmap`
|
||||
Persistent aggregate storage via memory-mapped files. Aggregate state survives process crashes without explicit flush. Zero serialization overhead.
|
||||
|
||||
---
|
||||
|
||||
## Storage Adapter Plugins
|
||||
|
||||
Plugins can register custom storage backends that users reference by name.
|
||||
|
|
|
|||
|
|
@ -426,7 +426,8 @@ brain.defineAggregate({
|
|||
count: { op: 'count' },
|
||||
average: { op: 'avg', field: 'amount' },
|
||||
highest: { op: 'max', field: 'amount' },
|
||||
lowest: { op: 'min', field: 'amount' }
|
||||
lowest: { op: 'min', field: 'amount' },
|
||||
spread: { op: 'stddev', field: 'amount' } // Welford's online algorithm
|
||||
},
|
||||
materialize: true // Optional: write results as NounType.Measurement entities
|
||||
})
|
||||
|
|
@ -441,7 +442,7 @@ brain.defineAggregate({
|
|||
| `source.where` | `Record<string, unknown>` | Metadata filter (same syntax as `find({ where })`) |
|
||||
| `source.service` | `string` | Multi-tenancy filter |
|
||||
| `groupBy` | `GroupByDimension[]` | Dimensions to group by — plain field names or `{ field, window }` for time bucketing |
|
||||
| `metrics` | `Record<string, AggregateMetricDef>` | Named metrics with `op` (`sum`, `count`, `avg`, `min`, `max`) and optional `field` |
|
||||
| `metrics` | `Record<string, AggregateMetricDef>` | Named metrics with `op` (`sum`, `count`, `avg`, `min`, `max`, `stddev`, `variance`) and optional `field` |
|
||||
| `materialize` | `boolean \| object` | Write results as `NounType.Measurement` entities (auto-visible in OData/Sheets/SSE) |
|
||||
|
||||
**Time window granularities:** `'hour'`, `'day'`, `'week'`, `'month'`, `'quarter'`, `'year'`, or `{ seconds: number }` for custom intervals.
|
||||
|
|
|
|||
242
docs/architecture/aggregation.md
Normal file
242
docs/architecture/aggregation.md
Normal file
|
|
@ -0,0 +1,242 @@
|
|||
# Aggregation Architecture
|
||||
|
||||
> Write-time incremental aggregation with O(1) reads
|
||||
|
||||
## Design Principles
|
||||
|
||||
1. **Write-time computation** — aggregates update on every `add()`, `update()`, and `delete()`, not as batch jobs
|
||||
2. **Incremental state** — running totals maintained per group, never rescanning the dataset
|
||||
3. **Provider interface** — TypeScript engine is the default; plugins can replace it with native implementations
|
||||
4. **Zero-allocation reads** — query results are computed from pre-aggregated state
|
||||
|
||||
## Component Overview
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ Brainy │
|
||||
│ │
|
||||
│ add() / update() / delete() │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ ┌──────────────────────┐ ┌────────────────────────┐ │
|
||||
│ │ AggregationIndex │ │ AggregateMaterializer │ │
|
||||
│ │ │───▶│ (debounced writes) │ │
|
||||
│ │ ├─ definitions Map │ └────────────────────────┘ │
|
||||
│ │ ├─ states Map │ │
|
||||
│ │ └─ staleMinMax Set │ ┌────────────────────────┐ │
|
||||
│ │ │ │ timeWindows.ts │ │
|
||||
│ │ Source filter ──────│───▶│ bucketTimestamp() │ │
|
||||
│ │ Group key ──────────│───▶│ parseBucketRange() │ │
|
||||
│ └──────────┬───────────┘ └────────────────────────┘ │
|
||||
│ │ │
|
||||
│ │ provider interface │
|
||||
│ ▼ │
|
||||
│ ┌──────────────────────┐ │
|
||||
│ │ AggregationProvider │ (optional, registered by │
|
||||
│ │ ├─ incrementalUpdate│ plugin like @soulcraft/cortex)│
|
||||
│ │ ├─ rebuildAggregate │ │
|
||||
│ │ ├─ queryAggregate │ │
|
||||
│ │ └─ serialize/restore│ │
|
||||
│ └──────────────────────┘ │
|
||||
└──────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## State Management
|
||||
|
||||
### Definitions
|
||||
|
||||
Registered via `brain.defineAggregate(def)`. Stored in a `Map<string, AggregateDefinition>` keyed by aggregate name. Persisted to storage under `__aggregation_definitions__` on flush.
|
||||
|
||||
### Group State
|
||||
|
||||
Each aggregate maintains a `Map<string, AggregateGroupState>` where keys are serialized group key values (e.g., `category=food|date=2024-01`). Each group holds per-metric `MetricState`:
|
||||
|
||||
```typescript
|
||||
interface MetricState {
|
||||
sum: number // Running total
|
||||
count: number // Entity count
|
||||
min: number // Minimum (Infinity if empty)
|
||||
max: number // Maximum (-Infinity if empty)
|
||||
m2?: number // Welford's M2 for stddev/variance
|
||||
}
|
||||
```
|
||||
|
||||
### Change Detection
|
||||
|
||||
On restart, definition hashes (FNV-1a 32-bit) are compared with the persisted hash. If a definition changed (different groupBy, metrics, or source), the aggregate state is reset and must be rebuilt.
|
||||
|
||||
## Write-Time Update Flow
|
||||
|
||||
When `brain.add(entity)` is called:
|
||||
|
||||
```
|
||||
1. For each registered aggregate:
|
||||
├─ Source filter check (type, service, where)
|
||||
│ └─ Skip if entity doesn't match
|
||||
├─ Aggregate entity check
|
||||
│ └─ Skip if entity.service === 'brainy:aggregation'
|
||||
│ or entity.metadata.__aggregate is set
|
||||
├─ Group key computation
|
||||
│ └─ Extract groupBy fields from metadata
|
||||
│ Apply time bucketing for windowed dimensions
|
||||
└─ Metric update
|
||||
└─ For each metric in the definition:
|
||||
├─ count: increment count
|
||||
├─ sum/avg: add value to sum, increment count
|
||||
├─ min/max: compare and update
|
||||
└─ stddev/variance: Welford's online update
|
||||
```
|
||||
|
||||
### Update Handling
|
||||
|
||||
On `brain.update(entity)`, the engine reverses the old entity's contribution and applies the new entity's contribution. This correctly handles:
|
||||
|
||||
- **Value changes**: old amount=10, new amount=20 — sum adjusts by +10
|
||||
- **Group key changes**: entity moves from category "food" to "drink" — both groups update
|
||||
- **Source filter changes**: entity type changes from Event to Document — removed from matching aggregates
|
||||
|
||||
### Delete Handling
|
||||
|
||||
On `brain.delete(id)`, the engine reverses the entity's contribution:
|
||||
|
||||
- `count` and `sum` are decremented
|
||||
- `min`/`max` may become stale (marked in `staleMinMax` for lazy recompute)
|
||||
- Welford's M2 is updated with the inverse formula
|
||||
- Empty groups (all metric counts at zero) are removed
|
||||
|
||||
## Algorithms
|
||||
|
||||
### Welford's Online Algorithm
|
||||
|
||||
Standard deviation and variance use Welford's numerically stable online algorithm with M2 tracking. This computes incrementally without storing individual values:
|
||||
|
||||
```
|
||||
On add(x):
|
||||
count += 1
|
||||
oldMean = (sum - x) / (count - 1) // mean before this value
|
||||
sum += x
|
||||
mean = sum / count // mean after this value
|
||||
M2 += (x - oldMean) * (x - mean)
|
||||
|
||||
On remove(x):
|
||||
oldMean = sum / count
|
||||
sum -= x
|
||||
count -= 1
|
||||
newMean = sum / count
|
||||
M2 = max(0, M2 - (x - oldMean) * (x - newMean))
|
||||
|
||||
Sample variance = M2 / (count - 1)
|
||||
Sample stddev = sqrt(variance)
|
||||
```
|
||||
|
||||
M2 is clamped to zero on remove to prevent floating-point drift from producing negative values.
|
||||
|
||||
### MIN/MAX Handling
|
||||
|
||||
The TypeScript engine uses simple comparison for add operations and marks MIN/MAX as potentially stale on delete (since removing the current min/max value requires a rescan). Stale values are lazily recomputed on the next query.
|
||||
|
||||
The Cortex native engine uses a `BTreeMap<OrderedFloat<f64>, u64>` that tracks the exact frequency of every value, providing precise MIN/MAX after any sequence of operations without rescanning.
|
||||
|
||||
### Time Window Bucketing
|
||||
|
||||
Timestamps (Unix milliseconds) are bucketed using UTC-based formatting:
|
||||
|
||||
| Granularity | Bucket Key | Algorithm |
|
||||
|------------|-----------|-----------|
|
||||
| `hour` | `2024-01-15T14` | UTC year-month-day-hour |
|
||||
| `day` | `2024-01-15` | UTC year-month-day |
|
||||
| `week` | `2024-W03` | ISO 8601 week (Monday start, week 1 contains first Thursday) |
|
||||
| `month` | `2024-01` | UTC year-month |
|
||||
| `quarter` | `2024-Q1` | `ceil((month) / 3)` |
|
||||
| `year` | `2024` | UTC year |
|
||||
| `{ seconds: N }` | ISO timestamp | `floor(timestamp / interval) * interval` |
|
||||
|
||||
Bucket keys can be parsed back into `{ start, end }` timestamp ranges via `parseBucketRange()`.
|
||||
|
||||
## Provider Interface
|
||||
|
||||
The `AggregationProvider` interface defines the contract between Brainy's `AggregationIndex` and plugin-provided native implementations:
|
||||
|
||||
```typescript
|
||||
interface AggregationProvider {
|
||||
defineAggregate?(def: AggregateDefinition): void
|
||||
removeAggregate?(name: string): void
|
||||
|
||||
incrementalUpdate(
|
||||
name: string,
|
||||
def: AggregateDefinition,
|
||||
entity: Record<string, unknown>,
|
||||
op: 'add' | 'update' | 'delete',
|
||||
prev?: Record<string, unknown>
|
||||
): AggregateGroupState[]
|
||||
|
||||
computeGroupKey(
|
||||
entity: Record<string, unknown>,
|
||||
groupBy: GroupByDimension[]
|
||||
): Record<string, string | number>
|
||||
|
||||
rebuildAggregate(
|
||||
def: AggregateDefinition,
|
||||
entities: Array<Record<string, unknown>>
|
||||
): Map<string, AggregateGroupState>
|
||||
|
||||
queryAggregate(
|
||||
state: Map<string, AggregateGroupState>,
|
||||
params: AggregateQueryParams
|
||||
): AggregateResult[]
|
||||
|
||||
restoreState?(data: string): void
|
||||
serializeState?(): string
|
||||
}
|
||||
```
|
||||
|
||||
When a native provider is registered:
|
||||
|
||||
1. `AggregationIndex` delegates `incrementalUpdate()` to the provider instead of running TypeScript logic
|
||||
2. Provider returns updated `AggregateGroupState[]` which are applied back into the state maps
|
||||
3. Query execution is delegated via `queryAggregate()`
|
||||
4. State serialization is delegated via `serializeState()`/`restoreState()`
|
||||
|
||||
Brainy retains ownership of the state maps and persistence. The provider handles computation.
|
||||
|
||||
## Materialization
|
||||
|
||||
The `AggregateMaterializer` converts aggregate group states into `NounType.Measurement` entities:
|
||||
|
||||
1. When an aggregate group is updated and `materialize` is enabled, `scheduleMaterialize()` is called
|
||||
2. Materialization is debounced (default: 1000ms) to batch rapid updates during ingestion
|
||||
3. On trigger, the materializer either creates or updates a `NounType.Measurement` entity
|
||||
4. Materialized entities include `service: 'brainy:aggregation'` and `metadata.__aggregate` to prevent infinite loops
|
||||
|
||||
Materialized entities are automatically visible through:
|
||||
- OData endpoints
|
||||
- Google Sheets integration
|
||||
- Server-Sent Events (SSE)
|
||||
- Webhook notifications
|
||||
|
||||
## Persistence
|
||||
|
||||
### Storage Keys
|
||||
|
||||
| Key | Content |
|
||||
|-----|---------|
|
||||
| `__aggregation_definitions__` | Array of all definitions with FNV-1a hashes |
|
||||
| `__aggregation_state_{name}__` | Per-aggregate group states (array of `AggregateGroupState`) |
|
||||
| `__aggregation_native_state__` | Serialized native provider state (JSON string) |
|
||||
|
||||
### Lifecycle
|
||||
|
||||
1. **`init()`** — Load definitions, compare hashes, load matching state, restore native provider state
|
||||
2. **Write operations** — Mark modified aggregates as dirty
|
||||
3. **`flush()`** — Persist all dirty aggregate states and native provider state
|
||||
4. **`close()`** — Flush and release resources
|
||||
|
||||
## Source Files
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `src/aggregation/AggregationIndex.ts` | Core engine: definitions, state, write hooks, query |
|
||||
| `src/aggregation/materializer.ts` | Debounced materialization of results as entities |
|
||||
| `src/aggregation/timeWindows.ts` | Time bucketing and bucket range parsing |
|
||||
| `src/aggregation/index.ts` | Module exports |
|
||||
| `src/types/brainy.types.ts` | Type definitions for all aggregation interfaces |
|
||||
524
docs/guides/aggregation.md
Normal file
524
docs/guides/aggregation.md
Normal file
|
|
@ -0,0 +1,524 @@
|
|||
# Aggregation Guide
|
||||
|
||||
> Real-time analytics on your entity data with incremental running totals
|
||||
|
||||
## Overview
|
||||
|
||||
Brainy's aggregation engine computes running totals at write time, so reading aggregate results is always O(1) regardless of dataset size. Define an aggregate once, and every `add()`, `update()`, and `delete()` automatically updates the running metrics.
|
||||
|
||||
No batch jobs. No scheduled recalculations. Aggregates stay current with every write.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```typescript
|
||||
import { Brainy, NounType } from '@soulcraft/brainy'
|
||||
|
||||
const brain = new Brainy()
|
||||
await brain.init()
|
||||
|
||||
// 1. Define an aggregate
|
||||
brain.defineAggregate({
|
||||
name: 'sales_by_category',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['category'],
|
||||
metrics: {
|
||||
revenue: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' },
|
||||
average: { op: 'avg', field: 'amount' }
|
||||
}
|
||||
})
|
||||
|
||||
// 2. Add entities — aggregates update automatically
|
||||
await brain.add({
|
||||
data: 'Coffee purchase',
|
||||
type: NounType.Event,
|
||||
metadata: { category: 'food', amount: 5.50 }
|
||||
})
|
||||
|
||||
await brain.add({
|
||||
data: 'Laptop purchase',
|
||||
type: NounType.Event,
|
||||
metadata: { category: 'electronics', amount: 1200 }
|
||||
})
|
||||
|
||||
await brain.add({
|
||||
data: 'Lunch purchase',
|
||||
type: NounType.Event,
|
||||
metadata: { category: 'food', amount: 12.00 }
|
||||
})
|
||||
|
||||
// 3. Query results
|
||||
const results = await brain.find({ aggregate: 'sales_by_category' })
|
||||
|
||||
// Results:
|
||||
// [
|
||||
// { groupKey: { category: 'food' }, metrics: { revenue: 17.50, count: 2, average: 8.75 } },
|
||||
// { groupKey: { category: 'electronics' }, metrics: { revenue: 1200, count: 1, average: 1200 } }
|
||||
// ]
|
||||
```
|
||||
|
||||
## Aggregation Operations
|
||||
|
||||
Brainy supports 7 aggregation operations:
|
||||
|
||||
### `sum` — Running Total
|
||||
|
||||
Adds up all values of a numeric field.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
total_revenue: { op: 'sum', field: 'amount' }
|
||||
}
|
||||
```
|
||||
|
||||
### `count` — Entity Count
|
||||
|
||||
Counts the number of matching entities. No `field` required.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
order_count: { op: 'count' }
|
||||
}
|
||||
```
|
||||
|
||||
### `avg` — Running Average
|
||||
|
||||
Computes `sum / count` incrementally.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
average_price: { op: 'avg', field: 'price' }
|
||||
}
|
||||
```
|
||||
|
||||
### `min` — Minimum Value
|
||||
|
||||
Tracks the minimum value across all entities in each group.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
lowest_price: { op: 'min', field: 'price' }
|
||||
}
|
||||
```
|
||||
|
||||
### `max` — Maximum Value
|
||||
|
||||
Tracks the maximum value across all entities in each group.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
highest_price: { op: 'max', field: 'price' }
|
||||
}
|
||||
```
|
||||
|
||||
### `stddev` — Sample Standard Deviation
|
||||
|
||||
Computes the sample standard deviation using Welford's numerically stable online algorithm. Updates incrementally without storing individual values.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
price_spread: { op: 'stddev', field: 'price' }
|
||||
}
|
||||
```
|
||||
|
||||
### `variance` — Sample Variance
|
||||
|
||||
Computes the sample variance (square of standard deviation) using Welford's online algorithm.
|
||||
|
||||
```typescript
|
||||
metrics: {
|
||||
price_variance: { op: 'variance', field: 'price' }
|
||||
}
|
||||
```
|
||||
|
||||
## GROUP BY Dimensions
|
||||
|
||||
Every aggregate requires at least one `groupBy` dimension. Results are grouped by the unique combinations of dimension values.
|
||||
|
||||
### Plain Fields
|
||||
|
||||
Group by a metadata field value:
|
||||
|
||||
```typescript
|
||||
groupBy: ['category']
|
||||
// Produces groups: { category: 'food' }, { category: 'electronics' }, ...
|
||||
```
|
||||
|
||||
### Multiple Fields
|
||||
|
||||
Group by multiple fields for composite keys:
|
||||
|
||||
```typescript
|
||||
groupBy: ['category', 'region']
|
||||
// Produces groups: { category: 'food', region: 'US' }, { category: 'food', region: 'EU' }, ...
|
||||
```
|
||||
|
||||
### Time Windows
|
||||
|
||||
Group by a timestamp field bucketed into time periods:
|
||||
|
||||
```typescript
|
||||
groupBy: [{ field: 'date', window: 'month' }]
|
||||
// Produces groups: { date: '2024-01' }, { date: '2024-02' }, ...
|
||||
```
|
||||
|
||||
Available time window granularities:
|
||||
|
||||
| Window | Format | Example |
|
||||
|--------|--------|---------|
|
||||
| `hour` | `YYYY-MM-DDThh` | `2024-01-15T14` |
|
||||
| `day` | `YYYY-MM-DD` | `2024-01-15` |
|
||||
| `week` | `YYYY-Wnn` | `2024-W03` |
|
||||
| `month` | `YYYY-MM` | `2024-01` |
|
||||
| `quarter` | `YYYY-Qn` | `2024-Q1` |
|
||||
| `year` | `YYYY` | `2024` |
|
||||
| `{ seconds: N }` | ISO 8601 | Custom interval |
|
||||
|
||||
### Combined Dimensions
|
||||
|
||||
Mix plain fields and time windows:
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'monthly_sales',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['region', { field: 'date', window: 'month' }],
|
||||
metrics: {
|
||||
revenue: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' }
|
||||
}
|
||||
})
|
||||
|
||||
// Produces groups like:
|
||||
// { region: 'US', date: '2024-01' }
|
||||
// { region: 'US', date: '2024-02' }
|
||||
// { region: 'EU', date: '2024-01' }
|
||||
```
|
||||
|
||||
## Querying Aggregates
|
||||
|
||||
Aggregate results are queried through the standard `find()` method.
|
||||
|
||||
### Basic Query
|
||||
|
||||
```typescript
|
||||
const results = await brain.find({ aggregate: 'sales_by_category' })
|
||||
```
|
||||
|
||||
### Filter by Group Key
|
||||
|
||||
Use `where` to filter on group key values:
|
||||
|
||||
```typescript
|
||||
const foodOnly = await brain.find({
|
||||
aggregate: 'sales_by_category',
|
||||
where: { category: 'food' }
|
||||
})
|
||||
```
|
||||
|
||||
### Sort and Paginate
|
||||
|
||||
Sort by any metric or group key field:
|
||||
|
||||
```typescript
|
||||
const topCategories = await brain.find({
|
||||
aggregate: {
|
||||
name: 'sales_by_category',
|
||||
orderBy: 'revenue',
|
||||
order: 'desc',
|
||||
limit: 10
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
### Combined Parameters
|
||||
|
||||
`where`, `orderBy`, `limit`, and `offset` from the outer `find()` call merge automatically with the aggregate query:
|
||||
|
||||
```typescript
|
||||
const recentTopSpenders = await brain.find({
|
||||
aggregate: 'monthly_sales',
|
||||
where: { region: 'US' },
|
||||
orderBy: 'revenue',
|
||||
order: 'desc',
|
||||
limit: 12,
|
||||
offset: 0
|
||||
})
|
||||
```
|
||||
|
||||
### Result Format
|
||||
|
||||
Each result is returned as a `Result<T>` with `type: NounType.Measurement`:
|
||||
|
||||
```typescript
|
||||
{
|
||||
id: string,
|
||||
score: 1.0,
|
||||
type: NounType.Measurement,
|
||||
metadata: {
|
||||
__aggregate: 'sales_by_category',
|
||||
category: 'food', // Group key values
|
||||
revenue: 17.50, // Computed metrics
|
||||
count: 2,
|
||||
average: 8.75
|
||||
},
|
||||
entity: Entity
|
||||
}
|
||||
```
|
||||
|
||||
## Source Filtering
|
||||
|
||||
Control which entities feed into an aggregate with the `source` property.
|
||||
|
||||
### Filter by Entity Type
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'event_stats',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['category'],
|
||||
metrics: { count: { op: 'count' } }
|
||||
})
|
||||
```
|
||||
|
||||
### Filter by Multiple Types
|
||||
|
||||
```typescript
|
||||
source: { type: [NounType.Event, NounType.Document] }
|
||||
```
|
||||
|
||||
### Filter by Metadata
|
||||
|
||||
Use the same `where` syntax as `find()`:
|
||||
|
||||
```typescript
|
||||
source: {
|
||||
type: NounType.Event,
|
||||
where: { domain: 'financial', subtype: 'transaction' }
|
||||
}
|
||||
```
|
||||
|
||||
### Filter by Service
|
||||
|
||||
For multi-tenant deployments:
|
||||
|
||||
```typescript
|
||||
source: { service: 'tenant-123' }
|
||||
```
|
||||
|
||||
Entities that don't match the source filter are silently skipped during incremental updates.
|
||||
|
||||
## Incremental Updates
|
||||
|
||||
The aggregation engine hooks into every write operation:
|
||||
|
||||
### On `add()`
|
||||
|
||||
When a new entity matches an aggregate's source filter:
|
||||
1. The group key is computed from the entity's metadata
|
||||
2. Each metric in the matching group is incremented
|
||||
3. New groups are created automatically
|
||||
|
||||
### On `update()`
|
||||
|
||||
When an existing entity is updated:
|
||||
1. The old entity's contribution is reversed from its group
|
||||
2. The new entity's contribution is applied to its (potentially different) group
|
||||
3. Handles group key changes — an entity moving from category "food" to "drink" updates both groups
|
||||
|
||||
### On `delete()`
|
||||
|
||||
When an entity is deleted:
|
||||
1. The entity's contribution is reversed from its group
|
||||
2. If a group becomes empty (all metric counts reach zero), it's removed
|
||||
|
||||
### Aggregate Entity Exclusion
|
||||
|
||||
Materialized `NounType.Measurement` entities are automatically excluded from all source matching, preventing infinite feedback loops. Entities with `service: 'brainy:aggregation'` or `metadata.__aggregate` are always skipped.
|
||||
|
||||
## Materialization
|
||||
|
||||
Materialization writes aggregate results as `NounType.Measurement` entities, making them automatically available through OData, Google Sheets, SSE, and webhook integrations.
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'daily_metrics',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: [{ field: 'date', window: 'day' }],
|
||||
metrics: {
|
||||
total: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' }
|
||||
},
|
||||
materialize: true
|
||||
})
|
||||
```
|
||||
|
||||
### Debounce Configuration
|
||||
|
||||
During high-throughput ingestion, materialization is debounced to avoid excessive writes:
|
||||
|
||||
```typescript
|
||||
materialize: {
|
||||
debounceMs: 2000, // Wait 2 seconds after last update before writing
|
||||
trackSources: true // Track which entities contributed
|
||||
}
|
||||
```
|
||||
|
||||
The default debounce interval is 1000ms.
|
||||
|
||||
## Multiple Aggregates
|
||||
|
||||
Define multiple aggregates that process the same entities:
|
||||
|
||||
```typescript
|
||||
// Revenue by category
|
||||
brain.defineAggregate({
|
||||
name: 'category_revenue',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['category'],
|
||||
metrics: { total: { op: 'sum', field: 'amount' } }
|
||||
})
|
||||
|
||||
// Monthly trends
|
||||
brain.defineAggregate({
|
||||
name: 'monthly_trends',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: [{ field: 'date', window: 'month' }],
|
||||
metrics: {
|
||||
revenue: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' },
|
||||
avg_order: { op: 'avg', field: 'amount' }
|
||||
}
|
||||
})
|
||||
|
||||
// Regional breakdown with statistical analysis
|
||||
brain.defineAggregate({
|
||||
name: 'regional_analysis',
|
||||
source: { type: NounType.Event },
|
||||
groupBy: ['region'],
|
||||
metrics: {
|
||||
revenue: { op: 'sum', field: 'amount' },
|
||||
spread: { op: 'stddev', field: 'amount' },
|
||||
variance: { op: 'variance', field: 'amount' }
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
Each `add()` call updates all matching aggregates automatically.
|
||||
|
||||
## Removing Aggregates
|
||||
|
||||
Remove an aggregate and clean up its state:
|
||||
|
||||
```typescript
|
||||
brain.removeAggregate('category_revenue')
|
||||
```
|
||||
|
||||
## Persistence
|
||||
|
||||
Aggregate definitions and running state are automatically persisted:
|
||||
|
||||
- **On `flush()`/`close()`**: All dirty aggregate state is written to storage
|
||||
- **On `init()`**: Definitions and state are restored from storage
|
||||
- **Change detection**: Definition changes are detected via FNV-1a hashing — only changed aggregates reset their state on restart
|
||||
|
||||
## Native Acceleration
|
||||
|
||||
When [Cortex](https://github.com/soulcraftlabs/cortex) is installed as a plugin, the aggregation engine automatically uses Rust-accelerated computation:
|
||||
|
||||
- Incremental updates run in Rust with BTreeMap-backed precise MIN/MAX
|
||||
- Welford's online stddev/variance computed natively
|
||||
- Rebuild uses Rayon parallel iterators across CPU cores (above 1,000 entities)
|
||||
- Time window bucketing uses integer arithmetic without `Date` object allocation
|
||||
|
||||
```typescript
|
||||
const brain = new Brainy({
|
||||
plugins: ['@soulcraft/cortex']
|
||||
})
|
||||
await brain.init()
|
||||
|
||||
// Aggregation automatically uses native engine
|
||||
brain.defineAggregate({ ... })
|
||||
```
|
||||
|
||||
Verify native acceleration is active:
|
||||
|
||||
```typescript
|
||||
const diag = brain.diagnostics()
|
||||
console.log(diag.providers.aggregation)
|
||||
// { source: 'plugin' }
|
||||
```
|
||||
|
||||
## Common Patterns
|
||||
|
||||
### Financial Analytics
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'monthly_spending',
|
||||
source: {
|
||||
type: NounType.Event,
|
||||
where: { domain: 'financial', subtype: 'transaction' }
|
||||
},
|
||||
groupBy: [
|
||||
'category',
|
||||
{ field: 'date', window: 'month' }
|
||||
],
|
||||
metrics: {
|
||||
total: { op: 'sum', field: 'amount' },
|
||||
count: { op: 'count' },
|
||||
average: { op: 'avg', field: 'amount' },
|
||||
highest: { op: 'max', field: 'amount' },
|
||||
lowest: { op: 'min', field: 'amount' }
|
||||
},
|
||||
materialize: true
|
||||
})
|
||||
```
|
||||
|
||||
### Time-Series Monitoring
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'hourly_metrics',
|
||||
source: { type: NounType.Event, where: { domain: 'monitoring' } },
|
||||
groupBy: [
|
||||
'service',
|
||||
{ field: 'timestamp', window: 'hour' }
|
||||
],
|
||||
metrics: {
|
||||
request_count: { op: 'count' },
|
||||
avg_latency: { op: 'avg', field: 'latency_ms' },
|
||||
max_latency: { op: 'max', field: 'latency_ms' },
|
||||
error_count: { op: 'sum', field: 'is_error' },
|
||||
latency_spread: { op: 'stddev', field: 'latency_ms' }
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
### Content Analytics
|
||||
|
||||
```typescript
|
||||
brain.defineAggregate({
|
||||
name: 'content_stats',
|
||||
source: { type: NounType.Document },
|
||||
groupBy: ['author', { field: 'publishedAt', window: 'month' }],
|
||||
metrics: {
|
||||
articles: { op: 'count' },
|
||||
total_words: { op: 'sum', field: 'wordCount' },
|
||||
avg_words: { op: 'avg', field: 'wordCount' }
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
## Performance
|
||||
|
||||
Aggregation complexity per write is O(A x G x M) where A = matching aggregates, G = groupBy dimensions, M = metrics. For typical configurations (2-5 aggregates, 1-3 dimensions, 3-5 metrics), this is effectively O(1).
|
||||
|
||||
With Cortex native acceleration:
|
||||
|
||||
| Operation | Throughput | Latency |
|
||||
|-----------|-----------|---------|
|
||||
| Incremental update (1K entities) | 809 ops/s | 1.2 ms |
|
||||
| Rebuild (10K entities) | 475 ops/s | 2.1 ms |
|
||||
| Rebuild (100K entities, Rayon) | 66 ops/s | 15.2 ms |
|
||||
| Query (1K groups, sort + paginate) | 986 ops/s | 1.0 ms |
|
||||
Loading…
Add table
Add a link
Reference in a new issue