SMA framework · system map

SMA — the master map

The agent stops trusting its own word — and v5 puts a fleet to work: a durable queue, headless workers, and an owner who only approves. The trust spine stays the only gate.

v5.7.1 6254 tests · 284 files p95 hook 155 ms zero LLM in the hot path plain files + git · source-available

The map

Eight version layers, newest on top. Click any node for its plain-language card — commands, files, honest numbers. One panel at a time; Esc or click-away closes; the URL remembers where you were.
amber = ships disabled, an owner turns it on · red = blocking · everything else observes and warns

V5.1–V5.6

The Daily Driver — the window runs the day

7 subsystems
Six shipped releases on top of the fleet: the app is the workplace, memory is measured and governed, a live session takes the wheel, and the docs cannot lie about the numbers.

V5

Orchestration — the 24/7 Fleet

8 subsystems
The fleet runs the work; the owner approves. The trust spine stays the only gate.

V4

Grade the Grader

6 subsystems
The judges themselves become audited instruments; every run is priced against the project's own spend.

V3.6

The One-Command Door

4 subsystems · 1 batch
Entering and leaving cost one command each; the newcomer sees their own repo first.

V3.5

Adoption & Trust Telemetry

15 subsystems · 15 plans
The release that makes trust visible to strangers.

V3

Trust Spine

10 subsystems · 10 plans
Every claim gets a receipt a stranger can re-run.

V2

Predictions · Reflexes · Coordination

6 subsystems
The agent starts betting on itself, in writing.

V1

Memory Foundation

5 subsystems
Durable memory that outlives the conversation.

The numbers

Every figure below is committed data or a verified release receipt. Nothing modeled, nothing projected.

Test-suite growth · measured points

0 2000 4000 6000 8000 v3.0.0 — 532 tests v3.5.0 — 744 tests, 70 files v3.6.0 — 792 tests, 74 files v4.0.0 — 876 tests, 78 files v5.0.1 — 1145 tests, 96 files (measured on main, 21.07.2026) main — 6254 tests, 284 files. NOT a release point: this is a working-tree measurement — the run receipt of 01.09.2026 at commit 236c2c1. 532 744 792 876 1145 6254 96 files 284 files v3.0.0 v3.5.0 v3.6.0 v4.0.0 v5.0.1 main*
Tests in the suite at the six measured points — four verified release tags, v5.0.1 measured on main, 21.07.2026, and the working-tree point marked with an asterisk: it is NOT a release measurement, it is the run receipt of the suite on main, 01.09.2026 (6254 tests / 284 files). The five release points are history and are never rewritten; the sixth is today's claim, so its counts, its stamp and the height it is drawn at are all re-derived from the run receipt by the numbers gate — a divergence turns npm test red. The lines are connectors, not a trend claim.

Hook overhead · log scale, lower is better

0.1 1 10 100 1000 ms V2 hooks · 3–4 spawns V2 baseline — 1268.6 ms per tool call 1268.6 ms V3 multiplexer p95 V3 multiplexer — p95 155 ms 155 ms Statusline render Statusline render — 0.43 ms 0.43 ms
Cost per tool call: the V2 hooks spawned 3–4 processes each time; the V3 multiplexer folds them into one call. Log scale — on a linear axis the 0.43 ms dot would sit on zero.

Prediction calibration by area · tiny n, shown on purpose

bridges 0% n=1 coordination 50% n=2 enforce 100% n=1 manifest 100% n=4 overall 16/20 collecting badge hidden until n=20 · a 0% or 100% at n=1 is noise — the chart says so
Per-area hit-rates from the committed calibration passport of SMA user #1. The point is honesty at small n: sample sizes are printed louder than the percentages, and the public badge stays hidden until n≥20.

Release gate receipts · v5.6.1

4043/4043test suite at the tag · 182 files
0package-check violations
0internal ids across 780 published files
62frozen web routes · grows only by recorded revision
0doc-number violations · every figure re-derived
The suite figure is the committed run receipt at the v5.6.1 tag; the package, leak, route and doc-number gates were re-run against this tree on 24.08.2026. Each one re-derives from the repository — none of them is a self-assessment.

Version story · what each layer added

  1. V1Memory + coordinationcorpus · sessions · claims
  2. V2Predictions + reflexesthe agent bets on itself
  3. V3Trust spine10 plans · receipts everywhere
  4. V3.5Adoption & telemetry15 plans · trust made public
  5. V3.6The one-command doornpx install · off-ramp · memory preview · rules block
  6. V4Grade the graderverdict scoring · economy meters · footprint ladder · quick-ship
  7. V5Orchestrationdurable queue · headless runners · roster front · the Creator
  8. V5.1The app and the memory foundation17-screen app · memory model 1.0 · import door
  9. V5.2Measured memorybenchmark · explain · reproducible proof
  10. V5.3Governance + hardened fleetlifecycle · state machine · fleet rules
  11. V5.4The whole working day, without the terminalthe window as daily driver · a QA department that uses the product
  12. V5.5The engine: steering a live sessiontranscript thread · live steer · return resumes the session
  13. V5.6The taskboard, and numbers that do not lieunits of work · doc-numbers gate · the worker's session, attempts that tell the truth
Everything through V5.6 shipped by 2026-08: the fleet runs the work, the window drives the day, and the trust spine stays the only gate. What comes next is named in the published roadmap — announced, not claimed.

The bets

The 10× scorecard — registered predictions, not results. Each target was frozen as a machine-scoreable record before the trust spine was built ("no measured base, no target"). Two baselines were forfeit when the measurement window was shortened; they are published as insufficient-data rather than dressed up. That honesty is the product.

< 1%S1 · false-"done" rate, with receipts on 100% of claims
100%S2 · destructive-gate firings preceded by a snapshot
n/aS3 · compaction survival — insufficient-data, window forfeit
0S4 · unverified subagent write claims in main
n/aS5 · time-to-context — insufficient-data, window forfeit
≥ 90%S6 · cross-machine collision warns, scored in a 2-machine drill
≤ 10%S7 · SMA self-cost share of session spend, p95 ≤ 300 ms
≥ 90%S8 · planted canary false-"dones" the blind verifier must catch

How one prompt flows

The path a single instruction takes through the framework. Each step opens the subsystem that owns it.