Add sssf skill, installable via the skills CLI

Port the sssf skill from ~/.agents/skills/sssf into this repo so it can be
distributed and installed with the skills CLI (skills add INDigitalStudio/skills
--skill sssf).

- Copy the skill (SKILL.md, cookbooks, references, scripts, templates, and the
  visualizer app source) into sssf/.
- Gitignore build/runtime artifacts: the visualizer's node_modules/ and dist/,
  Python bytecode, and the machine-specific repos.json.
- Make the skill location-independent: install.py now stamps the skill's real
  path into the stamped justfile's skill_dir (replacing the hardcoded
  ~/.agents/skills/sssf), so 'just obs' finds the visualizer wherever the CLI
  installed the skill.
- Update cookbooks to use <skill>/scripts/... instead of the hardcoded path,
  and document the skills CLI install command.
- Update the repo README with install instructions.
This commit is contained in:
INDigitalStudio 2026-08-09 21:00:28 +00:00
parent a42608f602
commit 2cc766aabe
98 changed files with 12508 additions and 0 deletions

View file

@ -0,0 +1,122 @@
# Create ADW
Compose a new ADW script — a thin, deterministic Python workflow over agents already in the config. Design the chain first, then generate or hand-write it.
## Step 1 — Design the chain
Answer four questions, in order:
1. **What agents, in what order?** Pick from the roster (`adws/adw_sssf_config/sssf.config.yaml`). The starter six cover most chains:
| Agent | Use when | Output type | Typical gates |
|---|---|---|---|
| `scout` | you need to FIND something first — read-only recon | `ScoutOutput` | `artifacts_exist` |
| `planner` | the work needs a plan before code changes | `PlanOutput` | `artifacts_exist`, `files_non_empty` |
| `builder` | code must change | `BuildOutput` | `diff_matches_claims` |
| `reviewer` | the change must be confirmed to BE what was asked for | `ReviewOutput` | `artifacts_exist`, `verdict_consistent` |
| *(no tester)* | verifying that it RUNS is a `kind="code"` phase over `quality.py`, not an agent | `QualityResult` → `as_envelope` | the exit code is the check |
| `documenter` | finished work needs a write-up (runs after a build, off the diff) | `DocumentOutput` | `artifacts_exist`, `files_non_empty` |
| any agent, generic ask | one-off prompt, no special shape | `GenericOutput` | as needed |
A new kind of agent needs a config entry + prompt pair + output type first — see `update_config.md`.
**The suite and the reviewer answer different questions.** "Does it run" is a test, and code can ask that. "Is this the thing that was asked for" is a review, and only an agent can. A green suite over a feature nobody requested is still a failed request, and neither one covers for the other.
2. **Where does code act?** Git branch/commit, migrations, deploys each get their own `kind="code"` phase — never buried inside an agent phase.
**Running the suite is one of these — there is no tester agent.** The command is written down in `quality.py`, so a `kind="code"` phase runs it (`quality.run_tests(run)` → `quality.as_envelope(result, "tests")` back into the builder) and the bounded repair loop is unchanged. An agent rediscovering `bun test` on every run buys nothing a subprocess does not already know. Capturing what changed is one of these: `changes.capture(run, ChangeCapture(base="main"))` diffs the working tree against a resolved base, writes `context_handoff/changes.diff`, and `changes.as_envelope(...)` hands it to the next agent. A diff is two git commands, not a judgement call.
3. **Does anything loop?** Test-fix cycles are bounded fix loops (see `update_adw.md`), not phase retries.
4. **What does each call need to prove?** Pick gates per call from `gates.py`: `artifacts_exist`, `files_non_empty`, `json_parses`, `diff_matches_claims`, `tests_pass("cmd")` — or an inline one-off.
## Step 2 — Ownership rules (the swim lanes depend on these)
- `kind="agent"` → `owner` MUST be an agent name from the config — it selects the harness (model, thinking, tools, prompts) AND the lane. `ph.call()` runs whoever owns the phase.
- `kind="engineer"` → `owner=run.engineer`. Every ADW opens with the engineer request phase — it is the system input record.
- `kind="code"` → `owner` is a short actor label (`"git"`, `"db"`); all code phases share the code lane.
- Phase `name` must be unique within the run (`plan`, `build`, `test_1`, `fix_1`, …) — the UI keys blocks on it.
- **`description` is required and must earn its place.** The name identifies the phase; the description explains it — what this phase does and why, in one sentence. It rides the `phase_start` event and is the only line of intent the trace, the console, and the phase block ever show. `PhaseParams` raises at construction on a blank description *or* one that merely restates the name (`commit_plan: "Commit the plan"`), so the rule fails before the phase opens rather than leaving an unreadable run in the db. Write `"Put the spec on record before any code exists to blur it"` instead.
- `retries=N` on an **agent** phase = extra gate-correction rounds re-sent into the same session (pi's `--session-id` creates-or-continues, so context stays intact). Code-phase re-execution is not implemented in v1.
## Step 3 — Generate or write it
```bash
uv run <skill>/scripts/make_adw.py --name review_docs --agents scout,builder
```
`<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf` or a repo's `.claude/skills/sssf`).
Writes `adws/adw_review_docs.py`: one agent phase per name, chained by `previous=`, starter agents mapped to their output types, unknown agents to `GenericOutput`. It does NOT create config entries or prompt files — do that first (`update_config.md`), or `agents.validate()` will stop the run and tell you what's missing.
## The canonical skeleton
Every `adw_*.py`, generated or hand-written, is a `uv` single-file script with this shape:
```python
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan Build — plan the request, then implement the plan."""
import argparse
import sys
from adw_modules import agents, gates, git_helper, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner", "builder"] # names, never models
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config) # 1. point to config
agents.validate(cfg, REQUIRED_AGENTS) # 2. fail fast — nothing spawns on a half-valid config
run = session.ensure(cfg, adw_id) # 3. pin-or-create the session → the Run object
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="build", kind="agent", owner="builder", retries=1,
description="Implement the plan exactly")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.diff_matches_claims]))
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Commit the working tree")) as ph:
message = build.commit_message or f"sssf({run.adw_id}): {build.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))
```
## Non-negotiables
- **`REQUIRED_AGENTS` + `agents.validate()`** — declare every agent name the script uses and validate before the first phase.
- **Every agent call declares a concrete output type** from `data_types.py`. No untyped handoffs.
- **`previous=` carries the chain** — the upstream envelope lands in the next agent's `user.md` as `{{previous_envelope}}`; bulky context moves through `context_handoff/` files the envelope references.
- **The engineer request phase comes first**, always.
- **Four-param rule** — `run.phase()` and `ph.call()` each take exactly one object; new helpers with >4 params get a data type.
- **Stay thin** — sequencing and acceptance only; real logic goes in `adw_modules/` (`update_modules.md`).
- **Committing is a code phase, and it needs a fallback.** `PlanOutput`, `BuildOutput`, and `DocumentOutput` each carry a `commit_message` the agent writes **for its own work product** — the spec, the code, the write-up. It defaults to empty, so always `envelope.commit_message or <fallback>`, and commit each product with the message of the agent that made it (`adw_simple_sdlc.py` commits three times and never crosses them). `git_helper.commit_all(message)` stages everything, commits, and returns the short sha; it raises a clear error when the cwd isn't a git repo or nothing changed, and that raise fails the phase.
## Before you ship it
1. `uv run adws/adw_<name>.py "a tiny real request"` — watch it go green end to end.
2. Check the trace: `sqlite3 adws/adw_data/sssf.db "select seq,name,kind,owner,status from phases where adw_id='<id>' order by seq;"`
3. Read the final `envelope.json` — is the output type earning its fields, or should it be sharper?

View file

@ -0,0 +1,63 @@
# Create Config
Generate `sssf.config.yaml` — the agent roster for a target repo.
## Generate it
```bash
uv run <skill>/scripts/make_config.py
```
`<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf` or a repo's `.claude/skills/sssf`).
Writes `adws/adw_sssf_config/sssf.config.yaml` — creating the directory if needed — with the starter agents (planner, builder, scout, reviewer, documenter) wired to the prompt files `/sssf install` stamped into `adws/adw_data/prompt_engineering/`. That path is the default every ADW and the justfile look for; `--config` overrides it. `make_config.py` refuses to overwrite an existing config unless you pass `--force`, so retuning an existing roster is a hand edit — see `update_config.md`.
## The rule
**One agent, one prompt, one purpose.** An entry defines who an agent *is*: its coding agent, model, thinking level, and exactly one system prompt plus one user prompt. How it gets *used* — the output type, a per-call user prompt override — lives at the ADW call site, never here.
## Schema
```yaml
defaults:
coding_agent: pi # v1: pi only (claude_code is specced, stubbed until v2)
model: google/gemini-3.6-flash # ALWAYS provider/model-id — a bare id is ambiguous
thinking: medium # off | minimal | low | medium | high | xhigh | max
harness_engineering: [] # pi extension names
data_dir: adws/adw_data # runtime home: {data_dir}/sessions/{adw_id}/{agent_name}/
observability:
db: adws/adw_data/sssf.db # tracer writes here; the UI polls it
poll_ms: 500 # visualizer live-poll cadence
agents:
- name: planner # ADW scripts name agents, never models
coding_agent: pi
model: google/gemini-3.6-flash
thinking: high
color: "#a78bfa" # optional hex — this agent's lane color in the visualizer
purpose: Turn a request into a plan the builder can implement without asking questions.
prompt_engineering:
system: adws/adw_data/prompt_engineering/planner/system.md
user: adws/adw_data/prompt_engineering/planner/user.md
- name: scout
thinking: high # unset keys fall through to defaults
purpose: Find and report where things live; change nothing.
prompt_engineering:
system: adws/adw_data/prompt_engineering/scout/system.md
user: adws/adw_data/prompt_engineering/scout/user.md
tools: # optional allowlist — omit the key entirely for all tools
- read
- bash
```
Every agent entry merges over `defaults`, so an entry only states what differs. Pi's builtin tools are `read`, `bash`, `edit`, `write` — a read-only recon agent gets `[read, bash]`; a builder omits `tools` altogether.
## After generating
1. Each agent needs its prompt pair to exist on disk: `adws/adw_data/prompt_engineering/{name}/system.md` and `user.md`. `agents.validate()` fails the run at startup if either is missing.
2. Write `purpose` as one sentence and make the system prompt say the same thing — the two should not drift.
3. Validate by running the smallest ADW that names your agents; a bad entry fails fast, before anything spawns.
Full field-by-field spec, thinking-level mapping, and model resolution: `references/config.md`. Retuning an existing roster: `update_config.md`.

View file

@ -0,0 +1,110 @@
# How to Prompt for the Engineering
Read this **before every ADW launch**. The prompt you pass is what the whole chain reads: the planner plans from it, the builder builds from it, the reviewer judges against it. Your prompt might run through 10s or 100s of agents. A sloppy prompt is not a small tax; it is paid again by every agent in the chain.
## Purpose
Turn what the engineer said into the prompt the ADW receives: **clearer, not different.** You are a translator, not a redesigner.
## The one rule
**The intent is theirs. The precision is yours.**
| You MAY | You MAY NOT |
|---|---|
| Carry every constraint forward, verbatim | Quietly drop a requirement because it looks hard or odd |
| Fix grammar, cut rambling, order the steps | Soften a strong ask ("rewrite" → "refactor a bit") |
| Change the language used to better communicate the idea | Research the codebase for exact file names, never go into the app |
If you catch yourself improving the *idea* rather than the *sentence*, stop. Raise the concern to the engineer in your own message and launch what they asked for.
## You never touch the application, you prompt, monitor, observe, and report.
Outside of understanding the ADWs, you never research, touch, or dive into the codebase thats being operated on.
Your role is to simply kick off the workflow. There are entire teams of agents inside these ADWs built to do the work.
Your job is to kick it off, monitor, observe, report. Not interact with the application layer. You operate only on the agentic layer, the ADWs, the software factory.
## The shape
Four lines. Nothing else earns its tokens.
```
<the ask — one imperative sentence, their words where they were specific>
Where: <files or dirs you verified>
Done means: <the observable result — a response shape, a passing test, a rendered element>
Out of scope: <what you were tempted to add, named so nobody adds it>
```
**Before** (what the engineer said):
> can we get tags on posts, sorted by popularity
**After** (what the ADW receives):
```
Add a GET /api/tags endpoint returning {tags: [{tag, count}]} — the distinct tags
across all posts with how many posts carry each, sorted by count descending then
tag ascending.
Where: src/server.ts (routes), src/server.test.ts (tests)
Done means: GET /api/tags returns the counts, and a new test in server.test.ts covers it.
Out of scope: tag editing UI, tag filtering on the post list.
```
Same idea, same scope. What changed is that "popularity" became a sort order, the files are named, and nobody has to guess where it stops.
## Which ADW
**If the engineer named one, launch that one.** Their call stands — no second-guessing, no "upgrading" them to a longer chain. If you think another fits better, say so in your own message and launch what they asked for.
**If they did not, read what this repo actually has and choose from that.**
```bash
ls adws/adw_*.py # the menu
head -20 adws/adw_<name>.py # every ADW opens with its `Phases:` line — the chain in one line
```
Chains are the engineer's to add, rename, and rewire, so **the files on disk are the only authority**. Never launch from memory or from a name you saw in a doc; read the docstrings, then match by shape:
| The work | Look for a chain that |
|---|---|
| Changes code, and the shape is not obvious — new behaviour, more than one file, anything you would want a plan for | goes end to end: plans, builds, verifies, reviews, and documents |
| Changes code, one well-understood edit | plans, builds, and verifies |
| Implements a plan this session already produced (`--adw-id`) | starts at build and verifies |
| Confirms built work is what was asked for | ends in a review phase |
| Writes up work already shipped | captures the diff and documents it |
| Is a question, and nothing should change | is a single read-only agent — the one case where one phase is right |
**Never a single-agent chain when the engineer asked for work to be done.** One-phase ADWs answer questions and run one-offs; they do not deliver.
**The more complex the ask, the more complete the chain.** Complexity means: more than one file, a behaviour you cannot describe in one sentence, anything touching data or an interface others call, or any request where you had to guess. When two chains both fit, take the longer one — a phase you did not need costs cents, while a change nobody planned, verified, reviewed, or wrote up costs an afternoon.
If nothing on disk fits the shape you need, say so and offer to compose one (`create_adw.md`) rather than forcing the work into a chain that skips the phase it needed.
## Workflow
1. **Read it twice.** Mark every noun that could point at two things.
2. **Verify before you write.** Every path, route, and symbol you put in the prompt must exist — check it. A wrong path costs a whole build phase.
3. **Draft the four lines.**
4. **Diff against the original.** Every specific thing they said, still there? Anything in your draft they did not say? Delete it.
5. **Ask at most one question**, only when two readings would produce different code. Otherwise state your assumption in the prompt and say so when you report.
6. **Launch** the chain from *Which ADW* above; `run_adw.md` covers the mechanics and the watching. Inline for a short ask; for anything longer, write `requests/<slug>.md` and pass the path — every ADW takes either.
## Rules that do not bend here
- **Do not write the plan.** Your prompt says WHAT and DONE MEANS. HOW belongs to the planner — unless the engineer specified how, and then you carry it word for word.
- **Do not address the harness in the prompt.** "Use the reviewer", "retry twice", "then commit" are chain choices, and the chain is chosen by which ADW you launch, not by prose the agents will read.
- **Do not pad.** No preamble, no restating the repo, no encouragement. Gates check claims, not prose.
- **Their exact words survive.** When the engineer was specific — a name, a number, a format, a file — quote it rather than paraphrasing.
## Report back
After launching, show the engineer three things so a bad translation dies in seconds rather than at the commit phase:
1. **The prompt you actually sent** — verbatim.
2. **The ADW you chose**, and the one-line reason — or that you used the one they named.
If they named a roster (a config, a model tier), say which one you ran on; if they did not, you ran the default, and switching that is their call, not yours (`run_adw.md`).
3. **The `adw_id`**, so they can watch it (`just phases <adw_id>`).
Then observe and report per `run_adw.md`. You run the system; you do not do the work inside it.

69
sssf/cookbooks/install.md Normal file
View file

@ -0,0 +1,69 @@
# Install
`/sssf install` — stamp the entire factory out of the skill and into the current working directory.
## Install the skill
The skill is distributed from the `INDigitalStudio/skills` repo. Install it with the skills CLI:
```bash
skills add INDigitalStudio/skills --skill sssf -g -y
```
That places the skill (with its `scripts/`, `templates/`, and the visualizer app) into your agent's skills directory. The rest of this cookbook assumes the skill is installed and you can point at its `scripts/install.py`.
## Run it
```bash
uv run <skill>/scripts/install.py
```
Run from the **target repo root** — the cwd is where everything lands. `<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf`, a repo's `.claude/skills/sssf`, or wherever the skills CLI placed it). The scripts resolve their own location, so any of those paths works.
`install.py` asks which **coding-agent harness** to use — `pi` or `omp` — and writes your choice into the stamped config's `defaults.coding_agent`. Pass `--harness pi|omp` to skip the prompt. It also **registers the repo** in the visualizer's `repos.json` (next to the skill), so the multi-repo trace UI can browse this repo's sessions.
## What gets stamped
`install.py` copies `templates/` into the cwd:
| Stamped | From | Tracked? |
|---|---|---|
| `adws/adw_sssf_config/sssf.config.yaml` | `templates/sssf.config.yaml` | yes — the agent roster |
| `.env.sample` | `templates/env.sample` | yes |
| `adws/adw_*.py` | `templates/adws/` | yes — the thirteen starter ADWs (incl. `adw_validate.py`) |
| `adws/adw_modules/` | `templates/adws/adw_modules/` | yes — all low-level logic |
| `adws/adw_data/prompt_engineering/{planner,builder,scout,reviewer,documenter}/` | `templates/prompt_engineering/` | yes — **the user-owned home for prompts** |
| `adws/adw_data/harness_engineering/` | `templates/harness_engineering/` | yes — **the user-owned home for pi extensions** |
| `justfile` | `templates/justfile` | yes — starter recipes: `just demo`, the workflows, the trace reads, `just obs` |
| `adws/adw_data/sessions/`, `adws/adw_data/sssf.db` | created at runtime | no — gitignored |
The two `*_engineering` dirs mirror the two config keys of the same name: `prompt_engineering` is what an agent is told, `harness_engineering` is what its harness can do. Both are yours the moment they are stamped. Edit them in `adws/adw_data/`, never back inside the skill.
`harness_engineering/` ships with `subagents.ts` — the pi extension backing `subagent_create` / `_continue` / `_list` / `_remove`, wired to the planner and scout in the starter roster.
## Idempotency
Re-running is safe. `install.py` skips **every** file that already exists — your config, your prompts, and previously stamped code alike — and reports what it skipped, so a second run doubles as a drift check. To refresh stamped code (`adw_modules/`, the starter `adw_*.py`) to the skill's current version, run with `--force` — but know that `--force` overwrites ALL existing stamped files, including `sssf.config.yaml` and `prompt_engineering/`, so commit or back up user-owned edits first.
## Post-install checklist
1. **Env** — if the repo has no `.env` yet, `cp .env.sample .env` and set the key(s) your harness needs. **If the repo already has a `.env`** (an app's own env, e.g. PocketBase keys), do NOT overwrite it — the ADWs read `.env` via dotenv and merge, so just append the SSSF keys you need. Pi needs `OPENROUTER_API_KEY`; omp reads its own model catalog (`omp models --json`). `ANTHROPIC_API_KEY` / `CLAUDE_CODE_PATH` are only needed once Claude Code lands in v2.
2. **The harness is installed and on PATH** — `pi --version` (or `omp --version`). Set `PI_PATH` / `OMP_PATH` in `.env` if not.
3. **The model resolves** — `just validate` checks the whole roster (names, prompt files, models) without running anything. If a model doesn't resolve, point the roster at one that does: set `defaults.model` and drop the per-agent `model:` overrides. See `references/config.md` for model resolution.
4. **Gitignore** — `install.py` appends `adws/adw_data/sessions/`, `adws/adw_data/sssf.db*`, and `.env` for you; confirm they landed. All three are runtime or secrets and must never be committed.
5. **Git repo** — ADWs that end in a commit phase call `git_helper.commit_all`, which raises if the cwd is not a git repository. Run `git init` and make a first commit before using `adw_plan_build.py`, `adw_plan_build_test.py`, or `adw_simple_sdlc.py`. `adw_document.py` needs one too: it measures the change with `git diff` against a base ref (`main` by default, `--base` to override).
6. **Smoke test** — first `just validate` (roster check, nothing runs), then `just demo` runs two cheap read-only workflows back to back, or run the smallest ADW directly:
```bash
just validate # roster check first
just demo # both, end to end
uv run adws/adw_prompt.py "reply with a one-line summary of this repo" # the raw form
```
Green means the whole path works: config validated, session minted, the coding agent ran, envelope parsed, events landed in `adws/adw_data/sssf.db`. Verify the trace exists before trusting anything larger:
```bash
sqlite3 adws/adw_data/sssf.db "select adw_id, status from sessions order by started_at desc limit 1;"
```
If the smoke test fails, fix it before composing chains — every multi-agent ADW rides on this exact path.

122
sssf/cookbooks/run_adw.md Normal file
View file

@ -0,0 +1,122 @@
# Run ADW
Run a workflow and report on it. **You run and observe — you never step into the process or do the work yourself.**
## Step 0 — translate the request
**Read [how_to_prompt_for_the_eng.md](how_to_prompt_for_the_eng.md) before you launch anything.** The prompt you pass is read by every agent in the chain, so it gets written deliberately: same intent, sharper words, verified paths, and a stated "done means". That cookbook is the whole procedure; this one starts once you have the prompt.
## The orchestrator's posture
The ADW is the worker. Your job is to launch it, watch the trace, and tell the engineer what happened. Do not read the agent's target files and "help", do not fix the code an agent was supposed to fix, do not edit an envelope. If a run fails, report the failing phase and its violations — the fix is a config, prompt, or ADW change, made deliberately, and then a re-run.
## Launch
Which chain to launch is decided in `how_to_prompt_for_the_eng.md`, and the short version is: **the ADW the engineer named, or else the most complete composed chain the work justifies — never a single-agent one.** Read `ls adws/adw_*.py` and the `Phases:` line in each docstring to see what this repo has; the names below are shape, not a menu.
```bash
uv run adws/<end-to-end-chain>.py "add a /health endpoint"
uv run adws/<plan-build-verify-chain>.py requests/health.md
uv run adws/<build-first-chain>.py "implement the plan" --adw-id a1b2c3d4
uv run adws/<recon-chain>.py "where is auth handled" --config path/to/other.config.yaml
```
The prompt is inline text or a file path. Launch in the background so you can poll while it works; the `adw_id` is printed on startup — capture it, everything else keys off it.
### Listen for the roster
The chain says *what runs*; the config says *who runs it*. **If the engineer references a roster, a config, or a model tier, pass it — do not fall through to the default.**
```bash
just rosters # every roster on disk, and the model each agent runs
```
That prints the path to pass and who is in it, in one read:
```
adws/adw_sssf_config/sssf.config.yaml
planner fireworks/accounts/fireworks/models/kimi-k3
builder google/gemini-3.6-flash (inherited)
adws/adw_sssf_config/sssf.frontier.config.yaml
planner anthropic/claude-opus-5
```
Read those from disk every time. Rosters are the engineer's to add, rename, and retune, so a name you remember from a doc is a guess.
They will rarely say `--config`. Treat any of these as naming a roster, then resolve it to a file:
| What they say | What it means |
|---|---|
| "run it on the frontier config", "use the frontier roster" | the roster file whose name matches |
| "run this with the big models", "use the sota roster" | the non-default roster — confirm which if there is more than one. Each config's header comment lists the names it answers to, so `head -3` on the file settles it |
| "have opus plan this one" | a roster whose planner is that model; if none exists, say so rather than editing the config mid-request |
| nothing about models at all | the default, `adws/adw_sssf_config/sssf.config.yaml` |
`--config` takes the path directly; the justfile recipes read `SSSF_CONFIG` instead:
```bash
uv run adws/<chain>.py "<prompt>" --config adws/adw_sssf_config/sssf.frontier.config.yaml
SSSF_CONFIG=adws/adw_sssf_config/sssf.frontier.config.yaml just <recipe> "<prompt>"
```
Two things that bite:
- **Never swap rosters on your own.** A different roster is a different cost and a different result. If the default's model looks wrong for the work, say so and let the engineer choose.
- **Switching rosters mid-session breaks resumption.** `agent_map.json` records the model each coding-agent session was created with, so a joined run (`--adw-id`) whose config now names a different model starts that agent **fresh** instead of resuming its context window. That is deliberate — a bad resume is worse — but it means "plan on the frontier roster, then build on the default" costs the builder its accumulated context. Say so when you report it.
`--adw-id` is optional on **every** ADW. Given one, the run joins that session if it exists or creates it pinned to exactly that id: same `sessions/{adw_id}/` dirs, same `context_handoff/`, envelopes appended, and each agent resumes its existing coding-agent context window via `agent_map.json`. That is how you chain ADWs — plan under one id, then build under the same id.
## Observe
The trace db is `adws/adw_data/sssf.db`. It is WAL, so reads never block the running writers — poll it as often as you like.
```bash
# where the run stands
sqlite3 adws/adw_data/sssf.db \
"select seq, name, kind, owner, status, attempt from phases where adw_id='a1b2c3d4' order by seq;"
# the live tail — cursor on rowid, same query the visualizer polls
sqlite3 adws/adw_data/sssf.db \
"select rowid, type, name, started_at from events where adw_id='a1b2c3d4' and rowid > 0 order by rowid limit 50;"
# why a phase failed
sqlite3 adws/adw_data/sssf.db \
"select attempt, gate, passed, checks_json from gate_results where adw_id='a1b2c3d4';"
# session-level status
sqlite3 adws/adw_data/sssf.db \
"select adw_id, request, status, total_tokens from sessions order by started_at desc limit 5;"
# what an agent actually did, slowest tool calls first
sqlite3 adws/adw_data/sssf.db \
"select name, tokens, started_at, ended_at from events
where adw_id='a1b2c3d4' and type='tool_call' order by ended_at desc limit 20;"
```
Poll on a cursor: keep the highest `rowid` you have seen and query `where rowid > ?`. Don't re-read the whole table each pass.
`tool_call` rows carry a real span, so durations come off the columns — see `references/observability.md` for which fields each event type populates.
The ADW also narrates to stdout, and every line it prints is written to the db as a `log` event — terminal and swim lane tell the same story by construction, so tailing the background process is a valid second view rather than a competing source of truth.
Files are the raw record if you need more than the db shows: `adws/adw_data/sessions/{adw_id}/{agent}/raw_output.jsonl` (full coding-agent stream), `envelope.json` (the parsed final response), `prompts/` (exactly what was sent), and `context_handoff/` (what agents wrote for each other).
## When a run is stuck
A hung coding agent produces no events at all, so the trace goes quiet rather than red. Read it in this order:
```bash
just phases <adw_id> # which phase is still `running`
just procs <adw_id> # what that phase is actually running, with pids
just kill <adw_id> # stop it — children first, then the workflow
```
`processes` rows with `ended_at IS NULL` are the live ones. If `procs` shows a pi child but the phase has produced no `tool_call` events and its `raw_output.jsonl` is empty, the agent never got started properly — check the model resolves and that nothing is blocking the subprocess, rather than waiting it out. `just kill` verifies each pid still matches the command that was recorded before signalling, because pids get recycled.
A killed run marks itself `fail` and closes its process rows, so the trace never claims work is in flight that is already dead.
## Report
Tell the engineer, in order: which chain and which roster you launched (name the config whenever it was not the default), which phase is running now (or which failed), phase statuses in sequence, and for a failure the gate violations or the error verbatim. Remember **every phase defaults to `fail`** — a phase showing `fail` may simply never have completed; `queued` means it never started. Don't dress up a partial run as a success.
For a visual live view, the visualizer app in the skill (`just obs`, or tmux sessions viz-api :4600 + viz-ui :4601) polls this same db — sessions as cards, runs as swim lanes, phases and tool calls drill-in. The sqlite queries above remain the headless equivalent.

View file

@ -0,0 +1,88 @@
# SSSF Overview
The system map the orchestrator reads on startup — what SSSF is, how a stamped repo is laid out, and which cookbook to load next.
## What SSSF is
Super Simple Software Factory builds repeatable **agents plus code** workflows. Deterministic Python (an ADW script) owns sequencing, retries, and acceptance; agents are bounded nodes inside that graph. Agent proposes, code disposes.
Your job as orchestrator: **run the system, observe the system, help the engineer interact with it.** You do not do the work an ADW exists to do.
## Layout of a stamped repo
```
adws/
├── adw_sssf_config/
│ └── sssf.config.yaml the agent roster — one agent, one prompt, one purpose
├── adw_prompt.py smallest ADW: one agent, one prompt, traced end-to-end
├── adw_plan.py, adw_scout.py, adw_build.py, adw_plan_build.py, adw_build_test.py, adw_plan_build_test.py
├── adw_build_review.py build → review: is this what was asked for? (not testing)
├── adw_document.py write up the work just done, from git diff vs main
├── adw_simple_sdlc.py plan → build → test → review → document; commits each product
├── adw_modules/ ALL low-level logic — ADW scripts stay thin
│ ├── data_types.py AgentCall, PhaseParams, Phase, Envelope + one output type per agent call
│ ├── agents.py load_config, validate, resolve entry → interface + model + thinking
│ ├── runner.py the Run object: run.phase(PhaseParams) → ph.call(AgentCall)
│ ├── agent_pi.py Pi interface (v1) · agent_cc.py Claude Code (v2, stubbed)
│ ├── gates.py gate(envelope, run) -> GateReport — one check per item verified
│ ├── changes.py git diff vs a resolved base → ChangeSet → envelope for the documenter
│ ├── prompts.py, session.py, tracer.py, console.py, git_helper.py, utils.py
└── adw_data/
├── prompt_engineering/{agent}/{system.md,user.md} tracked — edit prompts HERE, never in the skill
│ planner · builder · scout · reviewer · documenter
├── sessions/{adw_id}/ gitignored runtime
│ ├── agent_map.json agent → coding-agent session_id + model
│ ├── context_handoff/ the one place agents write files for the agents that follow
│ └── {agent}/{prompts/, raw_output.jsonl, envelope.json}
└── sssf.db gitignored SQLite trace db the visualizer polls
```
**v1 runs Pi only.** `coding_agent: pi`, default model `gemini-3.6-flash`, thinking `medium`. `claude_code` is specced in the config and stubbed in the interface — it lands in v2.
## The phase model
Every ADW run is a sequence of **phases**, each one `with run.phase(PhaseParams(...))`. Three kinds, three swim lanes:
- **engineer** — the human lane; today the system-input phase (who asked, and for what).
- **agent** — `ph.call(AgentCall(...))`: prompt in → typed envelope out → gates verified.
- **code** — deterministic steps that stand alone (git branch, git commit, migrate). Never buried inside an agent phase.
**Success must be earned — every phase defaults to `fail`.** A clean exit flips it to success; agent phases additionally require the envelope to parse and all gates to come back green. A raise keeps it failed, records an error event, and aborts the run. `retries=N` on an agent phase buys extra gate-correction rounds through the same session before that raise happens.
## Envelopes
Agents have exactly two output channels: reference files written into `context_handoff/`, and a **final valid-JSON response** parsed against the output type the call declared. Code persists it as `envelope.json` and injects it into the next agent's `user.md` via `{{previous_envelope}}`. Bad JSON is never a restart — the harness re-prompts the *same session, context intact*, until it parses (bounded). See `references/handoff.md`.
**The output contract is a synced triad**: the type in `data_types.py` ↔ the `## Report` JSON example in the agent's `user.md` ↔ `output_type=` at the call site. Editing any one of the three means editing all three in the same change — drift between them taxes every call with correction retries.
## Running an ADW
```bash
uv run adws/adw_plan.py "add a /health endpoint"
uv run adws/adw_plan_build.py requests/health.md --adw-id a1b2c3d4
```
The prompt is inline text or a file path. `--adw-id` is optional on every ADW: given one, the run joins that session (same dirs, same `context_handoff/`, agents resume their existing context windows); omitted, a fresh id is minted and printed.
## When you have finished reading this
You are done with startup. List the ADWs (`ls adws/adw_*.py`, plus each `Phases:` docstring line) as a table, and **wait for the engineer's request.**
Do not survey anything else — not the trace db, not the config, not past runs, not the repo tree. You do not yet know what the request is, so anything you gather now is a guess about what will matter, spent from the context the real work needs. Every cookbook and reference below is lazy-loaded, one per request, and that is the whole design.
## Where to go next
Load one cookbook per request — this overview is the only one you read up front.
| Request | Cookbook |
|---|---|
| Turn a request into the prompt an ADW gets | `how_to_prompt_for_the_eng.md` — **read before every launch** |
| Set the system up in a repo | `install.md` |
| Write a new ADW script | `create_adw.md` |
| Change an existing ADW chain | `update_adw.md` |
| Generate `sssf.config.yaml` | `create_config.md` |
| Add or retune an agent | `update_config.md` |
| Add low-level logic or a gate | `update_modules.md` |
| Run and monitor a workflow | `how_to_prompt_for_the_eng.md`, then `run_adw.md` |
References, loaded when you need the spec: `references/config.md` (full config schema), `references/handoff.md` (envelope + session layout), `references/observability.md` (events, db tables, polling).

View file

@ -0,0 +1,90 @@
# Update ADW
Modify an existing ADW chain — add phases, add gates, add a bounded fix loop.
## Add a phase
Insert a `with run.phase(...)` block where it belongs in the sequence. Pick the right `kind`: `agent` for a `ph.call(...)`, `code` for a deterministic step, `engineer` for a human touchpoint. If the new phase names an agent not already in `REQUIRED_AGENTS`, add it there too — otherwise validation passes and the run dies mid-flight instead of at startup.
```python
with run.phase(PhaseParams(name="scout", kind="agent", owner="scout",
description="Locate the code the request touches")) as ph:
found = ph.call(AgentCall(output_type=ScoutOutput, prompt=prompt))
```
Phase `name` must be unique within the run — that is what the UI keys blocks on. In a loop, suffix it (`f"test_{i}"`).
`description` is **required**, and `PhaseParams` rejects both a blank one and one that merely restates the name. It is the single line of intent the trace, the console, and the UI phase block show, so write what the phase does and why — `"Land the code only now: green suite, approved review"`, not `"Commit build"`.
A code phase does its work in the block body and logs what it did. The commit phase that closes `adw_plan_build.py` and `adw_plan_build_test.py` is the pattern:
```python
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Land the builder's changes, using the message it wrote")) as ph:
message = build.commit_message or f"sssf({run.adw_id}): {build.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
```
`commit_message` is a field on `PlanOutput`, `BuildOutput`, and `DocumentOutput` that the agent fills in **for its own work product**, so always pair it with a fallback — it defaults to empty. `commit_all` raises if the cwd is not a git repo or nothing changed, which fails the phase rather than committing nothing. A chain that commits more than once (`adw_simple_sdlc.py`) commits each product with its own author's message.
## Remove a phase
Delete the block, drop any now-unused agent from `REQUIRED_AGENTS`, and re-thread the chain: whatever the removed phase produced was probably somebody's `previous=`. Point that call at the surviving upstream envelope.
## Add gates
Gates are callables over the finished envelope — `gate(envelope, run) -> GateReport`, recording one `check(item, ok, note)` per thing they looked at, with violations derived from the failed ones. Compose them per call:
```python
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.artifacts_exist, gates.diff_matches_claims]))
```
On violations the harness does **not** restart the agent — it sends the violation list back into the **same session** as a correction (pi's `--session-id` creates-or-continues, so the context window is intact), bounded by that phase's `retries`. Every gate result is traced to the `gate_results` table. Exhausting the retries raises `GateFailure` and fails the phase.
Gate claims, not guesses: declared artifacts exist and are non-empty, declared JSON parses, declared changes appear in the diff, declared test commands pass. Never hardcode counts — express quantity as a property of the declared list ("at least one artifact", "ALL declared paths valid"). Plan quality and code taste are not gateable; that is a reviewer agent or a human. New reusable gates go in `adw_modules/gates.py` (`update_modules.md`).
## Add a bounded fix loop
The pattern from `adw_build_test.py` — always bounded by a module-level constant. The runner is a **code** phase, because the command is known; only repairing it needs an agent:
```python
MAX_FIX_LOOPS = 3
test = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"test_{i}", kind="code", owner="quality",
description="Run the suite — a known command, so code runs it")) as ph:
test = quality.run_tests(run) # QualityResult, not an envelope
ph.log(passed=test.passed, artifacts=", ".join(test.artifacts))
if test.passed:
break
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description="Repair what the suite reported, from its verbatim output")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(test, "tests"),
gates=[gates.diff_matches_claims]))
return run.finish(accepted=test is not None and test.passed,
reason=f"the suite still failed after {MAX_FIX_LOOPS} fix attempt(s)")
```
`run.finish()` ends every ADW, and it takes the acceptance criterion the phase
statuses cannot express. A test phase that ran a red suite **succeeded** — the
runner did its job — so phases alone would report a green run that never passed
its tests, in the db and the UI as well as the terminal. Pass `accepted=` and
the exit code, the session status, and the banner are decided together.
`quality.as_envelope` is the adapter: a deterministic result shaped as an envelope, so the builder cannot tell it came from code. Wire the real command in `quality.py` first — the stamped blocks are `echo` placeholders that announce themselves.
Three distinctions worth keeping straight:
- **Gate retries vs. JSON retries.** `retries` buys extra *gate*-correction rounds. Malformed final JSON is handled separately and always — `JSON_FIX_ATTEMPTS` in `adw_modules/agents.py` (2 by default) re-prompts the same session for a valid object even on a phase with `retries=0`. Raising the phase's `retries` does not buy more JSON attempts, and vice versa.
- **Phase retries vs. fix loops.** `retries=N` on `PhaseParams` re-attempts one agent phase's gate corrections, re-sent into the same session with its context intact. (Code-phase re-execution is not implemented in v1.) A fix loop is a *chain* of phases repeated — different agents, new envelopes each pass.
- **The test phase succeeds when it runs and reports correctly.** A failing suite does not fail that phase; it fails the run, checked at the end. The runner did its job; the code didn't.
## Keep scripts thin
An ADW is sequencing and acceptance — nothing else. The moment you are writing parsing, subprocess handling, retry mechanics, or a reusable predicate inside `adw_*.py`, it belongs in `adw_modules/`. See `update_modules.md`.

View file

@ -0,0 +1,103 @@
# Update Config
Add or retune agents in `sssf.config.yaml`.
## Retune model or thinking
Edit the agent's entry in place:
```yaml
- name: builder
model: google/gemini-3.6-flash # ALWAYS provider/model-id
thinking: high # was medium
```
Write the model as `provider/model-id`, never a bare id. The same model is usually carried by several providers, and an ambiguous pattern raises in `agents.validate()` — grounding every agent that inherits it. See `references/config.md`.
Thinking levels are Pi's reasoning effort: `off | minimal | low | medium | high | xhigh | max`. It only bites when the model is registered with `reasoning: true` in `~/.pi/agent/models.json`.
**A model change means a fresh session.** `agent_map.json` records the model each coding-agent session was created with. When a joined run (`--adw-id`) finds the config's model no longer matches the recorded one, that agent starts a **new** session rather than resuming — the map is updated, never a bad resume. Thinking changes do not invalidate a session; model changes do. Expect the agent to lose its accumulated context window on the first run after the change.
## Recolor an agent's lane
```yaml
- name: builder
color: "#22d3ee" # hex; the starter roster ships violet/cyan/amber/green
```
Purely cosmetic and safe to change mid-project: the color rides the `agent_start` event and the `agent_sessions` row, so the visualizer picks it up on the next run without touching past sessions. Omit the key to let the UI's fallback palette choose.
## Retune tools
Pi's seven builtins: `read`, `bash`, `edit`, `write`, `grep`, `find`, `ls`. The last three are **off in bare Pi**, so an agent that doesn't name them will shell out through `bash` to search and list.
Set the roster-wide floor in `defaults`, then narrow per agent:
```yaml
defaults:
tools: [read, bash, edit, write, grep, find, ls]
agents:
- name: reviewer
tools: # explicit list wins over defaults
- read
- grep
- find
- ls
- bash
- write
```
**Resolution:** the agent's own list wins → else it inherits `defaults.tools` → else `None`, meaning all tools. An empty list is not "all tools"; it is a tool-less agent, and it will stall.
Narrow by role, not by reflex:
- Any agent that must produce a `context_handoff/` artifact needs **`write`** — without it, it falls back to a `bash` heredoc to create the file the gate checks for.
- Withhold `edit`/`write` only where the restriction *is* the guarantee. The reviewer's contract is "change nothing", so withholding `edit` makes that structural instead of merely prompted.
- Recon agents should get the full read surface (`read`, `grep`, `find`, `ls`) — cheaper and more legible in the trace than the equivalent `bash` calls.
**Extension tools count against the allowlist.** `--tools` filters built-in, extension, and custom tools alike. Once an agent has a `tools` list — its own, or inherited from `defaults` — a tool registered by one of its `harness_engineering` extensions is dropped unless it is named there. Nothing errors: the extension loads, the run passes, the tool is just never offered. Any agent with a tool-registering extension must list that tool by name.
## Add harness extensions
```yaml
harness_engineering:
- .pi/extensions/json_guard.ts # a pi extension FILE PATH
```
Entries are pi extension **file paths**, passed through as `pi -e <path>`, applied to that agent only. Reach for an output-tightening extension when an agent keeps wrapping its envelope in prose and burning correction retries. The starter roster ships with none — this is an escape hatch, not a default.
**Adding a tool-registering extension is a two-part edit.** The extension path goes in `harness_engineering`, *and* the tool name it registers goes in that agent's `tools` list:
```yaml
- name: reviewer
harness_engineering:
- .pi/extensions/ast_query.ts # registers tool: ast_query
tools:
- read
- grep
- find
- ls
- bash
- ast_query # REQUIRED — or the extension loads and its tool is filtered out
```
Skip the second half and it fails silently: extension loaded, run green, tool never available to the model. Extensions that only shape output or register flags — no new tool — need no `tools` change.
## Add a new agent
Three steps, all required — skipping any one fails `agents.validate()` at ADW startup, before anything spawns:
1. **Prompts.** Create `adws/adw_data/prompt_engineering/{name}/system.md` (Purpose + Instructions — the agent's static identity, nothing else) and `user.md` (an h3 per incoming datum: `{{prompt}}`, `{{previous_envelope}}`, `{{context_handoff_dir}}`, then the task, then a `## Report` section showing the exact output JSON). Copy an existing pair as the shape.
2. **Config entry.** Name, purpose, prompt refs, plus anything that differs from `defaults`.
3. **An output type.** Every agent call parses against a concrete Pydantic model in `adw_modules/data_types.py`. If none of `PlanOutput`, `BuildOutput`, `ScoutOutput`, `ReviewOutput`, `DocumentOutput` fits the new agent's report, add one — see `update_modules.md`. The user prompt's `Report` section must show exactly that JSON shape.
Then name the agent in an ADW's `REQUIRED_AGENTS` and call it.
## Rules that do not bend
- ADW scripts name **agents**, never models. Swapping a model is a config edit and touches no Python.
- One agent, one prompt, one purpose. If an entry needs two purposes, it is two agents.
- Output types never appear in config — they live at the call site, paired with the user prompt.
Full spec: `references/config.md`.

View file

@ -0,0 +1,100 @@
# Update Modules
Extend `adws/adw_modules/` with new low-level logic.
## The rule
**ALL low-level logic lives in `adw_modules/`; ADW scripts stay thin.** An `adw_*.py` file declares agents, sequences phases, and returns an exit code. Anything else — subprocess handling, parsing, retry mechanics, git plumbing, reusable predicates — goes in a module.
## Where things go
| Module | Owns |
|---|---|
| `data_types.py` | Every Pydantic model: `AgentCall`, `PhaseParams`, `Phase`, `EnvelopeBase` + one output type per agent call, the config models (`AgentConfig`, `SSSFConfig`), `EventRecord`, and `PiRequest`/`PiResult` |
| `agents.py` | `load_config`, `validate`, resolving an entry → coding-agent interface + model + thinking + harness extensions |
| `runner.py` | the `Run` object; `run.phase(PhaseParams)` context manager; `ph.call(AgentCall)` |
| `agent_pi.py` | the Pi interface (v1) — non-interactive `pi -p --mode json`, JSONL stream tailed live, model resolved against `~/.pi/agent/models.json`; `--session-id` creates-or-continues, so running and continuing an agent are the same call |
| `agent_cc.py` | the Claude Code interface — stubbed in v1, lands in v2 |
| `gates.py` | validation gates over envelope claims |
| `changes.py` | deterministic change capture: resolve the base ref, `git diff` into `context_handoff/changes.diff`, adapt the `ChangeSet` into an envelope an agent can be handed |
| `prompts.py` | load system/user prompt refs from config, render placeholders |
| `session.py` | mint or join `adw_id`, maintain `agent_map.json`, create session dirs incl. `context_handoff/` |
| `tracer.py` | append JSONL **and** insert every event into `sssf.db` as it happens |
| `console.py` | the terminal narrative — every line printed also lands in the db as a `log` event, so the UI reads the same story; plain sequential lines, no spinners |
| `console.py` | the rich stdout reporter — every line printed is ALSO traced as a `log` event (`{message, level}`) so the terminal and the swim-lane UI tell the same story |
| `git_helper.py` | branch, status, diff, commit — the raw plumbing `changes.py` composes |
| `utils.py` | safe subprocess env, logging, `resolve_prompt` |
## Never `print()`
Modules report through `run.console` — never a bare `print()`. Each console method prints a rich line **and** writes it to `sssf.db` as a `log` event with payload `{message, level}`, both from one `_emit` helper, so the terminal narrative and the swim-lane UI can't drift. New output means a new method on `Console`, not a print at the call site.
## The four-param rule
**Any function taking more than 4 parameters gets them converted into a concrete data type in `data_types.py`.** `AgentCall` and `PhaseParams` are the pattern — `run.phase()` and `ph.call()` each take exactly one object. This is skill-wide: every module the factory generates obeys it.
```python
class ReviewParams(BaseModel):
"""Everything review_changes() needs. Passed as one object, never loose params."""
base_ref: str
paths: list[str]
max_diff_lines: int = 2000
ignore_generated: bool = True
reviewer: str = "scout"
```
## Adding an output type
Every agent call parses against a concrete type. Extend `EnvelopeBase` — `status`, `summary`, `artifacts`, `notes_for_next_agent` — with only the fields that call actually needs:
```python
class ReviewOutput(EnvelopeBase):
approved: bool
blocking: list[str] = []
```
**The output contract is a synced triad — one change means three edits, always together:**
1. The type in `data_types.py` (the enforcer).
2. The agent's `user.md` `## Report` section showing exactly that JSON (the ask).
3. Every call site passing `output_type=` (the binding) — `grep -rn "ReviewOutput" adws/` to find them all.
If the type and the Report example drift, the agent produces what the prompt asked for, the parser rejects what the type expects, and every call burns correction round-trips before landing — a slow, silent tax. Renaming or removing a field is the same triad edit. Schema details: `references/handoff.md`.
## Adding a gate
A gate is a callable — `gate(envelope, run) -> GateReport`. You record **one check per item you look at**, and the harness derives the verdict: any failed check is a violation, and no failed checks means pass.
```python
from adw_modules.data_types import GateReport
def tests_declared_passed(envelope, run) -> GateReport:
"""Verify the envelope's own test claims, after the fact."""
report = GateReport()
for f in envelope.failures:
report.check(f.test, False, f.error)
report.check("suite", envelope.passed,
"all declared tests passed" if envelope.passed
else f"{len(envelope.failures)} declared failure(s)")
return report
```
`report.check(item, ok, note)` appends and returns the report, so a single-item gate is one line: `return GateReport().check(command, ok, f"exit {code}")`.
**Write a note on passing checks too, not just failures.** The note is the evidence, and it is what makes a green gate worth reading — `artifacts_exist ✓ 1 checked · plan.md — exists, 454B` tells you what was verified, where a bare ✓ tells you nothing. Notes on failed checks double as the reason and are what the agent is told, so phrase them as the problem: `"claimed changed file does not exist"`.
Rules that keep gates honest:
- **Verify claims, never predict.** File names and counts are unknowable before the agent finishes; gates check what the envelope declared.
- **Quantity as properties, not counts.** "at least one artifact", "ALL declared paths exist" — never `len(artifacts) == 3`.
- **Record checks, don't raise.** The harness feeds the derived violations back into the same session as a correction — context intact, bounded by the phase's `retries` — and traces every check, passed or failed, to `gate_results.checks_json` and the `gate_pass`/`gate_fail` event payload.
- **Check every item, even after one fails.** Don't early-return on the first problem; the agent fixes more per correction round when it sees every failure at once, and the trace shows the full picture.
- **Don't gate the ungateable.** Plan quality and code taste are a reviewer agent's job or a human's.
A gate that returns a plain `list[str]` of violations still works — the harness adapts it — but it records no evidence for the items that passed, so prefer a `GateReport`.
Reusable gates live in `gates.py`; genuine one-offs can be defined inline at the ADW call site and passed in `gates=[...]`.
## Before you finish
Run the smoke ADW — `uv run adws/adw_prompt.py "ping"` — since every module change rides the same path a real run does.