Add sssf skill, installable via the skills CLI

Port the sssf skill from ~/.agents/skills/sssf into this repo so it can be
distributed and installed with the skills CLI (skills add INDigitalStudio/skills
--skill sssf).

- Copy the skill (SKILL.md, cookbooks, references, scripts, templates, and the
  visualizer app source) into sssf/.
- Gitignore build/runtime artifacts: the visualizer's node_modules/ and dist/,
  Python bytecode, and the machine-specific repos.json.
- Make the skill location-independent: install.py now stamps the skill's real
  path into the stamped justfile's skill_dir (replacing the hardcoded
  ~/.agents/skills/sssf), so 'just obs' finds the visualizer wherever the CLI
  installed the skill.
- Update cookbooks to use <skill>/scripts/... instead of the hardcoded path,
  and document the skills CLI install command.
- Update the repo README with install instructions.
This commit is contained in:
INDigitalStudio 2026-08-09 21:00:28 +00:00
parent a42608f602
commit 2cc766aabe
98 changed files with 12508 additions and 0 deletions

11
.gitignore vendored Normal file
View file

@ -0,0 +1,11 @@
# Build artifacts — the visualizer app is installed from source via `bun install`
sssf/apps/visualizer/node_modules/
sssf/apps/visualizer/dist/
# Python bytecode
__pycache__/
*.pyc
# Machine-specific: the visualizer's repo registry (absolute db paths),
# regenerated by install.py on each machine.
sssf/repos.json

View file

@ -3,3 +3,26 @@
General folder for the skills we create or modify.
This repository holds the skills we develop and maintain. Each skill lives in its own directory with a `SKILL.md` describing its purpose and usage.
## Installing a skill
Skills are distributed from this repo and installed with the [skills CLI](https://github.com/vercel-labs/skills):
```bash
# list what's available
skills add INDigitalStudio/skills --list
# install one skill (e.g. sssf) globally
skills add INDigitalStudio/skills --skill sssf -g -y
# install every skill in the repo
skills add INDigitalStudio/skills --all -g -y
```
The CLI discovers each skill by its `SKILL.md` (which must declare `name` and `description` frontmatter) and copies the whole skill directory into your agent's skills folder.
## Skills
### sssf
Super Simple Software Factory — deploy and operate repeatable agents+code workflows (ADWs) in any codebase. See [`sssf/SKILL.md`](sssf/SKILL.md) and its `cookbooks/` for usage.

76
sssf/SKILL.md Normal file
View file

@ -0,0 +1,76 @@
---
name: sssf
description: Super Simple Software Factory — deploy and operate repeatable agents+code workflows (ADWs) in any codebase. Use when the user says /sssf install, wants to create/run/update an ADW, manage the agent roster in sssf.config.yaml, or observe running agent workflows. Keywords - sssf, software factory, ADW, AI developer workflow, agent pipeline, install factory.
argument-hint: "[install | create adw | run adw | update config | ...]"
---
# Super Simple Software Factory (SSSF)
Reusable combination of **agents plus code**: deterministic Python ADW scripts own sequencing, retries, and acceptance; coding agents (Pi in v1) work inside bounded phases; typed JSON envelopes carry context between them; everything streams into SQLite for the polled visualizer. Agent proposes, code disposes.
## Startup
Three steps. Then stop.
1. Read [cookbooks/sssf_overview.md](cookbooks/sssf_overview.md) — the system map.
2. `ls adws/adw_*.py` and read each file's `Phases:` docstring line.
3. Print the ADWs as a table — name, the chain, one line on when to reach for it — and **wait for the engineer's request.**
```
| ADW | Chain | Use when |
|---|---|---|
| adw_scout | engineer → scout | read-only recon; nothing changes |
| adw_simple_sdlc | plan → build → test → review → document, 3 commits | the work is real and its shape is not obvious |
```
**Nothing else.** No trace-db queries, no reading the config or the ADW scripts' bodies, no repo inventory, no last-runs summary, no diagnosing an old failure, no "current state" dashboard. None of it was asked for, and it is not free:
- **Volunteered state is guessed state.** An orchestrator that improvised a status board queried a `runs` table and a `payload` column — neither exists (`sessions`, `payload_json`). The spec that would have said so is `references/observability.md`, one lazy read away. Probing to look prepared is how you end up confidently wrong in your first message.
- **It spends the context the real task needs**, before you know what the task is.
- **It is stale on arrival.** State printed before the request describes a system that the very next run changes.
Everything else — the db schema, the roster, the handoff contract — is lazy-loaded through the routing table below, when a request actually calls for it. Reading it early defeats the mechanism.
Two exceptions, both narrow: if the engineer's first message already contains a request, skip the waiting and route it; and if the factory is plainly not installed (no `adws/`, no config), say that in one line instead of the table.
## Orchestrator rules
You run the system, observe the system, and help the user interact with it. **You do no ADW work yourself:**
- Never implement, plan, or test in an agent's place — launch the ADW and watch it.
- Never edit files inside `adws/adw_data/sessions/` — that is the run record.
- Observe by querying `adws/adw_data/sssf.db` (WAL — reads never block writers) **when observing is the task**. This is a capability, not a startup step: query it to follow a run you launched or one the engineer asked about, never to volunteer a status report nobody requested.
- Report phase status plainly: name, owner, status, error if any.
## Request routing (lazy-load the cookbook, then follow it)
| Request | Cookbook |
|---|---|
| `/sssf install`, set up the factory in this repo | [cookbooks/install.md](cookbooks/install.md) |
| create a new ADW / workflow | [cookbooks/create_adw.md](cookbooks/create_adw.md) |
| modify an existing ADW chain | [cookbooks/update_adw.md](cookbooks/update_adw.md) |
| create the config / agent roster | [cookbooks/create_config.md](cookbooks/create_config.md) |
| add or retune an agent (model, thinking, tools, prompts) | [cookbooks/update_config.md](cookbooks/update_config.md) |
| extend adw_modules with new low-level logic | [cookbooks/update_modules.md](cookbooks/update_modules.md) |
| run / monitor an ADW | [cookbooks/how_to_prompt_for_the_eng.md](cookbooks/how_to_prompt_for_the_eng.md) **first**, then [cookbooks/run_adw.md](cookbooks/run_adw.md) |
| turn a request into an ADW prompt | [cookbooks/how_to_prompt_for_the_eng.md](cookbooks/how_to_prompt_for_the_eng.md) |
Deep specs, when needed: [references/config.md](references/config.md) · [references/handoff.md](references/handoff.md) · [references/observability.md](references/observability.md)
## Hard rules (enforced across everything the factory generates)
1. **Validate before running** — every ADW declares `REQUIRED_AGENTS` and calls `agents.validate()` first; a missing/misnamed agent fails before anything spawns.
2. **Typed outputs only** — every agent call pairs with a concrete `EnvelopeBase` subclass in `adw_modules/data_types.py`; parse failures re-prompt the same session (context intact), never restart.
**The output contract is a synced triad**: (a) the type in `data_types.py`, (b) the JSON example in the agent's `user.md` `## Report` section, (c) `output_type=` at every call site. These are ONE contract — change any one, update all three in the same edit (grep the type name to find every call site).
3. **Gates validate claims, not guesses** — `gate(envelope, run) -> list[str]` violations; failures return to the same session as corrections.
4. **Four-param rule** — any function with more than 4 parameters takes one concrete data type instead (`AgentCall`, `PhaseParams` are the pattern).
5. **One agent, one prompt, one purpose** — identity lives in `system.md`; task shape (user prompt + output type) lives at the call site.
6. **ADW scripts stay thin** — all low-level logic lives in `adw_modules/`.
7. **Every phase earns a description** — one sentence on what it does and why, never a restatement of its name. It is the only intent the trace, the console, and the UI ever show; `commit_plan: "Commit the plan"` is rejected at construction, blank is too.
8. **A known command is code, not an agent** — if you can write the invocation down (`bun test`, `ruff check`), it belongs in a `kind="code"` phase via `adw_modules/quality.py`. Agents are for the parts that need reading and deciding; failures come back to the builder as an envelope either way.
9. **`tools:` is a capability list, `writes:` is the boundary** — `bash` runs anything (including `git checkout`) and `write` reaches any path, so a tool list can never make "this agent changes nothing" true. `writes:` per agent and `protected_files` in defaults are enforced in `adw_modules/permissions.py` after every agent call: unauthorized changes are rolled back and the phase dies. The session runtime under `data_dir` is always writable — a read-only agent is read-only with respect to the REPO, never mute.
10. **Every ADW ends in `run.finish()`** — phases passing is not the same as the run being accepted. A test phase that ran a red suite succeeded at its job. Pass `accepted=` so the exit code, the session status, and the banner are decided together and cannot disagree.
## v1 scope
Pi or OMP coding agent (`coding_agent: pi` or `omp`), chosen at install time. `claude_code` is schema-valid but stubbed until v2. The visualizer app (`just obs`) ships with the skill — a multi-repo trace UI over each repo's `sssf.db`.

View file

@ -0,0 +1,18 @@
{
"$schema": "./node_modules/oxlint/configuration_schema.json",
"plugins": ["typescript", "unicorn", "oxc"],
"categories": {
"correctness": "error",
"suspicious": "warn",
"perf": "warn"
},
"env": {
"browser": true,
"es2024": true
},
"rules": {
"no-console": "off",
"typescript/no-explicit-any": "warn"
},
"ignorePatterns": ["dist/**", "node_modules/**"]
}

View file

@ -0,0 +1,263 @@
{
"lockfileVersion": 1,
"configVersion": 1,
"workspaces": {
"": {
"name": "sssf-visualizer",
"dependencies": {
"@fontsource/play": "^5.3.0",
"lucide-vue-next": "^1.0.0",
"vue": "^3.5.13",
},
"devDependencies": {
"@types/bun": "^1.1.14",
"@vitejs/plugin-vue": "^6",
"oxlint": "^1",
"typescript": "^5.7.2",
"vite": "^7",
"vue-tsc": "^3",
},
},
},
"packages": {
"@babel/helper-string-parser": ["@babel/helper-string-parser@7.29.7", "", {}, "sha512-Pb5ijPrZ89GDH8223L4UP8i6QApWxs04RbPQJTeWDV0/keR2E36MeKnyr6LYmUUvqRRI+Iv87SuF1W6ErINzYw=="],
"@babel/helper-validator-identifier": ["@babel/helper-validator-identifier@7.29.7", "", {}, "sha512-qehxGkRj55h/ff8EMaJ+cYhyaKlHIxqYDn682wQD7RNp9UujOQsHog2uS0r2vzr4pW+sXf90NeeayjcNaX3fFg=="],
"@babel/parser": ["@babel/parser@7.29.7", "", { "dependencies": { "@babel/types": "^7.29.7" }, "bin": "./bin/babel-parser.js" }, "sha512-hnORnjP/1P/zFEndoeX+n+t1RwWRJiJpM/jO7FW32Kn9r5+sJB2JWOdYo4L6k78j15eCwY3Gm/7364B1EMwtNg=="],
"@babel/types": ["@babel/types@7.29.7", "", { "dependencies": { "@babel/helper-string-parser": "^7.29.7", "@babel/helper-validator-identifier": "^7.29.7" } }, "sha512-4zBIxpPzowiZpusoFkyGVwakdRJUyuH5PxQ/PrqghfdFWWasvnCdPfQXHrenDai+gyLARulZjZowCOj6fjT4pA=="],
"@esbuild/aix-ppc64": ["@esbuild/aix-ppc64@0.28.1", "", { "os": "aix", "cpu": "ppc64" }, "sha512-Svl7tq8k/08+p6CXPpRjQ1fKX+1odH/BQbb48fV6fj3CWHhsoIOoY87w1oHXm0qEpkIK3ZfVgp0hed3XBXzXMQ=="],
"@esbuild/android-arm": ["@esbuild/android-arm@0.28.1", "", { "os": "android", "cpu": "arm" }, "sha512-0k2F129Xdio1TdJfzJ8sy1Q47vUD2NnwdhiAf7drUN1EBTfPf4hsFCtmMgu/6m8JSzsBrlmVjudMBQqOfG8usQ=="],
"@esbuild/android-arm64": ["@esbuild/android-arm64@0.28.1", "", { "os": "android", "cpu": "arm64" }, "sha512-34EGEbCIAgosYz6goLcopX6Mo7NyGv9tfwEM2/7Ce2VcVRk568iSvniGWcUXIy7wEDR1wzolcxcriFVrWYcwBg=="],
"@esbuild/android-x64": ["@esbuild/android-x64@0.28.1", "", { "os": "android", "cpu": "x64" }, "sha512-dbwY7ltSMDWsRatcRpCnES4F+im88OCUgGZjy52shC7GqHRE/cYlxNbB4Z4UpJswpcc4Qxd2oE/ufM0p61IKng=="],
"@esbuild/darwin-arm64": ["@esbuild/darwin-arm64@0.28.1", "", { "os": "darwin", "cpu": "arm64" }, "sha512-TZbWkQY7kvTAXbXUT7uVACR5cMHsDiSz9z7ZKAX/RTq/WJEk3QyRr0wZpNhBDX+/0CtdqUIJlOiodQcta6tY3Q=="],
"@esbuild/darwin-x64": ["@esbuild/darwin-x64@0.28.1", "", { "os": "darwin", "cpu": "x64" }, "sha512-zfdzgK9ACBNZLI/CyHTOx81SyNbM6YXn7rxSgX97VjyiPl9W1i4Ka4fgKECEoFCKGpvBj5qArWIGgQjOwkgskQ=="],
"@esbuild/freebsd-arm64": ["@esbuild/freebsd-arm64@0.28.1", "", { "os": "freebsd", "cpu": "arm64" }, "sha512-wG2EA8ENdEI0qhkSZMjfqrdY+ziCYCPMmtZjjIwOmXFjmyzEHn+UUxk5of+SYsjtfs3VpnlC7QLzSI5hY/rOAw=="],
"@esbuild/freebsd-x64": ["@esbuild/freebsd-x64@0.28.1", "", { "os": "freebsd", "cpu": "x64" }, "sha512-i7dZ9vQgnvSCzi/rYCXNgtF/U+eKZNJBzu3eTQbRgHnM7tNSizLOkRFAl3qzVc/Op/u5YkHHa4pf/3DOYHthLQ=="],
"@esbuild/linux-arm": ["@esbuild/linux-arm@0.28.1", "", { "os": "linux", "cpu": "arm" }, "sha512-qVXBOHQS+d5Y722GwJzJUtOLlX7km3CraOaGormF1pDtPd2C/l1SHRPgjLunLGe51Sh5YYWKMFDyV4SxgMQYTQ=="],
"@esbuild/linux-arm64": ["@esbuild/linux-arm64@0.28.1", "", { "os": "linux", "cpu": "arm64" }, "sha512-yHs+0uc8+nvEAfAfxrWQKK5peSNzBc4PegcMO0EJ2hT71uA7vB8Ihg2e77R2P7SG5uYjPbHlLLmve4LLLRCf0g=="],
"@esbuild/linux-ia32": ["@esbuild/linux-ia32@0.28.1", "", { "os": "linux", "cpu": "ia32" }, "sha512-d1z4ZuP0ajrfz/FhGT4vv278rX8KnPPJx8i5+AtK7TYbx9Le9F1hyzurZpkEyjkGa9dUGhQow4C1NmeGvqxN2w=="],
"@esbuild/linux-loong64": ["@esbuild/linux-loong64@0.28.1", "", { "os": "linux", "cpu": "none" }, "sha512-M5sRjUVZrkm1OAPR3dlOYzNmN+loZKGVi1VUQGrwuqLcbR6qeAz+famMhjASeH3YVKvZz+zT1jlh/keC3Rj/lg=="],
"@esbuild/linux-mips64el": ["@esbuild/linux-mips64el@0.28.1", "", { "os": "linux", "cpu": "none" }, "sha512-mRObBZeHh2OxcBFPWE/FjylkRgZdYuiTR3vaTozquCGOH14iP9oN4x4Ge81CoIDYQrXmIxpFumJBu5MtZpnQJQ=="],
"@esbuild/linux-ppc64": ["@esbuild/linux-ppc64@0.28.1", "", { "os": "linux", "cpu": "ppc64" }, "sha512-slScBsMAb3GFDcdrCgLwZtPYRoH2H/youv10QiZyRjmsP48fznoveWytSgCI/R0ZcUgpc0ZhIUEx6LHts8yrfQ=="],
"@esbuild/linux-riscv64": ["@esbuild/linux-riscv64@0.28.1", "", { "os": "linux", "cpu": "none" }, "sha512-kw0owk1o0GFETUJyW0jc0G4Yzs0BHZn0JDZ8JRT088vjJYX777BAs1fDGxAC+q831qOs2DTC96mNsG2opdfyyQ=="],
"@esbuild/linux-s390x": ["@esbuild/linux-s390x@0.28.1", "", { "os": "linux", "cpu": "s390x" }, "sha512-/lAIjX8aYFRByhh6L5rYtPEDRqa9de/4V/juOXcta5frjvzXO4/sqEtyytse0g3zZFuWu5cDN0MkLz2qRDD2Ag=="],
"@esbuild/linux-x64": ["@esbuild/linux-x64@0.28.1", "", { "os": "linux", "cpu": "x64" }, "sha512-u/anNYF2mmVOEDwLtnQ1wOr3EZ9sTNGLWrsYGYwHWzGA3Si84IOkHXlbWTD1NB+9/1lcnweYKO54uhxZydNzfA=="],
"@esbuild/netbsd-arm64": ["@esbuild/netbsd-arm64@0.28.1", "", { "os": "none", "cpu": "arm64" }, "sha512-oks0DYbLwWMmaakTsCb+zL4E+aHRVLom9IJZOAthMQEPiQmydXHkziYEsGYRx0uNV/IjEKGAV941JzH02pflqw=="],
"@esbuild/netbsd-x64": ["@esbuild/netbsd-x64@0.28.1", "", { "os": "none", "cpu": "x64" }, "sha512-aeL6lAnN89Hz43Mlh1G8ARasbuoYvSITDEx0tHh5b7jJnHcssqgjy9Yx430GDpmCa6OyrKoS0aNRjKundRizGg=="],
"@esbuild/openbsd-arm64": ["@esbuild/openbsd-arm64@0.28.1", "", { "os": "openbsd", "cpu": "arm64" }, "sha512-MEFJe5C3R8pwXdZ5Y21oo6m7ePiS0d9pWucn99O/wvyJZChoIQKrQDxKrGeW8F5+T0okTHesAmDeiHDTIq0V/Q=="],
"@esbuild/openbsd-x64": ["@esbuild/openbsd-x64@0.28.1", "", { "os": "openbsd", "cpu": "x64" }, "sha512-i/ZLIOafE0Z8cI/XANJAixoJL/uRAoS2xOA3rb0xN+KK0K177cMAsQYkzHtBrtMXAKuAc7HGgcWiZ/sRC1Nxgw=="],
"@esbuild/openharmony-arm64": ["@esbuild/openharmony-arm64@0.28.1", "", { "os": "none", "cpu": "arm64" }, "sha512-ge+Z7EXFNt2BO1oAMsVpiQ8EwndV9i1xXerAeTIK7AtPs3bKFXQM7nlRxDSIUIMeueR1CNXxqztLzdNeReKBJg=="],
"@esbuild/sunos-x64": ["@esbuild/sunos-x64@0.28.1", "", { "os": "sunos", "cpu": "x64" }, "sha512-BEjgtECkL3vY+SaSQ6nzVfiALUeFxpawyp8Jmf5PtYhf1Ug40N1h/hxlhts+f1FvSvarEigdxS3BlSMI2PJLcQ=="],
"@esbuild/win32-arm64": ["@esbuild/win32-arm64@0.28.1", "", { "os": "win32", "cpu": "arm64" }, "sha512-lCv9eK/H6ZJWbE7bh2nw54CZ9M2nupBxJcTsdk/QQnWkdSjKGuxmmH8/GWrlT1eMmZfn4dGcCjRte397WqfQXA=="],
"@esbuild/win32-ia32": ["@esbuild/win32-ia32@0.28.1", "", { "os": "win32", "cpu": "ia32" }, "sha512-zvb/mB2bSCoJOpoCBgYKKpX6YM6mJBlBUVUtVj41DlZJVEB6/0CKlRYxP5wWl1C1ILiCoAU5wZZ4q1P3qeS6Eg=="],
"@esbuild/win32-x64": ["@esbuild/win32-x64@0.28.1", "", { "os": "win32", "cpu": "x64" }, "sha512-bm4Mowrv+GXMlpWX++EcXw/iLyd1o3+bJkC2DkWXYVvgZCqD/bSj9ctZeAMC3cIxgjRVR2Dufaiu4YPxr5gW1A=="],
"@fontsource/play": ["@fontsource/play@5.3.0", "", {}, "sha512-VpI6fd/A3jT1J3bOuyHn01CioAqwGWwNReTGycrFF2D57An2EJhDAuG030up3Ho7jSy1sSGW85gtB+l/v3gXOw=="],
"@jridgewell/sourcemap-codec": ["@jridgewell/sourcemap-codec@1.5.5", "", {}, "sha512-cYQ9310grqxueWbl+WuIUIaiUaDcj7WOq5fVhEljNVgRfOUhY9fy2zTvfoqWsnebh8Sl70VScFbICvJnLKB0Og=="],
"@oxlint/binding-android-arm-eabi": ["@oxlint/binding-android-arm-eabi@1.76.0", "", { "os": "android", "cpu": "arm" }, "sha512-ZHIE5Zt9AsPDcY4nOlofXt0YfneEeo+QrKMPcPzLf2Z6Q8VtV2W73d7SFJ920WUwyik783u/doKCs3KXdwG+7w=="],
"@oxlint/binding-android-arm64": ["@oxlint/binding-android-arm64@1.76.0", "", { "os": "android", "cpu": "arm64" }, "sha512-shm/ngQilHK6bs+ElJWa4oHfNj5vL1Gl/iVEJldTQjpr0/67oSgr0KUpbmcnLig5Fo0v/l6j2567A7TOL89ONA=="],
"@oxlint/binding-darwin-arm64": ["@oxlint/binding-darwin-arm64@1.76.0", "", { "os": "darwin", "cpu": "arm64" }, "sha512-rvJmrAPKSQ9aWJ6wIS6CK2tJjwzfW0ApQH9qokq6sfDvmHwoyIHxHFMq7z7i7GiV6fdE6s8qvBqWKPTu8RmT6Q=="],
"@oxlint/binding-darwin-x64": ["@oxlint/binding-darwin-x64@1.76.0", "", { "os": "darwin", "cpu": "x64" }, "sha512-U/zYdb7VYKGY6pA9Vd2rYl9O/HlCylcOlb5PGPvVLtg+oLGsk6H3XGKEMHKyqD3nmmtmlmwb/8SwU2vfSAtvMw=="],
"@oxlint/binding-freebsd-x64": ["@oxlint/binding-freebsd-x64@1.76.0", "", { "os": "freebsd", "cpu": "x64" }, "sha512-WvKG9CAriuo0XNiFzpXjDngUZcRGFNpaK2kLyMUsnJlShxkT96u+BpJQ3KqdQwGOrvI14L6V8bAwXwAYNNY6Jg=="],
"@oxlint/binding-linux-arm-gnueabihf": ["@oxlint/binding-linux-arm-gnueabihf@1.76.0", "", { "os": "linux", "cpu": "arm" }, "sha512-qJ5+RH99TqFRq3UCDxkW0zJJu9c+OAHFY72vGlxZLEpuO+MpKo3POgqb8sYipL9KYm8XY6ofb0HsOuvY6hQNqQ=="],
"@oxlint/binding-linux-arm-musleabihf": ["@oxlint/binding-linux-arm-musleabihf@1.76.0", "", { "os": "linux", "cpu": "arm" }, "sha512-PvPCVptkgVARsucgIqFQQcSmJ6xc6GtnVB5bRBekRahTc9eObMtjHfMjy5M+C2tHt5UCMttWM9RuSk/H9NqYeg=="],
"@oxlint/binding-linux-arm64-gnu": ["@oxlint/binding-linux-arm64-gnu@1.76.0", "", { "os": "linux", "cpu": "arm64" }, "sha512-3KeFDx8Bu4HPAXbuHZOr/oHvN+QT+JQhMw/NYPz7Z071xLSsG27Jfh9PIQVEY7hk1I+jr43ExqRIeJ6VKk2yLw=="],
"@oxlint/binding-linux-arm64-musl": ["@oxlint/binding-linux-arm64-musl@1.76.0", "", { "os": "linux", "cpu": "arm64" }, "sha512-oPFkkKTgl0K/EIg9fQ8oA3IGcI05/Mq1en04iFa41mmNPT+6KEiByVazTOZZJiHMBBrbsns1YJ2e1Scqwzesjw=="],
"@oxlint/binding-linux-ppc64-gnu": ["@oxlint/binding-linux-ppc64-gnu@1.76.0", "", { "os": "linux", "cpu": "ppc64" }, "sha512-gN7yZ0eqflA5Fhf1wvHxGUltIV3FsvmB1zhNMDEK9vSHhc7E6qg9CuPeBgPZab66Tjzq6w6kHAtNEvnTHf4cyw=="],
"@oxlint/binding-linux-riscv64-gnu": ["@oxlint/binding-linux-riscv64-gnu@1.76.0", "", { "os": "linux", "cpu": "none" }, "sha512-S/HqMbn22mQrjtErUxEoS/a55u8kIeXvreIxiJu5G7Le3UecEd6SQZxrDIpuhtgaFnsY/nVra3ytP+pRljDilA=="],
"@oxlint/binding-linux-riscv64-musl": ["@oxlint/binding-linux-riscv64-musl@1.76.0", "", { "os": "linux", "cpu": "none" }, "sha512-ZIga3097VJZolGZk6SrIAUokIGfRkxRlhiHDUznZptGBfwrhD7pNfD1rzEzsCwvk/1DX0A1bLz+liuNh5QKIVQ=="],
"@oxlint/binding-linux-s390x-gnu": ["@oxlint/binding-linux-s390x-gnu@1.76.0", "", { "os": "linux", "cpu": "s390x" }, "sha512-ZGiiA7pFzMJSyMWYZTVlPgbTsx+Vl8ihLGMIujPwaslUF7kIPPWAbVmAlTc+9lWDV+DCiB8Ikixu+lSHeOIIWQ=="],
"@oxlint/binding-linux-x64-gnu": ["@oxlint/binding-linux-x64-gnu@1.76.0", "", { "os": "linux", "cpu": "x64" }, "sha512-JLiy5WuvEBFTT6ErIFV35SLzi0R7Iri6MKU6dZbTxfIx8pndbbPs3Mj780nMipBFcPkti+okAPOJ9POKkHFEgg=="],
"@oxlint/binding-linux-x64-musl": ["@oxlint/binding-linux-x64-musl@1.76.0", "", { "os": "linux", "cpu": "x64" }, "sha512-z7lgKQtbo/I1NIe8G5NHLesxJDv0tRSUWTpXKb9Pm3E9nKFKfO4IOSDtFroKgXtOYb0jQbcdH+0wzTyMXVes+A=="],
"@oxlint/binding-openharmony-arm64": ["@oxlint/binding-openharmony-arm64@1.76.0", "", { "os": "none", "cpu": "arm64" }, "sha512-JOjKymIpb9QcYfEhZsN6h4V9Ivd474W38cNIBRv6bg2TbIvogbMTH0Mg6YWW9TiRDqfcX+/Hyfsbo5vcSE5guQ=="],
"@oxlint/binding-win32-arm64-msvc": ["@oxlint/binding-win32-arm64-msvc@1.76.0", "", { "os": "win32", "cpu": "arm64" }, "sha512-pqDWZiwcmByWUEm1NFUBNiT6aentCcaoMWJv0HbXEmuYermJ4sg8ppVrshubYP2MZ6SHccJJcpr6x469PuDFIw=="],
"@oxlint/binding-win32-ia32-msvc": ["@oxlint/binding-win32-ia32-msvc@1.76.0", "", { "os": "win32", "cpu": "ia32" }, "sha512-Ba0O659kgMv6pwO3z9PdO+K3aMxQRaw9HnG+e6AtOfgwcKFvYilciQYBoUBmxfQvOCKZe1SwjMkuB542NkuDMQ=="],
"@oxlint/binding-win32-x64-msvc": ["@oxlint/binding-win32-x64-msvc@1.76.0", "", { "os": "win32", "cpu": "x64" }, "sha512-5qcirPHO8nKfkoowEVWtpAoVTcYDy6g0UT0NGic450Qv8J2NrOqg4uQ8QppRP4MDTC7Xx47lbZnmadTH03CGGA=="],
"@rolldown/pluginutils": ["@rolldown/pluginutils@1.0.1", "", {}, "sha512-2j9bGt5Jh8hj+vPtgzPtl72j0yRxHAyumoo6TNfAjsLB04UtpSvPbPcDcBMxz7n+9CYB0c1GxQFxYRg2jimqGw=="],
"@rollup/rollup-android-arm-eabi": ["@rollup/rollup-android-arm-eabi@4.62.3", "", { "os": "android", "cpu": "arm" }, "sha512-c0wdcekXtQvvn5Tsrk/+op/gUArrbWaFduBnTLP2l1cKLSQs4diMWjJw3m6A0DdzT8dAAX95KpkJ3qynCePbmw=="],
"@rollup/rollup-android-arm64": ["@rollup/rollup-android-arm64@4.62.3", "", { "os": "android", "cpu": "arm64" }, "sha512-3YjElDdWN+qXAFbJ/CzPV+0wspLqh54k/I6GfdYtEJRqg7buSgc1yPM3B+93j1M4neobtkATHZTmxK2AMVGfnA=="],
"@rollup/rollup-darwin-arm64": ["@rollup/rollup-darwin-arm64@4.62.3", "", { "os": "darwin", "cpu": "arm64" }, "sha512-Pch2pFNOxxz1hTjypIdPyRTR6riiwRl84+VcN9djS680fw+Co1nAJINrdpqp7KV0NvyuU8ilZXZCjd7ykJl1GQ=="],
"@rollup/rollup-darwin-x64": ["@rollup/rollup-darwin-x64@4.62.3", "", { "os": "darwin", "cpu": "x64" }, "sha512-LEuncFUHFiF8t4yZVZvvZA1wk0pjAscRnsrn1EfTEmN4HXotBi2YtcnLRyaK6UbuczW7xZS5ES+81Rdz8Z0T6g=="],
"@rollup/rollup-freebsd-arm64": ["@rollup/rollup-freebsd-arm64@4.62.3", "", { "os": "freebsd", "cpu": "arm64" }, "sha512-zvBUvsQUpOWALdDsk6qbS8bXf2VxmPisuudNDrY7x0p0jBdsoZl8HsHczIOgkQiZldmcacMKtBzpoGVNeIe2bQ=="],
"@rollup/rollup-freebsd-x64": ["@rollup/rollup-freebsd-x64@4.62.3", "", { "os": "freebsd", "cpu": "x64" }, "sha512-C2KmNrcSem/AMg984H/dev+si0lieQGdXdR/lYGJnuumXnFb9Y7QdiI62obFdLlxRYLBv4P0eUVIDbD4c1vVvw=="],
"@rollup/rollup-linux-arm-gnueabihf": ["@rollup/rollup-linux-arm-gnueabihf@4.62.3", "", { "os": "linux", "cpu": "arm" }, "sha512-ggXnsTAEzNQx74XpunRsiZ9aBZDsI7XIa0hm2nzR9f4WzH5/f/d73ZSDaC5ejJ8YLY4NW+V3wr0tjOaeCq8hqA=="],
"@rollup/rollup-linux-arm-musleabihf": ["@rollup/rollup-linux-arm-musleabihf@4.62.3", "", { "os": "linux", "cpu": "arm" }, "sha512-2vng+FlzNUhKZxtej3IUqJgbZoQk2M/dwQM20+ULV0R/E/8tr9/P6uEf2iiGIk4HL0zMKh5Jry7mUHdUOvyGgA=="],
"@rollup/rollup-linux-arm64-gnu": ["@rollup/rollup-linux-arm64-gnu@4.62.3", "", { "os": "linux", "cpu": "arm64" }, "sha512-LLLFZKt4/Nraf9rxDkhiU8QVgLF4WmCkfr0L4fj0fPfIZFBib0DeiFk1hhaYKd03LFAFJcxHslhDFlNJLylf5Q=="],
"@rollup/rollup-linux-arm64-musl": ["@rollup/rollup-linux-arm64-musl@4.62.3", "", { "os": "linux", "cpu": "arm64" }, "sha512-WJkdQCvS9sWNOUBJZfQRKpZGFBztRzcowI+nndmflKgU4XY+3a420FgTOSKTsVqJbnzSxeT4vaJalpOaPo2YCQ=="],
"@rollup/rollup-linux-loong64-gnu": ["@rollup/rollup-linux-loong64-gnu@4.62.3", "", { "os": "linux", "cpu": "none" }, "sha512-PwHXCCS2n64/1Ot6rP1YEYA02MGYBcQlr8CSZZyrUG2O7NH6NklYmvr9v3Jy+5e/eDeNchc/ukmKJi9LuflMIQ=="],
"@rollup/rollup-linux-loong64-musl": ["@rollup/rollup-linux-loong64-musl@4.62.3", "", { "os": "linux", "cpu": "none" }, "sha512-vUjxINQu3RC8NZS3ykk1gN65gIz8pAopOq2HXuZhiIxHdx7TFvDG+jgrdSgInu1Eza4/Rfi2VzZgyIgEH4WOaw=="],
"@rollup/rollup-linux-ppc64-gnu": ["@rollup/rollup-linux-ppc64-gnu@4.62.3", "", { "os": "linux", "cpu": "ppc64" }, "sha512-wzko4aJ13+0G3kGnviCg5gnXFKd40izKsrf2uOw12US4XqprkDrmwOpeW14aSNa37V8bfPcz5Fkob6LZ3BAPmA=="],
"@rollup/rollup-linux-ppc64-musl": ["@rollup/rollup-linux-ppc64-musl@4.62.3", "", { "os": "linux", "cpu": "ppc64" }, "sha512-8120ue0JUMSwy11stlwnfdX3pPd+WZYGCDBwEHWtIHi6pOpZmsEF5QKB7a/UN+XFdqvobxz98kv8RTqikyCEBw=="],
"@rollup/rollup-linux-riscv64-gnu": ["@rollup/rollup-linux-riscv64-gnu@4.62.3", "", { "os": "linux", "cpu": "none" }, "sha512-XLFHnR3tXMjbOCh2vtVJHmxt+995uJsTERQyseFDRA0xxMxyTZPLa3OIUlyFaO4mF/Lu0FjmWHCuPXJT1n/IOg=="],
"@rollup/rollup-linux-riscv64-musl": ["@rollup/rollup-linux-riscv64-musl@4.62.3", "", { "os": "linux", "cpu": "none" }, "sha512-se6yXvNGMIl0f+RQzyh7XAmia8/9kplQx424wnG2w0C1oi6XgO6Y8otKhdXFHbHs88Ihavzmvh1NWjuovE76BQ=="],
"@rollup/rollup-linux-s390x-gnu": ["@rollup/rollup-linux-s390x-gnu@4.62.3", "", { "os": "linux", "cpu": "s390x" }, "sha512-gNoxRefktVIiGflpONuxWWXZAzIQG++z9qHO3xKwk4WdDMuQja3JHGfE1u0i3PfPDyvhypdk+WrgIJqLhGG7sg=="],
"@rollup/rollup-linux-x64-gnu": ["@rollup/rollup-linux-x64-gnu@4.62.3", "", { "os": "linux", "cpu": "x64" }, "sha512-V4KtWtQfAFMU7+9/A/VDps/VI8CHd3cYz0L8sgJzz8qK7eY7wI4ruFD82UYIYvW9Z4DtlTfhQcsl4XyPHW5uSg=="],
"@rollup/rollup-linux-x64-musl": ["@rollup/rollup-linux-x64-musl@4.62.3", "", { "os": "linux", "cpu": "x64" }, "sha512-LBx9LYXvj2CBkMkjLdNAWLwH0MLMin7do2VcVo9kVPibGLkY0BQQut2fv7NVqkXqZ/CrAu9LqDHVV1xHCMpCPw=="],
"@rollup/rollup-openbsd-x64": ["@rollup/rollup-openbsd-x64@4.62.3", "", { "os": "openbsd", "cpu": "x64" }, "sha512-ABVf3Q0RCu7NcyCCOZQI0pJ3GuSdfSl8EXcy88QtdceIMIoCUdfhsJChZ64L9zVM2aJHjde1Bhn5uqSRcX9ySA=="],
"@rollup/rollup-openharmony-arm64": ["@rollup/rollup-openharmony-arm64@4.62.3", "", { "os": "none", "cpu": "arm64" }, "sha512-+2Cy/ldweGBLlPIKsQLF8U5N44a0KDdbrk1rAjHOM9M2K+kGdIVjHLmmrZIcx+9Ny3ke/1JomCsDI1ocb11+sg=="],
"@rollup/rollup-win32-arm64-msvc": ["@rollup/rollup-win32-arm64-msvc@4.62.3", "", { "os": "win32", "cpu": "arm64" }, "sha512-dtZvzc8BedpSaFNy75x6uiWwAGTH+aZHDtdrqP6qk+WcLJrfti6sGje1ZJ9UxyzDLF23d/mV+PaMwuC0hL7UVA=="],
"@rollup/rollup-win32-ia32-msvc": ["@rollup/rollup-win32-ia32-msvc@4.62.3", "", { "os": "win32", "cpu": "ia32" }, "sha512-Rj8Ra4noo+aYy7sKBggCx0407mws34kAb1ySyWuq5DAtFBQdkSwnsjCgPrhPe9cvgBKZIukpE+CVHvORCS93kQ=="],
"@rollup/rollup-win32-x64-gnu": ["@rollup/rollup-win32-x64-gnu@4.62.3", "", { "os": "win32", "cpu": "x64" }, "sha512-vp7N084ew/odXn2gi/mzm9mUkQu9l6AiN6dt4IeUM2Uvm9o+cVmP+YkqbMOteLbiGgqBBlJZjIMYVCfOOIVbVQ=="],
"@rollup/rollup-win32-x64-msvc": ["@rollup/rollup-win32-x64-msvc@4.62.3", "", { "os": "win32", "cpu": "x64" }, "sha512-MOG/3gTOn4Fwf574RVOaY61I5o6P90legkFADiTyn1hyjNydT+cerU2rLUwPdZkKKyJ+iT+K9p7WXK4LM1Ka6g=="],
"@types/bun": ["@types/bun@1.3.14", "", { "dependencies": { "bun-types": "1.3.14" } }, "sha512-h1hFqFVcvAvD9j9K7ZW7vd82aSA+rTdznZa+5bwvCwqSB1jmmfLcbIWhOLx1/+boy/xmjgCs/OMUL8hRJSmnPw=="],
"@types/estree": ["@types/estree@1.0.9", "", {}, "sha512-GhdPgy1el4/ImP05X05Uw4cw2/M93BCUmnEvWZNStlCzEKME4Fkk+YpoA5OiHNQmoS7Cafb8Xa3Pya8m1Qrzeg=="],
"@types/node": ["@types/node@26.1.2", "", { "dependencies": { "undici-types": "~8.3.0" } }, "sha512-Vu4a5UFA9rIIFJ7rB/Vaafh9lrCQszopTCx6KjFboXTGQbPNasehVR5TEiithSDGyd1DEiUByggTZsg8jukeIg=="],
"@vitejs/plugin-vue": ["@vitejs/plugin-vue@6.0.8", "", { "dependencies": { "@rolldown/pluginutils": "^1.0.1" }, "peerDependencies": { "vite": "^5.0.0 || ^6.0.0 || ^7.0.0 || ^8.0.0", "vue": "^3.2.25" } }, "sha512-0ZjgOg7oO6farnNGup7yvoM/YXZV84OZxHAwtflItNa/6zzQyVb5LNxyea3FEKEX2XlagIKzrlH7wwxkKgtiew=="],
"@volar/language-core": ["@volar/language-core@2.4.28", "", { "dependencies": { "@volar/source-map": "2.4.28" } }, "sha512-w4qhIJ8ZSitgLAkVay6AbcnC7gP3glYM3fYwKV3srj8m494E3xtrCv6E+bWviiK/8hs6e6t1ij1s2Endql7vzQ=="],
"@volar/source-map": ["@volar/source-map@2.4.28", "", {}, "sha512-yX2BDBqJkRXfKw8my8VarTyjv48QwxdJtvRgUpNE5erCsgEUdI2DsLbpa+rOQVAJYshY99szEcRDmyHbF10ggQ=="],
"@volar/typescript": ["@volar/typescript@2.4.28", "", { "dependencies": { "@volar/language-core": "2.4.28", "path-browserify": "^1.0.1", "vscode-uri": "^3.0.8" } }, "sha512-Ja6yvWrbis2QtN4ClAKreeUZPVYMARDYZl9LMEv1iQ1QdepB6wn0jTRxA9MftYmYa4DQ4k/DaSZpFPUfxl8giw=="],
"@vue/compiler-core": ["@vue/compiler-core@3.5.40", "", { "dependencies": { "@babel/parser": "^7.29.7", "@vue/shared": "3.5.40", "entities": "^7.0.1", "estree-walker": "^2.0.2", "source-map-js": "^1.2.1" } }, "sha512-39E8IgOhTbVDnoJFMKc2DvYnypcZwUqgUhQkccva/0m6FUwtIKSGV7n1hpVmYcFaoRAwf9pBcwnKlCEsN63ZEQ=="],
"@vue/compiler-dom": ["@vue/compiler-dom@3.5.40", "", { "dependencies": { "@vue/compiler-core": "3.5.40", "@vue/shared": "3.5.40" } }, "sha512-pwkx4vqlqOspFstrcmzwkKLePVMD3PT65imRzLhanU2V1Fj4K13g6OXjanOyzw3aTAuRk84BOmY8f3rEHqPaVA=="],
"@vue/compiler-sfc": ["@vue/compiler-sfc@3.5.40", "", { "dependencies": { "@babel/parser": "^7.29.7", "@vue/compiler-core": "3.5.40", "@vue/compiler-dom": "3.5.40", "@vue/compiler-ssr": "3.5.40", "@vue/shared": "3.5.40", "estree-walker": "^2.0.2", "magic-string": "^0.30.21", "postcss": "^8.5.19", "source-map-js": "^1.2.1" } }, "sha512-gIf497P4kpuALcvs5n3AEg1Vdn0pSY4XbjASIfHNYF1/MP3T2Mf2STERTubysBxCRxzJGJYtF/O7vwJrxFB3Vw=="],
"@vue/compiler-ssr": ["@vue/compiler-ssr@3.5.40", "", { "dependencies": { "@vue/compiler-dom": "3.5.40", "@vue/shared": "3.5.40" } }, "sha512-rrE5xiXG663+vHCHa3J9p2z5OcBRjXmoqenprJxAFQxg5pSshzeBiCE6pu46axapRJ2Adk0YDA2BRZVjiHXnhg=="],
"@vue/language-core": ["@vue/language-core@3.3.8", "", { "dependencies": { "@volar/language-core": "2.4.28", "@vue/compiler-dom": "^3.5.0", "@vue/shared": "^3.5.0", "alien-signals": "^3.2.1", "muggle-string": "^0.4.1", "path-browserify": "^1.0.1", "picomatch": "^4.0.4" } }, "sha512-ieGT8jJdhhy0mGzStZhsg/qPw5bQZJg5yF+3+XU6saf4sM7yo9ZXy3h+nCwrm2+b4qS/SypkNdR2jAF3uei9tA=="],
"@vue/reactivity": ["@vue/reactivity@3.5.40", "", { "dependencies": { "@vue/shared": "3.5.40" } }, "sha512-B7ot9UlUZOi1zbq61/LvE88ZLTV8IlajTdiZTAEiDQgrnIMIZoPr9kGw0Zw46ObW62O9+H/Be3kMbfb7kYPQZA=="],
"@vue/runtime-core": ["@vue/runtime-core@3.5.40", "", { "dependencies": { "@vue/reactivity": "3.5.40", "@vue/shared": "3.5.40" } }, "sha512-KAZLweuZ6uUJPK1PMSQPgBU5gCjgrrfjUhSglmU9NhH+Zjepa8cnwSydPWDWHDwOgY4g3VcZ+PljbiHlURNCbw=="],
"@vue/runtime-dom": ["@vue/runtime-dom@3.5.40", "", { "dependencies": { "@vue/reactivity": "3.5.40", "@vue/runtime-core": "3.5.40", "@vue/shared": "3.5.40", "csstype": "^3.2.3" } }, "sha512-ZfrX8ssZQds900L9pr8AuK05ddnMsR4MPMZr8cPN9GoqoPWcXLhjvvbIA2SMv+7a97sJ1vv9pj/zxK0Cq/eEFQ=="],
"@vue/server-renderer": ["@vue/server-renderer@3.5.40", "", { "dependencies": { "@vue/compiler-ssr": "3.5.40", "@vue/runtime-dom": "3.5.40", "@vue/shared": "3.5.40" } }, "sha512-XNJym9WpevhTVt1HuwOrCRJ5Q+9z4BjTMrDtjTrvx74SmUll8spNTw6whWJa9mEkO4PKn5TihI/bm/8ds2QVJw=="],
"@vue/shared": ["@vue/shared@3.5.40", "", {}, "sha512-WxnBtruIqOoV3rA4jeKDWzrYI5h7Cp4+pjwDi8kWGHz+IslhiN+wguLVVhtv2l8VoU02rzDCVfDjgCl1lNpZVg=="],
"alien-signals": ["alien-signals@3.2.1", "", {}, "sha512-I8FjmltrfnDFoZedi5CG8DghVYNhzb/Ijluz7tCSJH0xpd0484Kowhbb1XDYOxfJpU1p5wnM2X54dA+IfGyD1g=="],
"bun-types": ["bun-types@1.3.14", "", { "dependencies": { "@types/node": "*" } }, "sha512-4N0ig0fEomHt5R0KCFWjovxow98rIoRwKolrYdCcknNwMekCXRnWEUvgu5soYV8QXtVsrUD8B95MBOZGPvr6KQ=="],
"csstype": ["csstype@3.2.3", "", {}, "sha512-z1HGKcYy2xA8AGQfwrn0PAy+PB7X/GSj3UVJW9qKyn43xWa+gl5nXmU4qqLMRzWVLFC8KusUX8T/0kCiOYpAIQ=="],
"entities": ["entities@7.0.1", "", {}, "sha512-TWrgLOFUQTH994YUyl1yT4uyavY5nNB5muff+RtWaqNVCAK408b5ZnnbNAUEWLTCpum9w6arT70i1XdQ4UeOPA=="],
"esbuild": ["esbuild@0.28.1", "", { "optionalDependencies": { "@esbuild/aix-ppc64": "0.28.1", "@esbuild/android-arm": "0.28.1", "@esbuild/android-arm64": "0.28.1", "@esbuild/android-x64": "0.28.1", "@esbuild/darwin-arm64": "0.28.1", "@esbuild/darwin-x64": "0.28.1", "@esbuild/freebsd-arm64": "0.28.1", "@esbuild/freebsd-x64": "0.28.1", "@esbuild/linux-arm": "0.28.1", "@esbuild/linux-arm64": "0.28.1", "@esbuild/linux-ia32": "0.28.1", "@esbuild/linux-loong64": "0.28.1", "@esbuild/linux-mips64el": "0.28.1", "@esbuild/linux-ppc64": "0.28.1", "@esbuild/linux-riscv64": "0.28.1", "@esbuild/linux-s390x": "0.28.1", "@esbuild/linux-x64": "0.28.1", "@esbuild/netbsd-arm64": "0.28.1", "@esbuild/netbsd-x64": "0.28.1", "@esbuild/openbsd-arm64": "0.28.1", "@esbuild/openbsd-x64": "0.28.1", "@esbuild/openharmony-arm64": "0.28.1", "@esbuild/sunos-x64": "0.28.1", "@esbuild/win32-arm64": "0.28.1", "@esbuild/win32-ia32": "0.28.1", "@esbuild/win32-x64": "0.28.1" }, "bin": { "esbuild": "bin/esbuild" } }, "sha512-HrJrvZv5ayxBzPfwphOoNzkzOIIlifzk0KJrGK2c8R4+LKpMtpYLQeUdjnwjWv/LZlkH2laZk+4w78pi99D4Vw=="],
"estree-walker": ["estree-walker@2.0.2", "", {}, "sha512-Rfkk/Mp/DL7JVje3u18FxFujQlTNR2q6QfMSMB7AvCBx91NGj/ba3kCfza0f6dVDbw7YlRf/nDrn7pQrCCyQ/w=="],
"fdir": ["fdir@6.5.0", "", { "peerDependencies": { "picomatch": "^3 || ^4" }, "optionalPeers": ["picomatch"] }, "sha512-tIbYtZbucOs0BRGqPJkshJUYdL+SDH7dVM8gjy+ERp3WAUjLEFJE+02kanyHtwjWOnwrKYBiwAmM0p4kLJAnXg=="],
"fsevents": ["fsevents@2.3.3", "", { "os": "darwin" }, "sha512-5xoDfX+fL7faATnagmWPpbFtwh/R77WmMMqqHGS65C3vvB0YHrgF+B1YmZ3441tMj5n63k0212XNoJwzlhffQw=="],
"lucide-vue-next": ["lucide-vue-next@1.0.0", "", { "peerDependencies": { "vue": ">=3.0.1" } }, "sha512-V6SPvx1IHTj/UY+FrIYWV5faISsPSb8BnWSFDxAtezWKvWc9ZZ40PDrdu1/Qb5vg4lHWr1hs1BAMGVGm6V1Xdg=="],
"magic-string": ["magic-string@0.30.21", "", { "dependencies": { "@jridgewell/sourcemap-codec": "^1.5.5" } }, "sha512-vd2F4YUyEXKGcLHoq+TEyCjxueSeHnFxyyjNp80yg0XV4vUhnDer/lvvlqM/arB5bXQN5K2/3oinyCRyx8T2CQ=="],
"muggle-string": ["muggle-string@0.4.1", "", {}, "sha512-VNTrAak/KhO2i8dqqnqnAHOa3cYBwXEZe9h+D5h/1ZqFSTEFHdM65lR7RoIqq3tBBYavsOXV84NoHXZ0AkPyqQ=="],
"nanoid": ["nanoid@3.3.16", "", { "bin": { "nanoid": "bin/nanoid.cjs" } }, "sha512-bzlKTyNJ7+LdGIIwy8ijFpIqEQIvafahV7eYykJ8Cvh42EdJeODoJ6gUJXpQJvej1BddH8OqTXZNE/KfbWAu8Q=="],
"oxlint": ["oxlint@1.76.0", "", { "optionalDependencies": { "@oxlint/binding-android-arm-eabi": "1.76.0", "@oxlint/binding-android-arm64": "1.76.0", "@oxlint/binding-darwin-arm64": "1.76.0", "@oxlint/binding-darwin-x64": "1.76.0", "@oxlint/binding-freebsd-x64": "1.76.0", "@oxlint/binding-linux-arm-gnueabihf": "1.76.0", "@oxlint/binding-linux-arm-musleabihf": "1.76.0", "@oxlint/binding-linux-arm64-gnu": "1.76.0", "@oxlint/binding-linux-arm64-musl": "1.76.0", "@oxlint/binding-linux-ppc64-gnu": "1.76.0", "@oxlint/binding-linux-riscv64-gnu": "1.76.0", "@oxlint/binding-linux-riscv64-musl": "1.76.0", "@oxlint/binding-linux-s390x-gnu": "1.76.0", "@oxlint/binding-linux-x64-gnu": "1.76.0", "@oxlint/binding-linux-x64-musl": "1.76.0", "@oxlint/binding-openharmony-arm64": "1.76.0", "@oxlint/binding-win32-arm64-msvc": "1.76.0", "@oxlint/binding-win32-ia32-msvc": "1.76.0", "@oxlint/binding-win32-x64-msvc": "1.76.0" }, "peerDependencies": { "oxlint-tsgolint": ">=7.0.2001", "vite-plus": "*" }, "optionalPeers": ["oxlint-tsgolint", "vite-plus"], "bin": { "oxlint": "bin/oxlint" } }, "sha512-6QoFioEU4fNdiUx/2Eo6TRd6NG7H7njnRCz8rhB66cZmMHDTqcm1Rjvl8Wry+ZTQMBAmyb4Mlf62Mk5X+eHSOw=="],
"path-browserify": ["path-browserify@1.0.1", "", {}, "sha512-b7uo2UCUOYZcnF/3ID0lulOJi/bafxa1xPe7ZPsammBSpjSWQkjNxlt635YGS2MiR9GjvuXCtz2emr3jbsz98g=="],
"picocolors": ["picocolors@1.1.1", "", {}, "sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA=="],
"picomatch": ["picomatch@4.0.5", "", {}, "sha512-RvwwcruNjI1ncT5xRakeyS9Lf8lcItv34KD+aif+VH9kduAyfYBipGh12274xtenIPZ119/R9BdTBa8gAwSh0A=="],
"postcss": ["postcss@8.5.24", "", { "dependencies": { "nanoid": "^3.3.16", "picocolors": "^1.1.1", "source-map-js": "^1.2.1" } }, "sha512-8RyVklq0owXUTa4xlpzu4l9AaVKIdQvAcOHZWaMh98HgySsUtxRVf/chRe3dsSLqb6i40BzGRzEUddRaI+9TSw=="],
"rollup": ["rollup@4.62.3", "", { "dependencies": { "@types/estree": "1.0.9" }, "optionalDependencies": { "@rollup/rollup-android-arm-eabi": "4.62.3", "@rollup/rollup-android-arm64": "4.62.3", "@rollup/rollup-darwin-arm64": "4.62.3", "@rollup/rollup-darwin-x64": "4.62.3", "@rollup/rollup-freebsd-arm64": "4.62.3", "@rollup/rollup-freebsd-x64": "4.62.3", "@rollup/rollup-linux-arm-gnueabihf": "4.62.3", "@rollup/rollup-linux-arm-musleabihf": "4.62.3", "@rollup/rollup-linux-arm64-gnu": "4.62.3", "@rollup/rollup-linux-arm64-musl": "4.62.3", "@rollup/rollup-linux-loong64-gnu": "4.62.3", "@rollup/rollup-linux-loong64-musl": "4.62.3", "@rollup/rollup-linux-ppc64-gnu": "4.62.3", "@rollup/rollup-linux-ppc64-musl": "4.62.3", "@rollup/rollup-linux-riscv64-gnu": "4.62.3", "@rollup/rollup-linux-riscv64-musl": "4.62.3", "@rollup/rollup-linux-s390x-gnu": "4.62.3", "@rollup/rollup-linux-x64-gnu": "4.62.3", "@rollup/rollup-linux-x64-musl": "4.62.3", "@rollup/rollup-openbsd-x64": "4.62.3", "@rollup/rollup-openharmony-arm64": "4.62.3", "@rollup/rollup-win32-arm64-msvc": "4.62.3", "@rollup/rollup-win32-ia32-msvc": "4.62.3", "@rollup/rollup-win32-x64-gnu": "4.62.3", "@rollup/rollup-win32-x64-msvc": "4.62.3", "fsevents": "~2.3.2" }, "bin": { "rollup": "dist/bin/rollup" } }, "sha512-Gu0c0iH9FzgX1L1t7ByIbbS3Vmdz+6KHm/EsqmmC71gUQ82yvZRkTK6XzrFObSka91WUVdynqp6nsfilzr5k6Q=="],
"source-map-js": ["source-map-js@1.2.1", "", {}, "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA=="],
"tinyglobby": ["tinyglobby@0.2.17", "", { "dependencies": { "fdir": "^6.5.0", "picomatch": "^4.0.4" } }, "sha512-wXR/dYpcqKmfWpEdZjiKJOwCNFndD0DMnrW/cYjVGttEkBfVgcLFHoNrlj47mjOVic9yyNu65alsgF4NQyTa2g=="],
"typescript": ["typescript@5.9.3", "", { "bin": { "tsc": "bin/tsc", "tsserver": "bin/tsserver" } }, "sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw=="],
"undici-types": ["undici-types@8.3.0", "", {}, "sha512-j375ScV60dom+YkPFIfTLcOiPxkN/buHz5GobjLhixFuANaNs3C9l4GmrWqejgXWJ7BbJcFYpTEUkS1Ge8bpZQ=="],
"vite": ["vite@7.3.6", "", { "dependencies": { "esbuild": "^0.27.0 || ^0.28.0", "fdir": "^6.5.0", "picomatch": "^4.0.3", "postcss": "^8.5.6", "rollup": "^4.43.0", "tinyglobby": "^0.2.15" }, "optionalDependencies": { "fsevents": "~2.3.3" }, "peerDependencies": { "@types/node": "^20.19.0 || >=22.12.0", "jiti": ">=1.21.0", "less": "^4.0.0", "lightningcss": "^1.21.0", "sass": "^1.70.0", "sass-embedded": "^1.70.0", "stylus": ">=0.54.8", "sugarss": "^5.0.0", "terser": "^5.16.0", "tsx": "^4.8.1", "yaml": "^2.4.2" }, "optionalPeers": ["@types/node", "jiti", "less", "lightningcss", "sass", "sass-embedded", "stylus", "sugarss", "terser", "tsx", "yaml"], "bin": { "vite": "bin/vite.js" } }, "sha512-4XP60spRGjSZFf1qYH+dJIkK2znL3zQfl9KkOV9MkkRR/3Dls0dxaBsQPTloEc5BLXWPL9vsOxopxyKoMmDueg=="],
"vscode-uri": ["vscode-uri@3.1.0", "", {}, "sha512-/BpdSx+yCQGnCvecbyXdxHDkuk55/G3xwnC0GqY4gmQ3j+A+g8kzzgB4Nk/SINjqn6+waqw3EgbVF2QKExkRxQ=="],
"vue": ["vue@3.5.40", "", { "dependencies": { "@vue/compiler-dom": "3.5.40", "@vue/compiler-sfc": "3.5.40", "@vue/runtime-dom": "3.5.40", "@vue/server-renderer": "3.5.40", "@vue/shared": "3.5.40" }, "peerDependencies": { "typescript": "*" }, "optionalPeers": ["typescript"] }, "sha512-+8PJ4SJXdn/cHGImF4CKdxlWHIN5Dkt7DoufRREM6h6uVCx2m7QxgcEQmmzyOK8A9mcafg7sFbJFYsdFVubTig=="],
"vue-tsc": ["vue-tsc@3.3.8", "", { "dependencies": { "@volar/typescript": "2.4.28", "@vue/language-core": "3.3.8" }, "peerDependencies": { "typescript": ">=5.0.0" }, "bin": { "vue-tsc": "bin/vue-tsc.js" } }, "sha512-xXmYlVQpcwJDWyGlqbHrGVOl1h3UOsASymRibrHc+iy9j/UNnOrOn4u+fntHz4D6Cs74RtapeqVV6CzJeg+UlA=="],
}
}

View file

@ -0,0 +1,13 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<link rel="icon" type="image/svg+xml" href="/logo.svg" />
<title>Super Simple Software Factory</title>
</head>
<body>
<div id="app"></div>
<script type="module" src="/src/main.ts"></script>
</body>
</html>

View file

@ -0,0 +1,29 @@
{
"name": "sssf-visualizer",
"version": "0.1.0",
"private": true,
"type": "module",
"description": "Read-only observability UI for the Super Simple Software Factory — polls sssf.db in a target repo.",
"scripts": {
"dev": "vite",
"server": "bun run server/index.ts",
"dev:all": "bun run server/index.ts & vite",
"build": "vue-tsc --noEmit && vite build",
"preview": "bun run server/index.ts",
"typecheck": "vue-tsc --noEmit",
"lint": "oxlint ."
},
"dependencies": {
"@fontsource/play": "^5.3.0",
"lucide-vue-next": "^1.0.0",
"vue": "^3.5.13"
},
"devDependencies": {
"@types/bun": "^1.1.14",
"@vitejs/plugin-vue": "^6",
"oxlint": "^1",
"typescript": "^5.7.2",
"vite": "^7",
"vue-tsc": "^3"
}
}

View file

@ -0,0 +1,6 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32">
<!-- SSSF mark: three swim lanes. -->
<rect x="4" y="6" width="17" height="5" rx="2.5" fill="#e8b64a"/>
<rect x="8" y="13.5" width="20" height="5" rx="2.5" fill="#c89bff"/>
<rect x="4" y="21" width="13" height="5" rx="2.5" fill="#5ad2dd"/>
</svg>

After

Width:  |  Height:  |  Size: 316 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 15 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 96 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 51 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

View file

@ -0,0 +1,418 @@
/**
* SQLite reader over a target repo's sssf.db.
*
* The read connection is opened readonly and every query on it is a SELECT —
* the writers are the tracers of running ADW processes, and WAL lets us read
* straight through their inserts.
*
* ONE exception, opened lazily on its own connection: `setArchived`. Archiving
* is review triage — "I have looked at this run" — which has to outlive a
* browser, so it lives on the session row rather than in localStorage. It is
* the only write this process can make, it touches exactly one column, and it
* never runs unless a human clicks the button.
*/
import { Database } from "bun:sqlite";
import { existsSync } from "node:fs";
import { dirname, isAbsolute, resolve } from "node:path";
import type {
AgentSession,
AgentStartPayload,
Envelope,
Event,
EventsPage,
GateResult,
Phase,
Session,
SessionDetail,
SessionSummary,
SessionUsage,
} from "../shared/types.ts";
const DEFAULT_DB_RELATIVE = "adws/adw_data/sssf.db";
const MAX_LIMIT = 1000;
const DEFAULT_LIMIT = 500;
/**
* Resolve the db path: --db arg wins, then SSSF_DB, then <cwd>/adws/adw_data/sssf.db.
* The db lives in the TARGET repo, so cwd is the repo the visualizer is pointed at.
*/
export function resolveDbPath(argv: string[] = Bun.argv): string {
const flagIndex = argv.indexOf("--db");
const inline = argv.find((a) => a.startsWith("--db="));
const raw =
(flagIndex !== -1 ? argv[flagIndex + 1] : undefined) ??
inline?.slice("--db=".length) ??
process.env.SSSF_DB ??
DEFAULT_DB_RELATIVE;
return isAbsolute(raw) ? raw : resolve(process.cwd(), raw);
}
export class SssfDb {
readonly path: string;
/**
* Where the ADW session dirs live: `{data_dir}/sessions/{adw_id}/{agent}/`.
* The db sits in the same data_dir (config's `observability.db` defaults to
* `adws/adw_data/sssf.db`), so deriving it as a sibling of the db file keeps
* working when the whole data_dir is relocated.
*/
readonly sessionsDir: string;
readonly journalMode: string;
private readonly db: Database;
/** Opened on first archive and kept; null until then. */
private writer: Database | null = null;
/** Cache for optionalColumn(), keyed "table.column". Only ever false → true. */
private readonly columnCache = new Map<string, boolean>();
constructor(path: string) {
if (!existsSync(path)) {
throw new Error(
`sssf.db not found at ${path}\n` +
`Point the visualizer at a target repo: --db <path> or SSSF_DB=<path>, ` +
`or run it from a repo root containing ${DEFAULT_DB_RELATIVE}`,
);
}
this.path = path;
this.sessionsDir = resolve(dirname(path), "sessions");
this.db = new Database(path, { readonly: true });
// WAL is set by the tracer when it creates the db; a readonly connection
// cannot change it, so we assert rather than set, and always take the
// busy_timeout so a concurrent writer never turns into a failed request.
this.db.exec("PRAGMA busy_timeout = 5000");
this.db.exec("PRAGMA synchronous = NORMAL");
const mode = this.db
.query<{ journal_mode: string }, []>("PRAGMA journal_mode")
.get();
this.journalMode = mode?.journal_mode ?? "unknown";
if (this.journalMode.toLowerCase() !== "wal") {
console.warn(
`[db] journal_mode is "${this.journalMode}", expected "wal" — ` +
`live reads during agent writes may block`,
);
}
}
/**
* A SELECT fragment for a column the tracer adds by migration.
*
* We open readonly and cannot run those ALTERs ourselves, so selecting one
* blindly would throw "no such column" on every request against a db an older
* tracer wrote. Instead we probe and substitute NULL, which reads downstream
* as "this db predates the column" — the same thing the UI shows for a row
* the migration didn't backfill.
*
* The probe re-runs while the column is missing, because the tracer's ALTER
* can land while we're serving: a startup-only check would keep returning
* NULL for the rest of the process even after the data arrived. Once seen,
* a column never goes away, so it latches.
*/
private hasColumn(table: string, column: string): boolean {
const key = `${table}.${column}`;
if (!this.columnCache.get(key)) {
const cols = this.db
.query<{ name: string }, []>(`PRAGMA table_info(${table})`)
.all();
this.columnCache.set(key, cols.some((c) => c.name === column));
}
return this.columnCache.get(key) ?? false;
}
private optionalColumn(table: string, column: string): string {
return this.hasColumn(table, column) ? column : `NULL AS ${column}`;
}
close(): void {
this.writer?.close();
this.db.close();
}
/**
* Archive or restore a session — the only write in this process.
*
* busy_timeout matters: a run may be mid-insert on the same WAL db, and a
* click should wait its turn rather than fail. Returns false when the id
* does not exist, so the route can 404 instead of silently succeeding.
*/
setArchived(adwId: string, archived: boolean): boolean {
if (!this.hasColumn("sessions", "archived")) {
throw new Error("this db predates the archived column — run any ADW once to migrate it");
}
if (!this.writer) {
this.writer = new Database(this.path);
this.writer.exec("PRAGMA busy_timeout=5000;");
}
this.writer
.query("UPDATE sessions SET archived = ? WHERE adw_id = ?")
.run(archived ? 1 : 0, adwId);
return this.session(adwId) !== null;
}
/** Sessions, most recent first, each with its phase statuses for the progress dots. */
sessions(limit = 200): SessionSummary[] {
const rows = this.db
.query<Session, [number]>(
`SELECT adw_id, ${this.optionalColumn("sessions", "adw_name")}, request,
status, engineer, started_at, ended_at,
total_tokens, total_cost,
${this.optionalColumn("sessions", "archived")}
FROM sessions
WHERE COALESCE(${this.hasColumn("sessions", "archived") ? "archived" : "0"}, 0) = 0
ORDER BY started_at DESC, rowid DESC
LIMIT ?`,
)
.all(clamp(limit, 1, MAX_LIMIT));
if (rows.length === 0) return [];
// Embed each session's phases so the L1 progress dots cost no extra request.
const ids = rows.map((row) => row.adw_id);
const placeholders = ids.map(() => "?").join(", ");
const phaseRows = this.db
.query<Phase, string[]>(
`SELECT phase_id, adw_id, seq, name, kind, owner, description, status,
attempt, retries, error, started_at, ended_at
FROM phases WHERE adw_id IN (${placeholders}) ORDER BY seq, rowid`,
)
.all(...ids);
const byAdw = new Map<string, Phase[]>();
for (const phase of phaseRows) {
const list = byAdw.get(phase.adw_id);
if (list) list.push(phase);
else byAdw.set(phase.adw_id, [phase]);
}
// Agents come along too: an L1 card draws a per-agent dot timeline, and its
// dots are colored per agent — without this it would be one request per card.
const agentsByAdw = this.agentsFor(ids);
const summaries: SessionSummary[] = [];
for (const session of rows) {
const phases = byAdw.get(session.adw_id) ?? [];
summaries.push(
Object.assign(session, {
phases,
phase_count: phases.length,
agents: agentsByAdw.get(session.adw_id) ?? [],
}),
);
}
return summaries;
}
session(adwId: string): Session | null {
return (
this.db
.query<Session, [string]>(
`SELECT adw_id, ${this.optionalColumn("sessions", "adw_name")}, request,
status, engineer, started_at, ended_at,
total_tokens, total_cost
FROM sessions WHERE adw_id = ?`,
)
.get(adwId) ?? null
);
}
phases(adwId: string): Phase[] {
return this.db
.query<Phase, [string]>(
`SELECT phase_id, adw_id, seq, name, kind, owner, description, status,
attempt, retries, error, started_at, ended_at
FROM phases WHERE adw_id = ? ORDER BY seq, rowid`,
)
.all(adwId);
}
agentSessions(adwId: string): AgentSession[] {
return this.agentsFor([adwId]).get(adwId) ?? [];
}
/**
* Agents per session, for a set of ids at once: the agent_sessions rows plus
* anything that has started but not finished.
*
* agents.py writes the agent_sessions row only after the envelope persists, so
* a running agent has no row there — precisely the case the live view exists
* for. Its model, color and session_id are already on the agent_start event,
* so a lane is labelled and colored from the moment the agent spawns.
*/
private agentsFor(adwIds: string[]): Map<string, AgentSession[]> {
const byAdw = new Map<string, AgentSession[]>();
if (adwIds.length === 0) return byAdw;
const placeholders = adwIds.map(() => "?").join(", ");
const append = (adwId: string, agent: AgentSession) => {
const list = byAdw.get(adwId);
if (list) list.push(agent);
else byAdw.set(adwId, [agent]);
};
const color = this.optionalColumn("agent_sessions", "color");
const ctxUsed = this.optionalColumn("agent_sessions", "context_tokens");
const ctxWindow = this.optionalColumn("agent_sessions", "context_window");
const completed = this.db
.query<AgentSession, string[]>(
`SELECT adw_id, agent, coding_agent, model, session_id, ${color},
${ctxUsed}, ${ctxWindow}, created_at, last_used_at
FROM agent_sessions WHERE adw_id IN (${placeholders})
ORDER BY created_at, agent`,
)
.all(...adwIds);
for (const row of completed) append(row.adw_id, row);
const started = this.db
.query<
{
adw_id: string;
agent: string | null;
payload_json: string | null;
started_at: string | null;
},
string[]
>(
`SELECT e.adw_id, p.owner AS agent, e.payload_json, e.started_at
FROM events e JOIN phases p ON p.phase_id = e.phase_id
WHERE e.adw_id IN (${placeholders}) AND e.type = 'agent_start'
ORDER BY e.rowid`,
)
.all(...adwIds);
for (const row of started) {
if (!row.agent) continue;
// A finished row is authoritative; only fill genuine gaps.
if (byAdw.get(row.adw_id)?.some((a) => a.agent === row.agent)) continue;
let payload: AgentStartPayload = {};
try {
payload = JSON.parse(row.payload_json ?? "{}") as AgentStartPayload;
} catch {
// A malformed payload just means no label — never a failed request.
}
append(row.adw_id, {
adw_id: row.adw_id,
agent: row.agent,
coding_agent: null,
model: payload.model ?? null,
session_id: payload.session_id ?? null,
color: payload.color ?? null,
// Occupancy is only known once the agent's turn closes.
context_tokens: null,
context_window: null,
created_at: row.started_at,
last_used_at: row.started_at,
});
}
return byAdw;
}
/** Session + phases + agents in one shot — L2 needs all three to draw lanes. */
sessionDetail(adwId: string): SessionDetail | null {
const session = this.session(adwId);
if (!session) return null;
return {
session,
usage: this.usage(adwId),
phases: this.phases(adwId),
agents: this.agentSessions(adwId),
};
}
/**
* Raw tokens read and written, beside the billed headline.
*
* Derived from the `agent_end` payloads rather than stored, so every run
* already in the db gets the split without a migration or a re-run.
*
* `total_tokens` is a SPEND number: every turn re-sends the whole
* conversation, so an 86k conversation over 49 turns bills millions. These
* two say what actually moved — material read for the first time, and
* material generated. The gap between them and the headline is cached
* re-reads, which is usually most of it.
*/
usage(adwId: string): SessionUsage {
const rows = this.db
.query<{ payload_json: string | null }, [string]>(
"SELECT payload_json FROM events WHERE adw_id = ? AND type = 'agent_end'",
)
.all(adwId);
let read = 0;
let written = 0;
for (const row of rows) {
if (!row.payload_json) continue;
try {
const u = (JSON.parse(row.payload_json) as { usage?: Record<string, number> }).usage;
if (!u) continue;
// RAW reads only: material entering the context for the first time,
// billed either as uncached input or as a cache write. Cache reads are
// the same tokens served again on later turns — counting them here
// would rebuild the very inflation this split exists to expose.
read += (u.input_tokens ?? 0) + (u.cache_write_tokens ?? 0);
written += u.output_tokens ?? 0;
} catch {
/* a payload written by an older tracer simply contributes nothing */
}
}
return { read, written };
}
/**
* The polling query. Rowid cursor, insertion order, bounded page — the same
* mechanism serves the live tail and lazy-paged history.
*/
events(adwId: string, after = 0, limit = DEFAULT_LIMIT): EventsPage {
const cappedLimit = clamp(limit, 1, MAX_LIMIT);
const events = this.db
.query<Event, [string, number, number]>(
`SELECT rowid, event_id, adw_id, phase_id, parent_id, type, name,
payload_json, tokens, started_at, ended_at
FROM events
WHERE adw_id = ? AND rowid > ?
ORDER BY rowid
LIMIT ?`,
)
.all(adwId, Math.max(0, after), cappedLimit);
return {
events,
cursor: events.length > 0 ? events[events.length - 1]!.rowid : Math.max(0, after),
has_more: events.length === cappedLimit,
};
}
envelopes(adwId: string): Envelope[] {
return this.db
.query<Envelope, [string]>(
`SELECT envelope_id, adw_id, phase_id, agent, output_type, payload_json,
valid, attempt, created_at
FROM envelopes WHERE adw_id = ? ORDER BY created_at, rowid`,
)
.all(adwId);
}
gates(adwId: string): GateResult[] {
const checks = this.optionalColumn("gate_results", "checks_json");
return this.db
.query<GateResult, [string]>(
`SELECT id, adw_id, phase_id, attempt, gate, passed, violations_json,
${checks}, created_at
FROM gate_results WHERE adw_id = ? ORDER BY id`,
)
.all(adwId);
}
sessionCount(): number {
const row = this.db
.query<{ n: number }, []>("SELECT COUNT(*) AS n FROM sessions")
.get();
return row?.n ?? 0;
}
}
function clamp(value: number, min: number, max: number): number {
if (!Number.isFinite(value)) return min;
return Math.min(max, Math.max(min, Math.trunc(value)));
}

View file

@ -0,0 +1,261 @@
/**
* SSSF visualizer server — JSON API over one or more repos' sssf.db, plus the
* built UI when ./dist exists. Reads are read-only; the single write is
* POST /api/sessions/:adw_id/archive, which sets one review flag on a row.
*
* There is no ingest endpoint and no websocket. The data path is
* agents → sqlite → web ui, and the UI gets there by polling.
*
* Single repo (legacy):
* bun run server/index.ts
* bun run server/index.ts --db /path/to/repo/adws/adw_data/sssf.db
* SSSF_DB=/path/to/sssf.db PORT=4600 bun run server/index.ts
*
* Multi repo:
* SSSF_REPOS=/path/to/repos.json PORT=4600 bun run server/index.ts
*
* In multi-repo mode every route is prefixed with the repo slug:
* /api/repos
* /api/:repo/sessions
* /api/:repo/sessions/:adw_id
* ...
* The legacy unprefixed routes (/api/sessions, ...) are kept and resolve to
* the first repo, so a single-repo deployment and the SPA's default view
* keep working unchanged.
*/
import { existsSync, statSync } from "node:fs";
import { join, resolve, sep } from "node:path";
import { buildRepos, type RepoRuntime } from "./repos.ts";
import type { AgentPrompts, ApiError, HealthResponse } from "../shared/types.ts";
import type { Server } from "bun";
const PORT = Number(process.env.PORT ?? 4600);
const DIST_DIR = resolve(import.meta.dir, "..", "dist");
function json(data: unknown, status = 200): Response {
return new Response(JSON.stringify(data), {
status,
headers: {
"content-type": "application/json; charset=utf-8",
"cache-control": "no-store",
},
});
}
function notFound(message: string): Response {
return json({ error: message } satisfies ApiError, 404);
}
/** Guard every handler so a malformed query can't take the server down mid-run. */
function safely(
handler: (req: Request) => Response | Promise<Response>,
): (req: Request) => Promise<Response> {
return async (req) => {
try {
return await handler(req);
} catch (error) {
console.error(`[sssf] ${req.method} ${new URL(req.url).pathname}:`, error);
return json({ error: (error as Error).message } satisfies ApiError, 500);
}
};
}
/**
* adw_ids and agent names are path segments on disk, so anything that isn't a
* plain identifier is rejected outright rather than sanitized into something
* that might still escape the sessions directory.
*/
const SAFE_SEGMENT = /^[A-Za-z0-9._-]+$/;
function isSafeSegment(value: string): boolean {
return SAFE_SEGMENT.test(value) && value !== "." && value !== "..";
}
function param(req: Request, key: string): string {
return decodeURIComponent(
(req as Request & { params: Record<string, string> }).params[key] ?? "",
);
}
function intQuery(req: Request, key: string, fallback: number): number {
const raw = new URL(req.url).searchParams.get(key);
if (raw === null || raw.trim() === "") return fallback;
const parsed = Number.parseInt(raw, 10);
return Number.isFinite(parsed) ? parsed : fallback;
}
/** Serve the built SPA if it has been built; otherwise point at the dev server. */
async function serveStatic(req: Request): Promise<Response> {
const { pathname } = new URL(req.url);
if (!existsSync(DIST_DIR)) {
return new Response(
`SSSF visualizer API is running on :${PORT}.\n\n` +
`No ./dist build found. Run "bun run dev" for the Vite dev server ` +
`(it proxies /api here), or "bun run build" to serve the UI from this process.\n`,
{ status: 200, headers: { "content-type": "text/plain; charset=utf-8" } },
);
}
// Reject traversal before touching the filesystem.
const candidate = resolve(join(DIST_DIR, pathname));
if (candidate === DIST_DIR || candidate.startsWith(DIST_DIR + "/")) {
if (existsSync(candidate) && statSync(candidate).isFile()) {
return new Response(Bun.file(candidate));
}
}
// SPA fallback: breadcrumb routes are client-side.
const indexHtml = join(DIST_DIR, "index.html");
if (existsSync(indexHtml)) {
return new Response(Bun.file(indexHtml), {
headers: { "content-type": "text/html; charset=utf-8" },
});
}
return notFound("not found");
}
/** Build the route table for one repo, mounted at /api/:repo and /api (default). */
function repoRoutes(repo: RepoRuntime, prefix: string) {
const db = repo.db;
const p = (path: string) => `${prefix}${path}`;
return {
[p("/health")]: safely(
() =>
json({
ok: true,
repo: repo.slug,
db: db.path,
journal_mode: db.journalMode,
sessions: db.sessionCount(),
} satisfies HealthResponse),
),
[p("/sessions")]: safely((req) => json(db.sessions(intQuery(req, "limit", 200)))),
[p("/sessions/:adw_id")]: safely((req) => {
const detail = db.sessionDetail(param(req, "adw_id"));
return detail ? json(detail) : notFound(`no session ${param(req, "adw_id")}`);
}),
// The one write. Archiving is review triage — it belongs to the reader, not
// to the run — so it never touches anything a tracer wrote.
[p("/sessions/:adw_id/archive")]: {
POST: safely(async (req) => {
const adwId = param(req, "adw_id");
if (!isSafeSegment(adwId)) {
return json({ error: "invalid adw_id" } satisfies ApiError, 400);
}
const body = (await req.json().catch(() => ({}))) as { archived?: unknown };
const archived = body.archived === undefined ? true : Boolean(body.archived);
return db.setArchived(adwId, archived)
? json({ adw_id: adwId, archived })
: notFound(`no session ${adwId}`);
}),
},
[p("/sessions/:adw_id/events")]: safely((req) =>
json(
db.events(
param(req, "adw_id"),
intQuery(req, "after", 0),
intQuery(req, "limit", 500),
),
),
),
[p("/sessions/:adw_id/envelopes")]: safely((req) =>
json(db.envelopes(param(req, "adw_id"))),
),
[p("/sessions/:adw_id/gates")]: safely((req) => json(db.gates(param(req, "adw_id")))),
// The exact prompts an agent was sent, read from the session dir. Files are
// the raw record; the db has no copy of them.
[p("/sessions/:adw_id/agents/:agent/prompts")]: safely(async (req) => {
const adwId = param(req, "adw_id");
const agent = param(req, "agent");
if (!isSafeSegment(adwId) || !isSafeSegment(agent)) {
return json({ error: "invalid adw_id or agent" } satisfies ApiError, 400);
}
if (!db.session(adwId)) return notFound(`no session ${adwId}`);
const dir = resolve(db.sessionsDir, adwId, agent, "prompts");
// Defense in depth: the segment check already forbids traversal.
if (dir !== db.sessionsDir && !dir.startsWith(db.sessionsDir + sep)) {
return json({ error: "invalid path" } satisfies ApiError, 400);
}
// A prompt file is absent whenever the agent never ran in this session —
// a normal state, so it reads as null rather than an error.
const read = async (name: string): Promise<string | null> => {
const file = Bun.file(join(dir, `${name}.md`));
return (await file.exists()) ? await file.text() : null;
};
return json({
system: await read("system"),
user: await read("user"),
} satisfies AgentPrompts);
}),
};
}
// The registry is mutable: POST /api/reload re-reads repos.json and swaps the
// route table in place (Bun's server.reload), so new repos appear without a
// process restart. `server` is assigned below; the reload handler only runs on
// a request, by which point it is set.
let repos: RepoRuntime[] = (await buildRepos()).repos;
let server: Server<undefined>;
/** Build the route table for the current registry. */
function buildRoutes(repos: RepoRuntime[]): Record<string, unknown> {
const routes: Record<string, unknown> = {
"/api/repos": safely(() =>
json(repos.map((r) => ({ slug: r.slug, name: r.name, db: r.db.path }))),
),
// Re-read repos.json and swap the registry + routes in place. The UI calls
// this after a repo is added/removed, so it shows up without a restart.
"/api/reload": {
POST: safely(async () => {
const next = (await buildRepos()).repos;
repos = next;
server.reload({ routes: buildRoutes(next) as never });
return json({
ok: true,
repos: next.map((r) => ({ slug: r.slug, name: r.name, db: r.db.path })),
});
}),
},
};
for (const repo of repos) {
Object.assign(routes, repoRoutes(repo, `/api/${repo.slug}`));
}
Object.assign(routes, repoRoutes(repos[0], "/api"));
return routes;
}
const routes = buildRoutes(repos);
server = Bun.serve({
port: PORT,
routes: routes as never,
fetch(req) {
const { pathname } = new URL(req.url);
if (pathname.startsWith("/api/")) return notFound(`no route ${pathname}`);
return serveStatic(req);
},
});
console.log(`[sssf] visualizer api http://localhost:${server.port}`);
for (const repo of repos) {
console.log(`[sssf] repo ${repo.slug} ${repo.db.path} [journal_mode=${repo.db.journalMode}]`);
}
console.log(
existsSync(DIST_DIR)
? `[sssf] serving ui from ${DIST_DIR}`
: `[sssf] no ./dist — use "bun run dev" for the Vite dev server on :4601`,
);
process.on("SIGINT", () => {
server.stop();
process.exit(0);
});

View file

@ -0,0 +1,146 @@
/**
* Repo registry for the multi-repo visualizer.
*
* The visualizer can serve several repos' trace DBs from one process. A repo
* is a slug → sssf.db mapping, declared in a JSON file:
*
* [
* { "slug": "tailsandstays", "name": "Tails & Stays",
* "db": "/home/ima/projects/tailsandstays/adws/adw_data/sssf.db" },
* { "slug": "mempalace", "name": "MemPalace",
* "db": "/home/ima/projects/mempalace/adws/adw_data/sssf.db" }
* ]
*
* Point the server at it with SSSF_REPOS=/path/to/repos.json. When SSSF_REPOS
* is unset the server auto-discovers repos under a projects root (default
* ~/projects, override with SSSF_PROJECTS_ROOT) by scanning for a trace db at
* <root>/<dir>/adws/adw_data/sssf.db. If nothing is discovered it falls back to
* the single-repo behaviour (--db / SSSF_DB / <cwd>/adws/adw_data/sssf.db)
* under a slug derived from the repo dir name, so the original single-repo
* deployment keeps working.
*/
import { existsSync, readdirSync } from "node:fs";
import { homedir } from "node:os";
import { basename, dirname, isAbsolute, join, resolve } from "node:path";
import { SssfDb, resolveDbPath } from "./db.ts";
export interface RepoEntry {
slug: string;
name: string;
db: string;
}
export interface RepoRuntime {
slug: string;
name: string;
db: SssfDb;
}
const SAFE_SLUG = /^[A-Za-z0-9._-]+$/;
function isSafeSlug(slug: string): boolean {
return SAFE_SLUG.test(slug) && slug !== "." && slug !== "..";
}
/** Derive a display slug from a db path: <repo>/adws/adw_data/sssf.db → repo dir name. */
function slugFromDb(dbPath: string): string {
// adw_data → adws → repo root
const repoDir = dirname(dirname(dirname(dbPath)));
const name = basename(repoDir);
return name && isSafeSlug(name) ? name : "default";
}
async function loadReposFile(path: string): Promise<RepoEntry[]> {
const raw = Bun.file(path);
if (!existsSync(path)) {
throw new Error(`SSSF_REPOS file not found: ${path}`);
}
const data = JSON.parse(await raw.text()) as unknown;
if (!Array.isArray(data)) {
throw new Error(`SSSF_REPOS file must be a JSON array of { slug, name, db }`);
}
const entries: RepoEntry[] = [];
const seen = new Set<string>();
for (const item of data) {
const entry = item as Partial<RepoEntry>;
if (typeof entry.slug !== "string" || !isSafeSlug(entry.slug)) {
throw new Error(`invalid repo slug: ${String(entry.slug)}`);
}
if (typeof entry.db !== "string") {
throw new Error(`repo "${entry.slug}" is missing a db path`);
}
if (seen.has(entry.slug)) {
throw new Error(`duplicate repo slug: ${entry.slug}`);
}
seen.add(entry.slug);
const db = isAbsolute(entry.db) ? entry.db : resolve(dirname(path), entry.db);
entries.push({
slug: entry.slug,
name: typeof entry.name === "string" && entry.name ? entry.name : entry.slug,
db,
});
}
if (entries.length === 0) {
throw new Error("SSSF_REPOS file declares no repos");
}
return entries;
}
/**
* Auto-discover repos under a projects root by scanning for a trace db at
* <root>/<dir>/adws/adw_data/sssf.db. Used when SSSF_REPOS is unset so new
* repos appear without editing a registry file. The root defaults to
* ~/projects and can be overridden with SSSF_PROJECTS_ROOT.
*/
function discoverRepos(root: string): RepoEntry[] {
if (!existsSync(root)) {
return [];
}
const entries: RepoEntry[] = [];
for (const name of readdirSync(root, { withFileTypes: true })) {
if (!name.isDirectory() || name.name.startsWith(".")) {
continue;
}
const db = join(root, name.name, "adws", "adw_data", "sssf.db");
if (existsSync(db) && isSafeSlug(name.name)) {
entries.push({ slug: name.name, name: name.name, db });
}
}
return entries;
}
/**
* Build the repo registry. The first repo is served under the legacy
* unprefixed /api routes (the single repo in single-repo mode, or the first
* repo in multi-repo mode).
*/
export async function buildRepos(): Promise<{ repos: RepoRuntime[] }> {
const reposFile = process.env.SSSF_REPOS;
const entries: RepoEntry[] = reposFile
? await loadReposFile(reposFile)
: discoverRepos(process.env.SSSF_PROJECTS_ROOT ?? join(homedir(), "projects"));
if (entries.length === 0) {
// No discovered repos — fall back to the single-repo behaviour so the
// original deployment (--db / SSSF_DB / <cwd>/adws/adw_data/sssf.db) works.
entries.push({ slug: slugFromDb(resolveDbPath()), name: "", db: resolveDbPath() });
}
const repos: RepoRuntime[] = [];
for (const entry of entries) {
let db: SssfDb;
try {
db = new SssfDb(entry.db);
} catch (error) {
console.error(`[sssf] repo "${entry.slug}": ${(error as Error).message}`);
continue;
}
repos.push({ slug: entry.slug, name: entry.name, db });
}
if (repos.length === 0) {
console.error("[sssf] no repos could be opened — exiting");
process.exit(1);
}
return { repos };
}

View file

@ -0,0 +1,314 @@
/**
* Types shared by the read-only server and the Vue client.
*
* Every interface mirrors a table in sssf.db one-for-one (see
* references/observability.md). Nothing here is derived state: phase durations,
* session progress and lane layout are computed in the UI, never stored.
*/
/** sessions.status — a run is running until it earns success. */
export type SessionStatus = "running" | "success" | "fail";
/** phases.status — queued only for manifest-declared phases not yet entered. */
export type PhaseStatus = "queued" | "running" | "success" | "fail";
/** phases.kind — decides which lane a block renders in. */
export type PhaseKind = "engineer" | "code" | "agent";
/** events.type — the ten types tracer.py emits. */
export type EventType =
| "phase_start"
| "phase_end"
| "agent_start"
| "agent_end"
| "tool_call"
| "handoff"
| "gate_pass"
| "gate_fail"
| "log"
| "error";
export interface Session {
adw_id: string;
/** ADW script(s) that ran this session, e.g. "adw_plan + adw_build_test". */
adw_name: string | null;
request: string | null;
status: SessionStatus | null;
engineer: string | null;
started_at: string | null;
ended_at: string | null;
total_tokens: number | null;
total_cost: number | null;
/** 1 once archived out of the review list. Review state, not run state. */
archived: number | null;
}
/**
* A session row with its phases embedded, so the L1 table draws the
* mini-progress dots without a second request per row.
*/
export interface SessionSummary extends Session {
/** Full phase rows, ordered by seq — one dot each. */
phases: Phase[];
phase_count: number;
/**
* The session's agents, same shape and merge rules as SessionDetail.agents —
* so an L1 card can color its per-agent dots without a request per card.
*/
agents: AgentSession[];
}
export interface Phase {
phase_id: string;
adw_id: string;
seq: number | null;
name: string | null;
kind: PhaseKind | null;
owner: string | null;
description: string | null;
status: PhaseStatus | null;
attempt: number | null;
retries: number | null;
error: string | null;
started_at: string | null;
ended_at: string | null;
}
export interface Event {
/** SQLite rowid — the polling cursor. Monotonic, insertion-ordered. */
rowid: number;
event_id: string;
adw_id: string;
phase_id: string | null;
/** Span nesting: an agent phase expands into its tool-call children. */
parent_id: string | null;
type: EventType | null;
name: string | null;
/** Raw JSON string as written by the tracer; parse at the point of display. */
payload_json: string | null;
tokens: number | null;
started_at: string | null;
ended_at: string | null;
}
export interface Envelope {
envelope_id: string;
adw_id: string;
phase_id: string | null;
agent: string | null;
/** Name of the data_types model the response was parsed against. */
output_type: string | null;
payload_json: string | null;
/** SQLite integer boolean. */
valid: number | null;
attempt: number | null;
created_at: string | null;
}
export interface GateResult {
id: number;
adw_id: string;
phase_id: string | null;
attempt: number | null;
gate: string | null;
/** SQLite integer boolean. */
passed: number | null;
/** JSON array of violation strings; "[]" on a pass. */
violations_json: string | null;
/**
* JSON array of GateCheck — the per-item evidence behind the verdict, so a
* green gate can say WHAT it verified rather than only that it passed.
* Null on rows written before the tracer recorded checks; those are not
* backfilled, so fall back to the verdict alone.
*/
checks_json: string | null;
created_at: string | null;
}
/** One item a gate inspected — the parsed element of `checks_json`. */
export interface GateCheck {
item: string;
ok: boolean;
note: string;
}
/** agent_sessions — the queryable mirror of agent_map.json. Supplies lane labels (`name · model`). */
export interface AgentSession {
adw_id: string;
agent: string;
coding_agent: string | null;
model: string | null;
session_id: string | null;
/**
* The agent's lane color from sssf.config.yaml, e.g. "#a78bfa". Null on dbs
* written by a tracer predating the column, and on agents with no configured
* color — fall back to the UI's own palette.
*/
color: string | null;
/**
* How full the agent's context window was after its last turn, and the
* model's ceiling. Null on dbs predating the columns and on an agent still
* running — the lane draws no bar rather than a misleading empty one.
*/
context_tokens: number | null;
context_window: number | null;
created_at: string | null;
last_used_at: string | null;
}
// ── payload_json shapes ──────────────────────────────────────────────────────
// events.payload_json is stored as a string. These are the parsed shapes for
// the two payloads the UI renders; every field is optional because the tracer
// writes what the coding agent reported, which varies by agent and by version.
/** Parsed `agent_start` payload — the live source of a lane's label and color. */
export interface AgentStartPayload {
model?: string;
thinking?: string;
session_id?: string;
color?: string;
coding_agent?: string;
purpose?: string;
/** Tool allowlist; null means all tools. Absent on pre-config-payload rows. */
tools?: string[] | null;
harness_engineering?: string[];
}
/**
* Tokens and dollars per component for one agent phase, summed across every
* send it made (a retried phase paid more than once). Mirrors pi's `usage`:
* `input_tokens` EXCLUDES cache reads, which bill at their own rate.
*/
export interface UsageBreakdown {
input_tokens: number;
output_tokens: number;
cache_read_tokens: number;
cache_write_tokens: number;
/**
* Thinking tokens — the reasoning SHARE of `output_tokens`, not a fifth
* component. Billed at the output rate; adding it to the others would
* double-count. Absent (undefined) on runs predating the field.
*/
reasoning_tokens?: number;
total_tokens: number;
input_cost: number;
output_cost: number;
cache_read_cost: number;
cache_write_cost: number;
total_cost: number;
}
/** Parsed `agent_end` payload — closes out a call with its cost and context use. */
export interface AgentEndPayload {
cost?: number;
/** Absent on runs predating the breakdown; `cost` alone survives there. */
usage?: UsageBreakdown;
/** Window occupancy after the final turn, and the model's ceiling. */
context_tokens?: number;
context_window?: number;
}
/**
* Parsed `tool_call` payload — one event per real tool call, emitted when the
* tool returns. `result_snippet` and `duration_ms` are absent when the coding
* agent never reported a result.
*/
export interface ToolCallPayload {
tool?: string;
tool_call_id?: string;
args?: Record<string, unknown>;
result_snippet?: string;
ok?: boolean;
duration_ms?: number;
agent?: string;
}
// ── API responses ────────────────────────────────────────────────────────────
/** GET /api/sessions */
export type SessionsResponse = SessionSummary[];
/** GET /api/sessions/:adw_id */
/**
* What actually moved through a session, summed across every agent.
*
* Deliberately NOT the billed total: `sessions.total_tokens` also counts every
* cached re-read, which is the same context charged again on each turn.
*/
export interface SessionUsage {
/** Raw prompt tokens read for the first time: new input + cache writes. */
read: number;
/** Tokens generated. Each produced exactly once, so this needs no adjusting. */
written: number;
}
export interface SessionDetail {
session: Session;
/** Derived from agent_end payloads, so historical runs have it too. */
usage: SessionUsage;
/** Ordered by seq. */
phases: Phase[];
/**
* One entry per agent that has run OR is running under this adw_id — lane
* labels come from here. Finished agents come from the agent_sessions table;
* an agent still in flight has no row there yet, so its entry is built from
* its agent_start event (coding_agent is null until it finishes).
*/
agents: AgentSession[];
}
/**
* GET /api/sessions/:adw_id/events?after=<rowid>&limit=500
*
* Poll with `after` = the cursor from the previous response. `cursor` is the
* highest rowid in this page (or the `after` you sent, when the page is empty),
* so it can be fed straight back in. `has_more` means the page hit the limit.
*/
export interface EventsPage {
events: Event[];
cursor: number;
has_more: boolean;
}
/**
* GET /api/sessions/:adw_id/agents/:agent/prompts
*
* The exact compiled prompts sent to an agent, read from
* `{data_dir}/sessions/{adw_id}/{agent}/prompts/`. These live only as files —
* the db has no copy. Either field is null when that file isn't on disk, which
* is the normal state for an agent that never ran in this session, so a 200
* with two nulls is a valid answer rather than an error.
*/
export interface AgentPrompts {
system: string | null;
user: string | null;
}
/** Alias matching the naming of the other endpoint payloads. */
export type PromptsResponse = AgentPrompts;
/** GET /api/sessions/:adw_id/envelopes */
export type EnvelopesResponse = Envelope[];
/** GET /api/sessions/:adw_id/gates */
export type GatesResponse = GateResult[];
/** GET /api/repos */
export interface RepoInfo {
slug: string;
name: string;
db: string;
}
/** GET /api/health */
export interface HealthResponse {
ok: boolean;
repo: string;
db: string;
journal_mode: string;
sessions: number;
}
export interface ApiError {
error: string;
}

View file

@ -0,0 +1,264 @@
<script setup lang="ts">
import { onMounted, ref } from 'vue'
import { useRoute, hrefFor, phaseCrumb, navigate } from './lib/router'
import type { RepoInfo } from './lib/types'
import { fetchRepos, reloadRepos } from './lib/api'
import SessionsList from './components/SessionsList.vue'
import SessionTrace from './components/SessionTrace.vue'
const route = useRoute()
const repos = ref<RepoInfo[]>([])
const repoName = ref<string | null>(null)
const reloading = ref(false)
const reloadError = ref<string | null>(null)
onMounted(async () => {
try {
repos.value = await fetchRepos()
// The URL must always carry the repo slug — parse() reads parts[0] as the
// repo, so a bare #/<adw_id> would be misread as repo=adwId. If we loaded
// with no repo segment, redirect to the first repo.
if (!route.value.repo && repos.value.length > 0) {
navigate(repos.value[0].slug)
return
}
const current = repos.value.find((r) => r.slug === route.value.repo)
repoName.value = current?.name ?? current?.slug ?? null
} catch {
// Repo list is a nicety; the default repo still works without it.
}
})
function onRepoChange(event: Event) {
const slug = (event.target as HTMLSelectElement).value
navigate(slug || null)
}
/** Re-read the server's repos.json so newly added repos show up without a restart. */
async function reload() {
reloading.value = true
reloadError.value = null
try {
repos.value = await reloadRepos()
// If the current repo was removed, fall back to the first one.
if (route.value.repo && !repos.value.some((r) => r.slug === route.value.repo)) {
navigate(repos.value[0]?.slug ?? null)
}
const current = repos.value.find((r) => r.slug === route.value.repo)
repoName.value = current?.name ?? current?.slug ?? null
} catch (error) {
reloadError.value = (error as Error).message
} finally {
reloading.value = false
}
}
</script>
<template>
<div class="app">
<header class="topbar">
<nav class="crumbs">
<!-- Inline copy of public/logo.svg (the favicon) so the mark renders
crisply with no fetch; keep the two in sync. -->
<svg class="logo" viewBox="0 0 32 32" aria-hidden="true">
<rect x="4" y="6" width="17" height="5" rx="2.5" fill="#e8b64a" />
<rect x="8" y="13.5" width="20" height="5" rx="2.5" fill="#c89bff" />
<rect x="4" y="21" width="13" height="5" rx="2.5" fill="#5ad2dd" />
</svg>
<span class="brand">Super Simple Software Factory</span>
<span class="sep">›</span>
<a :href="hrefFor(route.repo)" :class="{ current: !route.adwId }">sessions</a>
<template v-if="route.adwId">
<span class="sep">›</span>
<a :href="hrefFor(route.repo, route.adwId)" :class="{ current: !route.phaseId }">{{
route.adwId
}}</a>
</template>
<template v-if="route.adwId && route.phaseId">
<span class="sep">›</span>
<span class="current">{{ phaseCrumb ?? route.phaseId }}</span>
</template>
</nav>
<div class="topbar-right">
<select
v-if="repos.length > 1"
class="repo-switcher"
:value="route.repo ?? repos[0]?.slug ?? ''"
@change="onRepoChange"
:title="repoName ?? undefined"
>
<option v-for="r in repos" :key="r.slug" :value="r.slug">{{ r.name || r.slug }}</option>
</select>
<button
class="reload-btn"
type="button"
:disabled="reloading"
:title="reloadError ?? 'Re-read repos.json to pick up newly added repos'"
@click="reload"
>
{{ reloading ? '…' : '↻' }}
</button>
<span class="live-hint"><span class="live-dot" /> live</span>
</div>
<div v-if="reloadError" class="reload-error">{{ reloadError }}</div>
</header>
<main>
<SessionsList v-if="!route.adwId" />
<SessionTrace v-else :key="route.adwId" :adw-id="route.adwId" :phase-id="route.phaseId" />
</main>
</div>
</template>
<style scoped>
.topbar {
display: flex;
align-items: center;
justify-content: space-between;
padding: 15px 28px;
background: rgba(11, 15, 24, 0.72);
backdrop-filter: blur(14px);
-webkit-backdrop-filter: blur(14px);
position: sticky;
top: 0;
z-index: 10;
}
/* Gradient hairline instead of a hard border — the brand colors, whispered. */
.topbar::after {
content: '';
position: absolute;
left: 0;
right: 0;
bottom: 0;
height: 1px;
background: linear-gradient(
90deg,
rgba(200, 155, 255, 0.45),
rgba(90, 210, 221, 0.35) 40%,
rgba(90, 210, 221, 0.06)
);
}
.crumbs {
display: flex;
align-items: center;
gap: 10px;
font-size: 17px;
min-width: 0;
}
.logo {
width: 28px;
height: 28px;
flex: none;
filter: drop-shadow(0 0 8px rgba(200, 155, 255, 0.35));
}
.brand {
background: linear-gradient(90deg, var(--purple), var(--cyan));
-webkit-background-clip: text;
background-clip: text;
color: transparent;
font-weight: 700;
letter-spacing: 0.05em;
white-space: nowrap;
}
.sep {
color: var(--faint);
}
.crumbs a {
color: var(--dim);
}
.crumbs a:hover {
color: var(--text);
}
.crumbs .current {
color: var(--text);
}
.live-hint {
display: inline-flex;
align-items: center;
gap: 8px;
color: var(--dim);
font-size: 16px;
white-space: nowrap;
}
.topbar-right {
display: flex;
align-items: center;
gap: 16px;
}
.repo-switcher {
background: rgba(11, 15, 24, 0.72);
color: var(--text);
border: 1px solid rgba(200, 155, 255, 0.25);
border-radius: 8px;
padding: 5px 10px;
font-size: 14px;
cursor: pointer;
}
.repo-switcher:hover {
border-color: rgba(200, 155, 255, 0.5);
}
.repo-switcher option {
background: #0b0f18;
color: var(--text);
}
.reload-btn {
background: rgba(11, 15, 24, 0.72);
color: var(--dim);
border: 1px solid rgba(200, 155, 255, 0.25);
border-radius: 8px;
width: 30px;
height: 30px;
font-size: 16px;
line-height: 1;
cursor: pointer;
display: inline-flex;
align-items: center;
justify-content: center;
}
.reload-btn:hover:not(:disabled) {
color: var(--text);
border-color: rgba(200, 155, 255, 0.5);
}
.reload-btn:disabled {
opacity: 0.5;
cursor: default;
}
.reload-error {
position: absolute;
right: 28px;
top: 100%;
margin-top: 6px;
padding: 6px 10px;
font-size: 13px;
color: #ff8a8a;
background: rgba(11, 15, 24, 0.9);
border: 1px solid rgba(255, 138, 138, 0.4);
border-radius: 8px;
z-index: 11;
}
.live-dot {
width: 9px;
height: 9px;
border-radius: 50%;
background: var(--green);
box-shadow: 0 0 10px rgba(74, 222, 128, 0.7);
animation: pulse 1.6s ease-in-out infinite;
}
</style>

View file

@ -0,0 +1,76 @@
<script setup lang="ts">
import type { Component } from 'vue'
defineProps<{
title: string
/** Lucide icon component rendered before the title. */
icon?: Component
/** Shown after the title; omit for sections without a natural count. */
count?: number | null
open: boolean
}>()
defineEmits<{ toggle: [] }>()
</script>
<template>
<section class="dsec">
<button class="dsec-head" @click="$emit('toggle')">
<span class="chev">{{ open ? '▾' : '▸' }}</span>
<component :is="icon" v-if="icon" class="dsec-icon" :size="19" :stroke-width="2" />
<span class="dsec-title">{{ title }}</span>
<span v-if="count != null" class="dsec-count dim">({{ count }})</span>
</button>
<div v-if="open" class="dsec-body">
<slot />
</div>
</section>
</template>
<style scoped>
.dsec {
margin-bottom: 14px;
}
.dsec-head {
display: flex;
align-items: center;
gap: 9px;
width: 100%;
padding: 6px 8px;
background: none;
border: none;
border-bottom: 1px solid var(--border-soft);
border-radius: 6px 6px 0 0;
color: var(--dim);
font-size: 16px;
font-weight: 700;
letter-spacing: 0.05em;
text-transform: lowercase;
cursor: pointer;
text-align: left;
}
.dsec-icon {
flex: none;
color: var(--faint);
}
.dsec-head:hover {
background: var(--panel-2);
color: var(--text);
}
.chev {
color: var(--faint);
flex: none;
}
.dsec-count {
font-weight: 500;
}
.dsec-body {
padding-top: 10px;
}
</style>

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,55 @@
<script setup lang="ts">
import { computed } from 'vue'
import type { Phase } from '../lib/types'
const props = defineProps<{ phases: Phase[] }>()
const ordered = computed(() => props.phases.toSorted((a, b) => (a.seq ?? 0) - (b.seq ?? 0)))
const glyph: Record<string, string> = {
success: '●',
running: '◐',
queued: '○',
fail: '✗',
}
</script>
<template>
<span class="dots">
<span
v-for="p in ordered"
:key="p.phase_id"
class="d"
:class="p.status"
:title="`${p.name} — ${p.status}`"
>{{ glyph[p.status ?? ''] ?? '○' }}</span
>
<span v-if="!ordered.length" class="faint">—</span>
</span>
</template>
<style scoped>
.dots {
display: inline-flex;
gap: 5px;
font-size: 16px;
letter-spacing: 0;
}
.d.success {
color: var(--green);
}
.d.fail {
color: var(--red);
}
.d.running {
color: var(--blue);
animation: pulse 1.2s ease-in-out infinite;
}
.d.queued {
color: var(--faint);
}
</style>

View file

@ -0,0 +1,473 @@
<script setup lang="ts">
import { computed, onMounted, onUnmounted, shallowRef, watch } from 'vue'
import type { EventRow, SessionSummary } from '../lib/types'
import { archiveSession, fetchEvents } from '../lib/api'
import { axisTicks, fmtDate, fmtOffset, ts } from '../lib/format'
import { agentColor, dotColor, eventLabel } from '../lib/events'
import { hrefFor, useRoute } from '../lib/router'
import StatusChip from './StatusChip.vue'
import StatChip from './StatChip.vue'
import PhaseDots from './PhaseDots.vue'
const props = defineProps<{ session: SessionSummary; nowMs: number }>()
const emit = defineEmits<{ archived: [adwId: string] }>()
const route = useRoute()
// The card is an <a>; the button lives inside it, so the click must not
// navigate. Told the parent optimistically — the poll would take up to half a
// second to drop the card, and a triage click should feel instant.
async function archive(event: MouseEvent) {
event.preventDefault()
event.stopPropagation()
emit('archived', props.session.adw_id)
try {
await archiveSession(route.value.repo, props.session.adw_id)
} catch {
emit('archived', '') // signals the parent to re-sync from the server
}
}
// Each card tails its own event stream: one full fetch on mount, then the
// same rowid-cursor poll as the trace view — but only while the run is live.
const events = shallowRef<EventRow[]>([])
let cursor = 0
let inflight = false
let timer: ReturnType<typeof setInterval> | undefined
function stopPolling() {
clearInterval(timer)
timer = undefined
}
async function pull() {
if (inflight) return
inflight = true
try {
const fresh: EventRow[] = []
let page
do {
// Cursor pagination is inherently sequential: each request needs the previous cursor.
// oxlint-disable-next-line no-await-in-loop
page = await fetchEvents(route.value.repo, props.session.adw_id, cursor, 1000)
cursor = Math.max(cursor, page.cursor)
fresh.push(...page.events)
} while (page.has_more)
if (fresh.length) events.value = [...events.value, ...fresh]
if (props.session.status !== 'running') stopPolling()
} catch {
/* the list view surfaces api errors; a card just retries next poll */
} finally {
inflight = false
}
}
onMounted(() => {
void pull()
if (props.session.status === 'running') timer = setInterval(() => void pull(), 500)
})
onUnmounted(stopPolling)
watch(
() => props.session.status,
(status) => {
if (status === 'running' && !timer) timer = setInterval(() => void pull(), 500)
// On the transition out of running, one last pull drains the tail and stops the timer.
else if (status !== 'running') void pull()
},
)
const running = computed(() => props.session.status === 'running')
const range = computed(() => {
const s = props.session
let t0 = ts(s.started_at)
if (!Number.isFinite(t0)) {
t0 = Math.min(...events.value.map((e) => ts(e.started_at)).filter(Number.isFinite))
}
if (!Number.isFinite(t0)) t0 = props.nowMs
let t1 = running.value ? props.nowMs : ts(s.ended_at)
if (!Number.isFinite(t1)) {
t1 = Math.max(...events.value.map((e) => ts(e.started_at)).filter(Number.isFinite))
}
if (!Number.isFinite(t1)) t1 = t0 + 1000
return { t0, span: Math.max(t1 - t0, 1000) }
})
const ticks = computed(() => axisTicks(range.value.span, 5))
interface TimelineDot {
id: string
xPct: number
color: string
title: string
latest: boolean
}
interface TimelineRow {
owner: string
color: string
title: string
dots: TimelineDot[]
}
// Per-agent rows: events attribute to an agent through their phase's owner.
const rows = computed<TimelineRow[]>(() => {
const owners: string[] = []
const ownerByPhase = new Map<string, string>()
for (const p of props.session.phases ?? []) {
if (p.kind !== 'agent' || !p.owner) continue
ownerByPhase.set(p.phase_id, p.owner)
if (!owners.includes(p.owner)) owners.push(p.owner)
}
if (!owners.length) return []
const { t0, span } = range.value
const byOwner = new Map<string, TimelineDot[]>(owners.map((o) => [o, []]))
let latest: TimelineDot | null = null
let latestT = -Infinity
for (const e of events.value) {
const owner = e.phase_id ? ownerByPhase.get(e.phase_id) : undefined
const color = dotColor(e.type)
if (!owner || !color) continue
const t = ts(e.started_at)
if (!Number.isFinite(t)) continue
const dot: TimelineDot = {
id: e.event_id,
xPct: Math.min(Math.max(((t - t0) / span) * 100, 0), 100),
color,
title: `${e.type} ${eventLabel(e)} at ${fmtOffset(t - t0)}`,
latest: false,
}
byOwner.get(owner)?.push(dot)
if (t >= latestT) {
latestT = t
latest = dot
}
}
if (running.value && latest) latest.latest = true
// /api/sessions embeds agents so the labels can use config colors with no
// extra request; historical sessions return color null → fallback palette.
return owners.map((owner, i) => {
const info = (props.session.agents ?? []).find((a) => a.agent === owner)
return {
owner,
color: agentColor(info?.color, null, i),
title: info?.model ? `${owner} ${info.model}` : owner,
dots: byOwner.get(owner) ?? [],
}
})
})
const durationMs = computed(() => {
const s = props.session
const start = ts(s.started_at)
if (!Number.isFinite(start)) return NaN
const end = running.value ? props.nowMs : ts(s.ended_at)
return (Number.isFinite(end) ? end : props.nowMs) - start
})
// Cards are a fixed size, so the timeline region fits exactly MAX_VISIBLE_ROWS
// row slots. A roster that overflows spends one slot on the "+N more" line and
// shows MIN_VISIBLE_ROWS agents in the rest — never fewer than three, so a
// five-agent chain still reads as a chain rather than as a pair and a count.
const MAX_VISIBLE_ROWS = 4
const MIN_VISIBLE_ROWS = 3
const overflowing = computed(() => rows.value.length > MAX_VISIBLE_ROWS)
const visibleRows = computed(() =>
overflowing.value ? rows.value.slice(0, MIN_VISIBLE_ROWS) : rows.value,
)
const hiddenRowCount = computed(() =>
overflowing.value ? rows.value.length - MIN_VISIBLE_ROWS : 0,
)
</script>
<template>
<a class="card" :class="session.status" :href="hrefFor(route.repo, session.adw_id)">
<button
class="card-archive"
type="button"
title="Archive — remove this run from review"
aria-label="Archive run"
@click="archive"
>
×
</button>
<span class="card-id">{{ session.adw_id }}</span>
<span class="card-adw" :title="session.adw_name ?? ''">{{ session.adw_name ?? '—' }}</span>
<span class="card-req" :title="session.request ?? ''">{{ session.request }}</span>
<div v-if="rows.length" class="tl">
<div class="tl-axis">
<span class="tl-gutter" />
<span class="tl-scale">
<span
v-for="(t, i) in ticks"
:key="i"
class="tl-tick"
:class="{ edge: t.pct === 0 }"
:style="{ left: `${t.pct}%` }"
>{{ t.label }}</span
>
</span>
</div>
<div v-for="row in visibleRows" :key="row.owner" class="tl-row">
<span class="tl-agent" :style="{ color: row.color }" :title="row.title">{{
row.owner
}}</span>
<span class="tl-track">
<span
v-for="dot in row.dots"
:key="dot.id"
class="tl-dot"
:class="{ latest: dot.latest }"
:style="{ left: `${dot.xPct}%`, background: dot.color }"
:title="dot.title"
/>
</span>
</div>
<div v-if="hiddenRowCount" class="tl-more dim">+{{ hiddenRowCount }} more agents</div>
</div>
<div v-else class="tl tl-empty faint">no agent activity yet</div>
<div class="card-foot">
<span class="foot-status">
<StatusChip :status="session.status ?? 'fail'" />
<PhaseDots :phases="session.phases ?? []" />
</span>
<span class="dim">{{ fmtDate(session.started_at) }}</span>
</div>
<div class="card-stats">
<StatChip kind="cost" :value="session.total_cost" />
<StatChip kind="runtime" :value="durationMs" />
<StatChip kind="tokens" :value="session.total_tokens" />
</div>
</a>
</template>
<style scoped>
.card {
/* Uniform size: the grid fixes the width, this fixes the height — content
clamps and truncates rather than resizing the card. Grew by one 40px row
slot when the timeline went from three to four. */
height: 420px;
display: flex;
flex-direction: column;
gap: 10px;
padding: 20px 22px;
position: relative; /* anchors the archive button */
border: 1px solid var(--border-soft);
border-radius: 16px;
background: var(--surface);
color: var(--text);
cursor: pointer;
overflow: hidden;
transition:
border-color 0.18s ease,
box-shadow 0.18s ease,
transform 0.18s ease;
}
.card-archive {
/* Top-right of the card, out of the text flow so nothing reflows around it. */
position: absolute;
top: 10px;
right: 12px;
width: 26px;
height: 26px;
padding: 0;
border: 0;
border-radius: 8px;
background: transparent;
color: var(--dim);
font-family: inherit;
font-size: 20px;
line-height: 1;
cursor: pointer;
opacity: 0;
transition:
opacity 0.15s ease,
background 0.15s ease,
color 0.15s ease;
}
/* Hidden until the card is hovered — 50 cards should read as runs, not as a
wall of close buttons. Focus reveals it too, so keyboards are not excluded. */
.card:hover .card-archive,
.card-archive:focus-visible {
opacity: 1;
}
.card-archive:hover {
background: rgba(255, 111, 103, 0.16);
color: #ff6f67;
}
.card:hover {
border-color: rgba(148, 163, 255, 0.45);
box-shadow: 0 10px 34px rgba(148, 163, 255, 0.12);
transform: translateY(-2px);
}
.card.running {
border-color: rgba(108, 182, 255, 0.6);
box-shadow: 0 0 22px rgba(108, 182, 255, 0.16);
}
.card.fail {
border-color: rgba(255, 111, 103, 0.6);
}
/* Text rows must never absorb flex shrink — the fixed-height card squeezes
overflow into .tl (which clips), not into the text. */
.card-id {
flex: none;
font-family: var(--mono);
font-size: 18px;
font-weight: 700;
color: var(--purple);
}
.card-adw {
flex: none;
font-family: var(--mono);
font-size: 16px;
color: var(--cyan);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.card-req {
flex: none;
font-size: 16px;
color: var(--text);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.tl {
display: flex;
flex-direction: column;
margin-top: 4px;
/* Fixed region: axis (28 + 6) + four 40px row slots, roster size or not.
Four slots is what lets three agents show alongside a "+N more" line. */
height: 194px;
flex: none;
overflow: hidden;
}
.tl-more {
display: flex;
align-items: center;
height: 40px;
padding-left: 96px;
font-size: 16px;
}
.tl-axis {
display: flex;
align-items: flex-end;
height: 28px;
margin-bottom: 6px;
}
.tl-gutter,
.tl-agent {
flex: none;
/* Wide enough for full agent names (planner, builder, documenter). */
width: 96px;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
padding-right: 8px;
}
.tl-scale {
position: relative;
flex: 1;
height: 100%;
border-bottom: 1px solid var(--border);
}
.tl-tick {
position: absolute;
bottom: 4px;
transform: translateX(-50%);
font-family: var(--mono);
font-size: 16px;
color: var(--faint);
white-space: nowrap;
}
.tl-tick.edge {
transform: none;
}
.tl-row {
display: flex;
align-items: center;
height: 40px;
}
.tl-agent {
font-size: 16px;
color: var(--dim);
}
.tl-track {
position: relative;
flex: 1;
height: 100%;
border-bottom: 1px solid var(--border-soft);
}
.tl-dot {
position: absolute;
top: 50%;
width: 9px;
height: 9px;
border-radius: 50%;
transform: translate(-50%, -50%);
}
.tl-dot.latest {
width: 13px;
height: 13px;
box-shadow: 0 0 10px currentColor;
animation: pulse 1.4s ease-in-out infinite;
}
.tl-empty {
align-items: center;
justify-content: center;
font-size: 16px;
}
.card-foot {
display: flex;
align-items: center;
justify-content: space-between;
gap: 16px;
margin-top: auto;
font-size: 16px;
}
.foot-status {
display: inline-flex;
align-items: center;
gap: 14px;
}
.card-stats {
display: flex;
align-items: center;
gap: 12px;
}
</style>

View file

@ -0,0 +1,862 @@
<script setup lang="ts">
import { computed, onMounted, onUnmounted, ref, watchEffect } from 'vue'
import type {
AgentSession,
AgentStartPayload,
Envelope,
EventRow,
GateResult,
Phase,
PhaseKind,
Session,
SessionUsage,
} from '../lib/types'
import { Bot, SquareTerminal, UserRound } from 'lucide-vue-next'
import { fetchEnvelopes, fetchEvents, fetchGates, fetchSession } from '../lib/api'
import { axisTicks, fmtDate, payloadOk, ts } from '../lib/format'
import { modelIcon, modelName } from '../lib/models'
import { agentColor, hexAlpha, parseAgentStart } from '../lib/events'
import { navigate, phaseCrumb, useRoute } from '../lib/router'
import StatusChip from './StatusChip.vue'
import StatChip from './StatChip.vue'
import PhaseDetail from './PhaseDetail.vue'
const props = defineProps<{ adwId: string; phaseId: string | null }>()
const route = useRoute()
const session = ref<Session | null>(null)
const phases = ref<Phase[]>([])
const agents = ref<AgentSession[]>([])
const usage = ref<SessionUsage>({ read: 0, written: 0 })
const events = ref<EventRow[]>([])
const envelopes = ref<Envelope[]>([])
const gates = ref<GateResult[]>([])
const apiError = ref<string | null>(null)
const loaded = ref(false)
const nowMs = ref(Date.now())
let cursor = 0
let inflight = false
let timer: ReturnType<typeof setInterval> | undefined
const SIDE_TABLE_TYPES = new Set(['gate_pass', 'gate_fail', 'handoff', 'agent_end', 'phase_end', 'error'])
async function tick() {
if (inflight) return
inflight = true
try {
const detail = await fetchSession(route.value.repo, props.adwId)
session.value = detail.session
phases.value = detail.phases.toSorted((a, b) => (a.seq ?? 0) - (b.seq ?? 0))
agents.value = detail.agents
usage.value = detail.usage
const fresh: EventRow[] = []
let page
do {
// Cursor pagination is inherently sequential: each request needs the previous cursor.
// oxlint-disable-next-line no-await-in-loop
page = await fetchEvents(route.value.repo, props.adwId, cursor, 1000)
cursor = Math.max(cursor, page.cursor)
fresh.push(...page.events)
} while (page.has_more)
if (fresh.length) events.value = [...events.value, ...fresh]
// Envelopes and gates only gain rows around phase/agent boundaries — refetch
// on those events instead of every tick.
if (!loaded.value || fresh.some((e) => e.type !== null && SIDE_TABLE_TYPES.has(e.type))) {
const [env, g] = await Promise.all([
fetchEnvelopes(route.value.repo, props.adwId),
fetchGates(route.value.repo, props.adwId),
])
envelopes.value = env
gates.value = g
}
nowMs.value = Date.now()
apiError.value = null
loaded.value = true
} catch (err) {
apiError.value = err instanceof Error ? err.message : String(err)
} finally {
inflight = false
}
}
onMounted(() => {
void tick()
timer = setInterval(() => void tick(), 500)
})
onUnmounted(() => {
clearInterval(timer)
phaseCrumb.value = null
})
const selectedPhase = computed(
() => phases.value.find((p) => p.phase_id === props.phaseId) ?? null,
)
watchEffect(() => {
phaseCrumb.value = selectedPhase.value?.name ?? null
})
// ── Lanes ────────────────────────────────────────────────────────────────────
const ENGINEER_COLOR = '#e8b64a'
const CODE_COLOR = '#5ad2dd'
const KIND_ICONS = { engineer: UserRound, code: SquareTerminal, agent: Bot }
interface Lane {
id: string
label: string
/** Model driving this lane's agent — rendered with its provider icon. */
model: string | null
/** Context-window occupancy, or null while unknown (running / old db). */
context: LaneContext | null
metaLines: string[]
color: string
kind: PhaseKind
phases: Phase[]
}
interface LaneContext {
used: number
window: number
/** 0–100, uncapped by the floor applied to the bar's width. */
pct: number
}
/** Occupancy for an agent lane. Null unless BOTH numbers are real — a bar
* against an unknown ceiling would be decoration, not data. */
function laneContext(info: AgentSession | undefined): LaneContext | null {
const used = info?.context_tokens ?? 0
const window = info?.context_window ?? 0
if (!used || !window) return null
return { used, window, pct: Math.min(100, (used / window) * 100) }
}
/** Sub-1% occupancy is common and real; round it away and the bar reads empty. */
function contextLabel(ctx: LaneContext): string {
return ctx.pct < 1 ? `${ctx.pct.toFixed(1)}%` : `${Math.round(ctx.pct)}%`
}
/** Keep a non-zero fill visible — the exact numbers ride in the label and title. */
function contextFill(ctx: LaneContext): string {
return `${Math.max(ctx.pct, 2)}%`
}
const NUM = new Intl.NumberFormat('en-US')
// A live agent's model/thinking/color arrive on its agent_start event before
// any agent_sessions row exists; attribute each start to its phase's owner.
const ownerStart = computed<Record<string, AgentStartPayload>>(() => {
const ownerByPhase = new Map<string, string | null>(
phases.value.map((p) => [p.phase_id, p.owner]),
)
const meta: Record<string, AgentStartPayload> = {}
for (const e of events.value) {
if (e.type !== 'agent_start') continue
const owner = (e.phase_id ? ownerByPhase.get(e.phase_id) : null) ?? e.name
if (!owner || meta[owner]) continue
const payload = parseAgentStart(e)
if (payload) meta[owner] = payload
}
return meta
})
const lanes = computed<Lane[]>(() => {
const ph = phases.value
const agentOwners: string[] = []
for (const p of ph) {
if (p.kind === 'agent' && p.owner && !agentOwners.includes(p.owner)) agentOwners.push(p.owner)
}
const codePhases = ph.filter((p) => p.kind === 'code')
const out: Lane[] = [
{
id: 'engineer',
label: session.value?.engineer ?? 'engineer',
model: null,
context: null,
metaLines: ['engineer'],
color: ENGINEER_COLOR,
kind: 'engineer' as const,
phases: ph.filter((p) => p.kind === 'engineer'),
},
]
if (codePhases.length) {
out.push({
id: 'code',
label: 'code',
model: null,
context: null,
metaLines: ['workspace'],
color: CODE_COLOR,
kind: 'code' as const,
phases: codePhases,
})
}
for (const [i, owner] of agentOwners.entries()) {
const info = agents.value.find((a) => a.agent === owner)
const start = ownerStart.value[owner]
out.push({
id: `agent:${owner}`,
label: owner,
// The model is the lane's whole story; thinking level lives in the
// phase detail's agent config section.
model: info?.model ?? start?.model ?? null,
context: laneContext(info),
metaLines: [],
color: agentColor(info?.color, start?.color, i),
kind: 'agent' as const,
phases: ph.filter((p) => p.kind === 'agent' && p.owner === owner),
})
}
return out
})
// ── Timeline geometry ────────────────────────────────────────────────────────
const range = computed(() => {
let t0 = Infinity
let t1 = -Infinity
const s = session.value
const sStart = ts(s?.started_at)
const sEnd = ts(s?.ended_at)
if (Number.isFinite(sStart)) t0 = Math.min(t0, sStart)
if (Number.isFinite(sEnd)) t1 = Math.max(t1, sEnd)
for (const p of phases.value) {
const a = ts(p.started_at)
const b = ts(p.ended_at)
if (Number.isFinite(a)) {
t0 = Math.min(t0, a)
t1 = Math.max(t1, a)
}
if (Number.isFinite(b)) t1 = Math.max(t1, b)
}
if (s?.status === 'running') t1 = Math.max(t1, nowMs.value)
if (!Number.isFinite(t0)) {
t0 = nowMs.value
t1 = t0 + 1000
}
if (t1 - t0 < 1000) t1 = t0 + 1000
return { t0, t1, span: t1 - t0 }
})
// The engineer's request opens the run and owns the start of the timeline: it
// gets an exclusive leading zone, and every later phase maps into the rest —
// nothing can render on top of it.
const REQ_ZONE_PCT = 16
const requestPhase = computed(
() => phases.value.find((p) => p.kind === 'engineer' && p.started_at) ?? null,
)
const zonePct = computed(() => (requestPhase.value ? REQ_ZONE_PCT : 0))
/**
* Where the post-request timeline begins, in ms.
*
* The earliest non-engineer phase start, not the request phase's end: a later
* ADW joining the session pushes the request row's ended_at forward, which
* would otherwise throw every already-finished phase behind the origin.
*/
const originMs = computed(() => {
const { t0 } = range.value
const req = requestPhase.value
if (!req) return t0
let earliest = Infinity
for (const p of phases.value) {
if (p.kind === 'engineer') continue
const s = ts(p.started_at)
if (Number.isFinite(s)) earliest = Math.min(earliest, s)
}
if (Number.isFinite(earliest)) return Math.max(earliest, t0)
const end = ts(req.ended_at ?? req.started_at)
return Number.isFinite(end) ? Math.max(end, t0) : t0
})
const postSpan = computed(() => Math.max(range.value.t1 - originMs.value, 1000))
const ticks = computed(() => {
const zone = zonePct.value
return axisTicks(postSpan.value, 7).map((t) => ({
pct: zone + (t.pct * (100 - zone)) / 100,
label: t.label,
}))
})
/**
* Adjusted layout for every timed phase, in track-%.
*
* Phases are sequential by doctrine, and the render must say so: when a
* near-zero phase (a git commit) is widened to a readable floor, every later
* block shifts right by the same amount instead of being overlapped, and the
* whole layout is normalized back into the track. Blocks may squeeze a hair;
* they never stack.
*/
const MIN_BLOCK_PCT = 3.5
const blockLayout = computed<Record<string, { left: number; width: number }>>(() => {
const zone = zonePct.value
const avail = 100 - zone - 0.4 // hair of right margin
const t0 = originMs.value
const span = postSpan.value
const reqId = requestPhase.value?.phase_id
const timed = phases.value
.filter((p) => p.phase_id !== reqId && Number.isFinite(ts(p.started_at)))
.map((p) => {
const start = ts(p.started_at)
let end = ts(p.ended_at)
if (!Number.isFinite(end)) end = p.status === 'running' ? nowMs.value : start
return {
id: p.phase_id,
start,
left: ((start - t0) / span) * avail,
width: ((Math.max(end, start) - start) / span) * avail,
}
})
.toSorted((a, b) => a.start - b.start)
let shift = 0
let prevEdge = 0
const rows: { id: string; left: number; width: number }[] = []
for (const b of timed) {
let left = b.left + shift
if (left < prevEdge) {
shift += prevEdge - left
left = prevEdge
}
const width = Math.max(b.width, MIN_BLOCK_PCT)
shift += width - b.width
prevEdge = left + width
rows.push({ id: b.id, left, width })
}
const scale = avail / Math.max(prevEdge, avail)
const out: Record<string, { left: number; width: number }> = {}
for (const r of rows) out[r.id] = { left: zone + r.left * scale, width: r.width * scale }
return out
})
function blockGeom(p: Phase): { left: string; width: string } | null {
// The request block fills its reserved zone, nothing else ever enters it.
if (p.phase_id === requestPhase.value?.phase_id && zonePct.value > 0) {
return { left: '0.4%', width: `${zonePct.value - 0.8}%` }
}
const geom = blockLayout.value[p.phase_id]
if (!geom) return null
return { left: `${geom.left}%`, width: `${geom.width}%` }
}
function blockStyle(p: Phase, lane: Lane): Record<string, string> | undefined {
const geom = blockGeom(p)
if (!geom) return undefined
return {
left: geom.left,
width: geom.width,
background: `linear-gradient(180deg, ${hexAlpha(lane.color, 0.2)}, ${hexAlpha(lane.color, 0.05)})`,
borderColor: p.status === 'fail' ? 'rgba(255, 111, 103, 0.8)' : hexAlpha(lane.color, 0.55),
'--lane-glow': hexAlpha(lane.color, 0.28),
}
}
function blockDurationMs(p: Phase): number {
const start = ts(p.started_at)
if (!Number.isFinite(start)) return NaN
const end = p.status === 'running' ? nowMs.value : ts(p.ended_at)
if (!Number.isFinite(end)) return NaN
return end - start
}
const STATUS_GLYPH: Record<string, string> = {
success: '✓',
fail: '✗',
running: '●',
queued: '○',
}
// Tool-call tick marks inside a phase block, positioned within the block's own span.
interface ToolTick {
t: number
ok: boolean
}
const toolTicks = computed(() => {
const map: Record<string, ToolTick[]> = {}
for (const e of events.value) {
if (e.type !== 'tool_call' || !e.phase_id) continue
map[e.phase_id] ??= []
map[e.phase_id]?.push({ t: ts(e.started_at), ok: payloadOk(e.payload_json) })
}
return map
})
function ticksFor(p: Phase): { x: number; ok: boolean }[] {
const start = ts(p.started_at)
if (!Number.isFinite(start)) return []
let end = ts(p.ended_at)
if (!Number.isFinite(end)) end = p.status === 'running' ? nowMs.value : start
const width = Math.max(end - start, 1)
return (toolTicks.value[p.phase_id] ?? [])
.filter((mark) => Number.isFinite(mark.t))
.map((mark) => ({
x: Math.min(Math.max(((mark.t - start) / width) * 100, 1), 99),
ok: mark.ok,
}))
}
const queuedByLane = computed(() => {
const map: Record<string, Phase[]> = {}
for (const lane of lanes.value) {
map[lane.id] = lane.phases.filter((p) => !p.started_at)
}
return map
})
const sessionDurationMs = computed(() => {
const s = session.value
if (!s) return NaN
const start = ts(s.started_at)
if (!Number.isFinite(start)) return NaN
const end = s.status === 'running' ? nowMs.value : ts(s.ended_at)
return (Number.isFinite(end) ? end : nowMs.value) - start
})
function selectPhase(p: Phase) {
navigate(route.value.repo, props.adwId, p.phase_id === props.phaseId ? null : p.phase_id)
}
</script>
<template>
<div class="trace">
<div v-if="apiError" class="error-bar">api unreachable — retrying {{ apiError }}</div>
<div v-if="session" class="run-strip">
<span class="request" :title="session.request ?? ''">{{ session.request }}</span>
<StatusChip :status="session.status ?? 'fail'" />
<span class="dim">started {{ fmtDate(session.started_at) }}</span>
<span class="run-stats">
<StatChip kind="cost" :value="session.total_cost" />
<StatChip kind="runtime" :value="sessionDurationMs" />
<StatChip kind="tokens" :value="session.total_tokens" />
<StatChip kind="read" :value="usage.read" />
<StatChip kind="written" :value="usage.written" />
</span>
</div>
<div v-if="phases.length" class="waterfall">
<div class="row axis-row">
<div class="label" />
<div class="track">
<span v-if="zonePct" class="zone-head" :style="{ width: `${zonePct}%` }">request</span>
<span
v-for="(t, i) in ticks"
:key="i"
class="axis-label"
:style="{ left: `${t.pct}%` }"
>{{ t.label }}</span
>
</div>
</div>
<div v-for="lane in lanes" :key="lane.id" class="row lane" :class="`kind-${lane.kind}`">
<div class="label">
<span class="lane-name" :style="{ color: lane.color }">
<component :is="KIND_ICONS[lane.kind]" class="lane-icon" :size="22" :stroke-width="2" />
{{ lane.label }}
</span>
<span v-if="lane.model" class="lane-meta lane-model" :title="lane.model">
<img v-if="modelIcon(lane.model)" class="model-icon" :src="modelIcon(lane.model)!" alt="" />
{{ modelName(lane.model) }}
</span>
<span
v-if="lane.context"
class="lane-ctx"
:title="`${NUM.format(lane.context.used)} / ${NUM.format(lane.context.window)} tokens used · ${NUM.format(lane.context.window - lane.context.used)} remaining`"
>
<span class="ctx-head">
<span class="ctx-label">Context</span>
<span class="ctx-pct">{{ contextLabel(lane.context) }}</span>
</span>
<span class="ctx-bar">
<span
class="ctx-fill"
:style="{
width: contextFill(lane.context),
background: `linear-gradient(90deg, ${hexAlpha(lane.color, 0.55)}, ${lane.color})`,
boxShadow: `0 0 10px ${hexAlpha(lane.color, 0.45)}`,
}"
/>
</span>
</span>
<span v-for="(line, i) in lane.metaLines" :key="i" class="lane-meta">{{ line }}</span>
</div>
<div class="track">
<span v-if="zonePct" class="zone-divider" :style="{ left: `${zonePct}%` }" />
<span v-for="(t, i) in ticks" :key="i" class="gridline" :style="{ left: `${t.pct}%` }" />
<template v-for="p in lane.phases" :key="p.phase_id">
<button
v-if="blockGeom(p)"
class="block"
:class="[p.status, { selected: p.phase_id === phaseId }]"
:style="blockStyle(p, lane)"
:title="`${p.name} — ${p.status}${p.description ? `\n${p.description}` : ''}`"
@click="selectPhase(p)"
>
<span class="b-top">
<span class="b-status" :class="p.status">{{
STATUS_GLYPH[p.status ?? ''] ?? '○'
}}</span>
<span class="b-name">{{ p.name }}</span>
<StatChip
v-if="Number.isFinite(blockDurationMs(p))"
class="b-dur"
kind="runtime"
compact
:value="blockDurationMs(p)"
/>
</span>
<span class="b-desc">{{ p.description }}</span>
<span
v-for="(tick, i) in ticksFor(p)"
:key="i"
class="tool-tick"
:class="{ err: !tick.ok }"
:style="{ left: `${tick.x}%` }"
/>
</button>
</template>
<button
v-for="(p, i) in queuedByLane[lane.id]"
:key="p.phase_id"
class="block queued"
:class="{ selected: p.phase_id === phaseId }"
:style="{ right: `${10 + i * 5}px`, width: '170px' }"
:title="`${p.name} — queued`"
@click="selectPhase(p)"
>
<span class="b-top">
<span class="b-status queued">○</span>
<span class="b-name">{{ p.name }}</span>
</span>
<span class="b-desc">queued</span>
</button>
</div>
</div>
</div>
<div v-else-if="loaded" class="empty-state">no phases recorded for this session</div>
<div v-else-if="!apiError" class="empty-state">loading trace…</div>
<PhaseDetail
v-if="selectedPhase"
:phase="selectedPhase"
:events="events"
:envelopes="envelopes"
:gates="gates"
@close="navigate(route.repo, props.adwId)"
/>
</div>
</template>
<style scoped>
.trace {
padding: 0 0 40px;
}
.run-strip {
display: flex;
align-items: center;
gap: 18px;
padding: 14px 24px;
border-bottom: 1px solid var(--border-soft);
flex-wrap: wrap;
}
.run-strip .request {
font-size: 17px;
color: var(--text);
max-width: 52ch;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.run-stats {
display: inline-flex;
gap: 12px;
flex-wrap: wrap;
}
.waterfall {
margin: 20px 28px;
border: 1px solid var(--border-soft);
border-radius: 16px;
background: var(--surface);
overflow: hidden;
}
.row {
display: grid;
grid-template-columns: 280px 1fr;
}
.axis-row {
border-bottom: 1px solid var(--border);
background: var(--panel-2);
}
.axis-row .track {
height: 40px;
overflow: hidden;
}
.zone-head {
position: absolute;
top: 0;
bottom: 0;
left: 0;
display: inline-flex;
align-items: center;
justify-content: center;
font-size: 16px;
color: var(--amber);
border-right: 1px solid var(--border);
}
.axis-label {
position: absolute;
bottom: 7px;
transform: translateX(-50%);
font-family: var(--mono);
font-size: 16px;
color: var(--dim);
white-space: nowrap;
}
.label {
padding: 12px 16px;
display: flex;
flex-direction: column;
justify-content: center;
gap: 2px;
border-right: 1px solid var(--border);
overflow: hidden;
white-space: nowrap;
}
.lane-name {
display: inline-flex;
align-items: center;
gap: 8px;
font-size: 17px;
font-weight: 700;
overflow: hidden;
text-overflow: ellipsis;
}
.lane-icon {
flex: none;
opacity: 0.85;
}
.lane-meta {
font-family: var(--mono);
font-size: 16px;
color: var(--dim);
overflow: hidden;
text-overflow: ellipsis;
}
.lane-model {
display: inline-flex;
align-items: center;
gap: 7px;
}
.model-icon {
width: 17px;
height: 17px;
flex: none;
object-fit: contain;
}
/* Context occupancy — label row over a thin track, under the model. */
.lane-ctx {
display: flex;
flex-direction: column;
gap: 4px;
margin-top: 2px;
max-width: 190px;
}
.ctx-head {
display: flex;
align-items: baseline;
justify-content: space-between;
gap: 8px;
}
.ctx-label {
font-size: 14px;
letter-spacing: 0.06em;
text-transform: uppercase;
color: var(--faint);
}
.ctx-pct {
font-family: var(--mono);
font-size: 14px;
color: var(--dim);
}
.ctx-bar {
height: 6px;
border-radius: 999px;
background: rgba(6, 8, 15, 0.75);
border: 1px solid var(--border-soft);
overflow: hidden;
}
.ctx-fill {
display: block;
height: 100%;
border-radius: 999px;
transition: width 300ms ease;
}
.lane {
border-bottom: 1px solid var(--border-soft);
}
.lane:last-child {
border-bottom: none;
}
.track {
position: relative;
height: 118px;
overflow: hidden;
}
.zone-divider {
position: absolute;
top: 0;
bottom: 0;
border-left: 1px solid var(--border);
}
.gridline {
position: absolute;
top: 0;
bottom: 0;
border-left: 1px dashed rgba(174, 191, 212, 0.14);
}
.block {
position: absolute;
top: 13px;
height: 92px;
display: flex;
flex-direction: column;
justify-content: flex-start;
gap: 4px;
padding: 10px 12px 16px;
border-radius: 10px;
border: 1px solid;
font-size: 16px;
color: var(--text);
cursor: pointer;
overflow: hidden;
white-space: nowrap;
text-align: left;
transition: box-shadow 0.16s ease;
}
.block:hover {
box-shadow: 0 0 18px var(--lane-glow, rgba(108, 182, 255, 0.2));
}
.b-top {
display: flex;
align-items: baseline;
gap: 10px;
min-width: 0;
}
.b-status {
flex: none;
font-size: 16px;
}
.b-status.success {
color: var(--green);
}
.b-status.fail {
color: var(--red);
}
.b-status.running {
color: var(--blue);
animation: pulse 1.2s ease-in-out infinite;
}
.b-status.queued {
color: var(--faint);
}
.block .b-name {
font-size: 17px;
font-weight: 700;
overflow: hidden;
text-overflow: ellipsis;
}
.block .b-dur {
margin-left: auto;
flex: none;
}
.block .b-desc {
color: var(--dim);
font-size: 16px;
overflow: hidden;
text-overflow: ellipsis;
min-width: 0;
}
.block.running {
animation: pulse 1.6s ease-in-out infinite;
}
.block.queued {
background: transparent;
border-style: dashed;
border-color: var(--faint);
color: var(--dim);
}
.block.selected {
outline: 2px solid var(--blue);
outline-offset: 2px;
box-shadow: 0 0 22px var(--lane-glow, rgba(108, 182, 255, 0.25));
}
.tool-tick {
position: absolute;
bottom: 4px;
width: 3px;
height: 9px;
background: currentColor;
opacity: 0.55;
border-radius: 1px;
}
.tool-tick.err {
background: var(--red);
opacity: 1;
}
</style>

View file

@ -0,0 +1,109 @@
<script setup lang="ts">
import { computed, onMounted, onUnmounted, ref, shallowRef, watch } from 'vue'
import type { SessionSummary } from '../lib/types'
import { fetchSessions } from '../lib/api'
import { ts } from '../lib/format'
import { useRoute } from '../lib/router'
import SessionCard from './SessionCard.vue'
const route = useRoute()
const sessions = shallowRef<SessionSummary[]>([])
const apiError = ref<string | null>(null)
const loaded = ref(false)
const nowMs = ref(Date.now())
let timer: ReturnType<typeof setInterval> | undefined
let inflight = false
async function tick() {
if (inflight) return
inflight = true
try {
sessions.value = await fetchSessions(route.value.repo)
nowMs.value = Date.now()
apiError.value = null
loaded.value = true
} catch (err) {
apiError.value = err instanceof Error ? err.message : String(err)
} finally {
inflight = false
}
}
onMounted(() => {
void tick()
timer = setInterval(() => void tick(), 500)
})
// Switching repo in the topbar changes the route but not the mounted view
// (adwId stays null), so refetch and reset the list.
watch(
() => route.value.repo,
() => {
sessions.value = []
loaded.value = false
void tick()
},
)
onUnmounted(() => clearInterval(timer))
/** Optimistic removal; an empty id means the write failed, so re-sync instead. */
function onArchived(adwId: string) {
if (!adwId) {
void tick()
return
}
sessions.value = sessions.value.filter((s) => s.adw_id !== adwId)
}
const ordered = computed(() =>
sessions.value.toSorted((a, b) => (ts(b.started_at) || 0) - (ts(a.started_at) || 0)),
)
</script>
<template>
<div class="sessions">
<div v-if="apiError" class="error-bar">api unreachable — retrying {{ apiError }}</div>
<div v-if="ordered.length" class="list-head dim">{{ ordered.length }} runs</div>
<div v-if="ordered.length" class="cards">
<SessionCard
v-for="s in ordered"
:key="s.adw_id"
:session="s"
:now-ms="nowMs"
@archived="onArchived"
/>
</div>
<div v-else-if="loaded" class="empty-state">no sessions yet — run an ADW to see it here</div>
<div v-else-if="!apiError" class="empty-state">loading sessions…</div>
</div>
</template>
<style scoped>
.sessions {
display: flex;
flex-direction: column;
}
.list-head {
padding: 16px 24px 0;
font-size: 16px;
}
.cards {
/* Uniform grid: every card the same width and (fixed in SessionCard) height,
independent of content. */
display: grid;
grid-template-columns: repeat(auto-fill, minmax(460px, 1fr));
gap: 18px;
padding: 16px 24px 28px;
}
</style>

View file

@ -0,0 +1,88 @@
<script setup lang="ts">
import { computed } from 'vue'
import { BookOpen, CircleDollarSign, Coins, PenLine, Timer } from 'lucide-vue-next'
import { fmtCost, fmtDuration, fmtTokens } from '../lib/format'
const props = defineProps<{
kind: 'cost' | 'tokens' | 'runtime' | 'read' | 'written'
/** Raw value — cost in dollars, tokens as a count, runtime in milliseconds. */
value: number | null | undefined
/** Bare value, no pill chrome — for tight spots like waterfall blocks. */
compact?: boolean
}>()
const ICONS = {
cost: CircleDollarSign,
tokens: Coins,
runtime: Timer,
read: BookOpen,
written: PenLine,
}
// Every chip explains itself on hover. The token numbers in particular are read
// wrong without one — the headline is billed volume, not distinct tokens.
const TITLES = {
cost: 'Cost — dollars billed for this run, all agents combined.',
tokens:
'Tokens exchanged (billed) — everything sent or generated, counted once per turn. ' +
'Each turn re-sends the whole conversation, so this is far larger than the ' +
'conversation itself: it is spend, not size. The gap between it and read + ' +
'written is cached context re-read on later turns.',
runtime: 'Duration — wall-clock from the first phase starting to the last one ending.',
read:
'Read — raw tokens the models took in: prompts, file contents and tool results, ' +
'counted the first time they enter the context. Excludes cached re-reads of ' +
'material already counted here.',
written:
'Written — tokens the models actually generated. Each one produced exactly ' +
'once, so this is a true count of output.',
}
const text = computed(() => {
if (props.kind === 'cost') return fmtCost(props.value)
if (props.kind === 'runtime') return fmtDuration(props.value ?? NaN)
return fmtTokens(props.value)
})
</script>
<template>
<span class="stat" :class="{ compact }" :title="TITLES[kind]">
<component :is="ICONS[kind]" class="stat-icon" :size="compact ? 17 : 19" :stroke-width="2" />
<span class="stat-value">{{ text }}</span>
</span>
</template>
<style scoped>
.stat {
display: inline-flex;
align-items: center;
gap: 7px;
padding: 3px 12px;
border: 1px solid var(--border-soft);
border-radius: 999px;
background: rgba(19, 26, 38, 0.6);
font-size: 16px;
white-space: nowrap;
}
.stat-icon {
color: var(--faint);
flex: none;
}
.stat-value {
color: var(--text);
font-family: var(--mono);
font-variant-numeric: tabular-nums;
}
.stat.compact {
padding: 0;
border: none;
background: transparent;
}
.stat.compact .stat-value {
color: var(--dim);
}
</style>

View file

@ -0,0 +1,73 @@
<script setup lang="ts">
import { Check, Circle, LoaderCircle, X } from 'lucide-vue-next'
defineProps<{ status: string }>()
const ICONS: Record<string, unknown> = {
success: Check,
fail: X,
running: LoaderCircle,
queued: Circle,
}
</script>
<template>
<span class="chip" :class="status">
<component :is="ICONS[status] ?? Circle" class="chip-icon" :size="18" :stroke-width="2.5" />
{{ status }}
</span>
</template>
<style scoped>
.chip {
display: inline-flex;
align-items: center;
gap: 7px;
padding: 3px 13px 3px 10px;
border-radius: 999px;
border: 1px solid var(--border);
font-size: 16px;
color: var(--dim);
white-space: nowrap;
}
.chip-icon {
flex: none;
}
.chip.success {
color: var(--green);
border-color: rgba(74, 222, 128, 0.45);
background: rgba(74, 222, 128, 0.09);
box-shadow: 0 0 12px rgba(74, 222, 128, 0.12);
}
.chip.fail {
color: var(--red);
border-color: rgba(255, 111, 103, 0.45);
background: rgba(255, 111, 103, 0.09);
box-shadow: 0 0 12px rgba(255, 111, 103, 0.12);
}
.chip.running {
color: var(--blue);
border-color: rgba(108, 182, 255, 0.45);
background: rgba(108, 182, 255, 0.09);
box-shadow: 0 0 12px rgba(108, 182, 255, 0.18);
}
.chip.running .chip-icon {
animation: spin 1.1s linear infinite;
}
@keyframes spin {
to {
transform: rotate(360deg);
}
}
.chip.queued {
color: var(--dim);
border-style: dashed;
}
</style>

View file

@ -0,0 +1,111 @@
import type {
Envelope,
EventRow,
EventsPage,
GateResult,
HealthResponse,
PromptsResponse,
RepoInfo,
SessionDetail,
SessionSummary,
} from './types'
async function getJson(url: string): Promise<unknown> {
const res = await fetch(url)
if (!res.ok) throw new Error(`GET ${url} → ${res.status}`)
return res.json()
}
/** Prefix an API path with the repo slug; null/empty → the server's default repo. */
function repoPath(repo: string | null, path: string): string {
return repo ? `/api/${encodeURIComponent(repo)}${path}` : `/api${path}`
}
export function fetchRepos(): Promise<RepoInfo[]> {
return getJson('/api/repos') as Promise<RepoInfo[]>
}
/** Re-read the server's repos.json and return the refreshed repo list. */
export async function reloadRepos(): Promise<RepoInfo[]> {
const res = await fetch('/api/reload', { method: 'POST' })
if (!res.ok) throw new Error(`POST /api/reload → ${res.status}`)
const data = (await res.json()) as { repos?: RepoInfo[] }
return data.repos ?? []
}
export function fetchSessions(repo: string | null): Promise<SessionSummary[]> {
return getJson(repoPath(repo, '/sessions')) as Promise<SessionSummary[]>
}
export async function fetchSession(repo: string | null, adwId: string): Promise<SessionDetail> {
const detail = (await getJson(
repoPath(repo, `/sessions/${encodeURIComponent(adwId)}`),
)) as SessionDetail
return {
session: detail.session,
usage: detail.usage ?? { read: 0, written: 0 },
phases: detail.phases ?? [],
agents: detail.agents ?? [],
}
}
export async function fetchEvents(
repo: string | null,
adwId: string,
after: number,
limit = 500,
): Promise<EventsPage> {
const page = (await getJson(
repoPath(repo, `/sessions/${encodeURIComponent(adwId)}/events?after=${after}&limit=${limit}`),
)) as EventsPage | EventRow[]
if (Array.isArray(page)) {
const cursor = page.reduce((max, e) => Math.max(max, e.rowid), after)
return { events: page, cursor, has_more: page.length === limit }
}
return { events: page.events ?? [], cursor: page.cursor ?? after, has_more: page.has_more ?? false }
}
/** Archive a run out of the review list (or restore it with archived=false). */
export async function archiveSession(
repo: string | null,
adwId: string,
archived = true,
): Promise<void> {
const url = repoPath(repo, `/sessions/${encodeURIComponent(adwId)}/archive`)
const res = await fetch(url, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ archived }),
})
if (!res.ok) throw new Error(`POST ${url} → ${res.status}`)
}
export function fetchHealth(repo: string | null): Promise<HealthResponse> {
return getJson(repoPath(repo, '/health')) as Promise<HealthResponse>
}
// PhaseDetail imports the prompts type from here alongside fetchPrompts.
export type { PromptsResponse }
export async function fetchPrompts(
repo: string | null,
adwId: string,
agent: string,
): Promise<PromptsResponse> {
const res = await fetch(
repoPath(repo, `/sessions/${encodeURIComponent(adwId)}/agents/${encodeURIComponent(agent)}/prompts`),
)
// Not recorded (or endpoint not deployed yet) renders as "no prompts", not an error.
if (res.status === 404) return { system: null, user: null }
if (!res.ok) throw new Error(`GET prompts → ${res.status}`)
const data = (await res.json()) as Partial<PromptsResponse>
return { system: data.system ?? null, user: data.user ?? null }
}
export function fetchEnvelopes(repo: string | null, adwId: string): Promise<Envelope[]> {
return getJson(repoPath(repo, `/sessions/${encodeURIComponent(adwId)}/envelopes`)) as Promise<Envelope[]>
}
export function fetchGates(repo: string | null, adwId: string): Promise<GateResult[]> {
return getJson(repoPath(repo, `/sessions/${encodeURIComponent(adwId)}/gates`)) as Promise<GateResult[]>
}

View file

@ -0,0 +1,131 @@
import type { AgentStartPayload, EventRow, ToolCallPayload } from './types'
// ── Event dot colors ─────────────────────────────────────────────────────────
// One color per event type, shared by the session-card timelines and the phase
// detail list. gate_fail reads as an error signal on purpose.
export const EVENT_DOT_COLORS: Record<string, string> = {
agent_start: '#c89bff',
tool_call: '#5ad2dd',
handoff: '#94a3ff',
agent_end: '#4ade80',
error: '#ff6f67',
gate_fail: '#ff6f67',
}
export function dotColor(type: string | null): string | null {
if (!type) return null
return EVENT_DOT_COLORS[type] ?? null
}
// ── Agent lane colors ────────────────────────────────────────────────────────
// Config color wins (agents[].color from the API, or the agent_start payload
// for in-flight agents); the palette below covers dbs written before the
// color column existed.
export const AGENT_FALLBACK_COLORS = ['#c89bff', '#5ad2dd', '#94a3ff', '#e8b64a', '#f2a2c4']
export function agentColor(
configColor: string | null | undefined,
payloadColor: string | null | undefined,
index: number,
): string {
return (
configColor ??
payloadColor ??
AGENT_FALLBACK_COLORS[index % AGENT_FALLBACK_COLORS.length] ??
'#c89bff'
)
}
/** "#c89bff" + alpha → rgba() usable in inline styles. Invalid input → transparent. */
export function hexAlpha(hex: string, alpha: number): string {
const m = /^#?([0-9a-f]{6})$/i.exec(hex.trim())
if (!m || !m[1]) return 'transparent'
const n = Number.parseInt(m[1], 16)
return `rgba(${(n >> 16) & 0xff}, ${(n >> 8) & 0xff}, ${n & 0xff}, ${alpha})`
}
// ── Payload parsing ──────────────────────────────────────────────────────────
export function parsePayload(raw: string | null | undefined): Record<string, unknown> | null {
if (!raw) return null
try {
const parsed: unknown = JSON.parse(raw)
if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) {
return parsed as Record<string, unknown>
}
} catch {
/* legacy or truncated payloads render raw */
}
return null
}
export function parseToolCall(e: EventRow): ToolCallPayload | null {
const payload = parsePayload(e.payload_json)
if (!payload || typeof payload.tool !== 'string') return null
return payload as ToolCallPayload
}
export function parseAgentStart(e: EventRow): AgentStartPayload | null {
return parsePayload(e.payload_json) as AgentStartPayload | null
}
// Args keys most likely to BE the call, in priority order — a bash command, a
// file path, a search pattern. Used to build the one-line label.
const ARG_LABEL_KEYS = [
'command',
'cmd',
'file_path',
'path',
'pattern',
'query',
'url',
'prompt',
'description',
]
const LABEL_MAX = 160
function oneLine(value: string): string {
const flat = value.replaceAll(/\s+/g, ' ').trim()
return flat.length > LABEL_MAX ? `${flat.slice(0, LABEL_MAX)}…` : flat
}
/** One-line summary of a tool's args: the call itself, not a JSON dump. */
export function argsSummary(args: Record<string, unknown> | undefined): string {
if (!args) return ''
for (const key of ARG_LABEL_KEYS) {
const v = args[key]
if (typeof v === 'string' && v.trim() !== '') return oneLine(v)
}
const parts: string[] = []
for (const [key, v] of Object.entries(args)) {
if (v == null) continue
parts.push(`${key}=${typeof v === 'string' ? v : JSON.stringify(v)}`)
}
return oneLine(parts.join(' '))
}
/**
* Exact-call row label for an event.
* Rich tool_call → "bash: bun test src". Legacy payloads fall back to
* the event name plus whatever hint the payload carries.
*/
export function eventLabel(e: EventRow): string {
if (e.type === 'tool_call') {
const call = parseToolCall(e)
if (call?.tool) {
// New-tracer rows already carry the human label in events.name
// ("bash: ls -la src") — prefer it over re-deriving.
if (e.name?.startsWith(call.tool)) return oneLine(e.name)
const summary = argsSummary(call.args)
return summary ? `${call.tool}: ${summary}` : call.tool
}
const legacy = parsePayload(e.payload_json)
if (legacy && typeof legacy.pi_event === 'string') {
return `${e.name ?? 'tool'} ${legacy.pi_event}`
}
}
return e.name ?? e.type ?? ''
}

View file

@ -0,0 +1,92 @@
export function ts(iso: string | null | undefined): number {
if (!iso) return NaN
return new Date(iso).getTime()
}
export function fmtDuration(ms: number): string {
if (!Number.isFinite(ms) || ms < 0) return '—'
if (ms < 1000) return `${(ms / 1000).toFixed(2)}s`
const s = ms / 1000
if (s < 60) return `${s.toFixed(1)}s`
const m = Math.floor(s / 60)
const rem = Math.round(s % 60)
if (m < 60) return `${m}m ${String(rem).padStart(2, '0')}s`
const h = Math.floor(m / 60)
return `${h}h ${String(m % 60).padStart(2, '0')}m`
}
export function fmtClock(iso: string | null | undefined): string {
const t = ts(iso)
if (!Number.isFinite(t)) return '—'
return new Date(t).toLocaleTimeString([], { hour12: false })
}
export function fmtDate(iso: string | null | undefined): string {
const t = ts(iso)
if (!Number.isFinite(t)) return '—'
const d = new Date(t)
return `${d.toLocaleDateString([], { month: 'short', day: 'numeric' })} ${fmtClock(iso)}`
}
export function fmtTokens(n: number | null | undefined): string {
if (n == null) return '—'
if (n < 1000) return String(n)
if (n < 1_000_000) return `${(n / 1000).toFixed(1)}k`
return `${(n / 1_000_000).toFixed(2)}M`
}
export function fmtCost(n: number | null | undefined): string {
if (n == null) return '—'
return n >= 1 ? `$${n.toFixed(2)}` : `$${n.toFixed(4)}`
}
// Compact offset label for time axes: 0s, 30s, 1m, 1m30s, 2m, 1h05m.
export function fmtOffset(ms: number): string {
const s = Math.round(ms / 1000)
if (s < 60) return `${s}s`
const m = Math.floor(s / 60)
if (m < 60) {
const rem = s % 60
return rem ? `${m}m${String(rem).padStart(2, '0')}s` : `${m}m`
}
const h = Math.floor(m / 60)
const mrem = m % 60
return mrem ? `${h}h${String(mrem).padStart(2, '0')}m` : `${h}h`
}
const TICK_STEPS_MS = [1, 2, 5, 10, 15, 30, 60, 120, 300, 600, 1200, 1800, 3600].map(
(s) => s * 1000,
)
/** Evenly-stepped time-axis ticks over a span, ≤ maxTicks of them. */
export function axisTicks(spanMs: number, maxTicks = 8): { pct: number; label: string }[] {
const span = Math.max(spanMs, 1)
const step = TICK_STEPS_MS.find((s) => span / s <= maxTicks) ?? 3_600_000
const out: { pct: number; label: string }[] = []
for (let t = 0; t <= span; t += step) {
out.push({ pct: (t / span) * 100, label: fmtOffset(t) })
}
return out
}
// New-tracer tool_call payloads carry ok:false on tool errors; older payloads
// have no ok key and count as ok.
export function payloadOk(raw: string | null | undefined): boolean {
if (!raw) return true
try {
const p: unknown = JSON.parse(raw)
if (p && typeof p === 'object' && 'ok' in p) return (p as { ok?: unknown }).ok !== false
} catch {
/* not JSON — treat as ok */
}
return true
}
export function prettyJson(raw: string | null | undefined): string {
if (!raw) return ''
try {
return JSON.stringify(JSON.parse(raw), null, 2)
} catch {
return raw
}
}

View file

@ -0,0 +1,54 @@
/**
* Dependency-free JSON syntax highlighting.
*
* Same safety model as markdown.ts: every character of input is HTML-escaped;
* the only tags in the output are the <span>s this module writes. Token
* colors live in style.css under the .j-* classes.
*/
export function escapeHtml(s: string): string {
return s
.replaceAll('&', '&amp;')
.replaceAll('<', '&lt;')
.replaceAll('>', '&gt;')
.replaceAll('"', '&quot;')
.replaceAll("'", '&#39;')
}
// Strings first so digits/keywords inside them are consumed as part of the
// string token. A string followed by a colon is an object key.
const TOKEN =
/("(?:\\.|[^"\\])*")(\s*:)?|\b(true|false)\b|\b(null)\b|(-?\d+(?:\.\d+)?(?:[eE][+-]?\d+)?)/g
/** Highlight text that is already (or claims to be) JSON. Escapes everything. */
export function highlightJsonText(text: string): string {
let out = ''
let last = 0
for (const m of text.matchAll(TOKEN)) {
out += escapeHtml(text.slice(last, m.index))
const [full, str, colon, bool, nil, num] = m
if (str !== undefined) {
const cls = colon !== undefined ? 'j-key' : 'j-str'
out += `<span class="${cls}">${escapeHtml(str)}</span>${escapeHtml(colon ?? '')}`
} else if (bool !== undefined) {
out += `<span class="j-bool">${bool}</span>`
} else if (nil !== undefined) {
out += `<span class="j-null">null</span>`
} else {
out += `<span class="j-num">${escapeHtml(num ?? full)}</span>`
}
last = (m.index ?? 0) + full.length
}
out += escapeHtml(text.slice(last))
return out
}
/** Pretty-print raw JSON and highlight it; non-JSON falls back to escaped raw. */
export function highlightJson(raw: string | null | undefined): string {
if (!raw) return ''
try {
return highlightJsonText(JSON.stringify(JSON.parse(raw), null, 2))
} catch {
return escapeHtml(raw)
}
}

View file

@ -0,0 +1,120 @@
/**
* Minimal, dependency-free markdown → HTML for the compiled prompt panels.
*
* Safety model: ALL input is HTML-escaped before any tags are produced, so
* the only HTML in the output is what this module writes. No raw-HTML
* passthrough, links restricted to http(s).
*
* Supported: #–#### headings, fenced code blocks, inline code, bold, links,
* unordered/ordered lists, blockquotes, horizontal rules, paragraphs
* (pre-wrap preserves intra-paragraph line breaks).
*/
import { escapeHtml, highlightJsonText } from './highlight'
/** Bold + links, applied only OUTSIDE inline-code spans. */
function inline(escaped: string): string {
const parts = escaped.split(/(`[^`\n]+`)/g)
return parts
.map((part, i) => {
if (i % 2 === 1) return `<code>${part.slice(1, -1)}</code>`
return part
.replaceAll(/\*\*([^*\n]+)\*\*/g, '<strong>$1</strong>')
.replaceAll(
/\[([^\]\n]+)\]\((https?:\/\/[^)\s]+)\)/g,
'<a href="$2" target="_blank" rel="noopener noreferrer">$1</a>',
)
})
.join('')
}
export function renderMarkdown(src: string): string {
const lines = src.replaceAll('\r\n', '\n').split('\n')
const out: string[] = []
let i = 0
const paragraph: string[] = []
function flushParagraph() {
if (paragraph.length) {
out.push(`<p>${paragraph.map(inline).join('\n')}</p>`)
paragraph.length = 0
}
}
while (i < lines.length) {
const raw = lines[i] ?? ''
const line = escapeHtml(raw)
// Fenced code block — verbatim until the closing fence. json fences get
// syntax highlighting (the Report contract in every user.md is one).
const fence = /^\s*```(\w*)/.exec(raw)
if (fence) {
flushParagraph()
const code: string[] = []
i += 1
while (i < lines.length && !/^\s*```/.test(lines[i] ?? '')) {
code.push(lines[i] ?? '')
i += 1
}
i += 1
const text = code.join('\n')
const body =
(fence[1] ?? '').toLowerCase() === 'json' ? highlightJsonText(text) : escapeHtml(text)
out.push(`<pre class="md-code"><code>${body}</code></pre>`)
continue
}
const heading = /^(#{1,4})\s+(.*)$/.exec(raw)
if (heading?.[1] && heading[2] !== undefined) {
flushParagraph()
const level = heading[1].length
out.push(`<h${level}>${inline(escapeHtml(heading[2]))}</h${level}>`)
i += 1
continue
}
if (/^\s*(---+|\*\*\*+)\s*$/.test(raw)) {
flushParagraph()
out.push('<hr>')
i += 1
continue
}
if (/^\s*&gt;\s?/.test(line)) {
flushParagraph()
const quote: string[] = []
while (i < lines.length && /^\s*>\s?/.test(lines[i] ?? '')) {
quote.push(inline(escapeHtml((lines[i] ?? '').replace(/^\s*>\s?/, ''))))
i += 1
}
out.push(`<blockquote>${quote.join('\n')}</blockquote>`)
continue
}
const ulItem = /^\s*[-*]\s+/.test(raw)
const olItem = /^\s*\d+\.\s+/.test(raw)
if (ulItem || olItem) {
flushParagraph()
const tag = ulItem ? 'ul' : 'ol'
const marker = ulItem ? /^\s*[-*]\s+/ : /^\s*\d+\.\s+/
const items: string[] = []
while (i < lines.length && marker.test(lines[i] ?? '')) {
items.push(`<li>${inline(escapeHtml((lines[i] ?? '').replace(marker, '')))}</li>`)
i += 1
}
out.push(`<${tag}>${items.join('')}</${tag}>`)
continue
}
if (raw.trim() === '') {
flushParagraph()
i += 1
continue
}
paragraph.push(line)
i += 1
}
flushParagraph()
return out.join('\n')
}

View file

@ -0,0 +1,29 @@
/**
* Model → provider icon, by contains-check on the model name.
*
* Icons live in public/models/ (served from the site root). First matching
* needle wins; unknown models render no icon.
*/
const MODEL_ICONS: [needles: string[], icon: string][] = [
[['claude', 'opus', 'sonnet', 'haiku'], '/models/claude.png'],
[['gemini'], '/models/gemini.png'],
[['kimi', 'moonshot'], '/models/kimi.png'],
[['gpt', 'openai', 'codex', 'o3', 'o4'], '/models/openai.png'],
[['glm', 'zai', 'z.ai'], '/models/zai.png'],
]
export function modelIcon(model: string | null | undefined): string | null {
if (!model) return null
const m = model.toLowerCase()
for (const [needles, icon] of MODEL_ICONS) {
if (needles.some((n) => m.includes(n))) return icon
}
return null
}
/** Keep provider-qualified IDs compact while preserving the full ID in titles. */
export function modelName(model: string | null | undefined): string {
if (!model) return ''
return model.split('/').filter(Boolean).at(-1) ?? model
}

View file

@ -0,0 +1,44 @@
import { ref } from 'vue'
// Hash routes: #/ → sessions · #/<repo>/<adw_id> → waterfall · #/<repo>/<adw_id>/<phase_id> → phase panel open
// The repo segment is optional and defaults to the first repo the server serves.
export interface Route {
repo: string | null
adwId: string | null
phaseId: string | null
}
function parse(): Route {
const parts = window.location.hash
.replace(/^#\/?/, '')
.split('/')
.filter(Boolean)
.map(decodeURIComponent)
return { repo: parts[0] ?? null, adwId: parts[1] ?? null, phaseId: parts[2] ?? null }
}
const route = ref<Route>(parse())
window.addEventListener('hashchange', () => {
route.value = parse()
})
export function useRoute() {
return route
}
// Display name for the phase crumb — set by the trace view once phases load,
// since the phase_id in the URL is not the display name.
export const phaseCrumb = ref<string | null>(null)
export function hrefFor(repo?: string | null, adwId?: string | null, phaseId?: string | null): string {
const parts: string[] = []
if (repo) parts.push(encodeURIComponent(repo))
if (adwId) parts.push(encodeURIComponent(adwId))
if (phaseId) parts.push(encodeURIComponent(phaseId))
return '#/' + parts.join('/')
}
export function navigate(repo?: string | null, adwId?: string | null, phaseId?: string | null): void {
window.location.hash = hrefFor(repo, adwId, phaseId)
}

View file

@ -0,0 +1,27 @@
// Single switch point onto the server's contract: everything UI-side imports
// table shapes from here, which re-exports shared/types.ts.
export type {
Session,
SessionSummary,
SessionUsage,
SessionDetail,
Phase,
Event as EventRow,
EventsPage,
Envelope,
GateResult,
GateCheck,
AgentSession,
AgentStartPayload,
AgentEndPayload,
UsageBreakdown,
ToolCallPayload,
AgentPrompts,
PromptsResponse,
HealthResponse,
RepoInfo,
SessionStatus,
PhaseStatus,
PhaseKind,
EventType,
} from '@shared/types'

View file

@ -0,0 +1,7 @@
import { createApp } from 'vue'
import '@fontsource/play/400.css'
import '@fontsource/play/700.css'
import App from './App.vue'
import './style.css'
createApp(App).mount('#app')

View file

@ -0,0 +1,206 @@
:root {
--bg: #06080f;
--panel: #0d1119;
--panel-2: #131a26;
--panel-3: #0a0e16;
--border: #232c3d;
--border-soft: #222b3d;
--text: #f2f5fa;
--dim: #aabdd5;
--faint: #8b9cb6;
--green: #4ade80;
--red: #ff6f67;
--blue: #6cb6ff;
--amber: #e8b64a;
--purple: #c89bff;
--cyan: #5ad2dd;
--violet: #94a3ff;
/* Play carries the UI voice; mono is reserved for data (ids, times, code). */
--sans: 'Play', 'Helvetica Neue', system-ui, sans-serif;
--mono: ui-monospace, 'SF Mono', SFMono-Regular, Menlo, Monaco, 'Cascadia Mono',
'Roboto Mono', monospace;
/* Surface gradient shared by cards and panels — one recipe, everywhere. */
--surface: linear-gradient(180deg, #10141f 0%, #0b0f18 100%);
}
* {
box-sizing: border-box;
}
html,
body {
margin: 0;
padding: 0;
}
body {
/* Deep space with a faint violet/cyan aurora — fixed so scrolling glides over it. */
background:
radial-gradient(1100px 700px at 8% -10%, rgba(148, 163, 255, 0.09), transparent 62%),
radial-gradient(1000px 720px at 102% 108%, rgba(90, 210, 221, 0.07), transparent 60%),
linear-gradient(180deg, #06080f 0%, #090d17 100%);
background-attachment: fixed;
color: var(--text);
font-family: var(--sans);
/* Readability floor: nothing in the app renders below 16px. */
font-size: 16px;
line-height: 1.5;
}
#app {
min-height: 100vh;
}
a {
color: var(--blue);
text-decoration: none;
}
pre {
margin: 0;
padding: 12px 14px;
background: var(--panel-3);
border: 1px solid var(--border-soft);
border-radius: 8px;
overflow-x: auto;
font-family: var(--mono);
font-size: 16px;
line-height: 1.55;
color: #cfdded;
white-space: pre-wrap;
word-break: break-word;
}
/* JSON token colors — spans emitted by lib/highlight.ts */
.j-key {
color: var(--purple);
}
.j-str {
color: var(--green);
}
.j-num {
color: var(--amber);
}
.j-bool {
color: var(--cyan);
}
.j-null {
color: var(--red);
}
.dim {
color: var(--dim);
}
.faint {
color: var(--faint);
}
.error-bar {
margin: 12px 24px;
padding: 10px 14px;
border: 1px solid rgba(255, 111, 103, 0.5);
background: rgba(255, 111, 103, 0.1);
color: var(--red);
border-radius: 8px;
font-size: 16px;
}
.empty-state {
padding: 56px 24px;
text-align: center;
color: var(--dim);
font-size: 16px;
}
/* Rendered-markdown blocks (compiled prompts) — global because the content is
injected via v-html and scoped styles can't reach it. */
.md {
font-size: 16px;
line-height: 1.6;
color: var(--text);
}
.md h1 {
font-size: 20px;
margin: 14px 0 8px;
}
.md h2 {
font-size: 18px;
margin: 14px 0 8px;
}
.md h3,
.md h4 {
font-size: 17px;
margin: 12px 0 6px;
}
.md h1:first-child,
.md h2:first-child,
.md h3:first-child {
margin-top: 0;
}
.md p {
margin: 8px 0;
white-space: pre-wrap;
}
.md code {
background: var(--panel-2);
border: 1px solid var(--border-soft);
border-radius: 4px;
padding: 1px 7px;
font-size: 16px;
}
.md pre.md-code {
margin: 10px 0;
white-space: pre;
}
.md pre.md-code code {
background: transparent;
border: none;
padding: 0;
}
.md ul,
.md ol {
margin: 8px 0;
padding-left: 28px;
}
.md li {
margin: 3px 0;
}
.md blockquote {
margin: 10px 0;
padding: 4px 0 4px 14px;
border-left: 3px solid var(--border);
color: var(--dim);
white-space: pre-wrap;
}
.md hr {
border: none;
border-top: 1px solid var(--border);
margin: 14px 0;
}
@keyframes pulse {
0%,
100% {
opacity: 1;
}
50% {
opacity: 0.35;
}
}

View file

@ -0,0 +1,30 @@
{
"compilerOptions": {
"target": "ESNext",
"module": "ESNext",
"moduleResolution": "bundler",
"lib": ["ESNext", "DOM", "DOM.Iterable"],
"types": ["bun", "vite/client"],
"strict": true,
"noUnusedLocals": true,
"noUnusedParameters": true,
"noFallthroughCasesInSwitch": true,
"verbatimModuleSyntax": true,
"isolatedModules": true,
"skipLibCheck": true,
"resolveJsonModule": true,
"allowImportingTsExtensions": true,
"esModuleInterop": true,
"jsx": "preserve",
"jsxImportSource": "vue",
"noEmit": true,
"baseUrl": ".",
"paths": {
"@/*": ["./src/*"],
"@shared/*": ["./shared/*"]
}
},
"include": ["shared/**/*.ts", "server/**/*.ts", "src/**/*.ts", "src/**/*.vue", "vite.config.ts"]
}

View file

@ -0,0 +1,28 @@
import { fileURLToPath, URL } from "node:url";
import { defineConfig } from "vite";
import vue from "@vitejs/plugin-vue";
const API_PORT = process.env.PORT ?? "4600";
export default defineConfig({
plugins: [vue()],
resolve: {
alias: {
"@": fileURLToPath(new URL("./src", import.meta.url)),
"@shared": fileURLToPath(new URL("./shared", import.meta.url)),
},
},
server: {
port: 4601,
proxy: {
"/api": {
target: `http://localhost:${API_PORT}`,
changeOrigin: true,
},
},
},
build: {
outDir: "dist",
emptyOutDir: true,
},
});

View file

@ -0,0 +1,122 @@
# Create ADW
Compose a new ADW script — a thin, deterministic Python workflow over agents already in the config. Design the chain first, then generate or hand-write it.
## Step 1 — Design the chain
Answer four questions, in order:
1. **What agents, in what order?** Pick from the roster (`adws/adw_sssf_config/sssf.config.yaml`). The starter six cover most chains:
| Agent | Use when | Output type | Typical gates |
|---|---|---|---|
| `scout` | you need to FIND something first — read-only recon | `ScoutOutput` | `artifacts_exist` |
| `planner` | the work needs a plan before code changes | `PlanOutput` | `artifacts_exist`, `files_non_empty` |
| `builder` | code must change | `BuildOutput` | `diff_matches_claims` |
| `reviewer` | the change must be confirmed to BE what was asked for | `ReviewOutput` | `artifacts_exist`, `verdict_consistent` |
| *(no tester)* | verifying that it RUNS is a `kind="code"` phase over `quality.py`, not an agent | `QualityResult` → `as_envelope` | the exit code is the check |
| `documenter` | finished work needs a write-up (runs after a build, off the diff) | `DocumentOutput` | `artifacts_exist`, `files_non_empty` |
| any agent, generic ask | one-off prompt, no special shape | `GenericOutput` | as needed |
A new kind of agent needs a config entry + prompt pair + output type first — see `update_config.md`.
**The suite and the reviewer answer different questions.** "Does it run" is a test, and code can ask that. "Is this the thing that was asked for" is a review, and only an agent can. A green suite over a feature nobody requested is still a failed request, and neither one covers for the other.
2. **Where does code act?** Git branch/commit, migrations, deploys each get their own `kind="code"` phase — never buried inside an agent phase.
**Running the suite is one of these — there is no tester agent.** The command is written down in `quality.py`, so a `kind="code"` phase runs it (`quality.run_tests(run)` → `quality.as_envelope(result, "tests")` back into the builder) and the bounded repair loop is unchanged. An agent rediscovering `bun test` on every run buys nothing a subprocess does not already know. Capturing what changed is one of these: `changes.capture(run, ChangeCapture(base="main"))` diffs the working tree against a resolved base, writes `context_handoff/changes.diff`, and `changes.as_envelope(...)` hands it to the next agent. A diff is two git commands, not a judgement call.
3. **Does anything loop?** Test-fix cycles are bounded fix loops (see `update_adw.md`), not phase retries.
4. **What does each call need to prove?** Pick gates per call from `gates.py`: `artifacts_exist`, `files_non_empty`, `json_parses`, `diff_matches_claims`, `tests_pass("cmd")` — or an inline one-off.
## Step 2 — Ownership rules (the swim lanes depend on these)
- `kind="agent"` → `owner` MUST be an agent name from the config — it selects the harness (model, thinking, tools, prompts) AND the lane. `ph.call()` runs whoever owns the phase.
- `kind="engineer"` → `owner=run.engineer`. Every ADW opens with the engineer request phase — it is the system input record.
- `kind="code"` → `owner` is a short actor label (`"git"`, `"db"`); all code phases share the code lane.
- Phase `name` must be unique within the run (`plan`, `build`, `test_1`, `fix_1`, …) — the UI keys blocks on it.
- **`description` is required and must earn its place.** The name identifies the phase; the description explains it — what this phase does and why, in one sentence. It rides the `phase_start` event and is the only line of intent the trace, the console, and the phase block ever show. `PhaseParams` raises at construction on a blank description *or* one that merely restates the name (`commit_plan: "Commit the plan"`), so the rule fails before the phase opens rather than leaving an unreadable run in the db. Write `"Put the spec on record before any code exists to blur it"` instead.
- `retries=N` on an **agent** phase = extra gate-correction rounds re-sent into the same session (pi's `--session-id` creates-or-continues, so context stays intact). Code-phase re-execution is not implemented in v1.
## Step 3 — Generate or write it
```bash
uv run <skill>/scripts/make_adw.py --name review_docs --agents scout,builder
```
`<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf` or a repo's `.claude/skills/sssf`).
Writes `adws/adw_review_docs.py`: one agent phase per name, chained by `previous=`, starter agents mapped to their output types, unknown agents to `GenericOutput`. It does NOT create config entries or prompt files — do that first (`update_config.md`), or `agents.validate()` will stop the run and tell you what's missing.
## The canonical skeleton
Every `adw_*.py`, generated or hand-written, is a `uv` single-file script with this shape:
```python
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan Build — plan the request, then implement the plan."""
import argparse
import sys
from adw_modules import agents, gates, git_helper, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner", "builder"] # names, never models
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config) # 1. point to config
agents.validate(cfg, REQUIRED_AGENTS) # 2. fail fast — nothing spawns on a half-valid config
run = session.ensure(cfg, adw_id) # 3. pin-or-create the session → the Run object
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="build", kind="agent", owner="builder", retries=1,
description="Implement the plan exactly")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.diff_matches_claims]))
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Commit the working tree")) as ph:
message = build.commit_message or f"sssf({run.adw_id}): {build.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))
```
## Non-negotiables
- **`REQUIRED_AGENTS` + `agents.validate()`** — declare every agent name the script uses and validate before the first phase.
- **Every agent call declares a concrete output type** from `data_types.py`. No untyped handoffs.
- **`previous=` carries the chain** — the upstream envelope lands in the next agent's `user.md` as `{{previous_envelope}}`; bulky context moves through `context_handoff/` files the envelope references.
- **The engineer request phase comes first**, always.
- **Four-param rule** — `run.phase()` and `ph.call()` each take exactly one object; new helpers with >4 params get a data type.
- **Stay thin** — sequencing and acceptance only; real logic goes in `adw_modules/` (`update_modules.md`).
- **Committing is a code phase, and it needs a fallback.** `PlanOutput`, `BuildOutput`, and `DocumentOutput` each carry a `commit_message` the agent writes **for its own work product** — the spec, the code, the write-up. It defaults to empty, so always `envelope.commit_message or <fallback>`, and commit each product with the message of the agent that made it (`adw_simple_sdlc.py` commits three times and never crosses them). `git_helper.commit_all(message)` stages everything, commits, and returns the short sha; it raises a clear error when the cwd isn't a git repo or nothing changed, and that raise fails the phase.
## Before you ship it
1. `uv run adws/adw_<name>.py "a tiny real request"` — watch it go green end to end.
2. Check the trace: `sqlite3 adws/adw_data/sssf.db "select seq,name,kind,owner,status from phases where adw_id='<id>' order by seq;"`
3. Read the final `envelope.json` — is the output type earning its fields, or should it be sharper?

View file

@ -0,0 +1,63 @@
# Create Config
Generate `sssf.config.yaml` — the agent roster for a target repo.
## Generate it
```bash
uv run <skill>/scripts/make_config.py
```
`<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf` or a repo's `.claude/skills/sssf`).
Writes `adws/adw_sssf_config/sssf.config.yaml` — creating the directory if needed — with the starter agents (planner, builder, scout, reviewer, documenter) wired to the prompt files `/sssf install` stamped into `adws/adw_data/prompt_engineering/`. That path is the default every ADW and the justfile look for; `--config` overrides it. `make_config.py` refuses to overwrite an existing config unless you pass `--force`, so retuning an existing roster is a hand edit — see `update_config.md`.
## The rule
**One agent, one prompt, one purpose.** An entry defines who an agent *is*: its coding agent, model, thinking level, and exactly one system prompt plus one user prompt. How it gets *used* — the output type, a per-call user prompt override — lives at the ADW call site, never here.
## Schema
```yaml
defaults:
coding_agent: pi # v1: pi only (claude_code is specced, stubbed until v2)
model: google/gemini-3.6-flash # ALWAYS provider/model-id — a bare id is ambiguous
thinking: medium # off | minimal | low | medium | high | xhigh | max
harness_engineering: [] # pi extension names
data_dir: adws/adw_data # runtime home: {data_dir}/sessions/{adw_id}/{agent_name}/
observability:
db: adws/adw_data/sssf.db # tracer writes here; the UI polls it
poll_ms: 500 # visualizer live-poll cadence
agents:
- name: planner # ADW scripts name agents, never models
coding_agent: pi
model: google/gemini-3.6-flash
thinking: high
color: "#a78bfa" # optional hex — this agent's lane color in the visualizer
purpose: Turn a request into a plan the builder can implement without asking questions.
prompt_engineering:
system: adws/adw_data/prompt_engineering/planner/system.md
user: adws/adw_data/prompt_engineering/planner/user.md
- name: scout
thinking: high # unset keys fall through to defaults
purpose: Find and report where things live; change nothing.
prompt_engineering:
system: adws/adw_data/prompt_engineering/scout/system.md
user: adws/adw_data/prompt_engineering/scout/user.md
tools: # optional allowlist — omit the key entirely for all tools
- read
- bash
```
Every agent entry merges over `defaults`, so an entry only states what differs. Pi's builtin tools are `read`, `bash`, `edit`, `write` — a read-only recon agent gets `[read, bash]`; a builder omits `tools` altogether.
## After generating
1. Each agent needs its prompt pair to exist on disk: `adws/adw_data/prompt_engineering/{name}/system.md` and `user.md`. `agents.validate()` fails the run at startup if either is missing.
2. Write `purpose` as one sentence and make the system prompt say the same thing — the two should not drift.
3. Validate by running the smallest ADW that names your agents; a bad entry fails fast, before anything spawns.
Full field-by-field spec, thinking-level mapping, and model resolution: `references/config.md`. Retuning an existing roster: `update_config.md`.

View file

@ -0,0 +1,110 @@
# How to Prompt for the Engineering
Read this **before every ADW launch**. The prompt you pass is what the whole chain reads: the planner plans from it, the builder builds from it, the reviewer judges against it. Your prompt might run through 10s or 100s of agents. A sloppy prompt is not a small tax; it is paid again by every agent in the chain.
## Purpose
Turn what the engineer said into the prompt the ADW receives: **clearer, not different.** You are a translator, not a redesigner.
## The one rule
**The intent is theirs. The precision is yours.**
| You MAY | You MAY NOT |
|---|---|
| Carry every constraint forward, verbatim | Quietly drop a requirement because it looks hard or odd |
| Fix grammar, cut rambling, order the steps | Soften a strong ask ("rewrite" → "refactor a bit") |
| Change the language used to better communicate the idea | Research the codebase for exact file names, never go into the app |
If you catch yourself improving the *idea* rather than the *sentence*, stop. Raise the concern to the engineer in your own message and launch what they asked for.
## You never touch the application, you prompt, monitor, observe, and report.
Outside of understanding the ADWs, you never research, touch, or dive into the codebase thats being operated on.
Your role is to simply kick off the workflow. There are entire teams of agents inside these ADWs built to do the work.
Your job is to kick it off, monitor, observe, report. Not interact with the application layer. You operate only on the agentic layer, the ADWs, the software factory.
## The shape
Four lines. Nothing else earns its tokens.
```
<the ask — one imperative sentence, their words where they were specific>
Where: <files or dirs you verified>
Done means: <the observable result — a response shape, a passing test, a rendered element>
Out of scope: <what you were tempted to add, named so nobody adds it>
```
**Before** (what the engineer said):
> can we get tags on posts, sorted by popularity
**After** (what the ADW receives):
```
Add a GET /api/tags endpoint returning {tags: [{tag, count}]} — the distinct tags
across all posts with how many posts carry each, sorted by count descending then
tag ascending.
Where: src/server.ts (routes), src/server.test.ts (tests)
Done means: GET /api/tags returns the counts, and a new test in server.test.ts covers it.
Out of scope: tag editing UI, tag filtering on the post list.
```
Same idea, same scope. What changed is that "popularity" became a sort order, the files are named, and nobody has to guess where it stops.
## Which ADW
**If the engineer named one, launch that one.** Their call stands — no second-guessing, no "upgrading" them to a longer chain. If you think another fits better, say so in your own message and launch what they asked for.
**If they did not, read what this repo actually has and choose from that.**
```bash
ls adws/adw_*.py # the menu
head -20 adws/adw_<name>.py # every ADW opens with its `Phases:` line — the chain in one line
```
Chains are the engineer's to add, rename, and rewire, so **the files on disk are the only authority**. Never launch from memory or from a name you saw in a doc; read the docstrings, then match by shape:
| The work | Look for a chain that |
|---|---|
| Changes code, and the shape is not obvious — new behaviour, more than one file, anything you would want a plan for | goes end to end: plans, builds, verifies, reviews, and documents |
| Changes code, one well-understood edit | plans, builds, and verifies |
| Implements a plan this session already produced (`--adw-id`) | starts at build and verifies |
| Confirms built work is what was asked for | ends in a review phase |
| Writes up work already shipped | captures the diff and documents it |
| Is a question, and nothing should change | is a single read-only agent — the one case where one phase is right |
**Never a single-agent chain when the engineer asked for work to be done.** One-phase ADWs answer questions and run one-offs; they do not deliver.
**The more complex the ask, the more complete the chain.** Complexity means: more than one file, a behaviour you cannot describe in one sentence, anything touching data or an interface others call, or any request where you had to guess. When two chains both fit, take the longer one — a phase you did not need costs cents, while a change nobody planned, verified, reviewed, or wrote up costs an afternoon.
If nothing on disk fits the shape you need, say so and offer to compose one (`create_adw.md`) rather than forcing the work into a chain that skips the phase it needed.
## Workflow
1. **Read it twice.** Mark every noun that could point at two things.
2. **Verify before you write.** Every path, route, and symbol you put in the prompt must exist — check it. A wrong path costs a whole build phase.
3. **Draft the four lines.**
4. **Diff against the original.** Every specific thing they said, still there? Anything in your draft they did not say? Delete it.
5. **Ask at most one question**, only when two readings would produce different code. Otherwise state your assumption in the prompt and say so when you report.
6. **Launch** the chain from *Which ADW* above; `run_adw.md` covers the mechanics and the watching. Inline for a short ask; for anything longer, write `requests/<slug>.md` and pass the path — every ADW takes either.
## Rules that do not bend here
- **Do not write the plan.** Your prompt says WHAT and DONE MEANS. HOW belongs to the planner — unless the engineer specified how, and then you carry it word for word.
- **Do not address the harness in the prompt.** "Use the reviewer", "retry twice", "then commit" are chain choices, and the chain is chosen by which ADW you launch, not by prose the agents will read.
- **Do not pad.** No preamble, no restating the repo, no encouragement. Gates check claims, not prose.
- **Their exact words survive.** When the engineer was specific — a name, a number, a format, a file — quote it rather than paraphrasing.
## Report back
After launching, show the engineer three things so a bad translation dies in seconds rather than at the commit phase:
1. **The prompt you actually sent** — verbatim.
2. **The ADW you chose**, and the one-line reason — or that you used the one they named.
If they named a roster (a config, a model tier), say which one you ran on; if they did not, you ran the default, and switching that is their call, not yours (`run_adw.md`).
3. **The `adw_id`**, so they can watch it (`just phases <adw_id>`).
Then observe and report per `run_adw.md`. You run the system; you do not do the work inside it.

69
sssf/cookbooks/install.md Normal file
View file

@ -0,0 +1,69 @@
# Install
`/sssf install` — stamp the entire factory out of the skill and into the current working directory.
## Install the skill
The skill is distributed from the `INDigitalStudio/skills` repo. Install it with the skills CLI:
```bash
skills add INDigitalStudio/skills --skill sssf -g -y
```
That places the skill (with its `scripts/`, `templates/`, and the visualizer app) into your agent's skills directory. The rest of this cookbook assumes the skill is installed and you can point at its `scripts/install.py`.
## Run it
```bash
uv run <skill>/scripts/install.py
```
Run from the **target repo root** — the cwd is where everything lands. `<skill>` is the directory this skill was installed into (e.g. `~/.agents/skills/sssf`, a repo's `.claude/skills/sssf`, or wherever the skills CLI placed it). The scripts resolve their own location, so any of those paths works.
`install.py` asks which **coding-agent harness** to use — `pi` or `omp` — and writes your choice into the stamped config's `defaults.coding_agent`. Pass `--harness pi|omp` to skip the prompt. It also **registers the repo** in the visualizer's `repos.json` (next to the skill), so the multi-repo trace UI can browse this repo's sessions.
## What gets stamped
`install.py` copies `templates/` into the cwd:
| Stamped | From | Tracked? |
|---|---|---|
| `adws/adw_sssf_config/sssf.config.yaml` | `templates/sssf.config.yaml` | yes — the agent roster |
| `.env.sample` | `templates/env.sample` | yes |
| `adws/adw_*.py` | `templates/adws/` | yes — the thirteen starter ADWs (incl. `adw_validate.py`) |
| `adws/adw_modules/` | `templates/adws/adw_modules/` | yes — all low-level logic |
| `adws/adw_data/prompt_engineering/{planner,builder,scout,reviewer,documenter}/` | `templates/prompt_engineering/` | yes — **the user-owned home for prompts** |
| `adws/adw_data/harness_engineering/` | `templates/harness_engineering/` | yes — **the user-owned home for pi extensions** |
| `justfile` | `templates/justfile` | yes — starter recipes: `just demo`, the workflows, the trace reads, `just obs` |
| `adws/adw_data/sessions/`, `adws/adw_data/sssf.db` | created at runtime | no — gitignored |
The two `*_engineering` dirs mirror the two config keys of the same name: `prompt_engineering` is what an agent is told, `harness_engineering` is what its harness can do. Both are yours the moment they are stamped. Edit them in `adws/adw_data/`, never back inside the skill.
`harness_engineering/` ships with `subagents.ts` — the pi extension backing `subagent_create` / `_continue` / `_list` / `_remove`, wired to the planner and scout in the starter roster.
## Idempotency
Re-running is safe. `install.py` skips **every** file that already exists — your config, your prompts, and previously stamped code alike — and reports what it skipped, so a second run doubles as a drift check. To refresh stamped code (`adw_modules/`, the starter `adw_*.py`) to the skill's current version, run with `--force` — but know that `--force` overwrites ALL existing stamped files, including `sssf.config.yaml` and `prompt_engineering/`, so commit or back up user-owned edits first.
## Post-install checklist
1. **Env** — if the repo has no `.env` yet, `cp .env.sample .env` and set the key(s) your harness needs. **If the repo already has a `.env`** (an app's own env, e.g. PocketBase keys), do NOT overwrite it — the ADWs read `.env` via dotenv and merge, so just append the SSSF keys you need. Pi needs `OPENROUTER_API_KEY`; omp reads its own model catalog (`omp models --json`). `ANTHROPIC_API_KEY` / `CLAUDE_CODE_PATH` are only needed once Claude Code lands in v2.
2. **The harness is installed and on PATH** — `pi --version` (or `omp --version`). Set `PI_PATH` / `OMP_PATH` in `.env` if not.
3. **The model resolves** — `just validate` checks the whole roster (names, prompt files, models) without running anything. If a model doesn't resolve, point the roster at one that does: set `defaults.model` and drop the per-agent `model:` overrides. See `references/config.md` for model resolution.
4. **Gitignore** — `install.py` appends `adws/adw_data/sessions/`, `adws/adw_data/sssf.db*`, and `.env` for you; confirm they landed. All three are runtime or secrets and must never be committed.
5. **Git repo** — ADWs that end in a commit phase call `git_helper.commit_all`, which raises if the cwd is not a git repository. Run `git init` and make a first commit before using `adw_plan_build.py`, `adw_plan_build_test.py`, or `adw_simple_sdlc.py`. `adw_document.py` needs one too: it measures the change with `git diff` against a base ref (`main` by default, `--base` to override).
6. **Smoke test** — first `just validate` (roster check, nothing runs), then `just demo` runs two cheap read-only workflows back to back, or run the smallest ADW directly:
```bash
just validate # roster check first
just demo # both, end to end
uv run adws/adw_prompt.py "reply with a one-line summary of this repo" # the raw form
```
Green means the whole path works: config validated, session minted, the coding agent ran, envelope parsed, events landed in `adws/adw_data/sssf.db`. Verify the trace exists before trusting anything larger:
```bash
sqlite3 adws/adw_data/sssf.db "select adw_id, status from sessions order by started_at desc limit 1;"
```
If the smoke test fails, fix it before composing chains — every multi-agent ADW rides on this exact path.

122
sssf/cookbooks/run_adw.md Normal file
View file

@ -0,0 +1,122 @@
# Run ADW
Run a workflow and report on it. **You run and observe — you never step into the process or do the work yourself.**
## Step 0 — translate the request
**Read [how_to_prompt_for_the_eng.md](how_to_prompt_for_the_eng.md) before you launch anything.** The prompt you pass is read by every agent in the chain, so it gets written deliberately: same intent, sharper words, verified paths, and a stated "done means". That cookbook is the whole procedure; this one starts once you have the prompt.
## The orchestrator's posture
The ADW is the worker. Your job is to launch it, watch the trace, and tell the engineer what happened. Do not read the agent's target files and "help", do not fix the code an agent was supposed to fix, do not edit an envelope. If a run fails, report the failing phase and its violations — the fix is a config, prompt, or ADW change, made deliberately, and then a re-run.
## Launch
Which chain to launch is decided in `how_to_prompt_for_the_eng.md`, and the short version is: **the ADW the engineer named, or else the most complete composed chain the work justifies — never a single-agent one.** Read `ls adws/adw_*.py` and the `Phases:` line in each docstring to see what this repo has; the names below are shape, not a menu.
```bash
uv run adws/<end-to-end-chain>.py "add a /health endpoint"
uv run adws/<plan-build-verify-chain>.py requests/health.md
uv run adws/<build-first-chain>.py "implement the plan" --adw-id a1b2c3d4
uv run adws/<recon-chain>.py "where is auth handled" --config path/to/other.config.yaml
```
The prompt is inline text or a file path. Launch in the background so you can poll while it works; the `adw_id` is printed on startup — capture it, everything else keys off it.
### Listen for the roster
The chain says *what runs*; the config says *who runs it*. **If the engineer references a roster, a config, or a model tier, pass it — do not fall through to the default.**
```bash
just rosters # every roster on disk, and the model each agent runs
```
That prints the path to pass and who is in it, in one read:
```
adws/adw_sssf_config/sssf.config.yaml
planner fireworks/accounts/fireworks/models/kimi-k3
builder google/gemini-3.6-flash (inherited)
adws/adw_sssf_config/sssf.frontier.config.yaml
planner anthropic/claude-opus-5
```
Read those from disk every time. Rosters are the engineer's to add, rename, and retune, so a name you remember from a doc is a guess.
They will rarely say `--config`. Treat any of these as naming a roster, then resolve it to a file:
| What they say | What it means |
|---|---|
| "run it on the frontier config", "use the frontier roster" | the roster file whose name matches |
| "run this with the big models", "use the sota roster" | the non-default roster — confirm which if there is more than one. Each config's header comment lists the names it answers to, so `head -3` on the file settles it |
| "have opus plan this one" | a roster whose planner is that model; if none exists, say so rather than editing the config mid-request |
| nothing about models at all | the default, `adws/adw_sssf_config/sssf.config.yaml` |
`--config` takes the path directly; the justfile recipes read `SSSF_CONFIG` instead:
```bash
uv run adws/<chain>.py "<prompt>" --config adws/adw_sssf_config/sssf.frontier.config.yaml
SSSF_CONFIG=adws/adw_sssf_config/sssf.frontier.config.yaml just <recipe> "<prompt>"
```
Two things that bite:
- **Never swap rosters on your own.** A different roster is a different cost and a different result. If the default's model looks wrong for the work, say so and let the engineer choose.
- **Switching rosters mid-session breaks resumption.** `agent_map.json` records the model each coding-agent session was created with, so a joined run (`--adw-id`) whose config now names a different model starts that agent **fresh** instead of resuming its context window. That is deliberate — a bad resume is worse — but it means "plan on the frontier roster, then build on the default" costs the builder its accumulated context. Say so when you report it.
`--adw-id` is optional on **every** ADW. Given one, the run joins that session if it exists or creates it pinned to exactly that id: same `sessions/{adw_id}/` dirs, same `context_handoff/`, envelopes appended, and each agent resumes its existing coding-agent context window via `agent_map.json`. That is how you chain ADWs — plan under one id, then build under the same id.
## Observe
The trace db is `adws/adw_data/sssf.db`. It is WAL, so reads never block the running writers — poll it as often as you like.
```bash
# where the run stands
sqlite3 adws/adw_data/sssf.db \
"select seq, name, kind, owner, status, attempt from phases where adw_id='a1b2c3d4' order by seq;"
# the live tail — cursor on rowid, same query the visualizer polls
sqlite3 adws/adw_data/sssf.db \
"select rowid, type, name, started_at from events where adw_id='a1b2c3d4' and rowid > 0 order by rowid limit 50;"
# why a phase failed
sqlite3 adws/adw_data/sssf.db \
"select attempt, gate, passed, checks_json from gate_results where adw_id='a1b2c3d4';"
# session-level status
sqlite3 adws/adw_data/sssf.db \
"select adw_id, request, status, total_tokens from sessions order by started_at desc limit 5;"
# what an agent actually did, slowest tool calls first
sqlite3 adws/adw_data/sssf.db \
"select name, tokens, started_at, ended_at from events
where adw_id='a1b2c3d4' and type='tool_call' order by ended_at desc limit 20;"
```
Poll on a cursor: keep the highest `rowid` you have seen and query `where rowid > ?`. Don't re-read the whole table each pass.
`tool_call` rows carry a real span, so durations come off the columns — see `references/observability.md` for which fields each event type populates.
The ADW also narrates to stdout, and every line it prints is written to the db as a `log` event — terminal and swim lane tell the same story by construction, so tailing the background process is a valid second view rather than a competing source of truth.
Files are the raw record if you need more than the db shows: `adws/adw_data/sessions/{adw_id}/{agent}/raw_output.jsonl` (full coding-agent stream), `envelope.json` (the parsed final response), `prompts/` (exactly what was sent), and `context_handoff/` (what agents wrote for each other).
## When a run is stuck
A hung coding agent produces no events at all, so the trace goes quiet rather than red. Read it in this order:
```bash
just phases <adw_id> # which phase is still `running`
just procs <adw_id> # what that phase is actually running, with pids
just kill <adw_id> # stop it — children first, then the workflow
```
`processes` rows with `ended_at IS NULL` are the live ones. If `procs` shows a pi child but the phase has produced no `tool_call` events and its `raw_output.jsonl` is empty, the agent never got started properly — check the model resolves and that nothing is blocking the subprocess, rather than waiting it out. `just kill` verifies each pid still matches the command that was recorded before signalling, because pids get recycled.
A killed run marks itself `fail` and closes its process rows, so the trace never claims work is in flight that is already dead.
## Report
Tell the engineer, in order: which chain and which roster you launched (name the config whenever it was not the default), which phase is running now (or which failed), phase statuses in sequence, and for a failure the gate violations or the error verbatim. Remember **every phase defaults to `fail`** — a phase showing `fail` may simply never have completed; `queued` means it never started. Don't dress up a partial run as a success.
For a visual live view, the visualizer app in the skill (`just obs`, or tmux sessions viz-api :4600 + viz-ui :4601) polls this same db — sessions as cards, runs as swim lanes, phases and tool calls drill-in. The sqlite queries above remain the headless equivalent.

View file

@ -0,0 +1,88 @@
# SSSF Overview
The system map the orchestrator reads on startup — what SSSF is, how a stamped repo is laid out, and which cookbook to load next.
## What SSSF is
Super Simple Software Factory builds repeatable **agents plus code** workflows. Deterministic Python (an ADW script) owns sequencing, retries, and acceptance; agents are bounded nodes inside that graph. Agent proposes, code disposes.
Your job as orchestrator: **run the system, observe the system, help the engineer interact with it.** You do not do the work an ADW exists to do.
## Layout of a stamped repo
```
adws/
├── adw_sssf_config/
│ └── sssf.config.yaml the agent roster — one agent, one prompt, one purpose
├── adw_prompt.py smallest ADW: one agent, one prompt, traced end-to-end
├── adw_plan.py, adw_scout.py, adw_build.py, adw_plan_build.py, adw_build_test.py, adw_plan_build_test.py
├── adw_build_review.py build → review: is this what was asked for? (not testing)
├── adw_document.py write up the work just done, from git diff vs main
├── adw_simple_sdlc.py plan → build → test → review → document; commits each product
├── adw_modules/ ALL low-level logic — ADW scripts stay thin
│ ├── data_types.py AgentCall, PhaseParams, Phase, Envelope + one output type per agent call
│ ├── agents.py load_config, validate, resolve entry → interface + model + thinking
│ ├── runner.py the Run object: run.phase(PhaseParams) → ph.call(AgentCall)
│ ├── agent_pi.py Pi interface (v1) · agent_cc.py Claude Code (v2, stubbed)
│ ├── gates.py gate(envelope, run) -> GateReport — one check per item verified
│ ├── changes.py git diff vs a resolved base → ChangeSet → envelope for the documenter
│ ├── prompts.py, session.py, tracer.py, console.py, git_helper.py, utils.py
└── adw_data/
├── prompt_engineering/{agent}/{system.md,user.md} tracked — edit prompts HERE, never in the skill
│ planner · builder · scout · reviewer · documenter
├── sessions/{adw_id}/ gitignored runtime
│ ├── agent_map.json agent → coding-agent session_id + model
│ ├── context_handoff/ the one place agents write files for the agents that follow
│ └── {agent}/{prompts/, raw_output.jsonl, envelope.json}
└── sssf.db gitignored SQLite trace db the visualizer polls
```
**v1 runs Pi only.** `coding_agent: pi`, default model `gemini-3.6-flash`, thinking `medium`. `claude_code` is specced in the config and stubbed in the interface — it lands in v2.
## The phase model
Every ADW run is a sequence of **phases**, each one `with run.phase(PhaseParams(...))`. Three kinds, three swim lanes:
- **engineer** — the human lane; today the system-input phase (who asked, and for what).
- **agent** — `ph.call(AgentCall(...))`: prompt in → typed envelope out → gates verified.
- **code** — deterministic steps that stand alone (git branch, git commit, migrate). Never buried inside an agent phase.
**Success must be earned — every phase defaults to `fail`.** A clean exit flips it to success; agent phases additionally require the envelope to parse and all gates to come back green. A raise keeps it failed, records an error event, and aborts the run. `retries=N` on an agent phase buys extra gate-correction rounds through the same session before that raise happens.
## Envelopes
Agents have exactly two output channels: reference files written into `context_handoff/`, and a **final valid-JSON response** parsed against the output type the call declared. Code persists it as `envelope.json` and injects it into the next agent's `user.md` via `{{previous_envelope}}`. Bad JSON is never a restart — the harness re-prompts the *same session, context intact*, until it parses (bounded). See `references/handoff.md`.
**The output contract is a synced triad**: the type in `data_types.py` ↔ the `## Report` JSON example in the agent's `user.md` ↔ `output_type=` at the call site. Editing any one of the three means editing all three in the same change — drift between them taxes every call with correction retries.
## Running an ADW
```bash
uv run adws/adw_plan.py "add a /health endpoint"
uv run adws/adw_plan_build.py requests/health.md --adw-id a1b2c3d4
```
The prompt is inline text or a file path. `--adw-id` is optional on every ADW: given one, the run joins that session (same dirs, same `context_handoff/`, agents resume their existing context windows); omitted, a fresh id is minted and printed.
## When you have finished reading this
You are done with startup. List the ADWs (`ls adws/adw_*.py`, plus each `Phases:` docstring line) as a table, and **wait for the engineer's request.**
Do not survey anything else — not the trace db, not the config, not past runs, not the repo tree. You do not yet know what the request is, so anything you gather now is a guess about what will matter, spent from the context the real work needs. Every cookbook and reference below is lazy-loaded, one per request, and that is the whole design.
## Where to go next
Load one cookbook per request — this overview is the only one you read up front.
| Request | Cookbook |
|---|---|
| Turn a request into the prompt an ADW gets | `how_to_prompt_for_the_eng.md` — **read before every launch** |
| Set the system up in a repo | `install.md` |
| Write a new ADW script | `create_adw.md` |
| Change an existing ADW chain | `update_adw.md` |
| Generate `sssf.config.yaml` | `create_config.md` |
| Add or retune an agent | `update_config.md` |
| Add low-level logic or a gate | `update_modules.md` |
| Run and monitor a workflow | `how_to_prompt_for_the_eng.md`, then `run_adw.md` |
References, loaded when you need the spec: `references/config.md` (full config schema), `references/handoff.md` (envelope + session layout), `references/observability.md` (events, db tables, polling).

View file

@ -0,0 +1,90 @@
# Update ADW
Modify an existing ADW chain — add phases, add gates, add a bounded fix loop.
## Add a phase
Insert a `with run.phase(...)` block where it belongs in the sequence. Pick the right `kind`: `agent` for a `ph.call(...)`, `code` for a deterministic step, `engineer` for a human touchpoint. If the new phase names an agent not already in `REQUIRED_AGENTS`, add it there too — otherwise validation passes and the run dies mid-flight instead of at startup.
```python
with run.phase(PhaseParams(name="scout", kind="agent", owner="scout",
description="Locate the code the request touches")) as ph:
found = ph.call(AgentCall(output_type=ScoutOutput, prompt=prompt))
```
Phase `name` must be unique within the run — that is what the UI keys blocks on. In a loop, suffix it (`f"test_{i}"`).
`description` is **required**, and `PhaseParams` rejects both a blank one and one that merely restates the name. It is the single line of intent the trace, the console, and the UI phase block show, so write what the phase does and why — `"Land the code only now: green suite, approved review"`, not `"Commit build"`.
A code phase does its work in the block body and logs what it did. The commit phase that closes `adw_plan_build.py` and `adw_plan_build_test.py` is the pattern:
```python
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Land the builder's changes, using the message it wrote")) as ph:
message = build.commit_message or f"sssf({run.adw_id}): {build.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
```
`commit_message` is a field on `PlanOutput`, `BuildOutput`, and `DocumentOutput` that the agent fills in **for its own work product**, so always pair it with a fallback — it defaults to empty. `commit_all` raises if the cwd is not a git repo or nothing changed, which fails the phase rather than committing nothing. A chain that commits more than once (`adw_simple_sdlc.py`) commits each product with its own author's message.
## Remove a phase
Delete the block, drop any now-unused agent from `REQUIRED_AGENTS`, and re-thread the chain: whatever the removed phase produced was probably somebody's `previous=`. Point that call at the surviving upstream envelope.
## Add gates
Gates are callables over the finished envelope — `gate(envelope, run) -> GateReport`, recording one `check(item, ok, note)` per thing they looked at, with violations derived from the failed ones. Compose them per call:
```python
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.artifacts_exist, gates.diff_matches_claims]))
```
On violations the harness does **not** restart the agent — it sends the violation list back into the **same session** as a correction (pi's `--session-id` creates-or-continues, so the context window is intact), bounded by that phase's `retries`. Every gate result is traced to the `gate_results` table. Exhausting the retries raises `GateFailure` and fails the phase.
Gate claims, not guesses: declared artifacts exist and are non-empty, declared JSON parses, declared changes appear in the diff, declared test commands pass. Never hardcode counts — express quantity as a property of the declared list ("at least one artifact", "ALL declared paths valid"). Plan quality and code taste are not gateable; that is a reviewer agent or a human. New reusable gates go in `adw_modules/gates.py` (`update_modules.md`).
## Add a bounded fix loop
The pattern from `adw_build_test.py` — always bounded by a module-level constant. The runner is a **code** phase, because the command is known; only repairing it needs an agent:
```python
MAX_FIX_LOOPS = 3
test = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"test_{i}", kind="code", owner="quality",
description="Run the suite — a known command, so code runs it")) as ph:
test = quality.run_tests(run) # QualityResult, not an envelope
ph.log(passed=test.passed, artifacts=", ".join(test.artifacts))
if test.passed:
break
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description="Repair what the suite reported, from its verbatim output")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(test, "tests"),
gates=[gates.diff_matches_claims]))
return run.finish(accepted=test is not None and test.passed,
reason=f"the suite still failed after {MAX_FIX_LOOPS} fix attempt(s)")
```
`run.finish()` ends every ADW, and it takes the acceptance criterion the phase
statuses cannot express. A test phase that ran a red suite **succeeded** — the
runner did its job — so phases alone would report a green run that never passed
its tests, in the db and the UI as well as the terminal. Pass `accepted=` and
the exit code, the session status, and the banner are decided together.
`quality.as_envelope` is the adapter: a deterministic result shaped as an envelope, so the builder cannot tell it came from code. Wire the real command in `quality.py` first — the stamped blocks are `echo` placeholders that announce themselves.
Three distinctions worth keeping straight:
- **Gate retries vs. JSON retries.** `retries` buys extra *gate*-correction rounds. Malformed final JSON is handled separately and always — `JSON_FIX_ATTEMPTS` in `adw_modules/agents.py` (2 by default) re-prompts the same session for a valid object even on a phase with `retries=0`. Raising the phase's `retries` does not buy more JSON attempts, and vice versa.
- **Phase retries vs. fix loops.** `retries=N` on `PhaseParams` re-attempts one agent phase's gate corrections, re-sent into the same session with its context intact. (Code-phase re-execution is not implemented in v1.) A fix loop is a *chain* of phases repeated — different agents, new envelopes each pass.
- **The test phase succeeds when it runs and reports correctly.** A failing suite does not fail that phase; it fails the run, checked at the end. The runner did its job; the code didn't.
## Keep scripts thin
An ADW is sequencing and acceptance — nothing else. The moment you are writing parsing, subprocess handling, retry mechanics, or a reusable predicate inside `adw_*.py`, it belongs in `adw_modules/`. See `update_modules.md`.

View file

@ -0,0 +1,103 @@
# Update Config
Add or retune agents in `sssf.config.yaml`.
## Retune model or thinking
Edit the agent's entry in place:
```yaml
- name: builder
model: google/gemini-3.6-flash # ALWAYS provider/model-id
thinking: high # was medium
```
Write the model as `provider/model-id`, never a bare id. The same model is usually carried by several providers, and an ambiguous pattern raises in `agents.validate()` — grounding every agent that inherits it. See `references/config.md`.
Thinking levels are Pi's reasoning effort: `off | minimal | low | medium | high | xhigh | max`. It only bites when the model is registered with `reasoning: true` in `~/.pi/agent/models.json`.
**A model change means a fresh session.** `agent_map.json` records the model each coding-agent session was created with. When a joined run (`--adw-id`) finds the config's model no longer matches the recorded one, that agent starts a **new** session rather than resuming — the map is updated, never a bad resume. Thinking changes do not invalidate a session; model changes do. Expect the agent to lose its accumulated context window on the first run after the change.
## Recolor an agent's lane
```yaml
- name: builder
color: "#22d3ee" # hex; the starter roster ships violet/cyan/amber/green
```
Purely cosmetic and safe to change mid-project: the color rides the `agent_start` event and the `agent_sessions` row, so the visualizer picks it up on the next run without touching past sessions. Omit the key to let the UI's fallback palette choose.
## Retune tools
Pi's seven builtins: `read`, `bash`, `edit`, `write`, `grep`, `find`, `ls`. The last three are **off in bare Pi**, so an agent that doesn't name them will shell out through `bash` to search and list.
Set the roster-wide floor in `defaults`, then narrow per agent:
```yaml
defaults:
tools: [read, bash, edit, write, grep, find, ls]
agents:
- name: reviewer
tools: # explicit list wins over defaults
- read
- grep
- find
- ls
- bash
- write
```
**Resolution:** the agent's own list wins → else it inherits `defaults.tools` → else `None`, meaning all tools. An empty list is not "all tools"; it is a tool-less agent, and it will stall.
Narrow by role, not by reflex:
- Any agent that must produce a `context_handoff/` artifact needs **`write`** — without it, it falls back to a `bash` heredoc to create the file the gate checks for.
- Withhold `edit`/`write` only where the restriction *is* the guarantee. The reviewer's contract is "change nothing", so withholding `edit` makes that structural instead of merely prompted.
- Recon agents should get the full read surface (`read`, `grep`, `find`, `ls`) — cheaper and more legible in the trace than the equivalent `bash` calls.
**Extension tools count against the allowlist.** `--tools` filters built-in, extension, and custom tools alike. Once an agent has a `tools` list — its own, or inherited from `defaults` — a tool registered by one of its `harness_engineering` extensions is dropped unless it is named there. Nothing errors: the extension loads, the run passes, the tool is just never offered. Any agent with a tool-registering extension must list that tool by name.
## Add harness extensions
```yaml
harness_engineering:
- .pi/extensions/json_guard.ts # a pi extension FILE PATH
```
Entries are pi extension **file paths**, passed through as `pi -e <path>`, applied to that agent only. Reach for an output-tightening extension when an agent keeps wrapping its envelope in prose and burning correction retries. The starter roster ships with none — this is an escape hatch, not a default.
**Adding a tool-registering extension is a two-part edit.** The extension path goes in `harness_engineering`, *and* the tool name it registers goes in that agent's `tools` list:
```yaml
- name: reviewer
harness_engineering:
- .pi/extensions/ast_query.ts # registers tool: ast_query
tools:
- read
- grep
- find
- ls
- bash
- ast_query # REQUIRED — or the extension loads and its tool is filtered out
```
Skip the second half and it fails silently: extension loaded, run green, tool never available to the model. Extensions that only shape output or register flags — no new tool — need no `tools` change.
## Add a new agent
Three steps, all required — skipping any one fails `agents.validate()` at ADW startup, before anything spawns:
1. **Prompts.** Create `adws/adw_data/prompt_engineering/{name}/system.md` (Purpose + Instructions — the agent's static identity, nothing else) and `user.md` (an h3 per incoming datum: `{{prompt}}`, `{{previous_envelope}}`, `{{context_handoff_dir}}`, then the task, then a `## Report` section showing the exact output JSON). Copy an existing pair as the shape.
2. **Config entry.** Name, purpose, prompt refs, plus anything that differs from `defaults`.
3. **An output type.** Every agent call parses against a concrete Pydantic model in `adw_modules/data_types.py`. If none of `PlanOutput`, `BuildOutput`, `ScoutOutput`, `ReviewOutput`, `DocumentOutput` fits the new agent's report, add one — see `update_modules.md`. The user prompt's `Report` section must show exactly that JSON shape.
Then name the agent in an ADW's `REQUIRED_AGENTS` and call it.
## Rules that do not bend
- ADW scripts name **agents**, never models. Swapping a model is a config edit and touches no Python.
- One agent, one prompt, one purpose. If an entry needs two purposes, it is two agents.
- Output types never appear in config — they live at the call site, paired with the user prompt.
Full spec: `references/config.md`.

View file

@ -0,0 +1,100 @@
# Update Modules
Extend `adws/adw_modules/` with new low-level logic.
## The rule
**ALL low-level logic lives in `adw_modules/`; ADW scripts stay thin.** An `adw_*.py` file declares agents, sequences phases, and returns an exit code. Anything else — subprocess handling, parsing, retry mechanics, git plumbing, reusable predicates — goes in a module.
## Where things go
| Module | Owns |
|---|---|
| `data_types.py` | Every Pydantic model: `AgentCall`, `PhaseParams`, `Phase`, `EnvelopeBase` + one output type per agent call, the config models (`AgentConfig`, `SSSFConfig`), `EventRecord`, and `PiRequest`/`PiResult` |
| `agents.py` | `load_config`, `validate`, resolving an entry → coding-agent interface + model + thinking + harness extensions |
| `runner.py` | the `Run` object; `run.phase(PhaseParams)` context manager; `ph.call(AgentCall)` |
| `agent_pi.py` | the Pi interface (v1) — non-interactive `pi -p --mode json`, JSONL stream tailed live, model resolved against `~/.pi/agent/models.json`; `--session-id` creates-or-continues, so running and continuing an agent are the same call |
| `agent_cc.py` | the Claude Code interface — stubbed in v1, lands in v2 |
| `gates.py` | validation gates over envelope claims |
| `changes.py` | deterministic change capture: resolve the base ref, `git diff` into `context_handoff/changes.diff`, adapt the `ChangeSet` into an envelope an agent can be handed |
| `prompts.py` | load system/user prompt refs from config, render placeholders |
| `session.py` | mint or join `adw_id`, maintain `agent_map.json`, create session dirs incl. `context_handoff/` |
| `tracer.py` | append JSONL **and** insert every event into `sssf.db` as it happens |
| `console.py` | the terminal narrative — every line printed also lands in the db as a `log` event, so the UI reads the same story; plain sequential lines, no spinners |
| `console.py` | the rich stdout reporter — every line printed is ALSO traced as a `log` event (`{message, level}`) so the terminal and the swim-lane UI tell the same story |
| `git_helper.py` | branch, status, diff, commit — the raw plumbing `changes.py` composes |
| `utils.py` | safe subprocess env, logging, `resolve_prompt` |
## Never `print()`
Modules report through `run.console` — never a bare `print()`. Each console method prints a rich line **and** writes it to `sssf.db` as a `log` event with payload `{message, level}`, both from one `_emit` helper, so the terminal narrative and the swim-lane UI can't drift. New output means a new method on `Console`, not a print at the call site.
## The four-param rule
**Any function taking more than 4 parameters gets them converted into a concrete data type in `data_types.py`.** `AgentCall` and `PhaseParams` are the pattern — `run.phase()` and `ph.call()` each take exactly one object. This is skill-wide: every module the factory generates obeys it.
```python
class ReviewParams(BaseModel):
"""Everything review_changes() needs. Passed as one object, never loose params."""
base_ref: str
paths: list[str]
max_diff_lines: int = 2000
ignore_generated: bool = True
reviewer: str = "scout"
```
## Adding an output type
Every agent call parses against a concrete type. Extend `EnvelopeBase` — `status`, `summary`, `artifacts`, `notes_for_next_agent` — with only the fields that call actually needs:
```python
class ReviewOutput(EnvelopeBase):
approved: bool
blocking: list[str] = []
```
**The output contract is a synced triad — one change means three edits, always together:**
1. The type in `data_types.py` (the enforcer).
2. The agent's `user.md` `## Report` section showing exactly that JSON (the ask).
3. Every call site passing `output_type=` (the binding) — `grep -rn "ReviewOutput" adws/` to find them all.
If the type and the Report example drift, the agent produces what the prompt asked for, the parser rejects what the type expects, and every call burns correction round-trips before landing — a slow, silent tax. Renaming or removing a field is the same triad edit. Schema details: `references/handoff.md`.
## Adding a gate
A gate is a callable — `gate(envelope, run) -> GateReport`. You record **one check per item you look at**, and the harness derives the verdict: any failed check is a violation, and no failed checks means pass.
```python
from adw_modules.data_types import GateReport
def tests_declared_passed(envelope, run) -> GateReport:
"""Verify the envelope's own test claims, after the fact."""
report = GateReport()
for f in envelope.failures:
report.check(f.test, False, f.error)
report.check("suite", envelope.passed,
"all declared tests passed" if envelope.passed
else f"{len(envelope.failures)} declared failure(s)")
return report
```
`report.check(item, ok, note)` appends and returns the report, so a single-item gate is one line: `return GateReport().check(command, ok, f"exit {code}")`.
**Write a note on passing checks too, not just failures.** The note is the evidence, and it is what makes a green gate worth reading — `artifacts_exist ✓ 1 checked · plan.md — exists, 454B` tells you what was verified, where a bare ✓ tells you nothing. Notes on failed checks double as the reason and are what the agent is told, so phrase them as the problem: `"claimed changed file does not exist"`.
Rules that keep gates honest:
- **Verify claims, never predict.** File names and counts are unknowable before the agent finishes; gates check what the envelope declared.
- **Quantity as properties, not counts.** "at least one artifact", "ALL declared paths exist" — never `len(artifacts) == 3`.
- **Record checks, don't raise.** The harness feeds the derived violations back into the same session as a correction — context intact, bounded by the phase's `retries` — and traces every check, passed or failed, to `gate_results.checks_json` and the `gate_pass`/`gate_fail` event payload.
- **Check every item, even after one fails.** Don't early-return on the first problem; the agent fixes more per correction round when it sees every failure at once, and the trace shows the full picture.
- **Don't gate the ungateable.** Plan quality and code taste are a reviewer agent's job or a human's.
A gate that returns a plain `list[str]` of violations still works — the harness adapts it — but it records no evidence for the items that passed, so prefer a `GateReport`.
Reusable gates live in `gates.py`; genuine one-offs can be defined inline at the ADW call site and passed in `gates=[...]`.
## Before you finish
Run the smoke ADW — `uv run adws/adw_prompt.py "ping"` — since every module change rides the same path a real run does.

202
sssf/references/config.md Normal file
View file

@ -0,0 +1,202 @@
# Config Reference
The full `sssf.config.yaml` spec: every field, how defaults merge, and how model / thinking / tools / extensions map onto the coding agent.
It lives at **`adws/adw_sssf_config/sssf.config.yaml`** — the default path every `adw_*.py` and the justfile resolve, and where `install.py` / `make_config.py` stamp it. Pass `--config <path>` to any ADW (or set `SSSF_CONFIG` for the justfile) to run against a different roster.
## Shape
```yaml
defaults:
coding_agent: pi
model: google/gemini-3.6-flash # ALWAYS provider/model-id
thinking: medium
harness_engineering: []
tools: [read, bash, edit, write, grep, find, ls]
data_dir: adws/adw_data
observability:
db: adws/adw_data/sssf.db
poll_ms: 500
agents:
- name: planner
coding_agent: pi
model: google/gemini-3.6-flash # ALWAYS provider/model-id
thinking: high
color: "#a78bfa"
purpose: Turn a request into a plan the builder can implement without asking questions.
prompt_engineering:
system: adws/adw_data/prompt_engineering/planner/system.md
user: adws/adw_data/prompt_engineering/planner/user.md
harness_engineering:
- json-enforcer
tools:
- read
- bash
```
## Fields
### `defaults`
| Field | Type | Meaning |
|---|---|---|
| `coding_agent` | `pi` \| `claude_code` | Which interface runs the agent. **v1 implements `pi` only**; `claude_code` is specced and stubbed in `agent_cc.py`, landing in v2. |
| `model` | string | Model id. For Pi, any id registered in `~/.pi/agent/models.json`. Default `gemini-3.6-flash`. |
| `thinking` | enum | Reasoning effort — see below. Default `medium`. |
| `color` | hex string | Lane color for every agent that does not set its own. Default empty — the visualizer falls back to its own palette. |
| `harness_engineering` | list[string] | Coding-agent extensions. Pi: extension names. Claude Code: reserved (MCP, hooks). |
| `tools` | list[string] | Roster-wide tool allowlist. Every agent that omits its own `tools` inherits this. Unset = all tools usable. |
| `protected_files` | list[string] | Paths **no** agent may modify unless it names them in its own `writes`. Default: `adws/adw_modules/`, `adws/adw_sssf_config/`, `adws/adw_*.py` — an agent must not be able to edit the machinery that decides whether its work passed. |
| `data_dir` | path | Runtime home. Sessions land at `{data_dir}/sessions/{adw_id}/{agent_name}/`. Default `adws/adw_data`. |
### `observability`
| Field | Type | Meaning |
|---|---|---|
| `db` | path | SQLite trace db. `tracer.py` writes it directly; the visualizer polls it. Default `adws/adw_data/sssf.db`. |
| `poll_ms` | int | Visualizer live-poll cadence in ms. History uses the same queries, lazy-paged. Default `500`. |
### `agents[]`
| Field | Required | Meaning |
|---|---|---|
| `name` | yes | The identifier ADW scripts use. **ADWs name agents, never models.** |
| `purpose` | yes | One sentence: what this agent is for. Should match its `system.md` Purpose. |
| `prompt_engineering.system` | yes | Path to the system prompt — who the agent is, its single purpose, its output contract. |
| `prompt_engineering.user` | yes | Path to the default user prompt — the task template with `{{prompt}}`, `{{previous_envelope}}`, `{{context_handoff_dir}}`. |
| `color` | no | Hex swatch (`"#a78bfa"`) for this agent's lane in the visualizer. Travels config → `agent_sessions.color` → `/api/sessions/:adw_id`, and rides the `agent_start` event so a lane is colored while the agent is still running. Unset = the UI's fallback palette. |
| `coding_agent`, `model`, `thinking`, `color`, `harness_engineering` | no | Override the corresponding `defaults` key. |
| `tools` | no | Allowlist. **Omitting the key means all tools usable.** A capability list, not a boundary — see `writes`. |
| `writes` | no | What this agent may modify **in the repo**, enforced after every call. Omitted = unrestricted (still barred from `protected_files`). `[]` = no repo writes at all. A list = only those paths: a trailing `/` is a directory prefix, `*` matches within one path segment, `**` crosses segments, anything else is an exact path. Naming a `protected_files` path here is what unlocks it. **The session runtime under `data_dir` is always writable** — `writes: []` means read-only with respect to the repo, not unable to write its own report. |
Output types are deliberately absent: config defines who an agent *is*; the ADW call site defines how it's *used*. One agent serves many calls — same system prompt, different user prompt + output type per call.
## Defaults merging
`agents.py` merges each entry **over** `defaults`, key by key. An entry states only what differs; anything unset inherits. `agents.validate(cfg, REQUIRED_AGENTS)` then confirms every name an ADW declares exists, resolves to a usable coding agent + model, and has both prompt files present on disk. Any miss fails the run immediately — **no agent is ever spawned against a half-valid config.**
## Thinking levels
Pi's reasoning-effort ladder, lowest to highest:
```
off | minimal | low | medium | high | xhigh | max
```
Mapped to Pi's reasoning effort control and honored when the model is registered with `reasoning: true` in `~/.pi/agent/models.json`. On a non-reasoning model the setting is inert — no error, no effect. Rough guidance: `high`/`xhigh` for planners and reviewers, `medium` for builders, `low` for mechanical read-and-report agents. (For Claude Code in v2, the same field maps to the thinking budget.)
## Model resolution
**Always write `model` as `provider/model-id`.** `agents.py` hands the string to the Pi interface, which resolves it against pi's merged catalog — `~/.pi/agent/models.json` plus pi's built-in providers. The same model is usually carried by more than one provider (`gemini-3.6-flash` lives under `google` *and* under `openrouter` as `google/gemini-3.6-flash`), and a bare id that matches several **raises at resolution**:
```
agent 'scout': model pattern 'gemini-3.6-flash' is ambiguous:
[('google', 'gemini-3.6-flash'), ('openrouter', 'google/gemini-3.6-flash'), ...]
```
That is `agents.validate()` doing its job — it fails before anything spawns rather than silently billing the wrong provider — but it means every agent in the roster inheriting that default is grounded until the pattern is qualified. Qualifying is the whole fix: `google/gemini-3.6-flash`, `openai/gpt-5.6-terra`, `fireworks/accounts/fireworks/models/kimi-k3`. The leading segment is matched against the provider list first, so the rest of the string can contain slashes.
Other consequences worth knowing:
- A model must be in the catalog before any agent can name it. An unknown id fails at resolution, before spawn. `pi --list-models` is the catalog the resolver actually reads.
- **Ambiguity can appear without you touching the config.** Registering a new provider that carries a model you already use turns a formerly-fine bare pattern ambiguous. If a roster stops validating and nobody edited it, that is why.
- Provider credentials come from the environment, not the config — the key that matches the provider you named (`GEMINI_API_KEY` for `google/...`, `OPENROUTER_API_KEY` for `openrouter/...`).
- The resolved model is recorded per session in `agent_map.json` and mirrored into the `agent_sessions` table. **Changing an agent's model invalidates its session**: a joined run starts that agent fresh instead of resuming a context window built by a different model.
## Tools
`tools` maps to `pi --tools`. Pi's seven builtin tool names:
| Tool | Purpose | Pi's own default |
|---|---|---|
| `read` | read file contents | on |
| `bash` | execute bash commands | on |
| `edit` | find/replace edits | on |
| `write` | create/overwrite files | on |
| `grep` | search file contents | **off** |
| `find` | find files by glob | **off** |
| `ls` | list directory contents | **off** |
`grep`, `find`, and `ls` are off in bare Pi, so an agent that does not name them will shell out through `bash` to do the same work. The starter roster therefore sets `defaults.tools` to all seven and lets each agent narrow from there.
**Resolution order:** an agent's own `tools` list wins; an agent that omits the key inherits `defaults.tools`; if neither is set, `tools` stays `None` and all tools are usable. An empty list is not "all tools" — it is a tool-less agent, and it will stall.
## Write permissions — `writes` and `protected_files`
`tools` cannot express a safety boundary, because two of the tools are general
purpose. `bash` runs anything, including `git checkout`, which discards an
engineer's uncommitted work; `write` reaches any path, not only the one report
file an agent was granted it for. So "this agent changes nothing" is a claim a
tool list can state but never keep.
`adw_modules/permissions.py` keeps it, the same way every other claim in this
system is kept — after the fact, against the repo. Before an agent's first
prompt the working tree's change-set is fingerprinted; after its last send
(including JSON retries and gate corrections) it is fingerprinted again. Any
path that appeared, vanished, or changed is attributed to that agent.
Comparing change-sets rather than watching writes is deliberate: a path that was
modified before the agent ran and is clean afterwards has been **reverted**, and
a reversion is a modification. That is what catches `git checkout`.
A breach is not a gate violation. Gates are for work an agent can be asked to
redo; a write has already happened, so re-prompting fixes nothing. Instead:
1. every unauthorized change the agent **introduced** is rolled back — tracked
files with `git checkout --`, untracked files by deletion;
2. a path that was **already dirty** before the agent ran is left untouched. The
operator had uncommitted work there, and discarding it to tidy up would be
the same harm this module exists to prevent;
3. the phase fails and names every path with what happened to it.
```yaml
defaults:
protected_files: [adws/adw_modules/, adws/adw_sssf_config/, "adws/adw_*.py"]
agents:
- name: builder # no `writes` key -> unrestricted, minus protected_files
- name: scout
writes: [] # no repo writes; its findings still land in context_handoff/
- name: planner
writes: [specs/]
- name: documenter
writes: [app_docs/, docs/, "**/*.md", "*.md"]
```
**The session runtime under `data_dir` is always writable, for every agent.**
`context_handoff/` is how agents hand work to each other, and each agent's
prompts, `raw_output.jsonl`, and `envelope.json` sit beside it. That grant comes
from `data_dir` rather than from `.gitignore`: the runtime is normally ignored,
so it never even appears in a snapshot, but an agent's ability to record its own
work must not depend on a gitignore line someone can delete.
Narrow by role, not by reflex. Anything that must produce a `context_handoff/` artifact needs `write`, or it will resort to a `bash` heredoc. Withhold `edit`/`write` only where the restriction *is* the guarantee — a reviewer that cannot edit cannot quietly fix what it was asked to report.
### Extension tools must be named explicitly
`pi --tools` is an allowlist over **built-in, extension, and custom tools alike** — not just builtins. So the moment an agent has a `tools` list at all (its own, or one inherited from `defaults`), any tool registered by its `harness_engineering` extensions is **excluded unless it appears in that list by name**.
This fails quietly. The extension still loads, the run still succeeds, and the tool the extension exists to provide is simply never offered to the model — you find out by noticing the agent never called it.
```yaml
- name: reviewer
harness_engineering:
- .pi/extensions/ast_query.ts # registers tool: ast_query
tools:
- read
- grep
- find
- ls
- bash
- ast_query # REQUIRED — the extension's tool, named or lost
```
Rule: **every entry in `harness_engineering` that registers a tool must have that tool name added to the agent's `tools` list.** Adding an extension is therefore a two-line change, never one. The alternative is dropping the `tools` key *and* leaving `defaults.tools` unset so the agent resolves to `None` (all tools) — but with a roster-wide `defaults.tools` in place, that escape hatch is closed; naming the tool is the only path.
## Harness engineering
`harness_engineering` entries are pi extension **file paths**, passed through as `pi -e <path>`, one flag per entry, scoped to that agent only. This is where per-agent harness changes live — e.g. an output-tightening extension for an agent that keeps wrapping its envelope in prose. The starter roster ships with none. On Claude Code the field is reserved for MCP config and hooks in v2.
**If the extension registers a tool, name that tool in the agent's `tools` list too** — `--tools` filters extension tools exactly like builtins, so an unnamed extension tool is silently unavailable no matter that the extension loaded fine. See [Extension tools must be named explicitly](#extension-tools-must-be-named-explicitly) above. Extensions that only shape output or add flags (no tool registration) need no `tools` change.

161
sssf/references/handoff.md Normal file
View file

@ -0,0 +1,161 @@
# Handoff Reference
The envelope schema, the two-channel output contract, and the session directory layout — how context transfers in code, not in conversation.
## Two output channels, exactly
An agent may produce output in two ways and no others:
1. **Reference files** written into `context_handoff/` — plans, notes, artifacts for the agents that follow.
2. **A final valid-JSON response** — the envelope, its direct response and nothing else.
Code does the rest: parse the response against the output type the call declared, persist it as `envelope.json`, and inject it into the next agent's user prompt.
## Envelope schema
Every output type extends `EnvelopeBase`:
```python
class EnvelopeBase(BaseModel):
status: Literal["success", "fail"] # the only required field
summary: str = "" # one sentence: what happened
artifacts: list[str] = [] # paths written, usually inside context_handoff/
notes_for_next_agent: str = "" # what the next agent must know
```
`status` is load-bearing: an envelope that parses but reports `status="fail"` raises, failing the phase. An agent declaring its own failure is not a successful phase.
The starter types in `adw_modules/data_types.py`:
```python
class GenericOutput(EnvelopeBase):
"""Fallback for an agent with no sharper contract yet."""
class PlanOutput(EnvelopeBase):
commit_message: str = "" # imperative git subject for the PLAN FILE itself
class BuildOutput(EnvelopeBase):
changed_files: list[str] = []
commit_message: str = "" # consumed by the git commit phase
class ScoutOutput(EnvelopeBase):
findings: list[ScoutFinding] = [] # ScoutFinding: {file: str, note: str}
class ReviewOutput(EnvelopeBase):
approved: bool = False # the verdict; status is only "did the review run"
findings: list[ReviewFinding] = [] # ReviewFinding: {requirement, met: bool, evidence}
blocking: list[str] = [] # what must change before approval
class DocumentOutput(EnvelopeBase):
document_path: str = "" # the write-up's home in the repo
documented_files: list[str] = []
commit_message: str = ""
```
`commit_message` defaults to empty, so a git phase consuming it always needs a fallback — see `cookbooks/create_adw.md`.
**Each `commit_message` describes its own agent's work product, never the next one's**: `PlanOutput`'s covers the spec file, `BuildOutput`'s the code, `DocumentOutput`'s the write-up. A chain that commits once can use whichever fits; a chain that commits per step (`adw_simple_sdlc.py`) needs all three, and reusing one agent's sentence for another's diff is how a commit log starts lying.
There is no test output type: running the suite is a `kind="code"` phase, and its `QualityResult` reaches the next agent through `quality.as_envelope`.
Two of these are adapters rather than agent reports — code shaped as an envelope so an agent can be handed a deterministic result through the same door: `VerifyOutput` (a lint/test block's result) and `ChangesOutput` (a captured `git diff`, from `changes.as_envelope`). The consuming agent cannot tell the difference, which is the point.
The envelope is a **manifest of claims**. Gates verify those claims after the fact — declared artifacts exist and are non-empty, declared changes appear in the diff, declared tests actually pass. See `cookbooks/update_modules.md`.
## The typed-output rule
**Every agent call passes a concrete output type**, and the agent's final JSON is parsed against exactly that type. No untyped handoffs.
```python
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist]))
```
The user prompt asks for the shape; the type enforces it. They always travel as a pair, which is what lets one agent serve many calls — same system prompt, different user prompt + output type per call site. Output types live in code, never in `sssf.config.yaml`.
**Parse failure is not a restart.** If the response doesn't parse or doesn't validate, the harness re-prompts the **same session** with a correction naming the required fields — bounded by `JSON_FIX_ATTEMPTS` in `agents.py` (2). Gate violations use the identical mechanism, bounded instead by the phase's `retries`. A cold restart would throw away the context that produced the near-miss.
In v1 there is no separate continue call to make: `agent_pi.run()` passes `--session-id`, which pi treats as create-or-continue, so running an agent and continuing it are the same call with the same id. Before parsing, the harness also tolerates a fenced `json` code block or prose wrapped around the object — but the prompt still asks for bare JSON, and every failed attempt is persisted as an invalid envelope row.
## Injecting the previous envelope
`prompts.py` renders the agent's `user.md`, substituting:
| Placeholder | Value |
|---|---|
| `{{prompt}}` | the engineer's ask (or the ADW's per-call prompt) |
| `{{previous_envelope}}` | the upstream envelope JSON, from `AgentCall(previous=...)` |
| `{{context_handoff_dir}}` | absolute path to this session's `context_handoff/` |
A `user.md` declares one h3 per incoming datum, then the task, then the output contract:
````markdown
# Scout Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Find what `prompt` asks about. Write findings into `context_handoff_dir`, then emit your `Report` JSON.
## Report
Respond with ONLY valid JSON matching `ScoutOutput` — no prose before or after:
```json
{
"status": "success",
"summary": "<one sentence on what you found>",
"findings": [
{ "file": "src/server.ts", "note": "<why this file matters>" }
],
"artifacts": ["<context_handoff_dir>/scout_findings.md"]
}
```
````
The `## Report` section shows the exact JSON shape of the declared output type — that is the agent's output contract, and it lives in `user.md` because the shape belongs to the *use*, not the identity. The matching `system.md` stays static: Purpose + Instructions only.
## Session directory layout
```
adws/adw_data/sessions/{adw_id}/
├── agent_map.json agent name → coding-agent session_id + model
├── context_handoff/ the ONE place agents write files for the agents that follow
└── {agent_name}/
├── prompts/ exact prompts sent (system.md + user.md), saved before execution
├── pi_sessions/ pi's own session state for this agent
├── raw_output.jsonl full JSONL stream from the coding agent, appended live
└── envelope.json the final valid-JSON response — captured, validated, persisted by code
```
`session.ensure(cfg, adw_id)` mints or joins the id and creates these dirs. One `context_handoff/` per session, shared by every agent — the single location for cross-agent files.
## agent_map.json and resuming
```json
{
"planner": {"session_id": "sssf-a1b2c3d4-planner-9f2e",
"model": "google/gemini-3.6-flash", "coding_agent": "pi"},
"builder": {"session_id": "sssf-a1b2c3d4-builder-71ac",
"model": "google/gemini-3.6-flash", "coding_agent": "pi"}
}
```
This map is the key that lets a later ADW rejoin each agent's **existing context window**. Run `adw_build.py --adw-id a1b2c3d4` after `adw_plan.py` and the builder resumes its own session rather than starting cold.
The map records the model each session was created with. If config drift changes an agent's model, that agent starts a **fresh** session and the map is updated — never a bad resume. `agent_sessions` in `sssf.db` is the queryable mirror of this file.
**Files are the raw record; the db is the queryable mirror.** Losing `sssf.db` loses nothing that can't be rebuilt from `raw_output.jsonl`, `envelope.json`, and `agent_map.json`.

View file

@ -0,0 +1,158 @@
# Observability Reference
The event schema, the seven SQLite tables, and the polling contract — the one data path is **agents → sqlite → web ui**.
## Two stores, one truth
**Files are the raw record** (`raw_output.jsonl` streams, `envelope.json`, `agent_map.json`); **SQLite (`sssf.db`) is the queryable mirror** the UI reads. `tracer.py` writes both. Losing the db loses nothing that can't be rebuilt from files.
Location comes from `observability.db` in `sssf.config.yaml`, default `adws/adw_data/sssf.db` — inside the **target** repo, gitignored.
## Event schema
`tracer.py` emits these types, every one logged against its `adw_id` **and** `phase_id`:
| Type | Emitted when |
|---|---|
| `phase_start` | a `run.phase(...)` block is entered |
| `agent_start` | a coding agent is spawned or resumed for `ph.call(...)` |
| `tool_call` | a tool (`read`, `bash`, `edit`, `write`) returns — **one event per real call**, named `bash: ls -la src`, payload `{tool, tool_call_id, args, result_snippet, ok, duration_ms, agent}` |
| `handoff` | an envelope crosses from one agent to the next |
| `gate_pass` | a gate found no failed checks — payload carries `attempt`, `checks` (the evidence), and an empty `violations` |
| `gate_fail` | a gate found at least one failed check — payload carries `attempt`, `checks`, and `violations` |
| `log` | an explicit `ph.log(...)` from the ADW script |
| `agent_end` | the agent's run completes; envelope parsed or not — payload carries `cost`, `usage` (the per-component breakdown), `context_tokens`, `context_window` |
| `phase_end` | the block exits; carries the resolved status |
| `error` | a raise inside a phase block |
`parent_id` nests spans, so an agent phase expands into its tool-call spans in the UI.
**Spend is itemised per phase.** `agent_end.usage` carries tokens *and* dollars for each component pi reports — `input`, `output`, `cache_read`, `cache_write` — summed across every send the phase made, so a phase that retried on a bad envelope or a failed gate shows what all its attempts cost, not just the last one. The four components sum to `total_tokens`, and their costs sum to `total_cost`; the visualizer's Cost panel renders them as a table you can add up by eye.
`reasoning_tokens` is the thinking share and is **inside** `output_tokens`, not a fifth component — measured across every session on disk, reasoning never exceeds output and the four components always reconcile to the total. It bills at the output rate, so the panel nests it under output rather than adding it. Runs predating the breakdown have no `usage` key at all; the lump `cost` and the event's own `tokens` still stand, and the UI says so rather than rendering zeroes.
**Context is occupancy, not spend.** `events.tokens` and `sessions.total_tokens` bill every turn, so they only grow — an agent that burned 100k tokens may be sitting in a 15k window. `context_tokens` is how full the window actually was when the agent stopped, which is what the visualizer's per-lane Context bar measures against `context_window`.
It is computed the way pi computes it for its own footer and its auto-compaction trigger (`calculateContextTokens` in the coding agent's `core/compaction/compaction.ts`): take the last *valid* assistant turn — skipping `aborted` and `error` turns — and read `usage.totalTokens`, falling back to `input + output + cacheRead + cacheWrite`. Cache reads count; cached prompt is still prompt. `context_window` is the same `contextWindow` pi reads from `~/.pi/agent/models.json`, so `context_tokens / context_window` is the number pi would show. Both are NULL on rows written before the columns existed, and the lane draws no bar rather than a misleading empty one.
Two caveats worth knowing. Pi adds an *estimate* for any messages trailing the last assistant usage; in a batch (`-p`) run the session ends on that message, so the two agree. And if auto-compaction fires as the very last act of a run, the recorded number is the pre-compaction size — pi itself reports `null` in that window rather than guessing.
**Gates record evidence, not just a verdict.** A gate returns one `{item, ok, note}` check per thing it looked at, and `violations` are derived from the failed ones. Both land in `gate_results` (`checks_json` + `violations_json`) and in the `gate_pass`/`gate_fail` payload, so a green gate can answer *what did you verify* — `{"item": "…/plan.md", "ok": true, "note": "exists, 454B"}` — rather than only *did it pass*. Rows written before this existed have `checks_json` NULL; treat that as "no evidence recorded", not "nothing checked".
The gate event payload carries `attempt` too, so the `gate_results` table and the event stream are equivalent sources — a live consumer can group gate results per correction round from events alone, without a second query.
**A `tool_call` is the one event that spans time**, so it fills both `started_at` and `ended_at` on the row — the tool's real start and return. Every other type is a point in time: `started_at` is when it was recorded and `ended_at` stays NULL. Lay tool calls out on a time axis from those columns, never by parsing `payload_json` (`duration_ms` is in the payload too, as pi's own number, but it is a convenience, not the source for layout).
**Streaming is solved by construction.** `agent_pi.py` tails pi's JSONL stdout line by line and the tracer inserts each event into `sssf.db` **while the agent is still working** — never batched at phase end (verified in the first smoke run: tool calls visible mid-run). Everything downstream is a poll → render.
## Tables
```sql
sessions (
adw_id TEXT PRIMARY KEY,
request TEXT, -- the engineer's ask
status TEXT, -- running | success | fail
engineer TEXT,
started_at TEXT, ended_at TEXT,
total_tokens INTEGER, total_cost REAL
);
phases (
phase_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
seq INTEGER,
name TEXT, kind TEXT, owner TEXT, description TEXT,
status TEXT DEFAULT 'fail', -- success must be earned
attempt INTEGER DEFAULT 0, retries INTEGER DEFAULT 0,
error TEXT,
started_at TEXT, ended_at TEXT
);
events (
event_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases, -- every event logs against adw + phase
parent_id TEXT, -- span nesting
type TEXT, -- phase_start | phase_end | agent_start | agent_end | tool_call
-- | handoff | gate_pass | gate_fail | log | error
name TEXT,
payload_json TEXT,
tokens INTEGER,
started_at TEXT, ended_at TEXT -- ended_at set only on events that span time
);
envelopes (
envelope_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases,
agent TEXT,
output_type TEXT, -- name of the data_types model it parsed against
payload_json TEXT,
valid INTEGER,
attempt INTEGER,
created_at TEXT
);
gate_results (
id INTEGER PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases,
attempt INTEGER,
gate TEXT,
passed INTEGER,
violations_json TEXT, -- derived: the failed checks, as "item: note"
checks_json TEXT, -- [{item, ok, note}] — everything the gate looked at
created_at TEXT
);
processes ( -- adw_id → pid, so a stuck run can be stopped
id INTEGER PRIMARY KEY AUTOINCREMENT,
adw_id TEXT REFERENCES sessions,
kind TEXT, -- 'adw' (the workflow process) | 'agent' (a coding-agent child)
name TEXT, -- '' for the adw, the agent name for a child
pid INTEGER,
command TEXT, -- what the pid WAS; pids get recycled, so verify before killing
started_at TEXT, ended_at TEXT -- ended_at NULL = believed alive
);
agent_sessions ( -- the queryable mirror of agent_map.json
adw_id TEXT REFERENCES sessions,
agent TEXT,
coding_agent TEXT, model TEXT, color TEXT, -- color: the config's lane swatch
session_id TEXT,
context_tokens INTEGER, -- window occupancy after the agent's last turn
context_window INTEGER, -- the model's ceiling, from the pi registry
created_at TEXT, last_used_at TEXT,
PRIMARY KEY (adw_id, agent)
);
```
**A hung agent emits nothing**, which is exactly when you need its pid: no events, no tokens, no output to read. `processes` is the only table that can answer "what is this run running, and how do I stop it" — `just procs <adw_id>` lists what is live, `just kill <adw_id>` stops children before the parent, and both verify the recorded `command` still matches the pid before signalling it. A killed run finalizes its own trace: SIGTERM and SIGINT are turned into `SystemExit` in `session.ensure`, so the session lands on `fail` with its process rows closed instead of reading `running` forever.
**Derived, never stored:** phase durations (`ended_at − started_at`), session phase-progress (query `phases` by `adw_id`), lane layout (`kind` + `owner`).
Phase status invariants: `queued` only for manifest-declared phases not yet entered (dashed in the UI); `running` on enter; only a clean exit writes `success` — agent phases additionally need the envelope parsed and gates green; everything else resolves to `fail`.
## WAL pragmas
Open **every** connection — writer and reader — with:
```sql
PRAGMA journal_mode=WAL;
PRAGMA synchronous=NORMAL;
PRAGMA busy_timeout=5000;
```
WAL allows readers during writes. Writers are the tracers of running ADW processes; concurrent writers are fine given one small transaction per event plus `busy_timeout`. The visualizer reads on a readonly connection with exactly one exception: archiving a session (`POST /api/sessions/:adw_id/archive`) opens a second connection to set `sessions.archived`. That flag is review triage — it says a human has looked at the run — so it is the reader's state living on the row, and no tracer ever writes or reads it.
## Polling contract
**The UI never receives pushes.** No ingest endpoint, no WebSocket, no backfill or dedup logic.
Live view polls on a rowid cursor every `observability.poll_ms` (default 500):
```sql
SELECT ... FROM events WHERE adw_id = ? AND rowid > ? ORDER BY rowid LIMIT 500;
```
Keep the highest `rowid` returned as the next cursor. History is **the same queries** with filters, lazy-paged as the engineer scrolls or drills in — one mechanism serves both live and past runs, which is why there is no separate replay path.

301
sssf/scripts/install.py Executable file
View file

@ -0,0 +1,301 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = []
# ///
"""/install — stamp the SSSF factory from the skill into the cwd. Idempotent.
Usage:
uv run <skill>/scripts/install.py [--force] [--harness pi|omp]
Stamps: adws/ (modules + starter ADWs), adws/adw_data/prompt_engineering/
(4 starter agents), adws/adw_sssf_config/sssf.config.yaml, .env.sample,
.gitignore entries.
Prompts for the coding-agent harness (pi or omp) unless --harness is given,
writes it into the stamped config's defaults, and registers the repo in the
visualizer's repos.json so the multi-repo UI can browse it.
Existing files are skipped unless --force.
"""
import argparse
import json
import shutil
import subprocess
import sys
from pathlib import Path
TEMPLATES = Path(__file__).resolve().parent.parent / "templates"
# The skill's own directory — where the visualizer app lives. install.py stamps
# it into the justfile's `skill_dir` so `just obs` finds the app no matter
# where the skill was installed (user scope, a repo, or via the skills CLI).
SKILL_DIR = Path(__file__).resolve().parent.parent
# The visualizer's repo registry lives next to the skill (machine-specific:
# it maps repo slugs to local db paths). install.py keeps it in sync.
REPOS_FILE = SKILL_DIR / "repos.json"
GITIGNORE_ENTRIES = [
"adws/adw_data/sessions/",
"adws/adw_data/sssf.db*",
".env",
# The ADWs are Python, so importing adw_modules writes bytecode next to it.
# Chains that end in a commit phase call `git add -A`, so without this a
# stamped repo commits its own .pyc files — 15 of them showed up in the
# first repo that was ever installed into from scratch.
"__pycache__/",
"*.pyc",
]
HARNESSES = ("pi", "omp")
def stamp(src: Path, dest: Path, force: bool, stamped: list, skipped: list) -> None:
if src.is_dir():
for child in sorted(src.iterdir()):
if child.name == "__pycache__":
continue
stamp(child, dest / child.name, force, stamped, skipped)
return
if dest.exists() and not force:
skipped.append(str(dest))
return
dest.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(src, dest)
stamped.append(str(dest))
def ensure_gitignore(root: Path, stamped: list) -> None:
gitignore = root / ".gitignore"
existing = gitignore.read_text().splitlines() if gitignore.exists() else []
missing = [e for e in GITIGNORE_ENTRIES if e not in existing]
if missing:
with gitignore.open("a") as f:
f.write("\n# sssf runtime\n" + "\n".join(missing) + "\n")
stamped.append(f"{gitignore} (+{len(missing)} entries)")
def stamp_justfile(root: Path, force: bool, stamped: list, skipped: list) -> None:
"""Stamp the justfile, substituting the skill's real path into `skill_dir`.
The `obs` recipe boots the visualizer app, which ships with the skill, so
the stamped justfile has to know where the skill lives. The template keeps
a `@SSSF_SKILL_DIR@` placeholder; this fills it with the actual install
location (user scope, a repo, or wherever the skills CLI put it).
"""
dest = root / "justfile"
if dest.exists() and not force:
skipped.append(str(dest))
return
text = (TEMPLATES / "justfile").read_text()
text = text.replace("@SSSF_SKILL_DIR@", str(SKILL_DIR))
dest.parent.mkdir(parents=True, exist_ok=True)
dest.write_text(text)
stamped.append(str(dest))
def choose_harness(arg: str | None) -> str:
"""Pick the coding-agent harness: --harness wins, else prompt."""
if arg:
if arg not in HARNESSES:
raise SystemExit(f"--harness must be one of {', '.join(HARNESSES)}, got {arg!r}")
return arg
while True:
choice = input(f"Coding-agent harness [{HARNESSES[0]}/{HARNESSES[1]}] (default {HARNESSES[0]}): ").strip().lower()
if not choice:
return HARNESSES[0]
if choice in HARNESSES:
return choice
print(f" {choice!r} is not a harness; pick {', '.join(HARNESSES)}.")
def set_harness_in_config(config_path: Path, harness: str) -> None:
"""Rewrite the stamped config's defaults.coding_agent to the chosen harness."""
text = config_path.read_text()
# Target the `defaults:` block's coding_agent, not the comment above it
# (which also contains the words "coding_agent"). The template ships
# ` coding_agent: pi` as the first line under `defaults:`.
defaults_marker = "defaults:"
d_idx = text.find(defaults_marker)
if d_idx == -1:
return
marker = "coding_agent:"
idx = text.find(marker, d_idx)
if idx == -1:
return
line_end = text.find("\n", idx)
text = text[:idx] + f"coding_agent: {harness}" + text[line_end:]
config_path.write_text(text)
def harness_catalog(harness: str) -> list[str]:
"""List the model selectors the harness can actually run, or [] if unknown.
The template config ships with a roster of models (google/gemini-3.6-flash,
fireworks/..., openai/...) that only resolve if the harness has them
registered. Most machines have a handful of local/cloud models, so the
stamped config usually fails validation until every agent is pointed at a
model that exists. Reading the harness's own catalog lets install offer to
fix that automatically.
"""
try:
if harness == "pi":
out = subprocess.run(
["pi", "--list-models"], capture_output=True, text=True,
timeout=30, check=False,
)
if out.returncode != 0:
return []
# provider model context max-out thinking images
return [f"{row[0]}/{row[1]}" for row in
(line.split() for line in out.stdout.splitlines()[1:])
if len(row) >= 2]
if harness == "omp":
out = subprocess.run(
["omp", "models", "--json"], capture_output=True, text=True,
timeout=30, check=False,
)
if out.returncode != 0:
return []
data = json.loads(out.stdout)
return [m.get("selector") or f"{m['provider']}/{m['id']}"
for m in data.get("models", []) if isinstance(m, dict)]
except (OSError, subprocess.TimeoutExpired, json.JSONDecodeError, KeyError):
return []
return []
def fix_models_in_config(config_path: Path, harness: str) -> None:
"""Point every agent at a model the harness can actually run.
The template roster names several providers; if none of them resolve in the
harness's catalog, offer to set `defaults.model` to the first registered
model and drop the per-agent `model:` overrides — the fastest way to a
validating roster on a single model. No-op when the catalog is unknown or
the template models already resolve.
"""
catalog = harness_catalog(harness)
if not catalog:
return
text = config_path.read_text()
# Collect the model patterns the template ships (defaults + per-agent).
import re
patterns = re.findall(r"^\s*model:\s*(\S+)", text, re.M)
if not patterns:
return
# Does any template model resolve? A pattern resolves if it's in the catalog
# or is a bare id that matches exactly one catalog entry.
def resolves(pattern: str) -> bool:
if pattern in catalog:
return True
matches = [c for c in catalog if pattern == c.split("/")[-1]]
return len(matches) == 1
if any(resolves(p) for p in patterns):
return
# None resolve — offer to point the whole roster at the first model.
target = catalog[0]
print(f" none of the template models resolve in {harness}'s catalog "
f"({len(catalog)} model(s) available)")
choice = input(f" point the whole roster at {target}? [Y/n]: ").strip().lower()
if choice in ("n", "no"):
return
# Set defaults.model and drop every per-agent `model:` line.
defaults_marker = "defaults:"
d_idx = text.find(defaults_marker)
if d_idx != -1:
m_idx = text.find("model:", d_idx)
if m_idx != -1:
line_end = text.find("\n", m_idx)
text = text[:m_idx] + f"model: {target}" + text[line_end:]
# Remove per-agent model overrides (indented `model:` lines in agent
# blocks, i.e. after the defaults block ends). The defaults.model line we
# just wrote stays.
lines = text.splitlines(keepends=True)
kept = []
in_defaults = False
for line in lines:
stripped = line.strip()
if stripped.startswith("defaults:"):
in_defaults = True
elif stripped and not stripped.startswith("#") and not line.startswith((" ", "\t")):
in_defaults = False
if not in_defaults and stripped.startswith("model:") and line.startswith((" ", "\t")):
continue
kept.append(line)
config_path.write_text("".join(kept))
print(f" defaults.model = {target}; per-agent model overrides removed")
def register_repo(root: Path) -> None:
"""Add/update this repo in the visualizer's repos.json (idempotent)."""
slug = root.resolve().name
db = str((root / "adws" / "adw_data" / "sssf.db").resolve())
entry = {"slug": slug, "name": slug, "db": db}
repos = []
if REPOS_FILE.exists():
try:
repos = json.loads(REPOS_FILE.read_text())
except (json.JSONDecodeError, OSError):
repos = []
if not isinstance(repos, list):
repos = []
# Replace an existing entry for this slug, else append.
repos = [r for r in repos if not (isinstance(r, dict) and r.get("slug") == slug)]
repos.append(entry)
REPOS_FILE.write_text(json.dumps(repos, indent=2) + "\n")
print(f" registered {slug} in {REPOS_FILE}")
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--force", action="store_true", help="overwrite existing files")
parser.add_argument("--harness", choices=HARNESSES, help="coding-agent harness (pi or omp)")
args = parser.parse_args()
harness = choose_harness(args.harness)
root = Path.cwd()
stamped, skipped = [], []
stamp(TEMPLATES / "adws", root / "adws", args.force, stamped, skipped)
stamp(TEMPLATES / "prompt_engineering",
root / "adws" / "adw_data" / "prompt_engineering", args.force, stamped, skipped)
stamp(TEMPLATES / "harness_engineering",
root / "adws" / "adw_data" / "harness_engineering", args.force, stamped, skipped)
stamp(TEMPLATES / "sssf.config.yaml",
root / "adws" / "adw_sssf_config" / "sssf.config.yaml",
args.force, stamped, skipped)
stamp(TEMPLATES / "env.sample", root / ".env.sample", args.force, stamped, skipped)
# The recipes are part of the operating experience, and several cookbooks
# plus the run banner tell you to use them, so a stamped repo has to have
# them. Skipped like any other file if the repo already has a justfile.
stamp_justfile(root, args.force, stamped, skipped)
ensure_gitignore(root, stamped)
config_path = root / "adws" / "adw_sssf_config" / "sssf.config.yaml"
if config_path.exists():
set_harness_in_config(config_path, harness)
print(f" defaults.coding_agent = {harness} in {config_path}")
fix_models_in_config(config_path, harness)
register_repo(root)
print(f"sssf installed into {root}")
print(f" stamped: {len(stamped)} file(s)")
for s in stamped:
print(f" + {s}")
if skipped:
print(f" skipped (already exist, use --force to overwrite): {len(skipped)}")
print("\nnext steps:")
print(" 1. just validate # roster check: names, prompts, models all resolve")
print(" 2. cp .env.sample .env # if no .env yet; else append the keys you need")
print(" 3. just demo # two cheap read-only runs, end to end")
print(" 4. just sessions # what just happened")
print(" 5. just obs # the trace UI, needs bun")
print("\n no just? the raw form of step 3 is:")
print(" uv run adws/adw_prompt.py \"say hello\" --agent scout")
return 0
if __name__ == "__main__":
sys.exit(main())

119
sssf/scripts/make_adw.py Normal file
View file

@ -0,0 +1,119 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = []
# ///
"""make_adw — generate a new one-shot ADW script from agents in the config.
Usage:
uv run <skill>/scripts/make_adw.py --name review_docs --agents scout,builder
Each named agent becomes one agent phase, chained by envelope. Starter agents
map to their concrete output types; unknown agents get GenericOutput (define a
concrete type in adw_modules/data_types.py and swap it in).
"""
import argparse
import sys
from pathlib import Path
OUTPUT_TYPES = {"planner": "PlanOutput", "builder": "BuildOutput",
"scout": "ScoutOutput",
"reviewer": "ReviewOutput", "documenter": "DocumentOutput"}
HEADER = '''#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml"]
# ///
"""ADW {title} — generated by make_adw.
Usage:
uv run adws/adw_{name}.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> {chain}
"""
import argparse
import sys
from adw_modules import agents, gates, session, utils
from adw_modules.data_types import AgentCall, PhaseParams, {imports}
REQUIRED_AGENTS = {agents_list}
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
previous = None
{phases}
return 0 if run.succeeded else 1
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))
'''
PHASE = ''' # TODO: replace this description — say what THIS phase does and why.
with run.phase(PhaseParams(name="{name}", kind="agent", owner="{agent}",
description="Run {agent} over the request and hand its envelope on")) as ph:
previous = ph.call(AgentCall(output_type={output_type}, prompt=prompt,
previous=previous,
gates=[gates.artifacts_exist]))
'''
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--name", required=True, help="snake_case adw name")
parser.add_argument("--agents", required=True, help="comma-separated agent names, in order")
parser.add_argument("--force", action="store_true")
args = parser.parse_args()
agent_names = [a.strip() for a in args.agents.split(",") if a.strip()]
if not agent_names:
print("no agents given")
return 1
types = [OUTPUT_TYPES.get(a, "GenericOutput") for a in agent_names]
seen: dict[str, int] = {}
phases = []
for agent, output_type in zip(agent_names, types):
seen[agent] = seen.get(agent, 0) + 1
phase_name = agent if seen[agent] == 1 else f"{agent}_{seen[agent]}"
phases.append(PHASE.format(name=phase_name, agent=agent, output_type=output_type))
body = HEADER.format(
title=args.name.replace("_", " ").title(),
name=args.name,
chain=" -> ".join(agent_names),
imports=", ".join(sorted(set(types))),
agents_list=repr(sorted(set(agent_names))),
phases="\n".join(phases),
)
dest = Path.cwd() / "adws" / f"adw_{args.name}.py"
if dest.exists() and not args.force:
print(f"{dest} already exists — use --force to overwrite")
return 1
dest.parent.mkdir(parents=True, exist_ok=True)
dest.write_text(body)
print(f"wrote {dest}")
print("next: replace each phase description — a generated one says nothing, "
"and the description is the only intent the trace ever shows")
print(f"run: uv run adws/adw_{args.name}.py \"your request\"")
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -0,0 +1,35 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = []
# ///
"""make_config — generate adws/adw_sssf_config/sssf.config.yaml with great defaults.
Usage:
uv run <skill>/scripts/make_config.py [--force]
"""
import argparse
import shutil
import sys
from pathlib import Path
TEMPLATE = Path(__file__).resolve().parent.parent / "templates" / "sssf.config.yaml"
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--force", action="store_true")
args = parser.parse_args()
dest = Path.cwd() / "adws" / "adw_sssf_config" / "sssf.config.yaml"
if dest.exists() and not args.force:
print(f"{dest} already exists — use --force to overwrite")
return 1
dest.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(TEMPLATE, dest)
print(f"wrote {dest}")
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -0,0 +1,45 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Build — one-shot implementation workflow.
Usage:
uv run adws/adw_build.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> builder
"""
import argparse
import sys
from adw_modules import agents, gates, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams
REQUIRED_AGENTS = ["builder"]
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="build", kind="agent", owner="builder", retries=1,
description="Implement the request")) as ph:
ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
gates=[gates.diff_matches_claims]))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,76 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Build Review — implement, then confirm it is what was asked for.
Usage:
uv run adws/adw_build_review.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> builder -> reviewer [-> builder(revise) -> reviewer ... bounded]
Review is not testing. Tests answer "does it run"; the reviewer answers "is this
the thing that was asked for" — it reads the spec (`plan.md` from a prior plan
phase if the session has one, else the prompt verbatim), reads the code that was
written, and rules on each requirement.
Like the tester, the reviewer's phase succeeds when it RUNS and REPORTS. A
rejection does not fail the phase; it fails the run, checked at the end, after
the bounded revise loop has had its chances.
"""
import argparse
import sys
from adw_modules import agents, gates, session, utils
from adw_modules.data_types import (AgentCall, BuildOutput, PhaseParams,
ReviewOutput)
REQUIRED_AGENTS = ["builder", "reviewer"]
MAX_REVISION_LOOPS = 3
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the request")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
gates=[gates.diff_matches_claims]))
review = None
for i in range(1, MAX_REVISION_LOOPS + 1):
with run.phase(PhaseParams(name=f"review_{i}", kind="agent", owner="reviewer",
description="Rule on every requirement in the spec, against the code on disk")) as ph:
review = ph.call(AgentCall(output_type=ReviewOutput, prompt=prompt,
previous=previous,
gates=[gates.artifacts_exist,
gates.verdict_consistent]))
if review.approved:
break
if i == MAX_REVISION_LOOPS:
break
with run.phase(PhaseParams(name=f"revise_{i}", kind="agent", owner="builder", retries=1,
description="Close every blocking finding the reviewer named")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=review,
gates=[gates.diff_matches_claims]))
return run.finish(accepted=review is not None and review.approved,
reason=f"the reviewer never approved after {MAX_REVISION_LOOPS} revision(s)")
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,79 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Build Test — implement, then verify; failures flow back into the builder.
Usage:
uv run adws/adw_build_test.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> builder -> code(test) [-> builder(fix) -> code(test) ... bounded]
Testing is CODE. The suite's command is written down in adw_modules/quality.py,
so running it needs no judgement — only repairing it does. Failures reach the
builder as an envelope through `quality.as_envelope`, which is the same door an
agent's report came through, so the repair loop is unchanged.
A failing suite does NOT fail its phase: the runner did its job, the code is
what failed. It fails the run, checked at the end, after the bounded fix loop
has had its chances.
"""
import argparse
import sys
from adw_modules import agents, gates, quality, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams
REQUIRED_AGENTS = ["builder"]
MAX_FIX_LOOPS = 3
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
def record(ph, result) -> None:
passed = sum(1 for check in result.checks if check.passed)
ph.log(passed=result.passed, checks=f"{passed}/{len(result.checks)}",
artifacts=", ".join(result.artifacts))
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the request")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
gates=[gates.diff_matches_claims]))
test = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"test_{i}", kind="code", owner="quality",
description="Run the suite — a known command, so code runs "
"it and no agent has to rediscover it")) as ph:
test = quality.run_tests(run)
record(ph, test)
if test.passed:
break
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description="Repair what the suite reported, from its "
"verbatim output")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(test, "tests"),
gates=[gates.diff_matches_claims]))
return run.finish(accepted=test is not None and test.passed,
reason=f"the suite still failed after {MAX_FIX_LOOPS} fix attempt(s)")
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,76 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Document — write up the work that was just done, from the diff.
Usage:
uv run adws/adw_document.py "<prompt or path/to/prompt.md>" [--base main] [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> code(changes) -> documenter
This runs AFTER a build, and the guard is structural rather than advisory: the
change capture is a code phase, and an empty diff raises there — before the
documenter is ever spawned. There is nothing to document until something was
built, and the phase says so instead of paying an agent to discover it.
`git diff` against `--base` (main by default) is what "the latest changes"
means here; see adw_modules/changes.py for how the base commit is resolved on a
branch, on main, and on a clean tree right after a chain committed.
"""
import argparse
import sys
from adw_modules import agents, changes, gates, session, utils
from adw_modules.data_types import (AgentCall, ChangeCapture, DocumentOutput,
PhaseParams)
REQUIRED_AGENTS = ["documenter"]
DOCUMENT_NOTES = ("Read diff_path in full before writing. Document only what the "
"diff shows, then copy the write-up into app_docs/ as your task "
"describes.")
def main(prompt: str, base: str = "main",
config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="changes", kind="code", owner="git",
description=f"Diff the working tree against {base} — the change to be written up")) as ph:
changeset = changes.capture(run, ChangeCapture(base=base))
ph.log(base=f"{changeset.base.label} @ {changeset.base.commit[:7]}",
reason=changeset.base.reason,
files=len(changeset.files) + len(changeset.untracked),
lines=f"+{changeset.insertions} -{changeset.deletions}",
diff=changeset.diff_path)
if changeset.empty:
raise RuntimeError(
f"nothing changed since {changeset.base.label} ({changeset.base.reason}) "
f"— documenting runs after a build. Build something first, or point "
f"--base at the ref the work should be measured from.")
with run.phase(PhaseParams(name="document", kind="agent", owner="documenter", retries=1,
description="Turn the captured diff into a write-up an engineer can read")) as ph:
ph.call(AgentCall(output_type=DocumentOutput, prompt=prompt,
previous=changes.as_envelope(changeset, DOCUMENT_NOTES),
gates=[gates.artifacts_exist, gates.files_non_empty]))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--base", default="main", help="ref the change is measured against")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.base, args.config, args.adw_id))

View file

@ -0,0 +1,15 @@
"""Claude Code interface — STUB in v1. The factory is Pi-only for now.
The config schema accepts `coding_agent: claude_code` so nothing breaks at the
schema level, but selecting it raises until v2 implements this interface
(`claude -p --output-format stream-json --resume <session_id>`).
"""
from __future__ import annotations
def run(*args, **kwargs):
raise NotImplementedError(
"coding_agent 'claude_code' is not implemented in v1 — SSSF v1 runs the "
"Pi coding agent only. Set coding_agent: pi (or omit it) in sssf.config.yaml."
)

View file

@ -0,0 +1,210 @@
"""OMP coding agent interface — an alternative to Pi for SSSF.
Runs `omp -p --mode json` and tails its JSONL stdout line by line, forwarding
each event to a callback WHILE the agent works. OMP's event stream is the same
shape as Pi's (session / message_start / message_end / turn_end /
tool_execution_start / tool_execution_end / agent_end), so the shared
ToolCallTracker folds tool calls identically.
The one real difference is session identity. Pi accepts `--session-id` and
creates-or-continues that exact id. OMP mints its own session id and resumes by
`--resume <id-prefix>`. To keep SSSF's retry/correction loop (which re-enters
the SAME context window) working, this adapter persists the real OMP session id
in a `session_id.txt` next to the session dir: the first send creates a session
and records its id; every later send with the same session_dir resumes it.
"""
from __future__ import annotations
import json
import os
import subprocess
from functools import lru_cache
from pathlib import Path
from typing import Callable, Optional
from .agent_pi import (ToolCallTracker, _context_tokens, _label, _text_of)
from .data_types import PiRequest, PiResult
from .utils import now_iso, operator_env
OMP_PATH = os.environ.get("OMP_PATH", "omp")
RESULT_SNIPPET_CHARS = 20_000 # tool output rides along whole; clip only guards pathological cases
ARG_VALUE_CHARS = 20_000 # args too — the UI scrolls, it must not be handed cut-off data
LABEL_CHARS = 80 # "bash: <command>" shown as the event name
# The arg that identifies a call at a glance, in the order tools tend to use.
PRIMARY_ARGS = ("command", "path", "file_path", "pattern", "query", "url")
# The config's `tools` list is written in pi's vocabulary. OMP's tool set
# overlaps but is not identical, so translate the names that differ and drop
# the ones OMP has no equivalent for (the `writes` boundary, enforced in code,
# is what actually keeps an agent read-only — the tool list is a hint).
TOOL_TRANSLATION = {
"find": "glob", # pi's find -> omp's glob
"ls": None, # omp reads directories via `read`
"subagent_create": "task", # omp's subagent tool
"subagent_continue": "task",
"subagent_list": "task",
"subagent_remove": "task",
}
def _translate_tools(tools: list[str]) -> list[str]:
"""Map pi tool names to omp's vocabulary, dropping unmappable ones."""
out = []
for tool in tools:
mapped = TOOL_TRANSLATION.get(tool, tool)
if mapped and mapped not in out:
out.append(mapped)
return out
@lru_cache(maxsize=1)
def _omp_catalog() -> list[dict]:
"""Read omp's merged model catalog (`omp models --json`)."""
try:
result = subprocess.run(
[OMP_PATH, "models", "--json"], capture_output=True, text=True,
timeout=30, env=operator_env(), check=False,
)
except (OSError, subprocess.TimeoutExpired):
return []
if result.returncode != 0:
return []
try:
return json.loads(result.stdout).get("models", [])
except json.JSONDecodeError:
return []
def resolve_model(pattern: str) -> tuple[str, str]:
"""Resolve a model pattern to an explicit ``(provider, model_id)`` pair.
OMP's catalog exposes a `selector` of the form ``provider/model-id``. A
pattern with a slash must match a selector exactly; a bare pattern matches
by model id (substring, then exact), failing on ambiguity.
"""
catalog = _omp_catalog()
selectors = {m.get("selector"): (m.get("provider", ""), m.get("id", ""))
for m in catalog if m.get("selector")}
if "/" in pattern:
if pattern in selectors:
return selectors[pattern]
raise ValueError(f"model pattern {pattern!r} not found in omp models — "
"run `omp models` to see available selectors")
matches = [(m.get("provider", ""), m.get("id", "")) for m in catalog
if pattern == m.get("id") or pattern in m.get("id", "")]
exact = [match for match in matches if match[1] == pattern]
if len(exact) == 1:
return exact[0]
if len(matches) == 1:
return matches[0]
if not matches:
raise ValueError(f"model pattern {pattern!r} not found in omp models — "
"run `omp models` to see available models")
raise ValueError(f"model pattern {pattern!r} is ambiguous: {matches}")
def context_window(provider: str, model_id: str) -> int:
"""The model's context ceiling from omp's merged model catalog."""
for model in _omp_catalog():
if model.get("provider") == provider and model.get("id") == model_id:
return int(model.get("contextWindow") or 0)
return 0
def _session_id_file(session_dir: Path) -> Path:
return session_dir / "session_id.txt"
def run(request: PiRequest, on_event: Optional[Callable[[dict], None]] = None,
on_spawn: Optional[Callable[[int], None]] = None,
on_exit: Optional[Callable[[int], None]] = None) -> PiResult:
"""Run one non-interactive omp turn.
`on_spawn(pid)` and `on_exit(pid)` bracket the child process so the caller
can record it as killable — a hung coding agent is otherwise a pid you have
to hunt for in `ps` while the run sits there.
"""
provider, model_id = resolve_model(request.model)
selector = f"{provider}/{model_id}"
session_dir = Path(request.session_dir)
session_dir.mkdir(parents=True, exist_ok=True)
sid_file = _session_id_file(session_dir)
cmd = [
OMP_PATH, "-p", "--mode", "json",
"--model", selector,
"--thinking", request.thinking,
"--system-prompt", request.system_prompt,
"--session-dir", str(session_dir),
"--auto-approve", # agents run tools autonomously; no approval prompts
]
# Create-or-continue: the first send mints a session and records its id;
# every later send with the same session_dir resumes that context window.
if sid_file.exists():
cmd += ["--resume", sid_file.read_text().strip()]
if request.tools:
cmd += ["--tools", ",".join(_translate_tools(request.tools))]
for extension in request.extensions:
cmd += ["-e", extension]
cmd.append(request.prompt)
raw_path = Path(request.raw_output_path)
raw_path.parent.mkdir(parents=True, exist_ok=True)
result = PiResult(session_id="", context_window=context_window(provider, model_id))
# stdin is DEVNULL, deliberately. The prompt travels in argv, so the child
# never needs stdin — but inheriting the parent's means omp sees a non-TTY
# and can sit forever waiting for piped input that will never arrive or EOF.
process = subprocess.Popen(cmd, stdin=subprocess.DEVNULL,
stdout=subprocess.PIPE, stderr=subprocess.PIPE,
text=True, bufsize=1, cwd=request.cwd,
env=operator_env())
if on_spawn:
on_spawn(process.pid)
with raw_path.open("a") as raw:
assert process.stdout is not None
for line in process.stdout:
raw.write(line)
raw.flush() # events land on disk as they happen
line = line.strip()
if not line:
continue
try:
event = json.loads(line)
except json.JSONDecodeError:
continue
if event.get("type") == "session":
sid = event.get("id")
if sid:
result.session_id = sid
sid_file.write_text(sid) # persist for later resumes
if event.get("type") == "message_end":
message = event.get("message", {})
if message.get("role") == "assistant":
text = _text_of(message)
if text:
result.text = text # last assistant message wins
usage = message.get("usage", {}) or {}
turn = _context_tokens(usage)
result.tokens += turn
result.usage.add_turn(usage, turn)
# Occupancy is read off the last VALID assistant turn, the
# way omp does it — an aborted or errored turn reports usage
# you can't trust, so it must not overwrite a good reading.
if turn and message.get("stopReason") not in ("aborted", "error"):
result.context_tokens = turn
result.cost += (usage.get("cost", {}) or {}).get("total", 0.0) or 0.0
if on_event:
on_event(event)
stderr = process.stderr.read() if process.stderr else ""
result.returncode = process.wait()
if on_exit:
on_exit(process.pid)
if result.returncode != 0 and not result.text:
raise RuntimeError(f"omp exited {result.returncode}: {stderr.strip()[-800:]}")
return result

View file

@ -0,0 +1,286 @@
"""Pi coding agent interface — v1's only coding agent.
Runs `pi -p --mode json` and tails its JSONL stdout line by line, forwarding
each event to a callback WHILE the agent works (the streaming crack, solved
by construction). `--session-id` creates-or-continues, so running and
continuing an agent are the same call: same session id = same context window.
"""
from __future__ import annotations
import json
import os
import subprocess
import time
from functools import lru_cache
from pathlib import Path
from typing import Callable, Optional
from .data_types import PiRequest, PiResult
from .utils import now_iso, operator_env
PI_PATH = os.environ.get("PI_PATH", "pi")
MODELS_JSON = os.environ.get("PI_MODELS_PATH",
str(Path.home() / ".pi" / "agent" / "models.json"))
RESULT_SNIPPET_CHARS = 20_000 # tool output rides along whole; clip only guards pathological cases
ARG_VALUE_CHARS = 20_000 # args too — the UI scrolls, it must not be handed cut-off data
LABEL_CHARS = 80 # "bash: <command>" shown as the event name
# The arg that identifies a call at a glance, in the order tools tend to use.
PRIMARY_ARGS = ("command", "path", "file_path", "pattern", "query", "url")
def _count(value: str) -> int:
"""Parse pi's compact model-list counts (`272K`, `1.0M`)."""
suffixes = {"K": 1_000, "M": 1_000_000}
suffix = value[-1:].upper()
if suffix in suffixes:
return int(float(value[:-1]) * suffixes[suffix])
return int(value)
@lru_cache(maxsize=1)
def _pi_catalog() -> list[tuple[str, str, int]]:
"""Read pi's merged catalog, including built-in providers and custom models."""
try:
result = subprocess.run(
[PI_PATH, "--list-models"], capture_output=True, text=True,
timeout=30, env=operator_env(), check=False,
)
except (OSError, subprocess.TimeoutExpired):
return []
if result.returncode != 0:
return []
rows = []
for line in result.stdout.splitlines()[1:]:
columns = line.split()
if len(columns) < 3:
continue
try:
rows.append((columns[0], columns[1], _count(columns[2])))
except ValueError:
continue
return rows
def resolve_model(pattern: str) -> tuple[str, str]:
"""Resolve a model pattern to an explicit ``(provider, model_id)`` pair.
Pi's catalog merges built-in models with ``~/.pi/agent/models.json``. Using
that same merged view lets SSSF target direct providers such as
``openai/gpt-5.6-terra`` without re-registering built-in models locally.
"""
catalog = [(provider, model_id) for provider, model_id, _ in _pi_catalog()]
if "/" in pattern:
provider, model_id = pattern.split("/", 1)
if (provider, model_id) in catalog:
return provider, model_id
matches = [(provider, model_id) for provider, model_id in catalog
if pattern == model_id or pattern in model_id]
exact = [match for match in matches
if match[1] == pattern or match[1].endswith("/" + pattern)]
if len(exact) == 1:
return exact[0]
if len(matches) == 1:
return matches[0]
if not matches:
raise ValueError(f"model pattern {pattern!r} not found in pi --list-models — "
"authenticate/register it or fix the config")
raise ValueError(f"model pattern {pattern!r} is ambiguous: {matches}")
def _context_tokens(usage: dict) -> int:
"""Tokens occupying the window after a turn.
Mirrors pi's own `calculateContextTokens` (coding-agent
`core/compaction/compaction.ts`), which is what pi compacts against and
shows in its footer: prefer the provider's `totalTokens`, else sum the
parts. Cache reads count — cached prompt is still prompt.
"""
total = usage.get("totalTokens") or 0
if total:
return int(total)
return int(sum(usage.get(part) or 0
for part in ("input", "output", "cacheRead", "cacheWrite")))
def context_window(provider: str, model_id: str) -> int:
"""The model's context ceiling from pi's merged model catalog."""
registry = json.loads(Path(MODELS_JSON).read_text())
for model in registry.get("providers", {}).get(provider, {}).get("models", []):
if model.get("id") == model_id:
return int(model.get("contextWindow") or 0)
for listed_provider, listed_model, window in _pi_catalog():
if listed_provider == provider and listed_model == model_id:
return window
return 0
def _text_of(container: dict) -> str:
"""Join the text blocks of anything pi shapes as {content: [...]} — a
message or a tool result."""
return "".join(part.get("text", "") for part in container.get("content", []) or []
if isinstance(part, dict) and part.get("type") == "text")
def _clip(text: str, limit: int) -> str:
return text if len(text) <= limit else text[:limit].rstrip() + "…"
def _label(tool: str, args: dict) -> str:
"""One-line human name for a tool call: `bash: ls -la src`."""
value = next((args[key] for key in PRIMARY_ARGS
if isinstance(args.get(key), str) and args[key].strip()), "")
if not value:
value = next((v for v in args.values() if isinstance(v, str) and v.strip()), "")
value = " ".join(str(value).split())
return f"{tool}: {_clip(value, LABEL_CHARS)}" if value else tool
class ToolCallTracker:
"""Folds pi's tool stream into ONE normalized record per completed call.
pi announces a call as a `toolCall` content block, then emits
tool_execution_start / _update / _end for it. Only the end carries the
result, so that is where a record is emitted — one trace event per real
tool call, the moment it returns, instead of three shapeless ones.
The record carries the call's real span (`started_at`/`ended_at`), which the
tracer writes to columns so the UI can lay tool calls on a time axis without
parsing every payload.
"""
def __init__(self) -> None:
self._open: dict[str, dict] = {}
def observe(self, event: dict) -> Optional[dict]:
"""Returns the record for a finished tool call, else None."""
etype = event.get("type", "")
if etype == "message_end":
for block in event.get("message", {}).get("content", []) or []:
if isinstance(block, dict) and block.get("type") == "toolCall":
self._announce(block.get("id"), block.get("name"),
block.get("arguments"))
return None
if etype == "tool_execution_start":
self._announce(event.get("toolCallId"), event.get("toolName"),
event.get("args"))
return None
if etype != "tool_execution_end":
return None
call_id = str(event.get("toolCallId") or "")
opened = self._open.pop(call_id, {})
tool = str(event.get("toolName") or opened.get("tool") or "tool")
args = event.get("args") or opened.get("args") or {}
record = {
"tool": tool,
"tool_call_id": call_id,
"args": {key: _clip(value, ARG_VALUE_CHARS) if isinstance(value, str) else value
for key, value in args.items()},
"ok": not event.get("isError", False),
"label": _label(tool, args),
}
result_text = _text_of(event.get("result") or {})
if result_text:
record["result_snippet"] = _clip(result_text, RESULT_SNIPPET_CHARS)
record["ended_at"] = now_iso()
if opened.get("clock"):
record["duration_ms"] = int((time.monotonic() - opened["clock"]) * 1000)
if opened.get("started_at"):
record["started_at"] = opened["started_at"]
return record
def _announce(self, call_id, tool, args) -> None:
"""First sighting starts the clock; a later sighting only fills gaps."""
if not call_id:
return
known = self._open.get(str(call_id), {})
self._open[str(call_id)] = {
"tool": tool or known.get("tool", ""),
"args": args or known.get("args", {}),
"started_at": known.get("started_at") or now_iso(), # wall clock, for the row
"clock": known.get("clock") or time.monotonic(), # monotonic, for duration
}
def run(request: PiRequest, on_event: Optional[Callable[[dict], None]] = None,
on_spawn: Optional[Callable[[int], None]] = None,
on_exit: Optional[Callable[[int], None]] = None) -> PiResult:
"""Run one non-interactive pi turn.
`on_spawn(pid)` and `on_exit(pid)` bracket the child process so the caller
can record it as killable — a hung coding agent is otherwise a pid you have
to hunt for in `ps` while the run sits there.
"""
provider, model_id = resolve_model(request.model)
cmd = [
PI_PATH, "-p", "--mode", "json",
"--provider", provider, "--model", model_id,
"--thinking", request.thinking,
"--session-id", request.session_id,
"--session-dir", request.session_dir,
"--system-prompt", request.system_prompt,
]
if request.tools:
cmd += ["--tools", ",".join(request.tools)]
for extension in request.extensions:
cmd += ["-e", extension]
cmd.append(request.prompt)
raw_path = Path(request.raw_output_path)
raw_path.parent.mkdir(parents=True, exist_ok=True)
result = PiResult(session_id=request.session_id,
context_window=context_window(provider, model_id))
# stdin is DEVNULL, deliberately. The prompt travels in argv, so the child
# never needs stdin — but inheriting the parent's means pi sees a non-TTY
# and can sit forever waiting for piped input that will never arrive or
# EOF. That failure is silent and total: no request goes out, no bytes come
# back, and the ADW blocks on a read loop with nothing to read. Observed as
# a run that sat idle at 0% CPU with an empty raw_output.jsonl.
process = subprocess.Popen(cmd, stdin=subprocess.DEVNULL,
stdout=subprocess.PIPE, stderr=subprocess.PIPE,
text=True, bufsize=1, cwd=request.cwd,
env=operator_env())
if on_spawn:
on_spawn(process.pid)
with raw_path.open("a") as raw:
assert process.stdout is not None
for line in process.stdout:
raw.write(line)
raw.flush() # events land on disk as they happen
line = line.strip()
if not line:
continue
try:
event = json.loads(line)
except json.JSONDecodeError:
continue
if event.get("type") == "message_end":
message = event.get("message", {})
if message.get("role") == "assistant":
text = _text_of(message)
if text:
result.text = text # last assistant message wins
usage = message.get("usage", {}) or {}
turn = _context_tokens(usage)
result.tokens += turn
result.usage.add_turn(usage, turn)
# Occupancy is read off the last VALID assistant turn, the
# way pi does it — an aborted or errored turn reports usage
# you can't trust, so it must not overwrite a good reading.
if turn and message.get("stopReason") not in ("aborted", "error"):
result.context_tokens = turn
result.cost += (usage.get("cost", {}) or {}).get("total", 0.0) or 0.0
if on_event:
on_event(event)
stderr = process.stderr.read() if process.stderr else ""
result.returncode = process.wait()
if on_exit:
on_exit(process.pid)
if result.returncode != 0 and not result.text:
raise RuntimeError(f"pi exited {result.returncode}: {stderr.strip()[-800:]}")
return result

View file

@ -0,0 +1,320 @@
"""Config loading/validation and agent execution.
Every ADW validates its agents before running (fail fast, nothing spawns
against a half-valid config). Every agent call parses against a concrete
output type; parse failures and gate violations re-prompt the SAME session
with a correction — context intact, bounded retries. Agent proposes, code
disposes.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import Optional
import yaml
from . import agent_omp, agent_pi, permissions, prompts
from .data_types import (AgentCall, AgentConfig, EnvelopeBase, EventRecord,
GateCheck, GateReport, Phase, PiRequest, SSSFConfig,
UsageBreakdown)
from .utils import new_id
JSON_FIX_ATTEMPTS = 2 # continue-with-correction attempts for malformed JSON
class GateFailure(RuntimeError):
pass
# ── config ───────────────────────────────────────────────────────────────────
def load_config(path: str = "adws/adw_sssf_config/sssf.config.yaml") -> SSSFConfig:
raw = yaml.safe_load(Path(path).read_text()) or {}
defaults = raw.get("defaults", {}) or {}
for agent in raw.get("agents", []) or []:
for key in ("coding_agent", "model", "thinking", "color", "tools", "writes"):
if key in defaults:
agent.setdefault(key, defaults[key])
agent.setdefault("harness_engineering", defaults.get("harness_engineering", []))
return SSSFConfig(**raw)
def resolve(cfg: SSSFConfig, name: str) -> AgentConfig:
for agent in cfg.agents:
if agent.name == name:
return agent
raise SystemExit(f"agent {name!r} is not defined in the config — "
f"available: {[a.name for a in cfg.agents]}")
def validate(cfg: SSSFConfig, required: list[str]) -> None:
"""Fail fast: every required name must resolve to a usable agent."""
problems = []
for name in required:
try:
agent = resolve(cfg, name)
except SystemExit as e:
problems.append(str(e))
continue
if agent.coding_agent not in ("pi", "omp"):
problems.append(f"agent {name!r}: coding_agent {agent.coding_agent!r} "
f"is not implemented (pi and omp are)")
for label, ref in (("system", agent.prompt_engineering.system),
("user", agent.prompt_engineering.user)):
if not Path(ref).is_file():
problems.append(f"agent {name!r}: {label} prompt not found: {ref}")
try:
_resolve_model(agent)
except ValueError as e:
problems.append(f"agent {name!r}: {e}")
if problems:
raise SystemExit("config validation failed:\n- " + "\n- ".join(problems))
# ── execution ────────────────────────────────────────────────────────────────
def execute(run, phase: Phase, call: AgentCall) -> EnvelopeBase:
"""One agent call: render prompts -> pi run -> typed parse -> gates -> envelope."""
agent = resolve(run.cfg, phase.params.owner)
agent_dir = run.session_dir / agent.name
agent_dir.mkdir(parents=True, exist_ok=True)
variables = {
"prompt": call.prompt,
"previous_envelope": call.previous.model_dump_json(indent=2) if call.previous else "(none)",
"context_handoff_dir": str(run.context_handoff_dir),
}
system_text = prompts.render(agent.prompt_engineering.system, variables)
user_text = prompts.render(agent.prompt_engineering.user, variables)
prompts.save(agent_dir / "prompts", "system.md", system_text)
prompts.save(agent_dir / "prompts", "user.md", user_text)
session_id = _agent_session_id(run, agent)
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="agent_start", name=agent.name,
payload={"model": agent.model, "thinking": agent.thinking,
"color": agent.color,
"session_id": session_id,
"coding_agent": agent.coding_agent,
"purpose": agent.purpose,
"tools": agent.tools, # None = all tools
"harness_engineering": agent.harness_engineering}))
run.console.agent_started(agent.name, agent.model, session_id)
# Parse retries and gate corrections re-enter the SAME pi session, so the
# last send is the one whose context occupancy is current — while spend is
# the opposite: every send costs, so usage accumulates across all of them.
latest: agent_pi.PiResult | None = None
spent = UsageBreakdown()
def send(prompt_text: str) -> agent_pi.PiResult:
nonlocal latest
request = PiRequest(
prompt=prompt_text,
system_prompt=system_text,
model=agent.model,
thinking=agent.thinking,
session_id=session_id,
# absolute: these are read by the coding-agent subprocess, which runs in repo_root
session_dir=_agent_session_dir(agent_dir, agent.coding_agent),
raw_output_path=str((agent_dir / "raw_output.jsonl").resolve()),
tools=agent.tools,
extensions=agent.harness_engineering,
cwd=str(run.repo_root),
)
result = _agent_runner(agent)(
request,
on_event=_event_forwarder(run, phase, agent.name),
on_spawn=lambda pid: run.tracer.process_start(
run.adw_id, "agent", agent.name, pid,
f"{agent.coding_agent} {agent.name} {agent.model}"),
on_exit=lambda pid: run.tracer.process_end(run.adw_id, pid))
run.add_usage(result.tokens, result.cost)
spent.merge(result.usage)
latest = result
return result
# What the tree looked like before this agent got its hands on it. Every
# send in this phase — first prompt, JSON retries, gate corrections — is
# measured against this one baseline.
tree_before = permissions.snapshot(run)
result = send(user_text)
envelope, attempt = _parse_with_retries(run, phase, call, result, send)
# claim gates — violations flow back into the SAME session as corrections
for gate_attempt in range(1, max(1, phase.params.retries + 1) + 1):
violations = []
for gate in call.gates:
report = _as_report(gate(envelope, run))
found = report.violations
run.tracer.gate_row(phase, gate.__name__, report, gate_attempt)
run.tracer.event(EventRecord(
adw_id=run.adw_id, phase_id=phase.phase_id,
type="gate_fail" if found else "gate_pass", name=gate.__name__,
payload={"attempt": gate_attempt, "violations": found,
"checks": [c.model_dump() for c in report.checks]}))
run.console.gate_result(gate.__name__, report)
violations.extend(found)
if not violations:
break
if gate_attempt > phase.params.retries:
raise GateFailure(f"{agent.name} failed gates after {gate_attempt} attempt(s):\n- "
+ "\n- ".join(violations))
phase.attempt = gate_attempt
run.console.retry(agent.name, gate_attempt, phase.params.retries,
f"{len(violations)} gate violation(s)")
correction = ("Your previous response failed validation:\n- "
+ "\n- ".join(violations)
+ "\n\nFix these problems, then re-emit ONLY your Report JSON.")
result = send(correction)
envelope, attempt = _parse_with_retries(run, phase, call, result, send)
# Permission is checked after every send is done, and before the envelope is
# accepted: an agent does not get to report success on a phase in which it
# wrote somewhere it was not allowed to.
try:
touched = permissions.enforce(run, phase, agent, tree_before)
except permissions.PermissionBreach as breach:
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="error", name="permission_breach",
payload={"agent": agent.name, "error": str(breach),
"writes": agent.writes,
"protected_files": run.cfg.defaults.protected_files}))
raise
if touched:
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="log", name="paths_touched",
payload={"agent": agent.name, "paths": touched}))
_persist_envelope(run, phase, agent.name, call, envelope, attempt, valid=True)
run.console.envelope_summary(envelope)
context = latest or result
run.tracer.agent_session_row(run.adw_id, agent, session_id,
context_tokens=context.context_tokens,
context_window=context.context_window)
run.save_agent_map(agent.name, {"session_id": session_id, "model": agent.model,
"coding_agent": agent.coding_agent})
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="handoff", name=agent.name,
payload={"artifacts": envelope.artifacts,
"summary": envelope.summary}))
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="agent_end", name=agent.name,
# Phase totals, not the last send's: a retried
# phase paid for every attempt.
tokens=spent.total_tokens,
payload={"cost": spent.total_cost,
"usage": spent.model_dump(),
"context_tokens": context.context_tokens,
"context_window": context.context_window}))
run.console.agent_finished(agent.name, spent.total_tokens, spent.total_cost)
if envelope.status != "success":
raise RuntimeError(f"{agent.name} reported status={envelope.status!r}: {envelope.summary}")
return envelope
# ── internals ────────────────────────────────────────────────────────────────
def _as_report(result) -> GateReport:
"""Accept a GateReport, or a legacy gate that returned a violations list."""
if isinstance(result, GateReport):
return result
return GateReport(checks=[GateCheck(item=str(v), ok=False) for v in (result or [])])
def _resolve_model(agent: AgentConfig) -> tuple[str, str]:
"""Resolve an agent's model pattern against its coding agent's catalog."""
if agent.coding_agent == "omp":
return agent_omp.resolve_model(agent.model)
return agent_pi.resolve_model(agent.model)
def _agent_runner(agent: AgentConfig):
"""The run() callable for an agent's coding agent."""
if agent.coding_agent == "omp":
return agent_omp.run
return agent_pi.run
def _agent_session_dir(agent_dir, coding_agent: str) -> str:
"""Absolute session dir for the coding agent's subprocess."""
sub = "omp_sessions" if coding_agent == "omp" else "pi_sessions"
return str((agent_dir / sub).resolve())
def _agent_session_id(run, agent: AgentConfig) -> str:
entry = run.agent_map.get(agent.name)
if entry and entry.get("model") == agent.model:
return entry["session_id"] # rejoin the existing context window
return f"sssf-{run.adw_id}-{agent.name}-{new_id(4)}"
def _event_forwarder(run, phase: Phase, agent_name: str):
"""One tool_call event per real tool call, with its exact args and result."""
tracker = agent_pi.ToolCallTracker()
def forward(event: dict) -> None:
record = tracker.observe(event)
if record is None:
return
# The call's span rides the columns; duration_ms stays in the payload as
# pi's own authoritative number.
run.tracer.event(EventRecord(adw_id=run.adw_id, phase_id=phase.phase_id,
type="tool_call", name=record.pop("label"),
started_at=record.pop("started_at", None),
ended_at=record.pop("ended_at", None),
payload={**record, "agent": agent_name}))
return forward
def _extract_json(text: str) -> dict:
candidate = text
if "```" in text:
for block in text.split("```")[1::2]:
block = block.removeprefix("json").strip()
if block.startswith("{"):
candidate = block
break
start, end = candidate.find("{"), candidate.rfind("}")
if start == -1 or end <= start:
raise ValueError("no JSON object found in the response")
return json.loads(candidate[start:end + 1])
def _parse_with_retries(run, phase: Phase, call: AgentCall, result, send):
"""Parse the final response against the declared output type; on failure,
continue the SAME session with a correction (bounded)."""
for attempt in range(1, JSON_FIX_ATTEMPTS + 2):
try:
payload = _extract_json(result.text)
return call.output_type.model_validate(payload), attempt
except Exception as error:
_persist_envelope(run, phase, phase.params.owner, call, None, attempt,
valid=False, raw=result.text)
if attempt > JSON_FIX_ATTEMPTS:
raise RuntimeError(
f"{phase.params.owner} never produced valid "
f"{call.output_type.__name__} JSON: {error}") from error
run.console.retry(phase.params.owner, attempt, JSON_FIX_ATTEMPTS,
f"invalid {call.output_type.__name__} JSON: {error}")
fields = ", ".join(call.output_type.model_fields.keys())
result = send(
f"Your response was not valid JSON for the required structure "
f"({error}). Respond again with ONLY a JSON object with these "
f"fields: {fields}. No prose, no code fences.")
def _persist_envelope(run, phase: Phase, agent_name: str, call: AgentCall,
envelope: Optional[EnvelopeBase], attempt: int,
valid: bool, raw: str = "") -> None:
payload_json = envelope.model_dump_json(indent=2) if envelope else json.dumps({"raw": raw[-2000:]})
run.tracer.envelope_row(phase, agent_name, call.output_type.__name__,
payload_json, valid, attempt)
if envelope:
record = {"agent_name": agent_name, "purpose": resolve(run.cfg, agent_name).purpose,
"output_type": call.output_type.__name__, "attempt": attempt,
**envelope.model_dump()}
(run.session_dir / agent_name / "envelope.json").write_text(json.dumps(record, indent=2))

View file

@ -0,0 +1,103 @@
"""Deterministic change capture: what was built, straight from git.
"What changed since main" is not a judgement call — it is two git commands and
a subtraction. So it is code, and an agent is only handed the result. The
capture writes the full diff into `context_handoff/` and returns a ChangeSet;
`as_envelope` adapts that into the one door every agent handoff uses.
The base is resolved, not assumed. Off the base branch the diff covers the
whole branch plus the working tree; on it, the uncommitted tree; and on a clean
tree, the last commit — because "document the work that was just done" still
has an answer right after a chain committed. Whichever it picked rides along in
`BaseRef.reason`, so the trace never leaves you guessing what a diff was
measured against.
"""
from __future__ import annotations
from . import git_helper
from .data_types import BaseRef, ChangeCapture, ChangeSet, ChangesOutput
DIFF_FILENAME = "changes.diff"
def resolve_base(ref: str) -> BaseRef:
"""Pick the commit the work is measured from, and record why that one."""
if not git_helper.is_repo():
raise RuntimeError(
"not a git repository — change capture needs one. Run `git init` in "
"the repo root before running an ADW that documents a change.")
if not git_helper.ref_exists(ref):
raise RuntimeError(
f"base ref {ref!r} does not exist in this repository — pass --base "
f"with a ref that does (e.g. --base master, --base HEAD~1).")
# Built first, then given its reason: BaseRef.label knows how to print a
# pinned sha, and the reason is the line a human reads in the trace.
base = BaseRef(ref=ref, commit=git_helper.merge_base(ref, "HEAD"))
if git_helper.short_sha(base.commit) != git_helper.short_sha("HEAD"):
base.reason = (f"HEAD is ahead of {base.label} — diffing every commit since, "
f"plus the working tree")
elif git_helper.is_dirty():
base.reason = f"HEAD is on {base.label} — diffing the uncommitted working tree"
elif git_helper.ref_exists("HEAD~1"):
base.commit = git_helper.rev("HEAD~1")
base.reason = (f"HEAD is on {base.label} with a clean tree — falling back to "
f"the last commit")
else:
base.reason = f"HEAD is on {base.label} with a clean tree and no parent commit"
return base
def capture(run, params: ChangeCapture) -> ChangeSet:
"""Diff the working tree against the resolved base and persist the evidence."""
base = resolve_base(params.base)
files = git_helper.diff_files(base.commit)
untracked = git_helper.untracked_files() if params.include_untracked else []
insertions, deletions = git_helper.diff_counts(base.commit)
stat = git_helper.diff_stat(base.commit)
text = git_helper.diff_text(base.commit)
lines = text.splitlines()
truncated = len(lines) > params.max_diff_lines
if truncated:
text = "\n".join(lines[:params.max_diff_lines])
text += (f"\n\n[truncated at {params.max_diff_lines} lines of "
f"{len(lines)} — run `git diff {base.commit}` for the rest]")
# Untracked files are absent from `git diff` by construction, so they are
# named here rather than silently missing from the record. The reader has
# `read` and can open any of them.
untracked_block = ("\n".join(f" {f}" for f in untracked) if untracked
else " (none)")
diff_path = run.context_handoff_dir / DIFF_FILENAME
diff_path.write_text(
f"# changes since {base.label} @ {git_helper.short_sha(base.commit)}\n"
f"# {base.reason}\n"
f"# +{insertions} -{deletions} across {len(files)} tracked file(s)\n\n"
f"## stat\n{stat or ' (no tracked changes)'}\n\n"
f"## untracked files\n{untracked_block}\n\n"
f"## diff\n{text}\n")
return ChangeSet(base=base, files=files, untracked=untracked,
insertions=insertions, deletions=deletions, stat=stat,
diff_path=str(diff_path), truncated=truncated)
def as_envelope(changes: ChangeSet, notes: str = "") -> ChangesOutput:
"""Wrap a captured change so an agent can be handed it directly."""
total = len(changes.files) + len(changes.untracked)
return ChangesOutput(
status="success",
summary=(f"{total} file(s) changed since {changes.base.label} "
f"(+{changes.insertions} -{changes.deletions})"),
artifacts=[changes.diff_path],
notes_for_next_agent=notes,
base=f"{changes.base.label} @ {git_helper.short_sha(changes.base.commit)} "
f"— {changes.base.reason}",
changed_files=changes.files + changes.untracked,
insertions=changes.insertions,
deletions=changes.deletions,
stat=changes.stat,
diff_path=changes.diff_path,
)

View file

@ -0,0 +1,131 @@
"""Console reporter: one narrative, two destinations.
Every line an ADW prints ALSO lands in the db as a `log` event, so the swim-lane
UI reads the same story the terminal does. Both go through `_emit` — print and
trace cannot drift. Plain sequential lines only: no spinners, no live displays,
so a CI log reads exactly like a terminal.
"""
from __future__ import annotations
from rich.console import Console as RichConsole
from rich.markup import escape
from rich.panel import Panel
from rich.text import Text
from .data_types import EnvelopeBase, EventRecord, Phase
KIND_COLOR = {"engineer": "cyan", "agent": "magenta", "code": "yellow"}
MAX_LINE = 160 # dynamic text (summaries, violations, errors) is clipped
def _clip(text: str, limit: int = MAX_LINE) -> str:
text = " ".join(str(text).split())
return text if len(text) <= limit else text[: limit - 1] + "…"
class Console:
"""Bound to one run's tracer. Reachable as `run.console` everywhere."""
def __init__(self, tracer, adw_id: str):
self.tracer = tracer
self.adw_id = adw_id
self.phase_id = "" # current lane — log events attach to it
self.phase_name = ""
self.results: list[str] = [] # phase statuses, for the summary
self._finished = False # the summary panel prints once
self._out = RichConsole(highlight=False, soft_wrap=True)
# ── the one helper: print AND trace, always together ────────────────────
def _emit(self, markup: str, level: str = "info", renderable=None) -> None:
text = Text.from_markup(markup)
self._out.print(renderable if renderable is not None else text)
self.tracer.event(EventRecord(
adw_id=self.adw_id, phase_id=self.phase_id, type="log",
name=self.phase_name or "console",
payload={"message": text.plain, "level": level}))
# ── session ─────────────────────────────────────────────────────────────
def session_started(self, adw_id: str, engineer: str) -> None:
self._emit(f"[bold cyan]adw_id:[/bold cyan] [bold]{escape(adw_id)}[/bold]"
f" [dim]engineer[/dim] {escape(engineer)}")
def session_finished(self, ok: bool, tokens: int, cost: float, db_path: str) -> None:
if self._finished:
return
self._finished = True
passed = sum(1 for r in self.results if r == "success")
status = "[green]✓ success[/green]" if ok else "[red]✗ fail[/red]"
rows = [f" [dim]status[/dim] {status}",
f" [dim]phases[/dim] {passed}/{len(self.results)} passed",
f" [dim]tokens[/dim] {tokens:,}",
f" [dim]cost[/dim] ${cost:.4f}",
f" [dim]adw_id[/dim] {escape(self.adw_id)}",
f" [dim]db[/dim] {escape(str(db_path))}",
f" [dim]next[/dim] [bold]just phases {escape(self.adw_id)}[/bold]"]
panel = Panel(Text.from_markup("\n".join(rows)),
title="[bold]ADW complete[/bold]",
border_style="green" if ok else "red", expand=False)
plain = (f"session {self.adw_id} {'success' if ok else 'fail'} · "
f"{passed}/{len(self.results)} phases · {tokens:,} tokens · ${cost:.4f}")
self._emit(escape(plain), level="info" if ok else "error", renderable=panel)
# ── phases ──────────────────────────────────────────────────────────────
def phase_started(self, phase: Phase) -> None:
self.phase_id, self.phase_name = phase.phase_id, phase.params.name
p = phase.params
color = KIND_COLOR.get(p.kind, "white")
line = (f"[bold {color}]▶ {phase.seq:02d} {escape(p.name)}[/bold {color}]"
f" [{color}]{p.kind}[/{color}] [dim]· {escape(p.owner)}[/dim]")
if p.description:
line += f" [dim]{escape(_clip(p.description))}[/dim]"
self._emit(line)
def phase_ended(self, phase: Phase, seconds: float) -> None:
ok = phase.status == "success"
self.results.append(phase.status)
line = (f" {'[green]✓[/green]' if ok else '[red]✗[/red]'} "
f"{escape(phase.params.name)} [dim]{seconds:.1f}s[/dim]")
if not ok and phase.error:
line += f" [red]{escape(_clip(phase.error))}[/red]"
self._emit(line, level="info" if ok else "error")
self.phase_id, self.phase_name = "", ""
def note(self, message: str) -> None:
"""Free-form detail inside the current phase — what `ph.log()` recorded."""
self._emit(f" [dim]· {escape(_clip(message))}[/dim]")
# ── agents ──────────────────────────────────────────────────────────────
def agent_started(self, name: str, model: str, session_id: str) -> None:
self._emit(f" [magenta]▸[/magenta] {escape(name)} [dim]{escape(model)}[/dim]"
f" [dim]session {escape(session_id)}[/dim]")
def agent_finished(self, name: str, tokens: int, cost: float) -> None:
self._emit(f" [dim]└ {escape(name)} used {tokens:,} tokens · ${cost:.4f}[/dim]")
def retry(self, name: str, attempt: int, limit: int, reason: str) -> None:
self._emit(f" [yellow]⟳[/yellow] {escape(name)} retry {attempt}/{limit} "
f"[dim]— same session · {escape(_clip(reason))}[/dim]", level="warn")
# ── verification ────────────────────────────────────────────────────────
def gate_result(self, name: str, report) -> None:
"""A gate reports WHAT it checked, not just whether it passed."""
ok = report.passed
mark = "[green]✓[/green]" if ok else "[red]✗[/red]"
summary = (f"{len(report.checks)} checked" if ok
else f"[red]{len(report.violations)} of {len(report.checks)} failed[/red]")
self._emit(f" {mark} gate [dim]{escape(name)}[/dim] [dim]{summary}[/dim]",
level="info" if ok else "error")
for check in report.checks:
style = "dim" if check.ok else "dim red"
detail = f" — {_clip(check.note)}" if check.note else ""
self._emit(f" [{style}]{'·' if check.ok else '✗'} {escape(_clip(check.item))}"
f"{escape(detail)}[/{style}]", level="info" if check.ok else "error")
def envelope_summary(self, envelope: EnvelopeBase) -> None:
ok = envelope.status == "success"
line = (f" {'[green]✓[/green]' if ok else '[red]✗[/red]'} "
f"{type(envelope).__name__} [dim]{escape(_clip(envelope.summary))}[/dim]")
self._emit(line, level="info" if ok else "error")
if envelope.artifacts:
self._emit(f" [dim]artifacts: {escape(_clip(', '.join(envelope.artifacts)))}[/dim]")

View file

@ -0,0 +1,447 @@
"""Concrete data types for the SSSF ADW system.
RULE (four-param rule): any function that takes more than 4 parameters takes
ONE of these objects instead. AgentCall and PhaseParams are the pattern.
Every agent call declares a concrete output type — an EnvelopeBase subclass —
that its final JSON response is parsed against. No untyped handoffs.
"""
from __future__ import annotations
from typing import Any, Callable, Literal, Optional, Type
from pydantic import BaseModel, Field, ValidationInfo, field_validator
PhaseKind = Literal["engineer", "agent", "code"]
PhaseStatus = Literal["queued", "running", "success", "fail"]
# ── Phases ────────────────────────────────────────────────────────────────────
class PhaseParams(BaseModel):
"""Everything run.phase() needs. Passed as one object, never loose params."""
name: str # short id, unique within the run: "plan", "build"
kind: PhaseKind # which lane the block renders in
owner: str # engineer's name, "git", or an agent name from config
description: str # REQUIRED: what this phase does and why — see below
retries: int = 0 # agent phases: gate-failure retries via continue
@field_validator("description")
@classmethod
def _description_must_be_earned(cls, value: str, info: ValidationInfo) -> str:
"""A phase name identifies; a description explains. Both are required.
The description is the only sentence the trace, the console, and the
phase block in the UI ever show about intent — everything else is ids,
statuses, and timings. `commit_plan: "Commit the plan"` tells a reader
nothing they could not already see, so an echo is rejected the same way
a blank one is. This is a construction-time error on purpose: it fires
before the phase opens, not after a run is already in the trace.
"""
text = " ".join(value.split())
name = str(info.data.get("name", "?"))
if not text:
raise ValueError(
f"phase {name!r}: description is required — one sentence on what this "
f"phase does and why. It is what the trace and the UI show.")
if text.rstrip(".").casefold() == name.replace("_", " ").casefold():
raise ValueError(
f"phase {name!r}: description {text!r} only restates the phase name — "
f"say what it does and why instead.")
return text
class Phase(BaseModel):
"""The persisted phase record — PhaseParams plus lifecycle."""
phase_id: str
adw_id: str
seq: int
params: PhaseParams
status: PhaseStatus = "fail" # success must be earned
attempt: int = 0
error: Optional[str] = None
started_at: Optional[str] = None
ended_at: Optional[str] = None
# ── Envelopes (agent output types) ───────────────────────────────────────────
class EnvelopeBase(BaseModel):
"""Base of every agent's final JSON response. Output types extend this."""
status: Literal["success", "fail"]
summary: str = ""
artifacts: list[str] = Field(default_factory=list)
notes_for_next_agent: str = ""
class GenericOutput(EnvelopeBase):
pass
class PlanOutput(EnvelopeBase):
# Subject for committing the PLAN — the spec file the planner wrote, not the
# implementation it describes. Each agent's commit_message covers its own
# work product, so a chain that commits per step never reuses one agent's
# words for another agent's diff.
commit_message: str = ""
class BuildOutput(EnvelopeBase):
changed_files: list[str] = Field(default_factory=list)
commit_message: str = "" # consumed by the git commit phase
class ScoutFinding(BaseModel):
file: str
note: str = ""
class ScoutOutput(EnvelopeBase):
findings: list[ScoutFinding] = Field(default_factory=list)
class ReviewFinding(BaseModel):
"""One thing the request (or plan) asked for, and whether it is there."""
requirement: str # the ask, in the requester's words
met: bool
evidence: str = "" # where it lives, or what is missing
class ReviewOutput(EnvelopeBase):
"""Confirmation that what was built is what was asked for — not a test run."""
approved: bool = False
findings: list[ReviewFinding] = Field(default_factory=list)
blocking: list[str] = Field(default_factory=list) # what must change before approval
class DocumentOutput(EnvelopeBase):
"""Where the write-up of a completed change landed."""
document_path: str = "" # the doc in the repo, e.g. app_docs/<adw_id>_<slug>.md
documented_files: list[str] = Field(default_factory=list)
commit_message: str = ""
# ── Deterministic quality blocks ─────────────────────────────────────────────
QualityArea = Literal["frontend", "backend"]
QualityOperation = Literal["lint", "typecheck", "build"]
class QualityCheckSpec(BaseModel):
"""One deterministic quality command."""
name: str
area: QualityArea
operation: QualityOperation
argv: list[str]
timeout_seconds: int = 120
class QualityCheckResult(BaseModel):
"""Captured evidence from one quality command."""
name: str
area: QualityArea
operation: QualityOperation
command: str
returncode: int
passed: bool
duration_seconds: float
output_artifact: str
# The tail of stdout+stderr, verbatim and unparsed. A failure has to travel
# back to the builder as an envelope, and the builder cannot open a log file
# it was never handed — so the evidence rides along. Deliberately raw: every
# runner formats failures differently and a generic parser would be
# confidently wrong. The full log is always at output_artifact.
output_tail: str = ""
class QualityResult(BaseModel):
"""Aggregate result from a quality block: every check it ran, and the verdict."""
passed: bool
checks: list[QualityCheckResult] = Field(default_factory=list)
failures: list[str] = Field(default_factory=list)
artifacts: list[str] = Field(default_factory=list)
# ── Change capture (git diff, deterministic) ─────────────────────────────────
class ChangeCapture(BaseModel):
"""Everything documentation.capture() needs. One object, never loose params."""
base: str = "main" # the ref the work is measured against
max_diff_lines: int = 2000 # the diff artifact is truncated past this
include_untracked: bool = True # a brand-new file is part of the change
class BaseRef(BaseModel):
"""The commit a change is measured from, and why that one.
`reason` is the line the trace shows. A diff is only as trustworthy as the
thing it was taken against, so the ADW records that choice instead of
leaving the reader to infer it.
"""
ref: str # what was asked for: "main", or a pinned sha
commit: str # the commit actually diffed against
reason: str = ""
@property
def label(self) -> str:
"""Display form — a named ref as itself, a pinned raw sha shortened."""
if len(self.ref) == 40 and all(c in "0123456789abcdef" for c in self.ref):
return self.ref[:7]
return self.ref
class ChangeSet(BaseModel):
"""What changed since the base commit — pure git facts, no judgement."""
base: BaseRef
files: list[str] = Field(default_factory=list)
untracked: list[str] = Field(default_factory=list)
insertions: int = 0
deletions: int = 0
stat: str = "" # `git diff --stat` output, verbatim
diff_path: str = "" # the full diff, written into context_handoff/
truncated: bool = False
@property
def empty(self) -> bool:
return not (self.files or self.untracked)
class ChangesOutput(EnvelopeBase):
"""A ChangeSet shaped as an envelope so an agent can be handed it directly.
Same adapter idea as VerifyOutput: code computes the diff, the documenter
consumes it through the one door every agent handoff uses.
"""
base: str = "" # "<ref> @ <commit> — <reason>"
changed_files: list[str] = Field(default_factory=list)
insertions: int = 0
deletions: int = 0
stat: str = ""
diff_path: str = "" # read this for the full diff
class VerifyOutput(EnvelopeBase):
"""A deterministic result, shaped as an envelope so an agent can consume it.
Agents hand each other typed envelopes; code blocks return QualityResult.
This is the adapter, so a failing lint or test run flows back into the
builder through exactly the same door a tester agent's report used to —
the ADW script is the only thing that knows the difference.
"""
passed: bool = False
failures: list[str] = Field(default_factory=list)
# ── Agent calls ──────────────────────────────────────────────────────────────
class GateCheck(BaseModel):
"""One thing a gate looked at, and what it found.
`note` is the evidence — "exists, 2.1KB", "exit 0", "not in the diff". On a
failed check it doubles as the reason, so it is what the agent is told.
"""
item: str # what was checked: a path, a command, a test
ok: bool
note: str = ""
class GateReport(BaseModel):
"""What every gate returns: the checks it ran. Violations are derived.
Authoring stays a one-liner per item — `report.check(...)` appends and
returns self, so a gate is a loop and a return.
"""
checks: list[GateCheck] = Field(default_factory=list)
def check(self, item: str, ok: bool, note: str = "") -> "GateReport":
self.checks.append(GateCheck(item=item, ok=ok, note=note))
return self
@property
def violations(self) -> list[str]:
return [f"{c.item}: {c.note or 'failed'}" for c in self.checks if not c.ok]
@property
def passed(self) -> bool:
return not self.violations
class AgentCall(BaseModel):
"""One agent invocation: prompt in, typed envelope out, gates verified."""
model_config = {"arbitrary_types_allowed": True}
output_type: Type[EnvelopeBase]
prompt: str
previous: Optional[EnvelopeBase] = None
gates: list[Callable] = Field(default_factory=list) # gate(envelope, run) -> list[str]
# ── Config ───────────────────────────────────────────────────────────────────
class PromptEngineering(BaseModel):
system: str # path to system.md
user: str # path to user.md
class AgentConfig(BaseModel):
name: str
coding_agent: Literal["pi", "omp", "claude_code"] = "pi"
model: str = "google/gemini-3.6-flash"
thinking: str = "medium" # off | minimal | low | medium | high | xhigh | max
color: str = "" # hex swatch for this agent's lane in the UI
purpose: str = ""
prompt_engineering: PromptEngineering
harness_engineering: list[str] = Field(default_factory=list)
tools: Optional[list[str]] = None # allowlist; None = all tools usable
# What this agent may MODIFY in the repo, enforced in code after every call
# (see adw_modules/permissions.py). `tools` cannot express this: `bash` runs
# anything and `write` reaches any path, so an agent's capability list is a
# statement of intent that nothing checks.
# None -> unrestricted, except the roster-wide `protected_files` paths
# [] -> read-only: may modify nothing tracked
# [...] -> only these. A trailing "/" means a directory prefix; a "*"
# makes it a glob; anything else is an exact path.
writes: Optional[list[str]] = None
class ConfigDefaults(BaseModel):
coding_agent: Literal["pi", "omp", "claude_code"] = "pi"
model: str = "google/gemini-3.6-flash"
thinking: str = "medium"
color: str = ""
harness_engineering: list[str] = Field(default_factory=list)
tools: Optional[list[str]] = None # roster-wide allowlist; None = all tools usable
# Off-limits to every agent that has not named them in its own `writes`.
# The factory's own code is the default: an agent must not be able to edit
# the machinery that decides whether its work passed.
protected_files: list[str] = Field(default_factory=lambda: [
"adws/adw_modules/", "adws/adw_sssf_config/", "adws/adw_*.py",
])
data_dir: str = "adws/adw_data"
class ObservabilityConfig(BaseModel):
db: str = "adws/adw_data/sssf.db"
poll_ms: int = 500
class SSSFConfig(BaseModel):
defaults: ConfigDefaults = Field(default_factory=ConfigDefaults)
observability: ObservabilityConfig = Field(default_factory=ObservabilityConfig)
agents: list[AgentConfig] = Field(default_factory=list)
# ── Tracing ──────────────────────────────────────────────────────────────────
class EventRecord(BaseModel):
"""One traced event, always logged against adw_id + phase."""
adw_id: str
phase_id: str = ""
type: str # phase_start | agent_start | tool_call | handoff | gate_pass | gate_fail | log | agent_end | phase_end | error
name: str = ""
payload: dict[str, Any] = Field(default_factory=dict)
parent_id: str = ""
tokens: Optional[int] = None
# Spans: set both when an event covers real elapsed time (a tool call), so
# the UI lays it out on a time axis without parsing payload JSON. Left unset,
# the tracer stamps started_at with the moment the event was recorded.
started_at: Optional[str] = None
ended_at: Optional[str] = None
# ── Pi coding agent interface ────────────────────────────────────────────────
class PiRequest(BaseModel):
"""Everything one non-interactive pi run needs."""
prompt: str
system_prompt: str
model: str # registry pattern, resolved to provider + id
thinking: str = "medium"
session_id: str # pi --session-id: creates or continues
session_dir: str
raw_output_path: str # JSONL stream lands here
tools: Optional[list[str]] = None
extensions: list[str] = Field(default_factory=list)
cwd: str = "." # set from run.repo_root — the codebase root agents work in
class UsageBreakdown(BaseModel):
"""Tokens and the dollars they cost, per component, summed over a call.
Mirrors pi's `usage` shape one-for-one so the numbers reconcile with what
pi itself reports: `input` EXCLUDES cache reads, which bill at their own
(cheaper) rate — add them to learn the size of the prompt that was sent.
"""
input_tokens: int = 0
output_tokens: int = 0
cache_read_tokens: int = 0
cache_write_tokens: int = 0
# Thinking tokens. NOT a fifth component: measured across every session on
# disk, reasoning is always <= output and the four components above always
# sum to totalTokens, so reasoning is the thinking SHARE of output, billed
# at the output rate. Report it nested under output, never added to it.
reasoning_tokens: int = 0
total_tokens: int = 0
input_cost: float = 0.0
output_cost: float = 0.0
cache_read_cost: float = 0.0
cache_write_cost: float = 0.0
total_cost: float = 0.0
def add_turn(self, usage: dict, total_tokens: int) -> None:
"""Fold in one pi `message_end` usage object.
`total_tokens` is passed in rather than re-derived: the caller already
computes it pi's way (totalTokens, else the sum of the parts).
"""
cost = usage.get("cost") or {}
self.input_tokens += usage.get("input") or 0
self.output_tokens += usage.get("output") or 0
self.cache_read_tokens += usage.get("cacheRead") or 0
self.cache_write_tokens += usage.get("cacheWrite") or 0
self.reasoning_tokens += usage.get("reasoning") or 0
self.total_tokens += total_tokens
self.input_cost += cost.get("input") or 0.0
self.output_cost += cost.get("output") or 0.0
self.cache_read_cost += cost.get("cacheRead") or 0.0
self.cache_write_cost += cost.get("cacheWrite") or 0.0
self.total_cost += cost.get("total") or 0.0
def merge(self, other: "UsageBreakdown") -> None:
"""Add another call's usage — a phase that retries spends more than once."""
for field in self.model_fields:
setattr(self, field, getattr(self, field) + getattr(other, field))
class PiResult(BaseModel):
text: str = ""
returncode: int = 0
session_id: str = ""
tokens: int = 0
cost: float = 0.0
usage: UsageBreakdown = Field(default_factory=UsageBreakdown)
# Context occupancy after the LAST turn — not a sum. `tokens` bills every
# turn; this is how full the window is right now, which is what the
# visualizer's context bar measures against `context_window`.
context_tokens: int = 0
context_window: int = 0 # 0 when the registry declares no ceiling

View file

@ -0,0 +1,108 @@
"""Validation gates: verify the envelope's CLAIMS, never guesses.
A gate is `gate(envelope, run) -> GateReport` — one check per item it looked at.
Violations are derived from the failed checks and sent back to the SAME agent
session as a correction. Every check is recorded either way, so a green gate
says WHAT it verified instead of only that it passed.
Gates check what is mechanically checkable; plan quality is a reviewer's job.
"""
from __future__ import annotations
import json
import subprocess
from pathlib import Path
from .data_types import EnvelopeBase, GateReport
TAIL_CHARS = 1000 # command output kept as evidence on a failure
def _size(path: Path) -> str:
n = path.stat().st_size
return f"{n}B" if n < 1024 else f"{n / 1024:.1f}KB"
def artifacts_exist(envelope: EnvelopeBase, run) -> GateReport:
report = GateReport()
for a in envelope.artifacts:
p = Path(a)
report.check(a, p.exists(),
f"exists, {_size(p)}" if p.exists() else "declared artifact does not exist")
return report
def files_non_empty(envelope: EnvelopeBase, run) -> GateReport:
report = GateReport()
for a in envelope.artifacts:
p = Path(a)
if not (p.exists() and p.is_file()):
continue # existence is artifacts_exist's job
empty = p.stat().st_size == 0
report.check(a, not empty, "declared artifact is empty" if empty else _size(p))
return report
def json_parses(envelope: EnvelopeBase, run) -> GateReport:
report = GateReport()
for a in envelope.artifacts:
p = Path(a)
if p.suffix != ".json" or not p.exists():
continue
try:
parsed = json.loads(p.read_text())
report.check(a, True, f"parses, {type(parsed).__name__}")
except json.JSONDecodeError as e:
report.check(a, False, f"declared JSON artifact does not parse: {e}")
return report
def diff_matches_claims(envelope: EnvelopeBase, run) -> GateReport:
"""Every file claimed changed must exist on disk."""
report = GateReport()
for f in getattr(envelope, "changed_files", []):
p = Path(f)
report.check(f, p.exists(),
f"exists, {_size(p)}" if p.exists() else "claimed changed file does not exist")
return report
def verdict_consistent(envelope: EnvelopeBase, run) -> GateReport:
"""A review's verdict must agree with the findings it just wrote down.
Nothing here judges the code — that is the reviewer's job. This checks the
envelope against itself: an approval that ships blocking items, or a
rejection that names no problem, is a claim the harness can refute without
reading a line of the diff.
"""
report = GateReport()
approved = bool(getattr(envelope, "approved", False))
blocking = list(getattr(envelope, "blocking", []))
unmet = [f.requirement for f in getattr(envelope, "findings", []) if not f.met]
report.check("approved vs blocking", not (approved and blocking),
"no blocking items" if not blocking
else f"{len(blocking)} blocking item(s) while approved=true"
if approved else f"{len(blocking)} blocking item(s), not approved")
report.check("approved vs findings", not (approved and unmet),
"every requirement met" if not unmet
else f"{len(unmet)} unmet requirement(s) while approved=true"
if approved else f"{len(unmet)} unmet requirement(s), not approved")
report.check("rejection names a problem", approved or bool(blocking or unmet),
"verdict is supported" if approved or blocking or unmet
else "approved=false but no blocking item or unmet requirement was given")
return report
def tests_pass(command: str):
"""Gate factory: the given shell command must exit 0."""
def gate(envelope: EnvelopeBase, run) -> GateReport:
result = subprocess.run(command, shell=True, capture_output=True, text=True)
ok = result.returncode == 0
note = f"exit {result.returncode}"
if not ok:
note += "\n" + (result.stdout + result.stderr)[-TAIL_CHARS:]
return GateReport().check(command, ok, note)
gate.__name__ = f"tests_pass({command})"
return gate

View file

@ -0,0 +1,120 @@
"""Low-level git operations for code phases. All low-level logic lives in adw_modules."""
from __future__ import annotations
import subprocess
from pathlib import Path
def _git(*args: str) -> str:
result = subprocess.run(["git", *args], capture_output=True, text=True)
if result.returncode != 0:
raise RuntimeError(f"git {' '.join(args)} failed: {result.stderr.strip()}")
return result.stdout.strip()
def current_branch() -> str:
return _git("rev-parse", "--abbrev-ref", "HEAD")
def create_branch(name: str) -> str:
_git("checkout", "-b", name)
return name
def is_repo() -> bool:
result = subprocess.run(["git", "rev-parse", "--git-dir"],
capture_output=True, text=True)
return result.returncode == 0
def repo_root() -> Path:
"""Absolute root of the codebase — where agents are spawned to work.
The git toplevel when there is one, else the process cwd (ADWs run fine in a
non-git dir; only a commit phase requires a repo). Always absolute, so it is
safe to hand to a subprocess regardless of where the ADW was launched from.
"""
if is_repo():
return Path(_git("rev-parse", "--show-toplevel")).resolve()
return Path.cwd().resolve()
def commit_all(message: str) -> str:
"""Stage the working tree and commit it. Returns the new short sha."""
if not is_repo():
raise RuntimeError(
"not a git repository — a commit phase needs one. Run `git init` in the "
"repo root (and make a first commit) before running an ADW that commits.")
_git("add", "-A")
if not _git("status", "--porcelain"):
raise RuntimeError("nothing to commit — the preceding phases changed no files")
_git("commit", "-m", message)
return _git("rev-parse", "--short", "HEAD")
def changed_files() -> list[str]:
out = _git("status", "--porcelain")
return [line[3:] for line in out.splitlines() if line]
# ── diff plumbing (composed into a ChangeSet by documentation.py) ────────────
def ref_exists(ref: str) -> bool:
"""True when `ref` resolves to a commit. Never raises — this is a question."""
result = subprocess.run(["git", "rev-parse", "--verify", "--quiet", f"{ref}^{{commit}}"],
capture_output=True, text=True)
return result.returncode == 0
def rev(ref: str = "HEAD") -> str:
return _git("rev-parse", ref)
def short_sha(ref: str = "HEAD") -> str:
return _git("rev-parse", "--short", ref)
def merge_base(ref: str, other: str = "HEAD") -> str:
"""The commit where `ref` and `other` diverged — the honest base of a branch.
On the base branch itself this returns HEAD, which makes the diff exactly
"what is not committed yet". Off it, the diff is the whole branch plus the
working tree. One command covers both cases, so no ADW has to branch on it.
"""
return _git("merge-base", ref, other)
def is_dirty() -> bool:
return bool(_git("status", "--porcelain"))
def untracked_files() -> list[str]:
out = _git("ls-files", "--others", "--exclude-standard")
return [line for line in out.splitlines() if line]
def diff_files(base: str) -> list[str]:
"""Tracked files that differ between `base` and the working tree."""
out = _git("diff", "--name-only", base)
return [line for line in out.splitlines() if line]
def diff_stat(base: str) -> str:
return _git("diff", "--stat", base)
def diff_counts(base: str) -> tuple[int, int]:
"""(insertions, deletions) across the diff. Binary files count as neither."""
insertions = deletions = 0
for line in _git("diff", "--numstat", base).splitlines():
added, removed, *_ = line.split("\t")
if added.isdigit():
insertions += int(added)
if removed.isdigit():
deletions += int(removed)
return insertions, deletions
def diff_text(base: str) -> str:
return _git("diff", base)

View file

@ -0,0 +1,185 @@
"""What an agent may CHANGE, enforced in code after the fact.
`tools:` is a capability list, not a sandbox, and two holes make it
unenforceable on its own:
* `bash` runs anything. A builder handed bash to run a test suite can also
run `git checkout adws/` — which is not hypothetical: one did, discarding
uncommitted changes to the very quality check it was about to be judged by.
* `write` reaches any path, not just the one report file an agent was given
it for. A reviewer configured with "no edit, so it cannot quietly fix"
could still rewrite the code it was reviewing.
So permission is verified the way every other claim in this system is —
after the fact, against the repo itself. `snapshot()` fingerprints the working
tree's change-set before an agent runs; `enforce()` compares it afterwards and
fails the phase if the agent touched anything outside its allowlist.
Comparing change-sets, rather than watching for writes, is what catches the
`git checkout` case: a path that was modified before the agent ran and is clean
afterwards has been reverted, and a reversion is a modification. Appearing,
disappearing, and changing all count.
A breach is NOT a gate violation. Gates are for work an agent can be asked to
redo; a breach cannot be corrected by re-prompting, because the write already
happened. It aborts the phase and names every offending path.
Two keys drive it, both in sssf.config.yaml:
defaults.protected_files paths no agent may touch unless it names them itself
agents[].writes None = unrestricted · [] = read-only · [...] = only these
"""
from __future__ import annotations
import re
import subprocess
from pathlib import Path
from .data_types import AgentConfig, SSSFConfig
class PermissionBreach(RuntimeError):
"""An agent modified a path it was not permitted to modify."""
def _git(args: list[str], cwd) -> str:
result = subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True)
return result.stdout if result.returncode == 0 else ""
def snapshot(run) -> dict[str, str]:
"""Fingerprint every path the working tree currently differs on.
Tracked files carry their numstat counts, so an edit to an already-dirty
file still registers as a change. Untracked files are listed by name.
Gitignored paths never appear, which is why the session runtime under
`data_dir` — where handoff files legitimately land — needs no special case.
"""
fingerprints: dict[str, str] = {}
for line in _git(["diff", "HEAD", "--numstat"], run.repo_root).splitlines():
fields = line.split("\t")
if len(fields) >= 3:
path = fields[-1].strip()
fingerprints[path] = f"{fields[0]},{fields[1]}"
for path in _git(["ls-files", "--others", "--exclude-standard"],
run.repo_root).splitlines():
if path.strip():
fingerprints[path.strip()] = "untracked"
return fingerprints
def changed_paths(before: dict[str, str], after: dict[str, str]) -> list[str]:
"""Every path whose state differs — appeared, vanished, or was rewritten."""
return sorted({p for p in set(before) | set(after)
if before.get(p) != after.get(p)})
def _glob(pattern: str) -> re.Pattern:
"""Translate a pattern, with `*` stopping at a path separator.
fnmatch would let `*` cross `/`, which quietly widens every pattern:
`adws/adw_*.py` would match `adws/adw_data/sessions/x/y.py` as well as the
ADW scripts it means. `**` is the way to say "cross directories".
"""
out, i = [], 0
while i < len(pattern):
char = pattern[i]
if pattern.startswith("**", i):
out.append(".*")
i += 2
elif char == "*":
out.append("[^/]*")
i += 1
elif char == "?":
out.append("[^/]")
i += 1
else:
out.append(re.escape(char))
i += 1
return re.compile("".join(out))
def _matches(path: str, pattern: str) -> bool:
if pattern.endswith("/"): # directory prefix
return path.startswith(pattern)
if "*" in pattern or "?" in pattern:
return _glob(pattern).fullmatch(path) is not None
return path == pattern
def always_writable(cfg: SSSFConfig) -> list[str]:
"""The session runtime, which EVERY agent must be able to write.
`context_handoff/` is the one place agents hand work to each other, and an
agent's own prompts, raw_output.jsonl, and envelope.json land beside it.
Scout writes its findings there, the reviewer its review, the planner its
plan — a read-only agent is read-only with respect to the REPO, never with
respect to its own report.
This is granted from `data_dir` rather than left to .gitignore. The runtime
is normally ignored, so it never even appears in a snapshot — but an agent's
ability to record its work must not hang on a gitignore entry that someone
can delete or that a changed `data_dir` can outgrow.
"""
return [cfg.defaults.data_dir.rstrip("/") + "/"]
def permitted(path: str, agent: AgentConfig, cfg: SSSFConfig) -> bool:
"""Session runtime first, then the agent's own list, then what is protected."""
if any(_matches(path, p) for p in always_writable(cfg)):
return True
if any(_matches(path, p) for p in (agent.writes or [])):
return True # naming a path is what unlocks a protected one
if any(_matches(path, p) for p in cfg.defaults.protected_files):
return False
return agent.writes is None # None = unrestricted, [] = no repo writes
def _roll_back(run, path: str, before: dict[str, str], after: dict[str, str]) -> str:
"""Undo one unauthorized change. Returns a word describing what happened.
Only changes the agent INTRODUCED are undone. A path that was already dirty
when the agent started is left exactly as it is: the operator had
uncommitted work there, and discarding it to tidy up would be the same harm
this module exists to prevent, committed by the cleanup instead of the agent.
"""
if path in before:
# Already dirty beforehand. If it is gone from the diff now, the agent
# reverted an engineer's uncommitted work and the content is not ours
# to reconstruct — say so loudly rather than pretend it was handled.
return "REVERTED-BY-AGENT (uncommitted work lost, cannot restore)" \
if path not in after else "left as-is (was already modified)"
if after.get(path) == "untracked":
try:
(Path(run.repo_root) / path).unlink()
return "deleted"
except OSError as error:
return f"could not delete ({error})"
result = subprocess.run(["git", "checkout", "--", path],
cwd=run.repo_root, capture_output=True, text=True)
return "rolled back" if result.returncode == 0 else "could not roll back"
def enforce(run, phase, agent: AgentConfig, before: dict[str, str]) -> list[str]:
"""Compare the tree against `before`; undo and raise if the agent overstepped.
Returns the paths it legitimately changed, so the trace records what an
agent actually touched rather than only what it claimed in its envelope.
Detection alone would leave the repo holding the unauthorized change while
reporting a failure, so anything the agent introduced outside its allowlist
is rolled back before the phase dies. What it cannot undo, it names.
"""
after = snapshot(run)
touched = changed_paths(before, after)
breaches = [p for p in touched if not permitted(p, agent, run.cfg)]
if not breaches:
return touched
outcomes = {p: _roll_back(run, p, before, after) for p in breaches}
scope = ("read-only" if agent.writes == []
else f"limited to {agent.writes}" if agent.writes
else f"barred from {run.cfg.defaults.protected_files}")
detail = "\n".join(f" - {p} — {outcome}" for p, outcome in outcomes.items())
raise PermissionBreach(
f"{agent.name} is {scope} but modified {len(breaches)} path(s):\n{detail}")

View file

@ -0,0 +1,21 @@
"""Prompt rendering: load system/user refs from config, replace {{placeholders}}."""
from __future__ import annotations
from pathlib import Path
def render(template_path: str | Path, variables: dict[str, str]) -> str:
text = Path(template_path).read_text()
for key, value in variables.items():
text = text.replace("{{" + key + "}}", value)
return text
def save(directory: str | Path, name: str, content: str) -> Path:
"""Save the exact prompt sent, before execution — the audit copy."""
directory = Path(directory)
directory.mkdir(parents=True, exist_ok=True)
path = directory / name
path.write_text(content)
return path

View file

@ -0,0 +1,240 @@
"""Deterministic lint, typecheck, build, and test blocks.
A known command is not a judgement call. Anything whose invocation you can write
down belongs here as code — it runs in milliseconds, costs nothing, and returns
the same answer every time. Agents are for the parts that need reading and
deciding.
╔══════════════════════════════════════════════════════════════════════════════╗
║ REPLACE THE PLACEHOLDER COMMANDS BELOW. ║
║ ║
║ Every block ships as an `echo` that exits 0 and announces it is fake. They ║
║ are placeholders on purpose: a stamped repo has no way to guess your test ║
║ runner, and a wrong-but-plausible command that silently passes is worse ║
║ than one that says so out loud. ║
║ ║
║ For each block you want: swap `_placeholder(...)` for the real argv, e.g. ║
║ argv=["bun", "test", "apps/web/server.test.ts"] ║
║ argv=["uv", "run", "pytest", "-q"] ║
║ argv=["npm", "run", "lint"] ║
║ Delete the blocks you don't need, and drop them from run_quality()'s list. ║
║ ║
║ Two rules when you write the real command: ║
║ 1. argv LIST, never a shell string — no quoting bugs, no shell injection. ║
║ 2. Call binaries by BARE NAME. These blocks inherit the operator's ║
║ environment (see utils.operator_env), so `bun`, `uv`, `pytest` resolve ║
║ exactly as they do in their terminal. Never hard-code an absolute path ║
║ like /Users/you/.bun/bin/bun — that bakes your machine into the trace. ║
╚══════════════════════════════════════════════════════════════════════════════╝
"""
from __future__ import annotations
import shlex
import subprocess
import time
from pathlib import Path
from typing import Callable
from .data_types import (EventRecord, QualityCheckResult, QualityCheckSpec, QualityResult,
VerifyOutput)
from .utils import now_iso, operator_env
# How much of a failing command's output rides back inside the envelope. Enough
# for a builder to act on without opening the artifact; bounded so a runaway
# stack trace can't swamp the next agent's context.
TAIL_CHARS = 4_000
def _placeholder(name: str) -> list[str]:
"""A command that does nothing and admits it. Replace every call to this."""
return ["echo", f"PLACEHOLDER {name}: edit adws/adw_modules/quality.py and "
f"replace this echo with the real {name} command"]
def _check_dir(run, name: str) -> Path:
seq = run.phases[-1].seq if run.phases else 0
path = run.context_handoff_dir / "quality" / f"{seq:02d}_{name}"
path.mkdir(parents=True, exist_ok=True)
return path
def _run(spec: QualityCheckSpec, run) -> QualityCheckResult:
phase = run.phases[-1]
output_dir = _check_dir(run, spec.name)
output_artifact = output_dir / "command.log"
command = shlex.join(spec.argv)
env = operator_env() # the engineer's own shell environment
run.console.note(f"quality {spec.name}: {command}")
started_at = now_iso()
clock = time.monotonic()
stdout = ""
stderr = ""
try:
completed = subprocess.run(
spec.argv,
cwd=run.repo_root,
env=env,
capture_output=True,
text=True,
timeout=spec.timeout_seconds,
)
returncode = completed.returncode
stdout = completed.stdout
stderr = completed.stderr
except subprocess.TimeoutExpired as error:
returncode = 124
stdout = error.stdout or ""
stderr = (error.stderr or "") + f"\nTimed out after {spec.timeout_seconds}s."
except OSError as error:
# A missing binary lands here as exit 127 with the real message — no
# pre-flight probe needed, and none wanted.
returncode = 127
stderr = str(error)
duration = time.monotonic() - clock
output_artifact.write_text(
f"$ {command}\nexit: {returncode}\nduration_seconds: {duration:.3f}\n"
f"\n--- stdout ---\n{stdout}\n--- stderr ---\n{stderr}\n"
)
passed = returncode == 0
run.tracer.event(EventRecord(
adw_id=run.adw_id,
phase_id=phase.phase_id,
type="tool_call",
name=f"quality:{spec.name}",
payload={
"area": spec.area,
"operation": spec.operation,
"command": command,
"returncode": returncode,
"passed": passed,
"output_artifact": str(output_artifact),
},
started_at=started_at,
ended_at=now_iso(),
))
run.console.note(
f"quality {spec.name}: {'passed' if passed else 'failed'} "
f"(exit {returncode}, {duration:.1f}s)"
)
return QualityCheckResult(
name=spec.name,
area=spec.area,
operation=spec.operation,
command=command,
returncode=returncode,
passed=passed,
duration_seconds=duration,
output_artifact=str(output_artifact),
output_tail=(stdout + stderr)[-TAIL_CHARS:],
)
# ── Blocks ────────────────────────────────────────────────────────────────────
# Replace every argv below. See the banner at the top of this file.
def test(run) -> QualityCheckResult:
"""Run the project's test suite. The highest-value block to wire up first."""
return _run(QualityCheckSpec(
name="test",
area="backend",
operation="build",
argv=_placeholder("test"), # e.g. ["bun", "test"] or ["uv", "run", "pytest", "-q"]
timeout_seconds=600,
), run)
def lint(run) -> QualityCheckResult:
return _run(QualityCheckSpec(
name="lint",
area="backend",
operation="lint",
argv=_placeholder("lint"), # e.g. ["bun", "x", "oxlint@1.36.0", "src"]
), run)
def typecheck(run) -> QualityCheckResult:
return _run(QualityCheckSpec(
name="typecheck",
area="backend",
operation="typecheck",
argv=_placeholder("typecheck"), # e.g. ["bun", "x", "tsc", "--noEmit"]
), run)
def build(run) -> QualityCheckResult:
output_dir = _check_dir(run, "build") / "bundle"
return _run(QualityCheckSpec(
name="build",
area="backend",
operation="build",
argv=_placeholder("build"), # e.g. ["bun", "build", "src/index.ts", "--outdir", str(output_dir)]
), run)
def run_tests(run) -> QualityResult:
"""The test suite alone, as a QualityResult — the deterministic test phase.
This is what replaces a `tester` agent once the command is written down. An
agent rediscovering the runner on every run costs a fortune to learn what a
subprocess already knows; the repair loop is unchanged, because a failure
still reaches the builder through `as_envelope` below.
"""
check = test(run)
failures = ([] if check.passed else
[f"{check.name}: `{check.command}` exited {check.returncode}\n"
f"{check.output_tail}".rstrip()])
return QualityResult(passed=check.passed, checks=[check], failures=failures,
artifacts=[check.output_artifact])
def as_envelope(result: QualityResult, what: str) -> VerifyOutput:
"""Wrap a deterministic result so an agent can be handed it directly.
Agents hand each other typed envelopes; code blocks return QualityResult.
This is the adapter, so a failing lint or test run flows back into the
builder through exactly the same door an agent's report would — the ADW
script is the only thing that knows the difference.
"""
return VerifyOutput(
status="success" if result.passed else "fail",
summary=(f"{what}: all {len(result.checks)} check(s) passed" if result.passed
else f"{what}: {len(result.failures)} of {len(result.checks)} check(s) failed"),
artifacts=result.artifacts,
notes_for_next_agent=("" if result.passed else
"Fix every failure below. The output is verbatim from the "
"command — trust it over any summary."),
passed=result.passed,
failures=result.failures,
)
def run_quality(run) -> QualityResult:
"""Run every block and collect ALL failures — one pass tells you everything.
Ordering contract for the caller: a failing block does NOT fail the phase.
The runner did its job; the CODE is what failed. Hand this result to the
builder and let the bounded repair loop decide the run's fate.
"""
blocks: list[Callable] = [
test,
lint,
typecheck,
build,
]
checks = [block(run) for block in blocks]
# A failure is the command, its exit code, and what it actually printed —
# everything a builder needs to repair without opening a log or being told
# what the error "means" by a parser that guessed.
failures = [
f"{check.name}: `{check.command}` exited {check.returncode}\n{check.output_tail}".rstrip()
for check in checks if not check.passed
]
return QualityResult(
passed=not failures,
checks=checks,
failures=failures,
artifacts=[check.output_artifact for check in checks],
)

View file

@ -0,0 +1,142 @@
"""The Run object: config + adw_id + agent_map + tracer + console, bound once.
`run.phase(PhaseParams(...))` is the ONE phase primitive — a context manager
for all three kinds (engineer, agent, code). Success must be earned: every
phase defaults to fail; only a clean exit flips it (agent phases additionally
require a parsed envelope + green gates, enforced inside ph.call).
"""
from __future__ import annotations
import json
import time
from contextlib import contextmanager
from pathlib import Path
from . import agents, git_helper
from .console import Console
from .data_types import AgentCall, EnvelopeBase, EventRecord, Phase, PhaseParams
from .utils import ensure_dir, now_iso
class PhaseHandle:
def __init__(self, run: "Run", phase: Phase):
self.run = run
self.phase = phase
def log(self, **payload) -> None:
self.run.tracer.event(EventRecord(adw_id=self.run.adw_id,
phase_id=self.phase.phase_id,
type="log", name=self.phase.params.name,
payload=payload))
self.run.console.note(", ".join(f"{k}: {v}" for k, v in payload.items()))
if self.phase.params.kind == "engineer" and "input" in payload:
self.run.tracer.session_request(self.run.adw_id, str(payload["input"]))
def call(self, call: AgentCall) -> EnvelopeBase:
if self.phase.params.kind != "agent":
raise RuntimeError("ph.call() is only valid inside an agent phase")
return agents.execute(self.run, self.phase, call)
class Run:
def __init__(self, cfg, adw_id: str, tracer, engineer: str):
self.cfg = cfg
self.adw_id = adw_id
self.tracer = tracer
self.console = Console(tracer, adw_id)
self.engineer = engineer
self.phases: list[Phase] = []
self.tokens = 0
self.cost = 0.0
self._seq = tracer.max_phase_seq(adw_id) # a joined run continues the sequence
self.repo_root = git_helper.repo_root() # where every agent is spawned to work
self.session_dir = ensure_dir(Path(cfg.defaults.data_dir) / "sessions" / adw_id)
self.context_handoff_dir = ensure_dir(self.session_dir / "context_handoff")
self._agent_map_path = self.session_dir / "agent_map.json"
self.agent_map: dict = (json.loads(self._agent_map_path.read_text())
if self._agent_map_path.exists() else {})
# ── agent map (adw_id -> per-agent coding-agent session ids) ────────────
def save_agent_map(self, agent: str, entry: dict) -> None:
self.agent_map[agent] = entry
self._agent_map_path.write_text(json.dumps(self.agent_map, indent=2))
# ── usage (run totals mirror what the tracer accumulates in sqlite) ─────
def add_usage(self, tokens: int, cost: float) -> None:
self.tokens += tokens
self.cost += cost
self.tracer.session_add_usage(self.adw_id, tokens, cost)
# ── the phase primitive ─────────────────────────────────────────────────
@contextmanager
def phase(self, params: PhaseParams):
self._seq += 1
phase = Phase(phase_id=f"{self.adw_id}_{self._seq:02d}_{params.name}",
adw_id=self.adw_id, seq=self._seq, params=params,
status="running", started_at=now_iso())
self.phases.append(phase)
self.tracer.phase_upsert(phase)
self.tracer.event(EventRecord(adw_id=self.adw_id, phase_id=phase.phase_id,
type="phase_start", name=params.name,
payload={"kind": params.kind, "owner": params.owner,
"description": params.description}))
self.console.phase_started(phase)
clock = time.monotonic()
try:
yield PhaseHandle(self, phase)
except BaseException as error:
phase.status = "fail" # success must be earned
phase.error = str(error)[:1000]
phase.ended_at = now_iso()
self.tracer.event(EventRecord(adw_id=self.adw_id, phase_id=phase.phase_id,
type="error", name=params.name,
payload={"error": phase.error}))
self.tracer.event(EventRecord(adw_id=self.adw_id, phase_id=phase.phase_id,
type="phase_end", name=params.name,
payload={"status": "fail"}))
self.tracer.phase_upsert(phase)
self.tracer.session_finish(self.adw_id, ok=False)
self.console.phase_ended(phase, time.monotonic() - clock)
self.console.session_finished(False, self.tokens, self.cost,
self.cfg.observability.db)
raise
else:
phase.status = "success"
phase.ended_at = now_iso()
self.tracer.event(EventRecord(adw_id=self.adw_id, phase_id=phase.phase_id,
type="phase_end", name=params.name,
payload={"status": "success"}))
self.tracer.phase_upsert(phase)
self.console.phase_ended(phase, time.monotonic() - clock)
# ── run outcome ─────────────────────────────────────────────────────────
def finish(self, accepted: bool = True, reason: str = "") -> int:
"""Finalize the run and return its exit code. Call this exactly once.
Two criteria, not one. Every phase must have passed, AND the ADW's own
acceptance test must hold. They are different questions on purpose: a
test phase that ran the suite did its job even when the suite came back
red, so the PHASE succeeds while the RUN must not.
This replaces a `succeeded` property that answered only the first
question — and, being a property with side effects, wrote the session
status and printed the banner before the caller's `and test.passed` was
ever evaluated. A run whose suite never passed was recorded green in the
db, on the terminal, and in the UI while exiting 1. Anyone reading the
trace saw success; only a CI job checking `$?` saw the truth. One call
now settles the db, the banner, and the exit code together, so the three
cannot disagree.
"""
phases_ok = bool(self.phases) and all(p.status == "success" for p in self.phases)
ok = phases_ok and accepted
if phases_ok and not accepted:
note = reason or "the run's acceptance criterion was not met"
self.tracer.event(EventRecord(
adw_id=self.adw_id,
phase_id=self.phases[-1].phase_id if self.phases else "",
type="error", name="not_accepted", payload={"reason": note}))
self.console.note(f"not accepted: {note}")
self.tracer.session_finish(self.adw_id, ok=ok)
self.console.session_finished(ok, self.tokens, self.cost, self.cfg.observability.db)
return 0 if ok else 1

View file

@ -0,0 +1,50 @@
"""Session lifecycle: pin-or-create an adw_id, build the Run object.
`ensure(cfg, adw_id)` joins the session if it exists or creates it under
exactly that id (pinned ids for repeatable runs); omitted, a fresh id is
minted and printed so the next ADW can pick it up.
"""
from __future__ import annotations
import os
import signal
import sys
from pathlib import Path
from .data_types import SSSFConfig
from .runner import Run
from .tracer import Tracer
from .utils import engineer_name, new_id
def _finalize_when_killed(run: Run) -> None:
"""A killed run still closes its own trace.
Python's default SIGTERM handling exits without unwinding, so `just kill`
(or any `kill <pid>`) would leave the session reading `running` forever and
its process rows open — the trace would claim work is in flight that is
already dead. Turning the signal into SystemExit both finalizes here and
lets the phase context manager record the phase as failed on the way out.
"""
def handler(signum, _frame):
run.tracer.session_finish(run.adw_id, ok=False) # also closes process rows
raise SystemExit(128 + signum)
for sig in (signal.SIGTERM, signal.SIGINT):
signal.signal(sig, handler)
def ensure(cfg: SSSFConfig, adw_id: str | None = None) -> Run:
adw_id = adw_id or new_id(8)
tracer = Tracer(cfg.observability.db,
f"{cfg.defaults.data_dir}/sessions/{adw_id}/events.jsonl")
run = Run(cfg=cfg, adw_id=adw_id, tracer=tracer, engineer=engineer_name())
tracer.session_start(adw_id, run.engineer, adw_name=Path(sys.argv[0]).stem)
# This process is the run. Record it before any phase opens, so a run that
# hangs in its first agent call is still killable by adw_id.
tracer.process_start(adw_id, "adw", "", os.getpid(),
" ".join([Path(sys.argv[0]).name, *sys.argv[1:]]))
_finalize_when_killed(run)
run.console.session_started(adw_id, run.engineer)
return run

View file

@ -0,0 +1,271 @@
"""Tracer: every event lands in JSONL and SQLite AS IT HAPPENS.
Files are the raw record; sssf.db is the queryable mirror the UI polls.
No push transport — the flow is always: agents -> sqlite -> web ui.
WAL mode so the UI can read while ADW processes write.
"""
from __future__ import annotations
import json
import sqlite3
from pathlib import Path
from .data_types import AgentConfig, EventRecord, GateReport, Phase
from .utils import ensure_dir, new_id, now_iso
SCHEMA = """
CREATE TABLE IF NOT EXISTS sessions (
adw_id TEXT PRIMARY KEY,
adw_name TEXT, -- ADW script(s) run, e.g. "adw_plan + adw_build_test"
request TEXT,
status TEXT,
engineer TEXT,
started_at TEXT, ended_at TEXT,
total_tokens INTEGER DEFAULT 0, total_cost REAL DEFAULT 0,
archived INTEGER DEFAULT 0 -- review triage, set by the UI; never by a run
);
CREATE TABLE IF NOT EXISTS phases (
phase_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
seq INTEGER,
name TEXT, kind TEXT, owner TEXT, description TEXT,
status TEXT DEFAULT 'fail',
attempt INTEGER DEFAULT 0, retries INTEGER DEFAULT 0,
error TEXT,
started_at TEXT, ended_at TEXT
);
CREATE TABLE IF NOT EXISTS events (
event_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases,
parent_id TEXT,
type TEXT,
name TEXT,
payload_json TEXT,
tokens INTEGER,
started_at TEXT, ended_at TEXT
);
CREATE TABLE IF NOT EXISTS envelopes (
envelope_id TEXT PRIMARY KEY,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases,
agent TEXT,
output_type TEXT,
payload_json TEXT,
valid INTEGER,
attempt INTEGER,
created_at TEXT
);
CREATE TABLE IF NOT EXISTS gate_results (
id INTEGER PRIMARY KEY AUTOINCREMENT,
adw_id TEXT REFERENCES sessions,
phase_id TEXT REFERENCES phases,
attempt INTEGER,
gate TEXT,
passed INTEGER,
violations_json TEXT,
checks_json TEXT, -- [{item, ok, note}] — WHAT the gate verified
created_at TEXT
);
CREATE TABLE IF NOT EXISTS processes (
id INTEGER PRIMARY KEY AUTOINCREMENT,
adw_id TEXT REFERENCES sessions,
kind TEXT, -- 'adw' (the workflow process) | 'agent' (a coding-agent child)
name TEXT, -- '' for the adw, the agent name for a child
pid INTEGER,
command TEXT, -- what the pid was, so a recycled pid is not killed by mistake
started_at TEXT, ended_at TEXT -- ended_at NULL = believed alive
);
CREATE TABLE IF NOT EXISTS agent_sessions (
adw_id TEXT REFERENCES sessions,
agent TEXT,
coding_agent TEXT, model TEXT, color TEXT,
session_id TEXT,
context_tokens INTEGER, -- window occupancy after the agent's last turn
context_window INTEGER, -- the model's ceiling; 0/NULL = unknown
created_at TEXT, last_used_at TEXT,
PRIMARY KEY (adw_id, agent)
);
"""
# Columns added after a schema shipped. CREATE TABLE IF NOT EXISTS never
# revisits an existing table, so additive changes need an explicit ALTER.
MIGRATIONS = [("agent_sessions", "color", "TEXT"),
("gate_results", "checks_json", "TEXT"),
("sessions", "adw_name", "TEXT"),
("agent_sessions", "context_tokens", "INTEGER"),
("agent_sessions", "context_window", "INTEGER"),
("sessions", "archived", "INTEGER DEFAULT 0")]
class Tracer:
def __init__(self, db_path: str | Path, events_jsonl: str | Path):
ensure_dir(Path(db_path).parent)
self.db_path = str(db_path)
self.events_jsonl = Path(events_jsonl)
ensure_dir(self.events_jsonl.parent)
self.conn = sqlite3.connect(self.db_path, isolation_level=None)
self.conn.execute("PRAGMA journal_mode=WAL;")
self.conn.execute("PRAGMA synchronous=NORMAL;")
self.conn.execute("PRAGMA busy_timeout=5000;")
self.conn.executescript(SCHEMA)
self._migrate()
def _migrate(self) -> None:
"""Additive column migrations, so a db from an older SSSF still opens."""
for table, column, decl in MIGRATIONS:
columns = {row[1] for row in self.conn.execute(f"PRAGMA table_info({table})")}
if column not in columns:
self.conn.execute(f"ALTER TABLE {table} ADD COLUMN {column} {decl}")
# ── events ──────────────────────────────────────────────────────────────
def event(self, record: EventRecord) -> str:
event_id = f"evt_{new_id(12)}"
ts = now_iso()
line = {"event_id": event_id, "ts": ts, **record.model_dump()}
with self.events_jsonl.open("a") as f:
f.write(json.dumps(line) + "\n")
self.conn.execute(
"INSERT INTO events (event_id, adw_id, phase_id, parent_id, type, name,"
" payload_json, tokens, started_at, ended_at) VALUES (?,?,?,?,?,?,?,?,?,?)",
(event_id, record.adw_id, record.phase_id, record.parent_id, record.type,
record.name, json.dumps(record.payload), record.tokens,
record.started_at or ts, record.ended_at),
)
return event_id
# ── sessions ────────────────────────────────────────────────────────────
def session_start(self, adw_id: str, engineer: str, adw_name: str | None = None) -> None:
self.conn.execute(
"INSERT INTO sessions (adw_id, status, engineer, started_at) VALUES (?,?,?,?) "
"ON CONFLICT(adw_id) DO UPDATE SET status='running'",
(adw_id, "running", engineer, now_iso()),
)
if not adw_name:
return
# A joined session chains ADWs — record each distinct one, in run order.
row = self.conn.execute("SELECT adw_name FROM sessions WHERE adw_id=?",
(adw_id,)).fetchone()
names = row[0].split(" + ") if row and row[0] else []
if adw_name not in names:
names.append(adw_name)
self.conn.execute("UPDATE sessions SET adw_name=? WHERE adw_id=?",
(" + ".join(names), adw_id))
def session_request(self, adw_id: str, request: str) -> None:
self.conn.execute("UPDATE sessions SET request=? WHERE adw_id=?",
(request[:500], adw_id))
def session_finish(self, adw_id: str, ok: bool) -> None:
self.conn.execute(
"UPDATE sessions SET status=?, ended_at=? WHERE adw_id=?",
("success" if ok else "fail", now_iso(), adw_id),
)
self.processes_end_all(adw_id) # nothing of this run is alive any more
def session_add_usage(self, adw_id: str, tokens: int, cost: float) -> None:
self.conn.execute(
"UPDATE sessions SET total_tokens=total_tokens+?, total_cost=total_cost+? WHERE adw_id=?",
(tokens, cost, adw_id),
)
# ── processes (adw_id → pid, so a hung run can be found and killed) ─────
def process_start(self, adw_id: str, kind: str, name: str, pid: int,
command: str) -> None:
"""Record a live process for this run.
A coding agent that hangs produces no events at all, which is exactly
when you need its pid — and `ps` cannot tell you which adw_id it
belongs to. Writing it here makes the trace the answer to "what is this
run running, and how do I stop it".
"""
self.conn.execute(
"INSERT INTO processes (adw_id, kind, name, pid, command, started_at)"
" VALUES (?,?,?,?,?,?)",
(adw_id, kind, name, pid, command[:500], now_iso()),
)
def process_end(self, adw_id: str, pid: int) -> None:
"""Mark the newest live row for this pid as finished."""
self.conn.execute(
"UPDATE processes SET ended_at=? WHERE id = ("
" SELECT id FROM processes WHERE adw_id=? AND pid=? AND ended_at IS NULL"
" ORDER BY id DESC LIMIT 1)",
(now_iso(), adw_id, pid),
)
def processes_end_all(self, adw_id: str) -> None:
"""Close out every live row for a run — called when the session ends."""
self.conn.execute(
"UPDATE processes SET ended_at=? WHERE adw_id=? AND ended_at IS NULL",
(now_iso(), adw_id),
)
# ── phases ──────────────────────────────────────────────────────────────
def max_phase_seq(self, adw_id: str) -> int:
"""Highest seq already recorded for this session; 0 when it is new.
A joined run continues the sequence instead of restarting at 1 — which
would collide with the first run's phases on both `seq` (breaking
ordering) and `phase_id` (silently overwriting a row through the
phase_upsert conflict clause).
"""
row = self.conn.execute("SELECT MAX(seq) FROM phases WHERE adw_id = ?",
(adw_id,)).fetchone()
return row[0] if row and row[0] is not None else 0
def phase_upsert(self, phase: Phase) -> None:
p = phase.params
self.conn.execute(
"INSERT INTO phases (phase_id, adw_id, seq, name, kind, owner, description,"
" status, attempt, retries, error, started_at, ended_at)"
" VALUES (?,?,?,?,?,?,?,?,?,?,?,?,?)"
" ON CONFLICT(phase_id) DO UPDATE SET status=excluded.status,"
" attempt=excluded.attempt, error=excluded.error, ended_at=excluded.ended_at",
(phase.phase_id, phase.adw_id, phase.seq, p.name, p.kind, p.owner,
p.description, phase.status, phase.attempt, p.retries, phase.error,
phase.started_at, phase.ended_at),
)
# ── envelopes / gates / agent sessions ──────────────────────────────────
def envelope_row(self, phase: Phase, agent: str, output_type: str,
payload_json: str, valid: bool, attempt: int) -> None:
self.conn.execute(
"INSERT INTO envelopes (envelope_id, adw_id, phase_id, agent, output_type,"
" payload_json, valid, attempt, created_at) VALUES (?,?,?,?,?,?,?,?,?)",
(f"env_{new_id(12)}", phase.adw_id, phase.phase_id, agent, output_type,
payload_json, int(valid), attempt, now_iso()),
)
def gate_row(self, phase: Phase, gate: str, report: GateReport, attempt: int) -> None:
"""The report carries both the verdict and the evidence behind it."""
self.conn.execute(
"INSERT INTO gate_results (adw_id, phase_id, attempt, gate, passed,"
" violations_json, checks_json, created_at) VALUES (?,?,?,?,?,?,?,?)",
(phase.adw_id, phase.phase_id, attempt, gate, int(report.passed),
json.dumps(report.violations),
json.dumps([c.model_dump() for c in report.checks]), now_iso()),
)
def agent_session_row(self, adw_id: str, agent: AgentConfig, session_id: str,
context_tokens: int = 0, context_window: int = 0) -> None:
"""The agent's config row is the source of truth for its label and color.
Context is carried here rather than derived from events because the lane
wants one number per agent — the latest — and a session that runs the
same agent twice overwrites it, exactly like model and session_id.
"""
ts = now_iso()
self.conn.execute(
"INSERT INTO agent_sessions (adw_id, agent, coding_agent, model, color,"
" session_id, context_tokens, context_window, created_at, last_used_at)"
" VALUES (?,?,?,?,?,?,?,?,?,?)"
" ON CONFLICT(adw_id, agent) DO UPDATE SET model=excluded.model,"
" color=excluded.color, session_id=excluded.session_id,"
" context_tokens=excluded.context_tokens,"
" context_window=excluded.context_window,"
" last_used_at=excluded.last_used_at",
(adw_id, agent.name, agent.coding_agent, agent.model, agent.color,
session_id, context_tokens, context_window, ts, ts),
)

View file

@ -0,0 +1,77 @@
"""Small shared helpers. Anything bigger belongs in its own module."""
from __future__ import annotations
import os
import secrets
import subprocess
from datetime import datetime, timezone
from pathlib import Path
from dotenv import load_dotenv
load_dotenv()
def operator_env() -> dict[str, str]:
"""The engineer's own environment, as their shell would hand it over.
Agents and quality blocks are meant to see exactly what the operator sees:
their PATH, their toolchains, their globally installed packages. Copying
os.environ gets almost all the way there — but ADWs launch under `uv run`,
which prepends its ephemeral venv's bin to PATH and sets VIRTUAL_ENV. That
venv holds the ADW's OWN dependencies (pydantic, pyyaml), not the
operator's, so anything a subprocess resolves through it — `python3`,
`pip`, every globally pip-installed CLI — silently becomes the wrong one.
Stripping the venv restores parity: `python3` in an agent's bash is the
same `python3` the engineer gets in their terminal. The ADW's own imports
are unaffected; this env is only ever handed to child processes.
"""
env = os.environ.copy()
venv = env.pop("VIRTUAL_ENV", "")
if not venv:
return env
venv_bin = str(Path(venv) / "bin")
parts = [p for p in env.get("PATH", "").split(os.pathsep) if p and p != venv_bin]
env["PATH"] = os.pathsep.join(parts)
return env
def new_id(length: int = 8) -> str:
return secrets.token_hex(length // 2)
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat(timespec="milliseconds")
def ensure_dir(path: str | Path) -> Path:
p = Path(path)
p.mkdir(parents=True, exist_ok=True)
return p
def resolve_prompt(arg: str) -> str:
"""CLI prompt arg: a file path resolves to its contents, else inline text."""
try:
p = Path(arg)
if p.is_file():
return p.read_text()
except OSError:
pass
return arg
def engineer_name() -> str:
name = os.environ.get("ENGINEER_NAME", "").strip()
if name:
return name
try:
out = subprocess.run(["git", "config", "user.name"],
capture_output=True, text=True, timeout=5)
if out.returncode == 0 and out.stdout.strip():
return out.stdout.strip()
except OSError:
pass
return os.environ.get("USER", "engineer")

View file

@ -0,0 +1,45 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan — one-shot planning workflow.
Usage:
uv run adws/adw_plan.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> planner
"""
import argparse
import sys
from adw_modules import agents, gates, session, utils
from adw_modules.data_types import AgentCall, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner"]
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,55 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan Build — two-agent chain: planner -> envelope -> builder.
Usage:
uv run adws/adw_plan_build.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> planner -> builder -> git(commit)
"""
import argparse
import sys
from adw_modules import agents, gates, git_helper, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner", "builder"]
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the plan exactly")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.diff_matches_claims]))
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Land the builder's changes, using the message it wrote")) as ph:
message = build.commit_message or f"sssf({run.adw_id}): {build.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,86 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan Build Test — the full starter chain.
Usage:
uv run adws/adw_plan_build_test.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> planner -> builder -> code(test) [-> builder(fix) -> code(test) ... bounded] -> git(commit)
Testing is CODE: the suite's command lives in adw_modules/quality.py, so no
agent spends a context window rediscovering it. Failures flow back to the
builder as an envelope, and only an exhausted fix loop fails the run.
"""
import argparse
import sys
from adw_modules import agents, gates, git_helper, quality, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner", "builder"]
MAX_FIX_LOOPS = 3
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
def record(ph, result) -> None:
passed = sum(1 for check in result.checks if check.passed)
ph.log(passed=result.passed, checks=f"{passed}/{len(result.checks)}",
artifacts=", ".join(result.artifacts))
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the plan exactly")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.artifacts_exist]))
test = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"test_{i}", kind="code", owner="quality",
description="Run the suite — a known command, so code runs "
"it and no agent has to rediscover it")) as ph:
test = quality.run_tests(run)
record(ph, test)
if test.passed:
break
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description="Repair what the suite reported, from its "
"verbatim output")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(test, "tests"),
gates=[gates.artifacts_exist]))
# Only tested work gets committed — a red suite leaves the tree uncommitted.
if test is not None and test.passed:
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Land the code only after the suite came back green")) as ph:
message = previous.commit_message or f"sssf({run.adw_id}): {previous.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
return run.finish(accepted=test is not None and test.passed,
reason=f"the suite still failed after {MAX_FIX_LOOPS} fix attempt(s)")
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,98 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Plan Build Test Quality — full agent chain plus deterministic quality.
Usage:
uv run adws/adw_plan_build_test_quality.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> planner -> builder -> [code(verify) -> code(test) -> builder(fix)] bounded -> git(commit)
Verify and test are CODE, not agents. Their commands are known, so running them
needs no judgement — only repairing them does. A failing block does not fail its
phase: the runner did its job, the code is what failed. The failure becomes an
envelope and flows back into the builder, and only an exhausted repair loop
fails the run.
"""
import argparse
import sys
from adw_modules import agents, gates, git_helper, quality, session, utils
from adw_modules.data_types import AgentCall, BuildOutput, PhaseParams, PlanOutput
REQUIRED_AGENTS = ["planner", "builder"]
MAX_FIX_LOOPS = 3
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the plan exactly")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.diff_matches_claims]))
def record(ph, result) -> None:
passed = sum(1 for check in result.checks if check.passed)
ph.log(passed=result.passed, checks=f"{passed}/{len(result.checks)}",
artifacts=", ".join(result.artifacts))
test_result = None
quality_result = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"verify_{i}", kind="code", owner="quality",
description="Lint, typecheck, and build before testing")) as ph:
quality_result = quality.run_quality(run)
record(ph, quality_result)
# run_quality() already includes the test block; a repo that wants tests
# in their own phase can split them out the way this comment does.
test_result = quality_result
if quality_result.passed and test_result.passed:
break
if i == MAX_FIX_LOOPS:
break
# Whichever block failed becomes the builder's spec — verbatim command
# output, no parser standing between the failure and the fix.
broken = quality_result if not quality_result.passed else test_result
what = "verification" if not quality_result.passed else "tests"
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description=f"Resolve the reported {what} failures")) as ph:
previous = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(broken, what),
gates=[gates.diff_matches_claims]))
verified = (quality_result is not None and quality_result.passed
and test_result is not None and test_result.passed)
if verified:
with run.phase(PhaseParams(name="commit", kind="code", owner="git",
description="Commit the tested and quality-verified working tree")) as ph:
message = previous.commit_message or f"sssf({run.adw_id}): {previous.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
return run.finish(accepted=verified,
reason=f"verify/test never came back clean after {MAX_FIX_LOOPS} fix attempt(s)")
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,44 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Prompt — the smallest ADW: one agent, one prompt, traced end-to-end.
Usage:
uv run adws/adw_prompt.py "<prompt or path/to/prompt.md>" [--agent builder] [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> <agent>
"""
import argparse
import sys
from adw_modules import agents, session, utils
from adw_modules.data_types import AgentCall, GenericOutput, PhaseParams
def main(prompt: str, agent: str = "builder",
config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, [agent])
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="prompt", kind="agent", owner=agent,
description=f"Send the request straight to {agent} and parse its envelope")) as ph:
ph.call(AgentCall(output_type=GenericOutput, prompt=prompt))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--agent", default="builder", help="agent name from the config")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.agent, args.config, args.adw_id))

View file

@ -0,0 +1,49 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Quality — lint, typecheck, and build the project.
Usage:
uv run adws/adw_quality.py "<reason for the quality run>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> code(quality)
"""
import argparse
import sys
from adw_modules import agents, quality, session, utils
from adw_modules.data_types import PhaseParams
REQUIRED_AGENTS: list[str] = []
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture why quality verification was requested")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="quality", kind="code", owner="quality",
description="Run the deterministic quality blocks")) as ph:
result = quality.run_quality(run)
passed = sum(1 for check in result.checks if check.passed)
ph.log(passed=result.passed, checks=f"{passed}/{len(result.checks)}",
artifacts=", ".join(result.artifacts))
if not result.passed:
raise RuntimeError("quality failed: " + "; ".join(result.failures))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,45 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Scout — read-only recon workflow. Just looking for stuff.
Usage:
uv run adws/adw_scout.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> scout
"""
import argparse
import sys
from adw_modules import agents, gates, session, utils
from adw_modules.data_types import AgentCall, PhaseParams, ScoutOutput
REQUIRED_AGENTS = ["scout"]
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt)
with run.phase(PhaseParams(name="scout", kind="agent", owner="scout",
description="Find and report where things live — change nothing")) as ph:
ph.call(AgentCall(output_type=ScoutOutput, prompt=prompt,
gates=[gates.artifacts_exist]))
return run.finish()
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,183 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Simple SDLC — plan, build, test, review, document, committing as it goes.
Usage:
uv run adws/adw_simple_sdlc.py "<prompt or path/to/prompt.md>" [--config adws/adw_sssf_config/sssf.config.yaml] [--adw-id a1b2c3d4]
Phases: engineer(request) -> planner -> git(commit_plan)
-> builder -> code(test) [-> builder(fix) -> code(test) ... bounded]
-> reviewer [-> builder(revise) -> reviewer ... bounded]
-> code(retest, only if a revision changed code)
-> git(commit_build) -> code(changes) -> documenter -> git(commit_docs)
Three commits, three work products, three authors. The plan, the code, and the
write-up each land in their own commit, and each commit message is the words of
the agent that produced it — `commit_message` on PlanOutput describes the spec,
on BuildOutput the code, on DocumentOutput the write-up. No agent's sentence is
ever reused for another agent's diff.
Testing is CODE, not an agent. `bun test` is a command, not a judgement call:
an agent rediscovering it every run costs a million tokens to learn what a
subprocess already knows. Failures travel back to the builder as an envelope,
so the repair loop is unchanged — only the runner became free and repeatable.
Two different questions still get asked, in order. The suite asks "does it
run"; the reviewer asks "is this what was asked for", against `plan.md` — and
neither can answer the other's. A revision that closes a review finding
re-enters the suite, so the tree that gets committed is the tree that was both
tested and approved.
The code commit lands after verification, not straight after the build: fixes
and revisions are part of the same work product, and red code has no business
on the branch. A run that fails verification therefore leaves the plan
committed and the working tree dirty — the spec is a real artifact either way,
and the unfinished code stays where the engineer can see it.
The documenter measures against the commit this run STARTED from, not against
`main`, because by then the run has moved `main` itself. That baseline is
pinned before the first commit phase and printed in the request phase.
"""
import argparse
import sys
from adw_modules import agents, changes, gates, git_helper, quality, session, utils
from adw_modules.data_types import (AgentCall, BuildOutput, ChangeCapture,
DocumentOutput, PhaseParams, PlanOutput,
ReviewOutput)
REQUIRED_AGENTS = ["planner", "builder", "reviewer", "documenter"]
MAX_FIX_LOOPS = 3
MAX_REVISION_LOOPS = 2
DOCUMENT_NOTES = ("Read diff_path in full before writing. Document only what the "
"diff shows, then copy the write-up into app_docs/ as your task "
"describes.")
def main(prompt: str, config: str = "adws/adw_sssf_config/sssf.config.yaml", adw_id: str | None = None) -> int:
cfg = agents.load_config(config)
agents.validate(cfg, REQUIRED_AGENTS)
run = session.ensure(cfg, adw_id)
baseline = git_helper.rev("HEAD") # pinned before this run commits anything
def commit(ph, envelope) -> None:
"""Commit what the preceding phase produced, in that agent's own words."""
message = envelope.commit_message or f"sssf({run.adw_id}): {envelope.summary}"
ph.log(sha=git_helper.commit_all(message), message=message)
def record(ph, result) -> None:
"""Log a deterministic block's verdict — the same shape every ADW uses."""
passed = sum(1 for check in result.checks if check.passed)
ph.log(passed=result.passed, checks=f"{passed}/{len(result.checks)}",
artifacts=", ".join(result.artifacts))
with run.phase(PhaseParams(name="request", kind="engineer", owner=run.engineer,
description="Capture the incoming ask")) as ph:
ph.log(input=prompt, baseline=git_helper.short_sha(baseline))
with run.phase(PhaseParams(name="plan", kind="agent", owner="planner",
description="Turn the request into an implementable plan")) as ph:
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="commit_plan", kind="code", owner="git",
description="Put the spec on record before any code exists to blur it")) as ph:
commit(ph, plan)
with run.phase(PhaseParams(name="build", kind="agent", owner="builder",
description="Implement the plan exactly")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=plan,
gates=[gates.diff_matches_claims]))
test = None
for i in range(1, MAX_FIX_LOOPS + 1):
with run.phase(PhaseParams(name=f"test_{i}", kind="code", owner="quality",
description="Run the suite — a known command, so code runs "
"it and no agent has to rediscover it")) as ph:
test = quality.run_tests(run)
record(ph, test)
if test.passed:
break
with run.phase(PhaseParams(name=f"fix_{i}", kind="agent", owner="builder", retries=1,
description="Repair what the suite reported, from its "
"verbatim output")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt,
previous=quality.as_envelope(test, "tests"),
gates=[gates.diff_matches_claims]))
review = None
revised = False
for i in range(1, MAX_REVISION_LOOPS + 1):
with run.phase(PhaseParams(name=f"review_{i}", kind="agent", owner="reviewer",
description="Confirm the build matches the plan")) as ph:
review = ph.call(AgentCall(output_type=ReviewOutput, prompt=prompt, previous=build,
gates=[gates.artifacts_exist, gates.verdict_consistent]))
if review.approved or i == MAX_REVISION_LOOPS:
break
with run.phase(PhaseParams(name=f"revise_{i}", kind="agent", owner="builder", retries=1,
description="Close the reviewer's blocking findings")) as ph:
build = ph.call(AgentCall(output_type=BuildOutput, prompt=prompt, previous=review,
gates=[gates.diff_matches_claims]))
revised = True
# A revision edited code after the suite last ran, so the green light is
# stale. Re-run it rather than commit on a result that predates the change.
if revised and review is not None and review.approved:
with run.phase(PhaseParams(name="retest", kind="code", owner="quality",
description="Re-run the suite — the revision changed code "
"after the last green result")) as ph:
test = quality.run_tests(run)
record(ph, test)
# Red tests or a rejected review stop the chain here: the code stays
# uncommitted and nothing is documented, because there is nothing worth
# describing yet. The plan commit stands — it is a record of what was asked.
verified = (test is not None and test.passed
and review is not None and review.approved)
if verified:
with run.phase(PhaseParams(name="commit_build", kind="code", owner="git",
description="Land the code only now: green suite, approved review")) as ph:
commit(ph, build)
with run.phase(PhaseParams(name="changes", kind="code", owner="git",
description="Diff the whole run against its pinned baseline, for the documenter")) as ph:
changeset = changes.capture(run, ChangeCapture(base=baseline))
ph.log(base=f"{changeset.base.label} @ {changeset.base.commit[:7]}",
reason=changeset.base.reason,
files=len(changeset.files) + len(changeset.untracked),
lines=f"+{changeset.insertions} -{changeset.deletions}",
diff=changeset.diff_path)
if changeset.empty:
raise RuntimeError(
f"nothing changed since {changeset.base.label} "
f"({changeset.base.reason}) — there is nothing to document.")
with run.phase(PhaseParams(name="document", kind="agent", owner="documenter", retries=1,
description="Write up the completed change")) as ph:
document = ph.call(AgentCall(output_type=DocumentOutput, prompt=prompt,
previous=changes.as_envelope(changeset, DOCUMENT_NOTES),
gates=[gates.artifacts_exist, gates.files_non_empty]))
with run.phase(PhaseParams(name="commit_docs", kind="code", owner="git",
description="Ship the write-up in its own commit, beside the code it describes")) as ph:
commit(ph, document)
return run.finish(accepted=verified,
reason="the suite or the review never came back clean")
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("prompt", help="inline text or a path to a prompt file")
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
parser.add_argument("--adw-id", default=None, help="join or pin an existing session")
args = parser.parse_args()
sys.exit(main(utils.resolve_prompt(args.prompt), args.config, args.adw_id))

View file

@ -0,0 +1,36 @@
#!/usr/bin/env -S uv run
# /// script
# dependencies = ["pydantic", "python-dotenv", "pyyaml", "rich"]
# ///
"""ADW Validate — check the config without running anything.
Loads the roster and validates every agent: the name resolves, the coding
agent is implemented, the prompt files exist, and the model resolves in the
harness's catalog. Exits 0 on success, 1 with a list of problems otherwise.
Usage:
uv run adws/adw_validate.py [--config adws/adw_sssf_config/sssf.config.yaml]
"""
import argparse
import sys
from adw_modules import agents
def main(config: str = "adws/adw_sssf_config/sssf.config.yaml") -> int:
cfg = agents.load_config(config)
names = [a.name for a in cfg.agents]
if not names:
print(f"config {config} declares no agents")
return 1
agents.validate(cfg, names)
print(f"config OK; agents: {', '.join(names)}")
return 0
if __name__ == "__main__":
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--config", default="adws/adw_sssf_config/sssf.config.yaml")
args = parser.parse_args()
sys.exit(main(args.config))

28
sssf/templates/env.sample Normal file
View file

@ -0,0 +1,28 @@
# SSSF environment. Copy to .env at the target repo root and fill in.
#
# WHICH KEYS YOU NEED DEPENDS ON YOUR ROSTER.
# Every `model:` in adws/adw_sssf_config/sssf.config.yaml is written as
# provider/model-id. The provider half decides which key must be set, and which
# key pi reads for that provider comes from ~/.pi/agent/models.json.
#
# The starter roster names three providers, so it needs three keys:
# google/gemini-3.6-flash served via openrouter -> OPENROUTER_API_KEY
# fireworks/accounts/fireworks/models/kimi-k3 fireworks -> FIREWORKS_API_KEY
# openai/gpt-5.6-terra, openai/gpt-5.6-luna openai -> OPENAI_API_KEY
#
# Point every agent at one provider and you only need that provider's key.
# Setting `defaults.model` and deleting the per-agent `model:` overrides is the
# fastest way to run the whole roster on a single key.
#
# Nothing validates this for you. `agents.validate()` checks that a model is
# written as provider/id, not that the provider is reachable or that its key is
# set, so a missing key shows up when that agent runs, not at startup.
OPENROUTER_API_KEY=
FIREWORKS_API_KEY=
OPENAI_API_KEY=
# Optional overrides
# PI_PATH=pi # pi binary if not on PATH
# PI_MODELS_PATH=~/.pi/agent/models.json
# ENGINEER_NAME= # engineer lane label, defaults to git user.name

View file

@ -0,0 +1,593 @@
/**
* Subagent Widget — /sub, /subclear, /subrm, /subcont commands with stacking live widgets
*
* Each /sub spawns a background Pi subagent with its own persistent session,
* enabling conversation continuations via /subcont.
*
* Usage: pi -e extensions/subagent-widget.ts
* Then:
* /sub list files and summarize — spawn using the parent model/thinking
* /sub --model openai/gpt-5 --thinking high review this code
* /subcont 1 --thinking xhigh now write tests for it
* /subrm 2 — remove subagent #2 widget
* /subclear — clear all subagent widgets
*/
import { StringEnum, type ThinkingLevel } from "@mariozechner/pi-ai";
import type { ExtensionAPI } from "@mariozechner/pi-coding-agent";
import { DynamicBorder } from "@mariozechner/pi-coding-agent";
import { Container, Text } from "@mariozechner/pi-tui";
import { Type } from "@sinclair/typebox";
const { spawn } = require("child_process") as any;
import * as fs from "fs";
import * as os from "os";
import * as path from "path";
import { applyExtensionDefaults } from "./themeMap.ts";
const FALLBACK_MODEL = "openrouter/google/gemini-3.5-flash";
const THINKING_OVERRIDES = ["low", "medium", "high", "xhigh"] as const;
type ThinkingOverride = (typeof THINKING_OVERRIDES)[number];
interface SpawnOptions {
model?: string;
thinking?: ThinkingOverride;
}
interface SubState {
id: number;
status: "running" | "done" | "error";
task: string;
textChunks: string[];
toolCount: number;
elapsed: number;
sessionFile: string; // persistent JSONL session path — used by /subcont to resume
turnCount: number; // increments each time /subcont continues this agent
model: string;
thinking: ThinkingLevel;
proc?: any; // active ChildProcess ref (for kill on /subrm)
}
interface ParsedCommand {
options: SpawnOptions;
rest: string;
error?: string;
}
function readCommandValue(input: string): { value?: string; rest: string } {
const trimmed = input.trimStart();
if (!trimmed) return { rest: "" };
const quote = trimmed[0];
if (quote === '"' || quote === "'") {
const end = trimmed.indexOf(quote, 1);
if (end === -1) return { rest: trimmed };
return { value: trimmed.slice(1, end), rest: trimmed.slice(end + 1) };
}
const end = trimmed.search(/\s/);
return end === -1
? { value: trimmed, rest: "" }
: { value: trimmed.slice(0, end), rest: trimmed.slice(end) };
}
function parseCommandOptions(input: string): ParsedCommand {
const options: SpawnOptions = {};
let rest = input.trimStart();
while (rest.startsWith("--")) {
const flagMatch = rest.match(/^--(model|thinking)(?:=([^\s]+))?(?:\s+|$)/);
if (!flagMatch) {
const flag = rest.match(/^\S+/)?.[0] || rest;
return { options, rest: "", error: `Unknown or malformed option: ${flag}` };
}
const flag = flagMatch[1];
let value = flagMatch[2];
rest = rest.slice(flagMatch[0].length);
if (!value) {
const parsed = readCommandValue(rest);
value = parsed.value;
rest = parsed.rest;
}
if (!value) return { options, rest: "", error: `Missing value for --${flag}` };
if (flag === "model") {
options.model = value;
rest = rest.trimStart();
continue;
}
const thinking = value.toLowerCase();
if (!THINKING_OVERRIDES.includes(thinking as ThinkingOverride)) {
return {
options,
rest: "",
error: "Thinking must be one of: low, medium, high, xhigh",
};
}
options.thinking = thinking as ThinkingOverride;
rest = rest.trimStart();
}
return { options, rest: rest.trim() };
}
export default function (pi: ExtensionAPI) {
const agents: Map<number, SubState> = new Map();
let nextId = 1;
let widgetCtx: any;
// ── Session file helpers ──────────────────────────────────────────────────
function makeSessionFile(id: number): string {
const dir = path.join(os.homedir(), ".pi", "agent", "sessions", "subagents");
fs.mkdirSync(dir, { recursive: true });
return path.join(dir, `subagent-${id}-${Date.now()}.jsonl`);
}
// ── Widget rendering ──────────────────────────────────────────────────────
function updateWidgets() {
if (!widgetCtx) return;
for (const [id, state] of Array.from(agents.entries())) {
const key = `sub-${id}`;
widgetCtx.ui.setWidget(key, (_tui: any, theme: any) => {
const container = new Container();
const borderFn = (s: string) => theme.fg("dim", s);
container.addChild(new Text("", 0, 0)); // top margin
container.addChild(new DynamicBorder(borderFn));
const content = new Text("", 1, 0);
container.addChild(content);
container.addChild(new DynamicBorder(borderFn));
return {
render(width: number): string[] {
const lines: string[] = [];
const statusColor = state.status === "running" ? "accent"
: state.status === "done" ? "success" : "error";
const statusIcon = state.status === "running" ? "●"
: state.status === "done" ? "✓" : "✗";
const taskPreview = state.task.length > 40
? state.task.slice(0, 37) + "..."
: state.task;
const turnLabel = state.turnCount > 1
? theme.fg("dim", ` · Turn ${state.turnCount}`)
: "";
lines.push(
theme.fg(statusColor, `${statusIcon} Subagent #${state.id}`) +
turnLabel +
theme.fg("dim", ` ${taskPreview}`) +
theme.fg("dim", ` (${Math.round(state.elapsed / 1000)}s)`) +
theme.fg("dim", ` | Tools: ${state.toolCount}`)
);
const fullText = state.textChunks.join("");
const lastLine = fullText.split("\n").filter((l: string) => l.trim()).pop() || "";
if (lastLine) {
const trimmed = lastLine.length > width - 10
? lastLine.slice(0, width - 13) + "..."
: lastLine;
lines.push(theme.fg("muted", ` ${trimmed}`));
}
content.setText(lines.join("\n"));
return container.render(width);
},
invalidate() {
container.invalidate();
},
};
});
}
}
// ── Streaming helpers ─────────────────────────────────────────────────────
function processLine(state: SubState, line: string) {
if (!line.trim()) return;
try {
const event = JSON.parse(line);
const type = event.type;
if (type === "message_update") {
const delta = event.assistantMessageEvent;
if (delta?.type === "text_delta") {
state.textChunks.push(delta.delta || "");
updateWidgets();
}
} else if (type === "tool_execution_start") {
state.toolCount++;
updateWidgets();
}
} catch {}
}
function spawnAgent(
state: SubState,
prompt: string,
ctx: any,
options: SpawnOptions = {},
): Promise<void> {
const parentProvider = ctx.model?.provider?.trim();
const parentModelId = ctx.model?.id?.trim();
const hasParentModel = parentProvider && parentModelId
&& parentProvider !== "unknown" && parentModelId !== "unknown";
const parentModel = hasParentModel
? `${parentProvider}/${parentModelId}`
: FALLBACK_MODEL;
const model = options.model?.trim() || parentModel;
const thinking = options.thinking || pi.getThinkingLevel();
state.model = model;
state.thinking = thinking;
return new Promise<void>((resolve) => {
const proc = spawn("pi", [
"--mode", "json",
"-p",
"--session", state.sessionFile, // persistent session for /subcont resumption
"--no-extensions",
"--model", model,
"--tools", "read,bash,grep,find,ls",
"--thinking", thinking,
prompt,
], {
stdio: ["ignore", "pipe", "pipe"],
env: { ...process.env },
});
state.proc = proc;
const startTime = Date.now();
const timer = setInterval(() => {
state.elapsed = Date.now() - startTime;
updateWidgets();
}, 1000);
let buffer = "";
proc.stdout!.setEncoding("utf-8");
proc.stdout!.on("data", (chunk: string) => {
buffer += chunk;
const lines = buffer.split("\n");
buffer = lines.pop() || "";
for (const line of lines) processLine(state, line);
});
proc.stderr!.setEncoding("utf-8");
proc.stderr!.on("data", (chunk: string) => {
if (chunk.trim()) {
state.textChunks.push(chunk);
updateWidgets();
}
});
proc.on("close", (code) => {
if (buffer.trim()) processLine(state, buffer);
clearInterval(timer);
state.elapsed = Date.now() - startTime;
state.status = code === 0 ? "done" : "error";
state.proc = undefined;
updateWidgets();
const result = state.textChunks.join("");
ctx.ui.notify(
`Subagent #${state.id} ${state.status} in ${Math.round(state.elapsed / 1000)}s`,
state.status === "done" ? "success" : "error"
);
pi.sendMessage({
customType: "subagent-result",
content: `Subagent #${state.id}${state.turnCount > 1 ? ` (Turn ${state.turnCount})` : ""} finished "${prompt}" in ${Math.round(state.elapsed / 1000)}s.\n\nResult:\n${result.slice(0, 8000)}${result.length > 8000 ? "\n\n... [truncated]" : ""}`,
display: true,
}, { deliverAs: "followUp", triggerTurn: true });
resolve();
});
proc.on("error", (err) => {
clearInterval(timer);
state.status = "error";
state.proc = undefined;
state.textChunks.push(`Error: ${err.message}`);
updateWidgets();
resolve();
});
});
}
// ── Tools for the Main Agent ──────────────────────────────────────────────
pi.registerTool({
name: "subagent_create",
description: "Spawn a background subagent. Thinking level is required and is the primary way to match the subagent to task complexity: low for lightweight/simple tasks, medium for routine tasks needing moderate reasoning, high for complex multi-step work, and xhigh for the hardest tasks or when accuracy and performance are critical. Unless the user explicitly requests a specific model, omit model and use the default inherited parent model. Returns immediately and delivers results as a follow-up message.",
parameters: Type.Object({
task: Type.String({ description: "The complete task description for the subagent to perform" }),
model: Type.Optional(Type.String({
description: "Leave blank or omit unless the user explicitly requests a specific model. Do not choose a different model autonomously. When explicitly requested, provide the override in provider/model form. The default reuses the parent caller's current model and falls back to openrouter/google/gemini-3.5-flash only if the parent has no model.",
})),
thinking: StringEnum([...THINKING_OVERRIDES], {
description: "Required thinking level. Use low for lightweight/simple tasks; medium for routine tasks needing moderate reasoning; high for complex, multi-step, or ambiguous work; and xhigh for the hardest tasks or when accuracy and performance are critical. Pi may clamp the value to the selected model's supported maximum.",
}),
}),
execute: async (callId, args, _signal, _onUpdate, ctx) => {
widgetCtx = ctx;
const id = nextId++;
const state: SubState = {
id,
status: "running",
task: args.task,
textChunks: [],
toolCount: 0,
elapsed: 0,
sessionFile: makeSessionFile(id),
turnCount: 1,
model: "",
thinking: pi.getThinkingLevel(),
};
agents.set(id, state);
updateWidgets();
// Fire-and-forget
spawnAgent(state, args.task, ctx, { model: args.model, thinking: args.thinking });
return {
content: [{ type: "text", text: `Subagent #${id} spawned with ${state.model} (${state.thinking} thinking) and is running in background.` }],
};
},
});
pi.registerTool({
name: "subagent_continue",
description: "Continue an existing subagent conversation. Thinking level is required and is the primary way to match this turn to task complexity: low for lightweight/simple tasks, medium for routine tasks needing moderate reasoning, high for complex multi-step work, and xhigh for the hardest tasks or when accuracy and performance are critical. Unless the user explicitly requests a specific model, omit model and use the default inherited parent model. Returns immediately while it runs in the background.",
parameters: Type.Object({
id: Type.Number({ description: "The ID of the subagent to continue" }),
prompt: Type.String({ description: "The follow-up prompt or new instructions" }),
model: Type.Optional(Type.String({
description: "Leave blank or omit unless the user explicitly requests a specific model. Do not choose a different model autonomously. When explicitly requested, provide the override in provider/model form for this turn. The default reuses the parent caller's current model.",
})),
thinking: StringEnum([...THINKING_OVERRIDES], {
description: "Required thinking level for this turn. Use low for lightweight/simple tasks; medium for routine tasks needing moderate reasoning; high for complex, multi-step, or ambiguous work; and xhigh for the hardest tasks or when accuracy and performance are critical. Pi may clamp the value to the selected model's supported maximum.",
}),
}),
execute: async (callId, args, _signal, _onUpdate, ctx) => {
widgetCtx = ctx;
const state = agents.get(args.id);
if (!state) {
return { content: [{ type: "text", text: `Error: No subagent #${args.id} found.` }] };
}
if (state.status === "running") {
return { content: [{ type: "text", text: `Error: Subagent #${args.id} is still running.` }] };
}
state.status = "running";
state.task = args.prompt;
state.textChunks = [];
state.elapsed = 0;
state.turnCount++;
updateWidgets();
ctx.ui.notify(`Continuing Subagent #${args.id} (Turn ${state.turnCount})…`, "info");
spawnAgent(state, args.prompt, ctx, { model: args.model, thinking: args.thinking });
return {
content: [{ type: "text", text: `Subagent #${args.id} continuing with ${state.model} (${state.thinking} thinking) in background.` }],
};
},
});
pi.registerTool({
name: "subagent_remove",
description: "Remove a specific subagent. Kills it if it's currently running.",
parameters: Type.Object({
id: Type.Number({ description: "The ID of the subagent to remove" }),
}),
execute: async (callId, args, _signal, _onUpdate, ctx) => {
widgetCtx = ctx;
const state = agents.get(args.id);
if (!state) {
return { content: [{ type: "text", text: `Error: No subagent #${args.id} found.` }] };
}
if (state.proc && state.status === "running") {
state.proc.kill("SIGTERM");
}
ctx.ui.setWidget(`sub-${args.id}`, undefined);
agents.delete(args.id);
return {
content: [{ type: "text", text: `Subagent #${args.id} removed successfully.` }],
};
},
});
pi.registerTool({
name: "subagent_list",
description: "List all active and finished subagents, showing their IDs, tasks, and status.",
parameters: Type.Object({}),
execute: async () => {
if (agents.size === 0) {
return { content: [{ type: "text", text: "No active subagents." }] };
}
const list = Array.from(agents.values()).map(s =>
`#${s.id} [${s.status.toUpperCase()}] (Turn ${s.turnCount}, ${s.model}, ${s.thinking}) - ${s.task}`
).join("\n");
return {
content: [{ type: "text", text: `Subagents:\n${list}` }],
};
},
});
// ── /sub [--model <model>] [--thinking <level>] <task> ────────────────────
pi.registerCommand("sub", {
description: "Spawn a subagent: /sub [--model provider/model] [--thinking low|medium|high|xhigh] <task>",
handler: async (args, ctx) => {
widgetCtx = ctx;
const parsed = parseCommandOptions(args || "");
if (parsed.error) {
ctx.ui.notify(parsed.error, "error");
return;
}
const task = parsed.rest;
if (!task) {
ctx.ui.notify("Usage: /sub [--model provider/model] [--thinking low|medium|high|xhigh] <task>", "error");
return;
}
const id = nextId++;
const state: SubState = {
id,
status: "running",
task,
textChunks: [],
toolCount: 0,
elapsed: 0,
sessionFile: makeSessionFile(id),
turnCount: 1,
model: "",
thinking: pi.getThinkingLevel(),
};
agents.set(id, state);
updateWidgets();
// Fire-and-forget
spawnAgent(state, task, ctx, parsed.options);
ctx.ui.notify(`Subagent #${id}: ${state.model} (${state.thinking} thinking)`, "info");
},
});
// ── /subcont <id> [--model <model>] [--thinking <level>] <prompt> ─────────
pi.registerCommand("subcont", {
description: "Continue a subagent: /subcont <id> [--model provider/model] [--thinking low|medium|high|xhigh] <prompt>",
handler: async (args, ctx) => {
widgetCtx = ctx;
const trimmed = args?.trim() ?? "";
const idMatch = trimmed.match(/^(\d+)(?:\s+|$)/);
if (!idMatch) {
ctx.ui.notify("Usage: /subcont <id> [--model provider/model] [--thinking low|medium|high|xhigh] <prompt>", "error");
return;
}
const num = parseInt(idMatch[1], 10);
const parsed = parseCommandOptions(trimmed.slice(idMatch[0].length));
if (parsed.error) {
ctx.ui.notify(parsed.error, "error");
return;
}
const prompt = parsed.rest;
if (!prompt) {
ctx.ui.notify("Usage: /subcont <id> [--model provider/model] [--thinking low|medium|high|xhigh] <prompt>", "error");
return;
}
const state = agents.get(num);
if (!state) {
ctx.ui.notify(`No subagent #${num} found. Use /sub to create one.`, "error");
return;
}
if (state.status === "running") {
ctx.ui.notify(`Subagent #${num} is still running — wait for it to finish first.`, "warning");
return;
}
// Resume: update state for a new turn
state.status = "running";
state.task = prompt;
state.textChunks = [];
state.elapsed = 0;
state.turnCount++;
updateWidgets();
ctx.ui.notify(`Continuing Subagent #${num} (Turn ${state.turnCount})…`, "info");
// Fire-and-forget — reuses the same sessionFile for conversation history
spawnAgent(state, prompt, ctx, parsed.options);
ctx.ui.notify(`Subagent #${num}: ${state.model} (${state.thinking} thinking)`, "info");
},
});
// ── /subrm <number> ───────────────────────────────────────────────────────
pi.registerCommand("subrm", {
description: "Remove a specific subagent widget: /subrm <number>",
handler: async (args, ctx) => {
widgetCtx = ctx;
const num = parseInt(args?.trim() ?? "", 10);
if (isNaN(num)) {
ctx.ui.notify("Usage: /subrm <number>", "error");
return;
}
const state = agents.get(num);
if (!state) {
ctx.ui.notify(`No subagent #${num} found.`, "error");
return;
}
// Kill the process if still running
if (state.proc && state.status === "running") {
state.proc.kill("SIGTERM");
ctx.ui.notify(`Subagent #${num} killed and removed.`, "warning");
} else {
ctx.ui.notify(`Subagent #${num} removed.`, "info");
}
ctx.ui.setWidget(`sub-${num}`, undefined);
agents.delete(num);
},
});
// ── /subclear ─────────────────────────────────────────────────────────────
pi.registerCommand("subclear", {
description: "Clear all subagent widgets",
handler: async (_args, ctx) => {
widgetCtx = ctx;
let killed = 0;
for (const [id, state] of Array.from(agents.entries())) {
if (state.proc && state.status === "running") {
state.proc.kill("SIGTERM");
killed++;
}
ctx.ui.setWidget(`sub-${id}`, undefined);
}
const total = agents.size;
agents.clear();
nextId = 1;
const msg = total === 0
? "No subagents to clear."
: `Cleared ${total} subagent${total !== 1 ? "s" : ""}${killed > 0 ? ` (${killed} killed)` : ""}.`;
ctx.ui.notify(msg, total === 0 ? "info" : "success");
},
});
// ── Session lifecycle ─────────────────────────────────────────────────────
pi.on("session_start", async (_event, ctx) => {
applyExtensionDefaults(import.meta.url, ctx);
for (const [id, state] of Array.from(agents.entries())) {
if (state.proc && state.status === "running") {
state.proc.kill("SIGTERM");
}
ctx.ui.setWidget(`sub-${id}`, undefined);
}
agents.clear();
nextId = 1;
widgetCtx = ctx;
});
}

View file

@ -0,0 +1,145 @@
/**
* themeMap.ts — Per-extension default theme assignments
*
* Themes live in .pi/themes/ and are mapped by extension filename (no extension).
* Each extension calls applyExtensionTheme(import.meta.url, ctx) in its session_start
* hook to automatically load its designated theme on boot.
*
* Available themes (.pi/themes/):
* catppuccin-mocha · cyberpunk · dracula · everforest · gruvbox
* midnight-ocean · nord · ocean-breeze · rose-pine
* synthwave · tokyo-night
*/
import type { ExtensionContext } from "@mariozechner/pi-coding-agent";
import { basename } from "path";
import { fileURLToPath } from "url";
// ── Theme assignments ──────────────────────────────────────────────────────
//
// Key = extension filename without extension (matches extensions/<key>.ts)
// Value = theme name from .pi/themes/<value>.json
//
export const THEME_MAP: Record<string, string> = {
"agent-chain": "midnight-ocean", // deep sequential pipeline
"agent-team": "dracula", // rich orchestration palette
"coms": "ocean-breeze", // peer-to-peer messaging, cross-boundary
"coms-net": "ocean-breeze", // peer-to-peer messaging, cross-boundary
"cross-agent": "ocean-breeze", // cross-boundary, connecting
"damage-control": "gruvbox", // grounded, earthy safety
"minimal": "synthwave", // synthwave by default now!
"pi-pi": "rose-pine", // warm creative meta-agent
"pure-focus": "everforest", // calm, distraction-free
"purpose-gate": "tokyo-night", // intentional, sharp focus
"session-replay": "catppuccin-mocha", // soft, reflective history
"subagent-widget": "cyberpunk", // multi-agent futuristic
"system-select": "catppuccin-mocha", // soft selection UI
"theme-cycler": "synthwave", // neon, it's a theme tool
"tilldone": "everforest", // task-focused calm
"tool-counter": "synthwave", // techy metrics
"tool-counter-widget":"synthwave", // same family
};
// ── Helpers ───────────────────────────────────────────────────────────────
/** Derive the extension name (e.g. "minimal") from its import.meta.url. */
function extensionName(fileUrl: string): string {
const filePath = fileUrl.startsWith("file://") ? fileURLToPath(fileUrl) : fileUrl;
return basename(filePath).replace(/\.[^.]+$/, "");
}
// ── Theme ──────────────────────────────────────────────────────────────────
/**
* Apply the mapped theme for an extension on session boot.
*
* @param fileUrl Pass `import.meta.url` from the calling extension file.
* @param ctx The ExtensionContext from the session_start handler.
* @returns true if the theme was applied successfully, false otherwise.
*/
export function applyExtensionTheme(fileUrl: string, ctx: ExtensionContext): boolean {
if (!ctx.hasUI) return false;
const name = extensionName(fileUrl);
// If there are multiple extensions stacked in 'ipi', they each fire session_start
// and try to apply their own mapped theme. The LAST one to fire wins.
// Since system-select is last in the ipi alias array, it was setting 'catppuccin-mocha'.
// We want to skip theme application for all secondary extensions if they are stacked,
// so the primary extension (first in the array) dictates the theme.
const primaryExt = primaryExtensionName();
if (primaryExt && primaryExt !== name) {
return true; // Pretend we succeeded, but don't overwrite the primary theme
}
let themeName = THEME_MAP[name];
if (!themeName) {
themeName = "synthwave";
}
const result = ctx.ui.setTheme(themeName);
if (!result.success && themeName !== "synthwave") {
return ctx.ui.setTheme("synthwave").success;
}
return result.success;
}
// ── Title ──────────────────────────────────────────────────────────────────
/**
* Read process.argv to find the first -e / --extension flag value.
*
* When Pi is launched as:
* pi -e extensions/subagent-widget.ts -e extensions/pure-focus.ts
*
* process.argv contains those paths verbatim. Every stacked extension calls
* this and gets the same answer ("subagent-widget"), so all setTitle calls
* are idempotent — no shared state or deduplication needed.
*
* Returns null if no -e flag is present (e.g. plain `pi` with no extensions).
*/
function primaryExtensionName(): string | null {
const argv = process.argv;
for (let i = 0; i < argv.length - 1; i++) {
if (argv[i] === "-e" || argv[i] === "--extension") {
return basename(argv[i + 1]).replace(/\.[^.]+$/, "");
}
}
return null;
}
/**
* Set the terminal title to "π - <first-extension-name>" on session boot.
* Reads the title from process.argv so all stacked extensions agree on the
* same value — no coordination or shared state required.
*
* Deferred 150 ms to fire after Pi's own startup title-set.
*/
function applyExtensionTitle(ctx: ExtensionContext): void {
if (!ctx.hasUI) return;
const name = primaryExtensionName();
if (!name) return;
setTimeout(() => ctx.ui.setTitle(`π - ${name}`), 150);
}
// ── Combined default ───────────────────────────────────────────────────────
/**
* Apply both the mapped theme AND the terminal title for an extension.
* Drop-in replacement for applyExtensionTheme — call this in every session_start.
*
* Usage:
* import { applyExtensionDefaults } from "./themeMap.ts";
*
* pi.on("session_start", async (_event, ctx) => {
* applyExtensionDefaults(import.meta.url, ctx);
* // ... rest of handler
* });
*/
export function applyExtensionDefaults(fileUrl: string, ctx: ExtensionContext): void {
applyExtensionTheme(fileUrl, ctx);
applyExtensionTitle(ctx);
}

97
sssf/templates/justfile Normal file
View file

@ -0,0 +1,97 @@
# SSSF starter recipes. Stamped by install.py, then yours to edit.
#
# Deliberately small. These are the handful you need on day one: run something,
# watch it, and open the trace. Add your own as your chains grow, and see the
# example branch for the fuller set (orchestrator agents, kill, rosters, ipi).
# `.env` reaches every ADW through this, so keys work without exporting them.
set dotenv-load
set positional-arguments
# Every recipe passes this through, so `SSSF_CONFIG=other.yaml just sdlc "..."`
# swaps the whole roster for one run.
config := env_var_or_default("SSSF_CONFIG", "adws/adw_sssf_config/sssf.config.yaml")
db := "adws/adw_data/sssf.db"
# Where the sssf skill lives — the visualizer app ships with it. install.py
# stamps the real path here at install time; edit it if you move the skill.
skill_dir := "@SSSF_SKILL_DIR@"
# list every recipe
default:
@just --list
# ── first run ───────────────────────────────────────────────────────────────
# Proves the whole path works: config validated, session minted, agent ran,
# envelope parsed, gates checked, trace written. Costs a few cents and changes
# nothing in your repo, because both workflows are read-only.
#
# (`just --list` shows only the LAST comment line, so that one is the summary.)
# start here: two cheap read-only runs, end to end
demo:
@echo "1/2 adw_prompt: one agent, one prompt"
uv run adws/adw_prompt.py --config {{config}} --agent scout "reply with a one-line summary of this repo"
@echo "\n2/2 adw_scout: read-only recon"
uv run adws/adw_scout.py --config {{config}} "list the top-level directories in this repo and what each is for. change nothing."
@echo "\nboth done. now run: just sessions (or: just obs)"
# check the roster without running anything: names, prompts, models all resolve
validate:
uv run adws/adw_validate.py --config {{config}}
# ── run a workflow ──────────────────────────────────────────────────────────
# Args pass straight through: "<prompt or path/to/prompt.md>" [--adw-id X]
# one agent, one prompt: just prompt "summarize this repo"
prompt *ARGS:
uv run adws/adw_prompt.py --config {{config}} "$@"
# read-only recon: just scout "where is auth handled"
scout *ARGS:
uv run adws/adw_scout.py --config {{config}} "$@"
# plan only: just plan "add a /health endpoint"
plan *ARGS:
uv run adws/adw_plan.py --config {{config}} "$@"
# planner, builder, commit: just plan-build "add a /health endpoint"
plan-build *ARGS:
uv run adws/adw_plan_build.py --config {{config}} "$@"
# plan, build, test, commit: just sdlc "add a /health endpoint"
sdlc *ARGS:
uv run adws/adw_plan_build_test.py --config {{config}} "$@"
# the full chain, plus review and docs: just simple-sdlc "add a /health endpoint"
simple-sdlc *ARGS:
uv run adws/adw_simple_sdlc.py --config {{config}} "$@"
# ── watch it ────────────────────────────────────────────────────────────────
# Reads never block a running workflow, the db is WAL. Poll as hard as you like.
# the last 10 runs
sessions:
@sqlite3 {{db}} "select adw_id, status, substr(request,1,50), total_tokens, round(total_cost,4) from sessions order by started_at desc limit 10;"
# phase status in sequence: just phases <adw_id>
phases ADW_ID:
@sqlite3 {{db}} "select seq, name, kind, owner, status, attempt from phases where adw_id='{{ADW_ID}}' order by seq;"
# the live event tail: just tail <adw_id>
tail ADW_ID:
@sqlite3 {{db}} "select rowid, type, name, started_at from events where adw_id='{{ADW_ID}}' order by rowid desc limit 25;"
# what a run has alive right now, with pids: just procs <adw_id>
procs ADW_ID:
@sqlite3 {{db}} "select kind, name, pid, command, started_at from processes where adw_id='{{ADW_ID}}' and ended_at is null order by id;"
# ── observability UI ────────────────────────────────────────────────────────
# Needs bun. The db path is passed explicitly because the server runs from the
# app dir and would otherwise look for a trace db sitting next to itself.
# boot the trace UI, http://localhost:4601 (api on :4600)
obs:
cd {{skill_dir}}/apps/visualizer && bun install && (SSSF_DB={{justfile_directory()}}/{{db}} bun run server/index.ts &) && bunx vite

View file

@ -0,0 +1,13 @@
# Builder Agent
## Purpose
Implement the plan (or request) exactly; report every file you changed.
## Instructions
- If `previous_envelope` references a plan or test failures, follow them — they are your spec.
- Make the smallest change that satisfies the request; do not refactor unrelated code.
- When fixing test failures, address every reported failure.
- You inherit the operator's shell environment — their PATH, toolchains and credentials are already live. Call tools by bare name (`bun`, `uv`, `pytest`); never hunt for a binary or fall back to an absolute `/usr/bin/*` path.
- Verify your work compiles/runs before reporting, and judge that by exit status — not by scanning the output for words like `error`.

View file

@ -0,0 +1,34 @@
# Build Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Implement the work described in `prompt`, guided by `previous_envelope` if present, then emit your `Report` JSON.
## Report
Respond with ONLY valid JSON matching `BuildOutput` — no prose before or after:
```json
{
"status": "success",
"summary": "<one sentence describing what you built>",
"changed_files": ["src/server.ts"],
"artifacts": [],
"commit_message": "<imperative one-line git subject for the code you changed — this is what the commit of your work will say>",
"notes_for_next_agent": "<how to verify this work>"
}
```

View file

@ -0,0 +1,17 @@
# Documenter Agent
## Purpose
Write up the change that was just made, from the diff, for the engineer who arrives next.
## Instructions
- `previous_envelope` carries the captured change: `base` (what it was measured against), `changed_files`, `stat`, and `diff_path`. **Read `diff_path`** — the full diff is the source of truth.
- Everything you write must be traceable to that diff. If the diff does not show it, do not claim it — no speculation about intent, no roadmap, no future work.
- **Name a file only if it is in `changed_files` or appears in the diff.** Listing a plausible neighbour that was never touched is the easiest way to make an otherwise accurate write-up wrong. Check the list before you write the sentence.
- Document what the change does, where it lives, and how to use or verify it. It is a write-up for a human, not a commit log and not a replay of the diff.
- Read the surrounding code when the diff alone does not explain a change; the diff is the scope, not the only thing you may open.
- Write documentation only. Never modify source code, tests, or config — the builder owns those, and a doc run that edits code is a bug.
- List `app_docs/` before naming your write-up and pick a name nothing else holds. Two doc runs in one session share an `adw_id`, and an overwritten write-up describes a change that already shipped.
- Keep it tight. A reader should understand the change in under two minutes.
- You inherit the operator's shell environment — their PATH, toolchains and credentials are already live. Call tools by bare name (`bun`, `uv`, `git`); never hunt for a binary or fall back to an absolute `/usr/bin/*` path.

View file

@ -0,0 +1,48 @@
# Document Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Document the completed work described by `previous_envelope`, using `prompt` for what was originally asked.
1. Read the full diff at `previous_envelope.diff_path`, plus any changed file that needs context.
2. Write the write-up to `<context_handoff_dir>/document.md`. Cover: what changed and why it matters, the files that carry it, and how to use or verify it.
3. Copy that file into the repo under `app_docs/`:
- **List `app_docs/` before you pick the name.** A session that documents more than once reuses its `<adw_id>`, so the obvious name may already be taken.
- Base name: `app_docs/<adw_id>_<slug>.md`, where `<adw_id>` is the session directory name inside `context_handoff_dir` (`.../sessions/<adw_id>/context_handoff`) and `<slug>` is two to four kebab-case words naming the work.
- If a file with that name already exists, use `app_docs/<adw_id>_<slug>_v2.md`, then `_v3`, and so on until the name is free. **Never overwrite an existing write-up** — it describes a change that already shipped.
- **Copy it, do not retype it.** One bash call does the whole step:
`mkdir -p app_docs && cp "<context_handoff_dir>/document.md" "app_docs/<adw_id>_<slug>.md"`
Writing the document a second time through `write` re-emits every line you already wrote, which costs the whole write-up again in output tokens and lets the two copies drift.
4. Emit your `Report` JSON, declaring BOTH paths in `artifacts`.
## Report
Respond with ONLY valid JSON matching `DocumentOutput` — no prose before or after:
```json
{
"status": "success",
"summary": "<one sentence describing what you documented>",
"document_path": "app_docs/<adw_id>_<slug>.md",
"documented_files": ["src/server.ts"],
"artifacts": ["<context_handoff_dir>/document.md", "app_docs/<adw_id>_<slug>.md"],
"commit_message": "<imperative one-line git subject for committing THIS WRITE-UP, not the change it describes — e.g. 'Document the /health endpoint'>",
"notes_for_next_agent": "<anything the diff left unexplained>"
}
```
`document_path` and the `app_docs/` entry in `artifacts` are the path you ACTUALLY wrote, `_v2` suffix and all. Gates open these files — a name you meant to use fails them.

View file

@ -0,0 +1,21 @@
# Planner Agent
## Purpose
Turn a request into a plan the builder can implement without asking questions.
## Instructions
- Read only what you need to understand the request.
- Write the full plan to `<context_handoff_dir>/plan.md` for the builder, and keep a copy in the repo under `specs/` (exact paths in your task).
- List `specs/` before naming that copy and pick a name nothing else holds. Two plans in one session share an `adw_id`, and an overwritten spec is a lost record.
- Keep the plan concrete: files to touch, changes to make, how to verify.
- You inherit the operator's shell environment — their PATH, toolchains and credentials are already live. Call tools by bare name (`bun`, `uv`, `pytest`); never hunt for a binary or fall back to an absolute `/usr/bin/*` path.
- Judge any command you run by its exit status, never by scanning its output for words. `error` or `not found` inside passing output is text, not a failure.
- Do not implement anything.
## Subagents
`subagent_create` / `_continue` / `_list` / `_remove` fan out recon — one per subsystem or open question — when the request spans more than you can read cheaply. Give each a self-contained task; omit `model`.
They run in the background. **Wait for every one you spawned to report before writing `plan.md` or your Report JSON.** Skip them when a few reads would do.

View file

@ -0,0 +1,45 @@
# Plan Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Plan the work described in `prompt`.
1. Write the full plan to `<context_handoff_dir>/plan.md` — this is the copy the builder reads.
2. Copy that file into the repo under `specs/`:
- **List `specs/` before you pick the name.** A session that plans more than once reuses its `<adw_id>`, so the obvious name may already be taken.
- Base name: `specs/<adw_id>_<slug>.md`, where `<adw_id>` is the session directory name inside `context_handoff_dir` (`.../sessions/<adw_id>/context_handoff`) and `<slug>` is two to four kebab-case words naming the work.
- If a file with that name already exists, use `specs/<adw_id>_<slug>_v2.md`, then `_v3`, and so on until the name is free. **Never overwrite an existing spec** — the earlier plan is the record of what was asked for then.
- **Copy it, do not retype it.** One bash call does the whole step:
`mkdir -p specs && cp "<context_handoff_dir>/plan.md" "specs/<adw_id>_<slug>.md"`
Writing the plan a second time through `write` re-emits every line you already wrote, which costs the whole document again in output tokens and lets the two copies drift.
3. Emit your `Report` JSON, declaring BOTH paths in `artifacts`.
## Report
Respond with ONLY valid JSON matching `PlanOutput` — no prose before or after:
```json
{
"status": "success",
"summary": "<one sentence describing the plan>",
"artifacts": ["<context_handoff_dir>/plan.md", "specs/<adw_id>_<slug>.md"],
"commit_message": "<imperative one-line git subject for committing THIS PLAN DOCUMENT, not the work it describes — e.g. 'Add spec for the /health endpoint'>",
"notes_for_next_agent": "<what the builder must know>"
}
```
Both `artifacts` entries are the paths you ACTUALLY wrote, `_v2` suffix and all. Gates open these files — a name you meant to use fails them.

View file

@ -0,0 +1,16 @@
# Reviewer Agent
## Purpose
Confirm that what was built is what was asked for. This is not testing.
## Instructions
- Your spec is `<context_handoff_dir>/plan.md` when that file exists — the plan is the refined ask. Otherwise the spec is `prompt`, verbatim.
- Judge the code on disk, never the builder's summary of it. Start from `previous_envelope.changed_files`, read them, and use `git diff` for anything the envelope did not mention.
- Break the spec into concrete requirements and rule on each one: met, or not met with the evidence — a `file:line`, or exactly what is missing.
- Not your job: running tests, style opinions, refactors, or anything the request did not ask for. Work the request never asked for is not blocking on its own; work the request DID ask for and is missing always is.
- Change nothing. Findings go back to the builder — that is the only repair path.
- `approved` is true ONLY when every requirement is met and `blocking` is empty. Every blocking item names the specific gap, so the builder can fix it without guessing.
- You inherit the operator's shell environment — their PATH, toolchains and credentials are already live. Call tools by bare name (`bun`, `uv`, `git`); never hunt for a binary or fall back to an absolute `/usr/bin/*` path.
- Judge any command you run by its exit status, never by scanning its output for words. `error` or `not found` inside passing output is text, not a failure.

View file

@ -0,0 +1,44 @@
# Review Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Confirm that the work reported in `previous_envelope` is what was asked for.
1. Establish the spec: read `<context_handoff_dir>/plan.md` if it exists, else use `prompt`.
2. Read the code that was actually written, starting from `previous_envelope.changed_files`.
3. Rule on every requirement in the spec — one `findings` entry each, with evidence.
4. Write the review to `<context_handoff_dir>/review.md`, then emit your `Report` JSON.
## Report
Respond with ONLY valid JSON matching `ReviewOutput` — no prose before or after:
```json
{
"status": "success",
"approved": false,
"summary": "<one sentence: N of M requirements met>",
"findings": [
{ "requirement": "<the ask, in the requester's words>", "met": true, "evidence": "src/server.ts:42 — handler registered" }
],
"blocking": ["<what must change before this can be approved>"],
"artifacts": ["<context_handoff_dir>/review.md"],
"notes_for_next_agent": "<what the builder must fix, or how to verify if approved>"
}
```
`status` is `success` when the review itself completed — it is not the verdict. The verdict is `approved`, and it is true only when `findings` has no unmet entry and `blocking` is empty.

View file

@ -0,0 +1,20 @@
# Scout Agent
## Purpose
Find and report where things live. Change nothing.
## Instructions
- Read-only: search, read, and report — never write to the codebase.
- Cite exact file paths (with line hints where useful).
- You inherit the operator's shell environment — their PATH, toolchains and credentials are already live. Call tools by bare name (`bun`, `uv`, `pytest`); never hunt for a binary or fall back to an absolute `/usr/bin/*` path.
- Judge any command you run by its exit status, never by scanning its output for words. `error` or `not found` inside passing output is text, not a failure.
- Write your findings to `<context_handoff_dir>/scout_findings.md` for agents that follow.
- If you find nothing, say so plainly — an empty finding is a valid finding.
## Subagents
`subagent_create` / `_continue` / `_list` / `_remove` search several directions at once — one per lead or directory — instead of walking the codebase serially. Give each a self-contained task and hold it to read-only work; omit `model`.
They run in the background. **Wait for every one you spawned to report before writing `scout_findings.md` or your Report JSON.** Skip them when a couple of greps would do.

View file

@ -0,0 +1,34 @@
# Scout Task
## Variables
### prompt
{{prompt}}
### previous_envelope
{{previous_envelope}}
### context_handoff_dir
{{context_handoff_dir}}
## Task
Find what `prompt` asks about. Write findings into `context_handoff_dir`, then emit your `Report` JSON.
## Report
Respond with ONLY valid JSON matching `ScoutOutput` — no prose before or after:
```json
{
"status": "success",
"summary": "<one sentence on what you found>",
"findings": [
{ "file": "src/server.ts", "note": "<why this file matters>" }
],
"artifacts": ["<context_handoff_dir>/scout_findings.md"]
}
```

View file

@ -0,0 +1,143 @@
# sssf.config.yaml — the factory's agent roster. One agent, one prompt, one purpose.
# v1 runs the Pi or OMP coding agent; coding_agent: claude_code arrives in v2.
defaults:
coding_agent: pi
model: google/gemini-3.6-flash # provider/id — a bare pattern is ambiguous across providers
thinking: medium # off | minimal | low | medium | high | xhigh | max
harness_engineering: [] # pi extensions loaded into the harness (-e)
# Roster-wide allowlist; any agent may override with its own list.
# NOTE: --tools filters extension and custom tools too, not just builtins. An agent
# whose harness_engineering extension registers a tool MUST name that tool in its own
# tools list — otherwise the extension loads and its tool is silently filtered out.
tools:
- read # read file contents
- bash # execute bash commands
- edit # find/replace edits
- write # create/overwrite files
- grep # search file contents (pi default: OFF)
- find # find files by glob (pi default: OFF)
- ls # list directories (pi default: OFF)
# Off-limits to every agent that does not name them in its own `writes`.
# `tools` alone cannot protect these: bash runs `git checkout`, and write
# reaches any path. An agent must not be able to edit the machinery that
# decides whether its own work passed. Enforced in adw_modules/permissions.py.
#
# `writes:` per agent says what it may change IN THE REPO. It never restricts
# the session runtime under data_dir — context_handoff/, envelopes, prompts,
# raw output. Every agent can always write its own report; `writes: []` means
# read-only with respect to the repo, not mute.
protected_files:
- adws/adw_modules/
- adws/adw_sssf_config/
- adws/adw_*.py
data_dir: adws/adw_data # runtime home: {data_dir}/sessions/{adw_id}/{agent_name}/
observability:
db: adws/adw_data/sssf.db # tracer writes here directly; the UI polls it
poll_ms: 500 # visualizer live-poll cadence
agents:
- name: planner
model: fireworks/accounts/fireworks/models/kimi-k3
thinking: high
color: "#a78bfa" # optional hex — the agent's lane color in the visualizer
purpose: Turn a request into a plan the builder can implement without asking questions.
prompt_engineering:
system: adws/adw_data/prompt_engineering/planner/system.md
user: adws/adw_data/prompt_engineering/planner/user.md
harness_engineering:
- adws/adw_data/harness_engineering/subagents.ts # registers the four subagent_* tools below
writes: # the plan is the only thing it may leave in the repo
- specs/
tools: # full recon + write for plan.md; no edit — the planner never touches repo files
- read
- grep
- find
- ls
- bash
- write
- subagent_create # extension tools MUST be named here or they are filtered out
- subagent_continue
- subagent_list
- subagent_remove
- name: builder
color: "#22d3ee"
purpose: Implement the plan exactly; report every changed file in the envelope.
prompt_engineering:
system: adws/adw_data/prompt_engineering/builder/system.md
user: adws/adw_data/prompt_engineering/builder/user.md
# No `writes` key: unrestricted, and the only agent that is. It still cannot
# touch defaults.protected_files — the builder does not get to edit its own grader.
tools: # the only agent that mutates the repo — everything on
- read
- grep
- find
- ls
- bash
- edit
- write
- name: scout
color: "#fbbf24"
purpose: Find and report where things live; change nothing.
prompt_engineering:
system: adws/adw_data/prompt_engineering/scout/system.md
user: adws/adw_data/prompt_engineering/scout/user.md
harness_engineering:
- adws/adw_data/harness_engineering/subagents.ts # registers the four subagent_* tools below
writes: [] # read-only, and now actually read-only: its findings
# go to context_handoff/, which is runtime, not the repo
tools: # search-heavy recon; write only so scout_findings.md lands without a bash heredoc
- read
- grep
- find
- ls
- bash
- write
- subagent_create # extension tools MUST be named here or they are filtered out
- subagent_continue
- subagent_list
- subagent_remove
# No tester agent: running the suite is a known command, so it is a kind="code"
# phase over adw_modules/quality.py. See SKILL.md hard rule 8.
- name: reviewer
model: openai/gpt-5.6-terra
thinking: high
color: "#fb7185"
purpose: Confirm that what was built is what was asked for; change nothing.
prompt_engineering:
system: adws/adw_data/prompt_engineering/reviewer/system.md
user: adws/adw_data/prompt_engineering/reviewer/user.md
writes: [] # a reviewer that cannot fix cannot quietly fix — the
# claim the tool list only implied, now enforced
tools: # full read surface; write only for review.md, no edit
- read
- grep
- find
- ls
- bash
- write
- name: documenter
model: openai/gpt-5.6-luna
color: "#e879f9"
purpose: Write up the change that was just made, from the diff; document only.
prompt_engineering:
system: adws/adw_data/prompt_engineering/documenter/system.md
user: adws/adw_data/prompt_engineering/documenter/user.md
writes: # documentation only — "document only" is a rule now,
- app_docs/ # not a line in a prompt the model may drift from.
- docs/ # Markdown anywhere, because docs live next to the
- "**/*.md" # code they describe as often as in a docs folder.
- "*.md"
tools: # reads the diff and the code; writes/edits documentation only
- read
- grep
- find
- ls
- bash
- write
- edit