Add sssf skill, installable via the skills CLI
Port the sssf skill from ~/.agents/skills/sssf into this repo so it can be distributed and installed with the skills CLI (skills add INDigitalStudio/skills --skill sssf). - Copy the skill (SKILL.md, cookbooks, references, scripts, templates, and the visualizer app source) into sssf/. - Gitignore build/runtime artifacts: the visualizer's node_modules/ and dist/, Python bytecode, and the machine-specific repos.json. - Make the skill location-independent: install.py now stamps the skill's real path into the stamped justfile's skill_dir (replacing the hardcoded ~/.agents/skills/sssf), so 'just obs' finds the visualizer wherever the CLI installed the skill. - Update cookbooks to use <skill>/scripts/... instead of the hardcoded path, and document the skills CLI install command. - Update the repo README with install instructions.
This commit is contained in:
parent
a42608f602
commit
2cc766aabe
98 changed files with 12508 additions and 0 deletions
202
sssf/references/config.md
Normal file
202
sssf/references/config.md
Normal file
|
|
@ -0,0 +1,202 @@
|
|||
# Config Reference
|
||||
|
||||
The full `sssf.config.yaml` spec: every field, how defaults merge, and how model / thinking / tools / extensions map onto the coding agent.
|
||||
|
||||
It lives at **`adws/adw_sssf_config/sssf.config.yaml`** — the default path every `adw_*.py` and the justfile resolve, and where `install.py` / `make_config.py` stamp it. Pass `--config <path>` to any ADW (or set `SSSF_CONFIG` for the justfile) to run against a different roster.
|
||||
|
||||
## Shape
|
||||
|
||||
```yaml
|
||||
defaults:
|
||||
coding_agent: pi
|
||||
model: google/gemini-3.6-flash # ALWAYS provider/model-id
|
||||
thinking: medium
|
||||
harness_engineering: []
|
||||
tools: [read, bash, edit, write, grep, find, ls]
|
||||
data_dir: adws/adw_data
|
||||
|
||||
observability:
|
||||
db: adws/adw_data/sssf.db
|
||||
poll_ms: 500
|
||||
|
||||
agents:
|
||||
- name: planner
|
||||
coding_agent: pi
|
||||
model: google/gemini-3.6-flash # ALWAYS provider/model-id
|
||||
thinking: high
|
||||
color: "#a78bfa"
|
||||
purpose: Turn a request into a plan the builder can implement without asking questions.
|
||||
prompt_engineering:
|
||||
system: adws/adw_data/prompt_engineering/planner/system.md
|
||||
user: adws/adw_data/prompt_engineering/planner/user.md
|
||||
harness_engineering:
|
||||
- json-enforcer
|
||||
tools:
|
||||
- read
|
||||
- bash
|
||||
```
|
||||
|
||||
## Fields
|
||||
|
||||
### `defaults`
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `coding_agent` | `pi` \| `claude_code` | Which interface runs the agent. **v1 implements `pi` only**; `claude_code` is specced and stubbed in `agent_cc.py`, landing in v2. |
|
||||
| `model` | string | Model id. For Pi, any id registered in `~/.pi/agent/models.json`. Default `gemini-3.6-flash`. |
|
||||
| `thinking` | enum | Reasoning effort — see below. Default `medium`. |
|
||||
| `color` | hex string | Lane color for every agent that does not set its own. Default empty — the visualizer falls back to its own palette. |
|
||||
| `harness_engineering` | list[string] | Coding-agent extensions. Pi: extension names. Claude Code: reserved (MCP, hooks). |
|
||||
| `tools` | list[string] | Roster-wide tool allowlist. Every agent that omits its own `tools` inherits this. Unset = all tools usable. |
|
||||
| `protected_files` | list[string] | Paths **no** agent may modify unless it names them in its own `writes`. Default: `adws/adw_modules/`, `adws/adw_sssf_config/`, `adws/adw_*.py` — an agent must not be able to edit the machinery that decides whether its work passed. |
|
||||
| `data_dir` | path | Runtime home. Sessions land at `{data_dir}/sessions/{adw_id}/{agent_name}/`. Default `adws/adw_data`. |
|
||||
|
||||
### `observability`
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `db` | path | SQLite trace db. `tracer.py` writes it directly; the visualizer polls it. Default `adws/adw_data/sssf.db`. |
|
||||
| `poll_ms` | int | Visualizer live-poll cadence in ms. History uses the same queries, lazy-paged. Default `500`. |
|
||||
|
||||
### `agents[]`
|
||||
|
||||
| Field | Required | Meaning |
|
||||
|---|---|---|
|
||||
| `name` | yes | The identifier ADW scripts use. **ADWs name agents, never models.** |
|
||||
| `purpose` | yes | One sentence: what this agent is for. Should match its `system.md` Purpose. |
|
||||
| `prompt_engineering.system` | yes | Path to the system prompt — who the agent is, its single purpose, its output contract. |
|
||||
| `prompt_engineering.user` | yes | Path to the default user prompt — the task template with `{{prompt}}`, `{{previous_envelope}}`, `{{context_handoff_dir}}`. |
|
||||
| `color` | no | Hex swatch (`"#a78bfa"`) for this agent's lane in the visualizer. Travels config → `agent_sessions.color` → `/api/sessions/:adw_id`, and rides the `agent_start` event so a lane is colored while the agent is still running. Unset = the UI's fallback palette. |
|
||||
| `coding_agent`, `model`, `thinking`, `color`, `harness_engineering` | no | Override the corresponding `defaults` key. |
|
||||
| `tools` | no | Allowlist. **Omitting the key means all tools usable.** A capability list, not a boundary — see `writes`. |
|
||||
| `writes` | no | What this agent may modify **in the repo**, enforced after every call. Omitted = unrestricted (still barred from `protected_files`). `[]` = no repo writes at all. A list = only those paths: a trailing `/` is a directory prefix, `*` matches within one path segment, `**` crosses segments, anything else is an exact path. Naming a `protected_files` path here is what unlocks it. **The session runtime under `data_dir` is always writable** — `writes: []` means read-only with respect to the repo, not unable to write its own report. |
|
||||
|
||||
Output types are deliberately absent: config defines who an agent *is*; the ADW call site defines how it's *used*. One agent serves many calls — same system prompt, different user prompt + output type per call.
|
||||
|
||||
## Defaults merging
|
||||
|
||||
`agents.py` merges each entry **over** `defaults`, key by key. An entry states only what differs; anything unset inherits. `agents.validate(cfg, REQUIRED_AGENTS)` then confirms every name an ADW declares exists, resolves to a usable coding agent + model, and has both prompt files present on disk. Any miss fails the run immediately — **no agent is ever spawned against a half-valid config.**
|
||||
|
||||
## Thinking levels
|
||||
|
||||
Pi's reasoning-effort ladder, lowest to highest:
|
||||
|
||||
```
|
||||
off | minimal | low | medium | high | xhigh | max
|
||||
```
|
||||
|
||||
Mapped to Pi's reasoning effort control and honored when the model is registered with `reasoning: true` in `~/.pi/agent/models.json`. On a non-reasoning model the setting is inert — no error, no effect. Rough guidance: `high`/`xhigh` for planners and reviewers, `medium` for builders, `low` for mechanical read-and-report agents. (For Claude Code in v2, the same field maps to the thinking budget.)
|
||||
|
||||
## Model resolution
|
||||
|
||||
**Always write `model` as `provider/model-id`.** `agents.py` hands the string to the Pi interface, which resolves it against pi's merged catalog — `~/.pi/agent/models.json` plus pi's built-in providers. The same model is usually carried by more than one provider (`gemini-3.6-flash` lives under `google` *and* under `openrouter` as `google/gemini-3.6-flash`), and a bare id that matches several **raises at resolution**:
|
||||
|
||||
```
|
||||
agent 'scout': model pattern 'gemini-3.6-flash' is ambiguous:
|
||||
[('google', 'gemini-3.6-flash'), ('openrouter', 'google/gemini-3.6-flash'), ...]
|
||||
```
|
||||
|
||||
That is `agents.validate()` doing its job — it fails before anything spawns rather than silently billing the wrong provider — but it means every agent in the roster inheriting that default is grounded until the pattern is qualified. Qualifying is the whole fix: `google/gemini-3.6-flash`, `openai/gpt-5.6-terra`, `fireworks/accounts/fireworks/models/kimi-k3`. The leading segment is matched against the provider list first, so the rest of the string can contain slashes.
|
||||
|
||||
Other consequences worth knowing:
|
||||
|
||||
- A model must be in the catalog before any agent can name it. An unknown id fails at resolution, before spawn. `pi --list-models` is the catalog the resolver actually reads.
|
||||
- **Ambiguity can appear without you touching the config.** Registering a new provider that carries a model you already use turns a formerly-fine bare pattern ambiguous. If a roster stops validating and nobody edited it, that is why.
|
||||
- Provider credentials come from the environment, not the config — the key that matches the provider you named (`GEMINI_API_KEY` for `google/...`, `OPENROUTER_API_KEY` for `openrouter/...`).
|
||||
- The resolved model is recorded per session in `agent_map.json` and mirrored into the `agent_sessions` table. **Changing an agent's model invalidates its session**: a joined run starts that agent fresh instead of resuming a context window built by a different model.
|
||||
|
||||
## Tools
|
||||
|
||||
`tools` maps to `pi --tools`. Pi's seven builtin tool names:
|
||||
|
||||
| Tool | Purpose | Pi's own default |
|
||||
|---|---|---|
|
||||
| `read` | read file contents | on |
|
||||
| `bash` | execute bash commands | on |
|
||||
| `edit` | find/replace edits | on |
|
||||
| `write` | create/overwrite files | on |
|
||||
| `grep` | search file contents | **off** |
|
||||
| `find` | find files by glob | **off** |
|
||||
| `ls` | list directory contents | **off** |
|
||||
|
||||
`grep`, `find`, and `ls` are off in bare Pi, so an agent that does not name them will shell out through `bash` to do the same work. The starter roster therefore sets `defaults.tools` to all seven and lets each agent narrow from there.
|
||||
|
||||
**Resolution order:** an agent's own `tools` list wins; an agent that omits the key inherits `defaults.tools`; if neither is set, `tools` stays `None` and all tools are usable. An empty list is not "all tools" — it is a tool-less agent, and it will stall.
|
||||
|
||||
## Write permissions — `writes` and `protected_files`
|
||||
|
||||
`tools` cannot express a safety boundary, because two of the tools are general
|
||||
purpose. `bash` runs anything, including `git checkout`, which discards an
|
||||
engineer's uncommitted work; `write` reaches any path, not only the one report
|
||||
file an agent was granted it for. So "this agent changes nothing" is a claim a
|
||||
tool list can state but never keep.
|
||||
|
||||
`adw_modules/permissions.py` keeps it, the same way every other claim in this
|
||||
system is kept — after the fact, against the repo. Before an agent's first
|
||||
prompt the working tree's change-set is fingerprinted; after its last send
|
||||
(including JSON retries and gate corrections) it is fingerprinted again. Any
|
||||
path that appeared, vanished, or changed is attributed to that agent.
|
||||
|
||||
Comparing change-sets rather than watching writes is deliberate: a path that was
|
||||
modified before the agent ran and is clean afterwards has been **reverted**, and
|
||||
a reversion is a modification. That is what catches `git checkout`.
|
||||
|
||||
A breach is not a gate violation. Gates are for work an agent can be asked to
|
||||
redo; a write has already happened, so re-prompting fixes nothing. Instead:
|
||||
|
||||
1. every unauthorized change the agent **introduced** is rolled back — tracked
|
||||
files with `git checkout --`, untracked files by deletion;
|
||||
2. a path that was **already dirty** before the agent ran is left untouched. The
|
||||
operator had uncommitted work there, and discarding it to tidy up would be
|
||||
the same harm this module exists to prevent;
|
||||
3. the phase fails and names every path with what happened to it.
|
||||
|
||||
```yaml
|
||||
defaults:
|
||||
protected_files: [adws/adw_modules/, adws/adw_sssf_config/, "adws/adw_*.py"]
|
||||
|
||||
agents:
|
||||
- name: builder # no `writes` key -> unrestricted, minus protected_files
|
||||
- name: scout
|
||||
writes: [] # no repo writes; its findings still land in context_handoff/
|
||||
- name: planner
|
||||
writes: [specs/]
|
||||
- name: documenter
|
||||
writes: [app_docs/, docs/, "**/*.md", "*.md"]
|
||||
```
|
||||
|
||||
**The session runtime under `data_dir` is always writable, for every agent.**
|
||||
`context_handoff/` is how agents hand work to each other, and each agent's
|
||||
prompts, `raw_output.jsonl`, and `envelope.json` sit beside it. That grant comes
|
||||
from `data_dir` rather than from `.gitignore`: the runtime is normally ignored,
|
||||
so it never even appears in a snapshot, but an agent's ability to record its own
|
||||
work must not depend on a gitignore line someone can delete.
|
||||
|
||||
Narrow by role, not by reflex. Anything that must produce a `context_handoff/` artifact needs `write`, or it will resort to a `bash` heredoc. Withhold `edit`/`write` only where the restriction *is* the guarantee — a reviewer that cannot edit cannot quietly fix what it was asked to report.
|
||||
|
||||
### Extension tools must be named explicitly
|
||||
|
||||
`pi --tools` is an allowlist over **built-in, extension, and custom tools alike** — not just builtins. So the moment an agent has a `tools` list at all (its own, or one inherited from `defaults`), any tool registered by its `harness_engineering` extensions is **excluded unless it appears in that list by name**.
|
||||
|
||||
This fails quietly. The extension still loads, the run still succeeds, and the tool the extension exists to provide is simply never offered to the model — you find out by noticing the agent never called it.
|
||||
|
||||
```yaml
|
||||
- name: reviewer
|
||||
harness_engineering:
|
||||
- .pi/extensions/ast_query.ts # registers tool: ast_query
|
||||
tools:
|
||||
- read
|
||||
- grep
|
||||
- find
|
||||
- ls
|
||||
- bash
|
||||
- ast_query # REQUIRED — the extension's tool, named or lost
|
||||
```
|
||||
|
||||
Rule: **every entry in `harness_engineering` that registers a tool must have that tool name added to the agent's `tools` list.** Adding an extension is therefore a two-line change, never one. The alternative is dropping the `tools` key *and* leaving `defaults.tools` unset so the agent resolves to `None` (all tools) — but with a roster-wide `defaults.tools` in place, that escape hatch is closed; naming the tool is the only path.
|
||||
|
||||
## Harness engineering
|
||||
|
||||
`harness_engineering` entries are pi extension **file paths**, passed through as `pi -e <path>`, one flag per entry, scoped to that agent only. This is where per-agent harness changes live — e.g. an output-tightening extension for an agent that keeps wrapping its envelope in prose. The starter roster ships with none. On Claude Code the field is reserved for MCP config and hooks in v2.
|
||||
|
||||
**If the extension registers a tool, name that tool in the agent's `tools` list too** — `--tools` filters extension tools exactly like builtins, so an unnamed extension tool is silently unavailable no matter that the extension loaded fine. See [Extension tools must be named explicitly](#extension-tools-must-be-named-explicitly) above. Extensions that only shape output or add flags (no tool registration) need no `tools` change.
|
||||
161
sssf/references/handoff.md
Normal file
161
sssf/references/handoff.md
Normal file
|
|
@ -0,0 +1,161 @@
|
|||
# Handoff Reference
|
||||
|
||||
The envelope schema, the two-channel output contract, and the session directory layout — how context transfers in code, not in conversation.
|
||||
|
||||
## Two output channels, exactly
|
||||
|
||||
An agent may produce output in two ways and no others:
|
||||
|
||||
1. **Reference files** written into `context_handoff/` — plans, notes, artifacts for the agents that follow.
|
||||
2. **A final valid-JSON response** — the envelope, its direct response and nothing else.
|
||||
|
||||
Code does the rest: parse the response against the output type the call declared, persist it as `envelope.json`, and inject it into the next agent's user prompt.
|
||||
|
||||
## Envelope schema
|
||||
|
||||
Every output type extends `EnvelopeBase`:
|
||||
|
||||
```python
|
||||
class EnvelopeBase(BaseModel):
|
||||
status: Literal["success", "fail"] # the only required field
|
||||
summary: str = "" # one sentence: what happened
|
||||
artifacts: list[str] = [] # paths written, usually inside context_handoff/
|
||||
notes_for_next_agent: str = "" # what the next agent must know
|
||||
```
|
||||
|
||||
`status` is load-bearing: an envelope that parses but reports `status="fail"` raises, failing the phase. An agent declaring its own failure is not a successful phase.
|
||||
|
||||
The starter types in `adw_modules/data_types.py`:
|
||||
|
||||
```python
|
||||
class GenericOutput(EnvelopeBase):
|
||||
"""Fallback for an agent with no sharper contract yet."""
|
||||
|
||||
class PlanOutput(EnvelopeBase):
|
||||
commit_message: str = "" # imperative git subject for the PLAN FILE itself
|
||||
|
||||
class BuildOutput(EnvelopeBase):
|
||||
changed_files: list[str] = []
|
||||
commit_message: str = "" # consumed by the git commit phase
|
||||
|
||||
class ScoutOutput(EnvelopeBase):
|
||||
findings: list[ScoutFinding] = [] # ScoutFinding: {file: str, note: str}
|
||||
|
||||
class ReviewOutput(EnvelopeBase):
|
||||
approved: bool = False # the verdict; status is only "did the review run"
|
||||
findings: list[ReviewFinding] = [] # ReviewFinding: {requirement, met: bool, evidence}
|
||||
blocking: list[str] = [] # what must change before approval
|
||||
|
||||
class DocumentOutput(EnvelopeBase):
|
||||
document_path: str = "" # the write-up's home in the repo
|
||||
documented_files: list[str] = []
|
||||
commit_message: str = ""
|
||||
```
|
||||
|
||||
`commit_message` defaults to empty, so a git phase consuming it always needs a fallback — see `cookbooks/create_adw.md`.
|
||||
|
||||
**Each `commit_message` describes its own agent's work product, never the next one's**: `PlanOutput`'s covers the spec file, `BuildOutput`'s the code, `DocumentOutput`'s the write-up. A chain that commits once can use whichever fits; a chain that commits per step (`adw_simple_sdlc.py`) needs all three, and reusing one agent's sentence for another's diff is how a commit log starts lying.
|
||||
|
||||
There is no test output type: running the suite is a `kind="code"` phase, and its `QualityResult` reaches the next agent through `quality.as_envelope`.
|
||||
|
||||
Two of these are adapters rather than agent reports — code shaped as an envelope so an agent can be handed a deterministic result through the same door: `VerifyOutput` (a lint/test block's result) and `ChangesOutput` (a captured `git diff`, from `changes.as_envelope`). The consuming agent cannot tell the difference, which is the point.
|
||||
|
||||
The envelope is a **manifest of claims**. Gates verify those claims after the fact — declared artifacts exist and are non-empty, declared changes appear in the diff, declared tests actually pass. See `cookbooks/update_modules.md`.
|
||||
|
||||
## The typed-output rule
|
||||
|
||||
**Every agent call passes a concrete output type**, and the agent's final JSON is parsed against exactly that type. No untyped handoffs.
|
||||
|
||||
```python
|
||||
plan = ph.call(AgentCall(output_type=PlanOutput, prompt=prompt,
|
||||
gates=[gates.artifacts_exist]))
|
||||
```
|
||||
|
||||
The user prompt asks for the shape; the type enforces it. They always travel as a pair, which is what lets one agent serve many calls — same system prompt, different user prompt + output type per call site. Output types live in code, never in `sssf.config.yaml`.
|
||||
|
||||
**Parse failure is not a restart.** If the response doesn't parse or doesn't validate, the harness re-prompts the **same session** with a correction naming the required fields — bounded by `JSON_FIX_ATTEMPTS` in `agents.py` (2). Gate violations use the identical mechanism, bounded instead by the phase's `retries`. A cold restart would throw away the context that produced the near-miss.
|
||||
|
||||
In v1 there is no separate continue call to make: `agent_pi.run()` passes `--session-id`, which pi treats as create-or-continue, so running an agent and continuing it are the same call with the same id. Before parsing, the harness also tolerates a fenced `json` code block or prose wrapped around the object — but the prompt still asks for bare JSON, and every failed attempt is persisted as an invalid envelope row.
|
||||
|
||||
## Injecting the previous envelope
|
||||
|
||||
`prompts.py` renders the agent's `user.md`, substituting:
|
||||
|
||||
| Placeholder | Value |
|
||||
|---|---|
|
||||
| `{{prompt}}` | the engineer's ask (or the ADW's per-call prompt) |
|
||||
| `{{previous_envelope}}` | the upstream envelope JSON, from `AgentCall(previous=...)` |
|
||||
| `{{context_handoff_dir}}` | absolute path to this session's `context_handoff/` |
|
||||
|
||||
A `user.md` declares one h3 per incoming datum, then the task, then the output contract:
|
||||
|
||||
````markdown
|
||||
# Scout Task
|
||||
|
||||
## Variables
|
||||
|
||||
### prompt
|
||||
|
||||
{{prompt}}
|
||||
|
||||
### previous_envelope
|
||||
|
||||
{{previous_envelope}}
|
||||
|
||||
### context_handoff_dir
|
||||
|
||||
{{context_handoff_dir}}
|
||||
|
||||
## Task
|
||||
|
||||
Find what `prompt` asks about. Write findings into `context_handoff_dir`, then emit your `Report` JSON.
|
||||
|
||||
## Report
|
||||
|
||||
Respond with ONLY valid JSON matching `ScoutOutput` — no prose before or after:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "success",
|
||||
"summary": "<one sentence on what you found>",
|
||||
"findings": [
|
||||
{ "file": "src/server.ts", "note": "<why this file matters>" }
|
||||
],
|
||||
"artifacts": ["<context_handoff_dir>/scout_findings.md"]
|
||||
}
|
||||
```
|
||||
````
|
||||
|
||||
The `## Report` section shows the exact JSON shape of the declared output type — that is the agent's output contract, and it lives in `user.md` because the shape belongs to the *use*, not the identity. The matching `system.md` stays static: Purpose + Instructions only.
|
||||
|
||||
## Session directory layout
|
||||
|
||||
```
|
||||
adws/adw_data/sessions/{adw_id}/
|
||||
├── agent_map.json agent name → coding-agent session_id + model
|
||||
├── context_handoff/ the ONE place agents write files for the agents that follow
|
||||
└── {agent_name}/
|
||||
├── prompts/ exact prompts sent (system.md + user.md), saved before execution
|
||||
├── pi_sessions/ pi's own session state for this agent
|
||||
├── raw_output.jsonl full JSONL stream from the coding agent, appended live
|
||||
└── envelope.json the final valid-JSON response — captured, validated, persisted by code
|
||||
```
|
||||
|
||||
`session.ensure(cfg, adw_id)` mints or joins the id and creates these dirs. One `context_handoff/` per session, shared by every agent — the single location for cross-agent files.
|
||||
|
||||
## agent_map.json and resuming
|
||||
|
||||
```json
|
||||
{
|
||||
"planner": {"session_id": "sssf-a1b2c3d4-planner-9f2e",
|
||||
"model": "google/gemini-3.6-flash", "coding_agent": "pi"},
|
||||
"builder": {"session_id": "sssf-a1b2c3d4-builder-71ac",
|
||||
"model": "google/gemini-3.6-flash", "coding_agent": "pi"}
|
||||
}
|
||||
```
|
||||
|
||||
This map is the key that lets a later ADW rejoin each agent's **existing context window**. Run `adw_build.py --adw-id a1b2c3d4` after `adw_plan.py` and the builder resumes its own session rather than starting cold.
|
||||
|
||||
The map records the model each session was created with. If config drift changes an agent's model, that agent starts a **fresh** session and the map is updated — never a bad resume. `agent_sessions` in `sssf.db` is the queryable mirror of this file.
|
||||
|
||||
**Files are the raw record; the db is the queryable mirror.** Losing `sssf.db` loses nothing that can't be rebuilt from `raw_output.jsonl`, `envelope.json`, and `agent_map.json`.
|
||||
158
sssf/references/observability.md
Normal file
158
sssf/references/observability.md
Normal file
|
|
@ -0,0 +1,158 @@
|
|||
# Observability Reference
|
||||
|
||||
The event schema, the seven SQLite tables, and the polling contract — the one data path is **agents → sqlite → web ui**.
|
||||
|
||||
## Two stores, one truth
|
||||
|
||||
**Files are the raw record** (`raw_output.jsonl` streams, `envelope.json`, `agent_map.json`); **SQLite (`sssf.db`) is the queryable mirror** the UI reads. `tracer.py` writes both. Losing the db loses nothing that can't be rebuilt from files.
|
||||
|
||||
Location comes from `observability.db` in `sssf.config.yaml`, default `adws/adw_data/sssf.db` — inside the **target** repo, gitignored.
|
||||
|
||||
## Event schema
|
||||
|
||||
`tracer.py` emits these types, every one logged against its `adw_id` **and** `phase_id`:
|
||||
|
||||
| Type | Emitted when |
|
||||
|---|---|
|
||||
| `phase_start` | a `run.phase(...)` block is entered |
|
||||
| `agent_start` | a coding agent is spawned or resumed for `ph.call(...)` |
|
||||
| `tool_call` | a tool (`read`, `bash`, `edit`, `write`) returns — **one event per real call**, named `bash: ls -la src`, payload `{tool, tool_call_id, args, result_snippet, ok, duration_ms, agent}` |
|
||||
| `handoff` | an envelope crosses from one agent to the next |
|
||||
| `gate_pass` | a gate found no failed checks — payload carries `attempt`, `checks` (the evidence), and an empty `violations` |
|
||||
| `gate_fail` | a gate found at least one failed check — payload carries `attempt`, `checks`, and `violations` |
|
||||
| `log` | an explicit `ph.log(...)` from the ADW script |
|
||||
| `agent_end` | the agent's run completes; envelope parsed or not — payload carries `cost`, `usage` (the per-component breakdown), `context_tokens`, `context_window` |
|
||||
| `phase_end` | the block exits; carries the resolved status |
|
||||
| `error` | a raise inside a phase block |
|
||||
|
||||
`parent_id` nests spans, so an agent phase expands into its tool-call spans in the UI.
|
||||
|
||||
**Spend is itemised per phase.** `agent_end.usage` carries tokens *and* dollars for each component pi reports — `input`, `output`, `cache_read`, `cache_write` — summed across every send the phase made, so a phase that retried on a bad envelope or a failed gate shows what all its attempts cost, not just the last one. The four components sum to `total_tokens`, and their costs sum to `total_cost`; the visualizer's Cost panel renders them as a table you can add up by eye.
|
||||
|
||||
`reasoning_tokens` is the thinking share and is **inside** `output_tokens`, not a fifth component — measured across every session on disk, reasoning never exceeds output and the four components always reconcile to the total. It bills at the output rate, so the panel nests it under output rather than adding it. Runs predating the breakdown have no `usage` key at all; the lump `cost` and the event's own `tokens` still stand, and the UI says so rather than rendering zeroes.
|
||||
|
||||
**Context is occupancy, not spend.** `events.tokens` and `sessions.total_tokens` bill every turn, so they only grow — an agent that burned 100k tokens may be sitting in a 15k window. `context_tokens` is how full the window actually was when the agent stopped, which is what the visualizer's per-lane Context bar measures against `context_window`.
|
||||
|
||||
It is computed the way pi computes it for its own footer and its auto-compaction trigger (`calculateContextTokens` in the coding agent's `core/compaction/compaction.ts`): take the last *valid* assistant turn — skipping `aborted` and `error` turns — and read `usage.totalTokens`, falling back to `input + output + cacheRead + cacheWrite`. Cache reads count; cached prompt is still prompt. `context_window` is the same `contextWindow` pi reads from `~/.pi/agent/models.json`, so `context_tokens / context_window` is the number pi would show. Both are NULL on rows written before the columns existed, and the lane draws no bar rather than a misleading empty one.
|
||||
|
||||
Two caveats worth knowing. Pi adds an *estimate* for any messages trailing the last assistant usage; in a batch (`-p`) run the session ends on that message, so the two agree. And if auto-compaction fires as the very last act of a run, the recorded number is the pre-compaction size — pi itself reports `null` in that window rather than guessing.
|
||||
|
||||
**Gates record evidence, not just a verdict.** A gate returns one `{item, ok, note}` check per thing it looked at, and `violations` are derived from the failed ones. Both land in `gate_results` (`checks_json` + `violations_json`) and in the `gate_pass`/`gate_fail` payload, so a green gate can answer *what did you verify* — `{"item": "…/plan.md", "ok": true, "note": "exists, 454B"}` — rather than only *did it pass*. Rows written before this existed have `checks_json` NULL; treat that as "no evidence recorded", not "nothing checked".
|
||||
|
||||
The gate event payload carries `attempt` too, so the `gate_results` table and the event stream are equivalent sources — a live consumer can group gate results per correction round from events alone, without a second query.
|
||||
|
||||
**A `tool_call` is the one event that spans time**, so it fills both `started_at` and `ended_at` on the row — the tool's real start and return. Every other type is a point in time: `started_at` is when it was recorded and `ended_at` stays NULL. Lay tool calls out on a time axis from those columns, never by parsing `payload_json` (`duration_ms` is in the payload too, as pi's own number, but it is a convenience, not the source for layout).
|
||||
|
||||
**Streaming is solved by construction.** `agent_pi.py` tails pi's JSONL stdout line by line and the tracer inserts each event into `sssf.db` **while the agent is still working** — never batched at phase end (verified in the first smoke run: tool calls visible mid-run). Everything downstream is a poll → render.
|
||||
|
||||
## Tables
|
||||
|
||||
```sql
|
||||
sessions (
|
||||
adw_id TEXT PRIMARY KEY,
|
||||
request TEXT, -- the engineer's ask
|
||||
status TEXT, -- running | success | fail
|
||||
engineer TEXT,
|
||||
started_at TEXT, ended_at TEXT,
|
||||
total_tokens INTEGER, total_cost REAL
|
||||
);
|
||||
|
||||
phases (
|
||||
phase_id TEXT PRIMARY KEY,
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
seq INTEGER,
|
||||
name TEXT, kind TEXT, owner TEXT, description TEXT,
|
||||
status TEXT DEFAULT 'fail', -- success must be earned
|
||||
attempt INTEGER DEFAULT 0, retries INTEGER DEFAULT 0,
|
||||
error TEXT,
|
||||
started_at TEXT, ended_at TEXT
|
||||
);
|
||||
|
||||
events (
|
||||
event_id TEXT PRIMARY KEY,
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
phase_id TEXT REFERENCES phases, -- every event logs against adw + phase
|
||||
parent_id TEXT, -- span nesting
|
||||
type TEXT, -- phase_start | phase_end | agent_start | agent_end | tool_call
|
||||
-- | handoff | gate_pass | gate_fail | log | error
|
||||
name TEXT,
|
||||
payload_json TEXT,
|
||||
tokens INTEGER,
|
||||
started_at TEXT, ended_at TEXT -- ended_at set only on events that span time
|
||||
);
|
||||
|
||||
envelopes (
|
||||
envelope_id TEXT PRIMARY KEY,
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
phase_id TEXT REFERENCES phases,
|
||||
agent TEXT,
|
||||
output_type TEXT, -- name of the data_types model it parsed against
|
||||
payload_json TEXT,
|
||||
valid INTEGER,
|
||||
attempt INTEGER,
|
||||
created_at TEXT
|
||||
);
|
||||
|
||||
gate_results (
|
||||
id INTEGER PRIMARY KEY,
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
phase_id TEXT REFERENCES phases,
|
||||
attempt INTEGER,
|
||||
gate TEXT,
|
||||
passed INTEGER,
|
||||
violations_json TEXT, -- derived: the failed checks, as "item: note"
|
||||
checks_json TEXT, -- [{item, ok, note}] — everything the gate looked at
|
||||
created_at TEXT
|
||||
);
|
||||
|
||||
processes ( -- adw_id → pid, so a stuck run can be stopped
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
kind TEXT, -- 'adw' (the workflow process) | 'agent' (a coding-agent child)
|
||||
name TEXT, -- '' for the adw, the agent name for a child
|
||||
pid INTEGER,
|
||||
command TEXT, -- what the pid WAS; pids get recycled, so verify before killing
|
||||
started_at TEXT, ended_at TEXT -- ended_at NULL = believed alive
|
||||
);
|
||||
|
||||
agent_sessions ( -- the queryable mirror of agent_map.json
|
||||
adw_id TEXT REFERENCES sessions,
|
||||
agent TEXT,
|
||||
coding_agent TEXT, model TEXT, color TEXT, -- color: the config's lane swatch
|
||||
session_id TEXT,
|
||||
context_tokens INTEGER, -- window occupancy after the agent's last turn
|
||||
context_window INTEGER, -- the model's ceiling, from the pi registry
|
||||
created_at TEXT, last_used_at TEXT,
|
||||
PRIMARY KEY (adw_id, agent)
|
||||
);
|
||||
```
|
||||
|
||||
**A hung agent emits nothing**, which is exactly when you need its pid: no events, no tokens, no output to read. `processes` is the only table that can answer "what is this run running, and how do I stop it" — `just procs <adw_id>` lists what is live, `just kill <adw_id>` stops children before the parent, and both verify the recorded `command` still matches the pid before signalling it. A killed run finalizes its own trace: SIGTERM and SIGINT are turned into `SystemExit` in `session.ensure`, so the session lands on `fail` with its process rows closed instead of reading `running` forever.
|
||||
|
||||
**Derived, never stored:** phase durations (`ended_at − started_at`), session phase-progress (query `phases` by `adw_id`), lane layout (`kind` + `owner`).
|
||||
|
||||
Phase status invariants: `queued` only for manifest-declared phases not yet entered (dashed in the UI); `running` on enter; only a clean exit writes `success` — agent phases additionally need the envelope parsed and gates green; everything else resolves to `fail`.
|
||||
|
||||
## WAL pragmas
|
||||
|
||||
Open **every** connection — writer and reader — with:
|
||||
|
||||
```sql
|
||||
PRAGMA journal_mode=WAL;
|
||||
PRAGMA synchronous=NORMAL;
|
||||
PRAGMA busy_timeout=5000;
|
||||
```
|
||||
|
||||
WAL allows readers during writes. Writers are the tracers of running ADW processes; concurrent writers are fine given one small transaction per event plus `busy_timeout`. The visualizer reads on a readonly connection with exactly one exception: archiving a session (`POST /api/sessions/:adw_id/archive`) opens a second connection to set `sessions.archived`. That flag is review triage — it says a human has looked at the run — so it is the reader's state living on the row, and no tracer ever writes or reads it.
|
||||
|
||||
## Polling contract
|
||||
|
||||
**The UI never receives pushes.** No ingest endpoint, no WebSocket, no backfill or dedup logic.
|
||||
|
||||
Live view polls on a rowid cursor every `observability.poll_ms` (default 500):
|
||||
|
||||
```sql
|
||||
SELECT ... FROM events WHERE adw_id = ? AND rowid > ? ORDER BY rowid LIMIT 500;
|
||||
```
|
||||
|
||||
Keep the highest `rowid` returned as the next cursor. History is **the same queries** with filters, lazy-paged as the engineer scrolls or drills in — one mechanism serves both live and past runs, which is why there is no separate replay path.
|
||||
Loading…
Add table
Add a link
Reference in a new issue