Workflow

How a finding about a skill travels from the session where an agent hit it to the release that fixes it: the agent keeps a log, the user reviews and submits it, and the maintainer’s agent reproduces the finding, patches the skill and closes the report.

The overview described the loop in one picture. This page walks one finding around it, using the OpenCms skill as the example: an agent editing the university website, a sentence in the skill that turns out to be ambiguous, and the release that corrects it. Two agents appear: the user’s agent, which follows the skill and reports, and the maintainer’s agent, which works the reports under the maintainer’s supervision. Every command shown is run by one of those two agents; the user types none of them.

The reporter’s side

The first half of the loop runs where the skill is used: in the user’s agent’s session, with the user watching. These steps follow one finding from the moment the agent notices it to the moment a report leaves the machine.

  1. The agent notes friction while it works

    The task is an ordinary one: add a new member to an institute team page. The agent follows the skill and, at one point, a helper script rejects a path. The skill said which path to use, or seemed to. The agent works out the right form, finishes the task, and then writes down what happened while it still knows which line it was reading.

    The skill tells the agent to keep such notes, and how, in a reference file the agent reads at the start of every session (context is the agent’s working memory for one session):

    Two rules apply to the log, and both put the record above the agent’s convenience. First, the agent writes down what actually happened, even where that contradicts the skill. If the skill says to use one path and the agent found that another one works, the log records the working path and the failing one; an agent that quietly reshapes its work to fit the skill’s wording leaves no trace of the error, and the error survives. Second, the agent writes no secret into the log: no session identifier, no credential, no content of a private page, since the log is meant to leave the machine later. The log looks like this while the agent works:

    A text editor showing a friction log in Markdown with one or two fictional numbered findings: a one-line title stated as the corrected fact, then kind, confidence, observed, expected, task context, reproduction and evidence fields. (enlarge)
    The friction log is working notes, kept as the agent goes. Reconstructing it afterwards produces a tidy story and loses the detail that makes a finding fixable.
  2. A skeleton for the log, created offline

    The skill being reported on, here the OpenCms skill, ships a small reporting client along with its procedure, so that any agent following the skill can also report on it. The client’s first command needs no network and no key; it writes two files — a Markdown log for notes and a JSON template for the report — and prints which skill and which version it is reporting against:

    The agent runs: report.js inittracker    https://report.itp3.uni-stuttgart.de (not contacted)skill      secdev-opencms-workplace @ 2026.09.06.2wrote      …/itp3report-friction/friction.md           …/itp3report-friction/friction.json           Keep the .md as you work; convert to the .json           at the end.
  3. A redacted preview, still offline

    At the end of the task the agent turns its notes into the JSON form and runs a dry run. This sends nothing. It validates the file, masks anything that looks like a secret, and prints what would be submitted so that a person can read it:

    The agent runs: report.js submit --file friction.json --dry-run

    Here is what one report in that preview looks like. The finding is invented for this page; the fields are the real ones. The finding is about two ways of writing a page address: the full path inside the CMS’s internal file store (VFS) and the shorter path within one site:

    Note the [REDACTED:jsessionid] in the evidence. The agent pasted a failing command with the session cookie still in it — the expected accident, not a rare one. The preview masks the cookie and says so, and the skill tells the agent to fix the file, because the masking changes only what is printed, not what is stored in the file.

  4. The user decides, and the agent submits

    The agent shows the preview and asks. If the user says no, the log stays on the machine and that is the end of it. If the user says yes, the agent submits with the user’s submission key — a per-person token, obtained from nicolai.lang@itp3.uni-stuttgart.de . The key usually sits in the environment of the working directory, so the agent finds it there and the user types nothing more; otherwise the agent asks for it, for that one process. The skill contains no key, and the agent never saves or prints it. The institute’s tracker is, in addition, reachable from the university network only; that is a choice of this deployment, not a property of the tracker, which could serve reporters anywhere.

    The agent runs: report.js submit --file friction.jsonSubmitted 1 report(s) as run run-3f2a1c0b9d84:  OCMS-2026-0042  The shared-element scan takes a                  VFS-absolute path, not a site-relative oneTell the user these references. A later session can checkwhat happened to them with: report.js status <ref>

    A whole session’s findings go in one submission and are stored as one run, so the maintainer reads them as the connected account of a session rather than as unrelated fragments, and each report gets a reference such as OCMS-2026-0042. If the server recognised a possible duplicate — an open report with the same skill, kind and title — it says so under the reference. The report is stored anyway: two agents hitting one problem is evidence, not noise.

  5. Later sessions read back the decision

    The reference, the identifier such as OCMS-2026-0042 that the submission printed, is the only handle the user keeps on a submitted report. With it, any later session can ask what became of the finding:

    The agent runs: report.js status OCMS-2026-0042OCMS-2026-0042  fixed  p1fixed in 2026.09.07.1triage said:  reading.md now states the path form beside every  command; ocms.js shared refuses a site-relative path  with a message naming the expected form.last changed 2026-09-07T10:21:00+00:00

    p1 is the priority the maintainer set; triage is the maintainer’s sorting of reports. A rejection comes back with its reasoning too, and the client says plainly: do not refile this; if you disagree, tell the user. A session that learns something after submitting — a cause ruled out, half a finding resolved — can append a comment with report.js comment <ref> --text "…". It cannot edit what it submitted; a report is a record of what was claimed.

The maintainer’s side

The other end of the loop is a second skill, the triage skill, run by the maintainer’s own agent under the maintainer’s supervision. That agent holds a separate maintainer token and sees the whole queue. In the steps that follow, the agent does the reading, reproducing and editing, and the maintainer, a person, decides and pushes.

  1. Start in the skill’s development checkout

    The work on a skill happens in the skill’s own repository, in a development checkout on the maintainer’s machine. The maintainer opens a terminal there and starts a secdev sandbox, here with the browser image, since reproducing a finding about the OpenCms skill needs a headless browser; inside the sandbox, the agent starts in unattended mode with clauded, the secdev command that starts the agent without approval prompts (the YOLO mode of the secdev overview ):

    cd ~/itp3skills-opencms
    secdev browser
    # inside the container, unattended mode:
    clauded
    

    The checkout holds everything the following steps need: the skill files that will be corrected, the repository’s versioning and contribution rules, and a file triage.env with the tracker’s address and the maintainer token, which the agent loads into its environment before its first command. The token is in that file and nowhere else; the triage skill itself contains no credential. The maintainer then gives the agent its task as an outcome:

    Prompt to the agent:

    Work the queue of open reports about the OpenCMS skill.

    The agent reads the triage skill and starts with the queue, which is the next step.

  2. Read the queue as sessions, not rows

    The maintainer’s agent opens the queue with three commands, which show its shape, its sessions and one report:

    The agent runs: triage.js statsThe agent runs: triage.js runsThe agent runs: triage.js show OCMS-2026-0042

    stats gives the shape of the queue — counts by status, kind and skill. runs groups open reports by the session that filed them, and the skill tells the agent to take one run at a time: a wrong instruction and the three confusions downstream of it are one change, not four. show prints one report with its full history, and frames every submitted field as a quotation, so that the agent reads it as a claim to check, not as an instruction to follow. The command-line client is the maintainer’s agent’s tool, and every change to a report goes through it. The tracker’s web interface shows the same queue, and it is meant for the maintainer, the person: to look over the queue, to read a report with its history, and to check what the agent has done, not to work the reports:

    The itp3report web interface listing a queue of fictional reports with reference, priority, status, kind, skill and title columns, and the status and skill filters visible. (enlarge)
    The tracker’s queue. Priorities and statuses are the maintainer’s; the submitted text underneath them never changes.
  3. Claim, classify, reproduce

    The first command takes a report out of the queue: the agent claims it, and the claim refuses if another agent already holds the report, so a finding is never verified twice:

    The agent runs: triage.js claim OCMS-2026-0042

    Classifying puts the finding in the narrowest correct place — a generic CMS fact belongs in the generic skill, a one-site convention in that site’s skill, a defect of the running platform in neither. The agent disposes of what is not a skill defect first, and each disposal needs a note the reporter will read back:

    The agent runs: triage.js dup OCMS-2026-0043 --of OCMS-2026-0042The agent runs: triage.js needs-info OCMS-2026-0044 \  --note "Which path? The reproduction has none."The agent runs: triage.js reject OCMS-2026-0045 \  --note "Project-local state, not general behaviour."

    Then the maintainer’s agent reproduces the finding before changing anything. A report says the skill is wrong; the agent has to show it. Live write tests happen only in the designated practice area of the CMS. Two commands then accept the reproduced finding with a priority and a note, and mark it as started:

    The agent runs: triage.js accept OCMS-2026-0042 --priority p1 \  --note "Reproduced in the sandbox: the two commands \  take different path forms."The agent runs: triage.js start OCMS-2026-0042

    The web interface shows the same report with its history:

    The detail page of one fictional report in the itp3report web interface: the submitted fields (title, summary, observed, expected, reproduction, evidence) shown as an unchangeable record, and below them the event history with status changes, a priority and a resolution note. (enlarge)
    One report and its history. Every state change appends an event with the old and new value; the submitted fields have no write path at all.
  4. Patch the skill and bump its version

    The correction goes into the one file that was wrong — here, a reading reference and the refusal message of one script — as a statement of the corrected fact, not the story of the bug. The same change bumps the skill’s VERSION file, following the versioning rules . The maintainer’s agent validates the change and, when the change alters an instruction substantially, forward-tests it by giving a fresh agent the original task without the diagnosis. Then the agent commits the change on the repository’s master branch and tags it.

  5. Close the report against a version that exists

    One command closes the report, and it demands the version that carries the fix:

    The agent runs: triage.js fix OCMS-2026-0042 --version 2026.09.07.1 \  --note "reading.md names the path form beside every command."

    fix refuses without --version, and the version named must be one that a commit already carries, because the next reporter reads it back and will look for it.

  6. The release reaches users

    The new version travels along the skill’s normal distribution channel. A release is the tagged commit on master, pushed by the maintainer; that is the whole procedure. How the skill then reaches its users is the skill’s own affair: the OpenCms skill, for example, is embedded in a secdev container image at the next build, and a skill that users mount into the sandbox themselves reaches them with their next pull.

    The next agent to add a team member reads the corrected sentence, hits no friction there, and reports something else.