Technical Details

How the report tracker is built: one process, one database file, one configuration file, and no language model anywhere in it.

The workflow followed one finding around the loop. This page is the machinery under it: what the tracker is built from, who may reach which of its three entry points, how it scrubs what arrives, how a report is classified and moved, and the versioning rule that connects a closed report to a release the reporter can obtain.

Architecture

The tracker, itp3report, is a Python web application with a single SQLite file as its storage, run as two containers: the service and a nightly backup loop. Everything reaches the service through the host’s reverse proxy, the web server in front of the application that handles the encryption (TLS) and decides which network may reach which entry point. The entry points are called surfaces below. The following diagram shows the three surfaces and what sits behind them:

Three surfaces behind one reverse proxy. nginx decides who may reach a surface; the application decides whether the credential fits it.

Surfaces and credentials

Each surface has its own network zone and its own credential, and a credential opens only its own surface. The following table lists what each accepts and who may reach it:

SurfacePrefixReachable fromCredentialCan do
Intake API/api/v1/any university addressper-user submission tokenread the allowed report kinds and limits; submit reports; comment on and read the status of its own reports
Triage API/maint/the institute’s addressesmaintainer tokenlist, claim, change state, comment, delete, export
Web UI/the institute’s addressesHTTP basic auththe same queue, in a browser

The intake surface offers no listing and no search: a token reaches only the reports it filed itself, and a foreign reference gets the same answer as one that does not exist.

Per-token skill permissions

Every submission token is bound to the skills it may report against, and the server rejects the whole batch if any report names another skill, and stores nothing. The binding is attribution, not secrecy: it gives the tracker a skill field the submitter cannot get wrong by accident, and it limits a leaked token to filling the queues of the skills it is bound to.

Redaction, on both ends

The skill tells agents never to include a session identifier, a credential or private page content. The server does not trust that. On arrival, the server scrubs every free-text field with a fixed list of rules for cookies, tokens, passwords, private keys and any bare random-looking string (high entropy); a match becomes a marker, and the server flags the report as redacted so the client can tell the user.

Redaction is blunt on purpose: a false positive costs a maintainer one follow-up question, while a false negative puts a live session identifier into every export, backup and rendered page. The rules spare paths and resource names, because they are how a report says where. The reporting client carries the same rule list, so its offline preview masks exactly what the server would mask.

Report kinds and the state machine

The classification is defined once, in the server, and served to clients so a reporting script can validate a report before posting it. Free-text categories drift; these stay comparable across years. A report’s kind is about the skill, not about the system the skill drives:

kindMeaning
doc-wrongThe skill stated something untrue
doc-ambiguousIt permitted a reading that fails
doc-missingThe answer was not there at all
script-bugA bundled script failed or misbehaved
optimizationIt works, but a cheaper or safer route exists
feature-requestThe skill should be able to do something and has no path to it
platform-defectThe system the skill drives is broken; the skill is right

A report also carries a confidence, the affected skill’s version, what was observed and expected, and optional evidence. The server assigns the reference, the timestamp and the submitting token, so a caller cannot backdate a report or choose its own reference.

Triage moves a report through a fixed set of states with fixed allowed moves between them (a state machine), which the server enforces on every change. The following diagram shows the main path:

The main path. A report reaches ‘fixed’ only by having been accepted and worked, which is what makes its history a readable account and not a list of assertions.

The server also allows the ways back: a wrongly closed report, or one whose missing information has arrived, returns to the queue without being filed again. Internal comments are private; the note given when a report is closed or set to needs-info is what the reporter reads back. Closing a report is a deliberate act of telling the reporter why, because a silent rejection is how the same finding arrives twice. Submitted fields have no write path from any surface, and every state change appends an event, so the history is not editable either.

Duplicates and limits

On submission the server looks for open reports with the same skill, kind and title and returns them as possible duplicates; the report is stored regardless, and the hint lets a person drop an obvious repeat in the offline preview.

The server also enforces limits: how many reports one submission may hold, how long each field may be, how far back the duplicate search looks, and how many submissions one token may make in a period. All of them are settings in the server’s configuration file, and the intake API publishes them, so a reporting client can check a report against them before sending it. The rate limit is counted against the token’s name, the label the administrator gave the token in the configuration, not against the secret value. The administrator can therefore replace a token’s value, for example after it leaked, without resetting the holder’s limit or losing their history, as long as the name stays the same.

Configuration

The server reads its settings from a single YAML file, and a broken file, an unknown field or an unreplaced placeholder credential stops the service from starting. Each submission token is one entry:

intake_tokens:
  - name: alice
    # private; never in reports, APIs or logs
    note: "A. Example — uses the team-page skill"
    token: CHANGE-ME
    skills:
      - team-page-skill

The note records who received the token and why, and stays in the file; clients receive only the token value. The maintainer token and the web credentials are separate sections of the same file.

Calendar versions

Each skill has its own VERSION file holding a calendar version of the form YYYY.MM.DD.N in UTC, for example 2026.09.22.10, where N starts at 1 each day. Versions are independent between skills, so the generic skill and a site skill never share a number. Two rules carry the scheme:

  • Any change to a skill’s instructions, references or scripts bumps its VERSION in the same commit. Several edits in one change share one bump; changes outside the skill need none.
  • A released version is never reused for different content. A release is a commit on master carrying a tag of the form <skill-name>/v<VERSION>.

The reporting client reads that VERSION file, so every report names the exact release it was filed against; a development copy with uncommitted changes reports itself as +dirty instead.

What “fixed” means

A report is closed as fixed once the fix is committed on master with the affected skill’s VERSION bumped in that same commit. The report then records that version, and that value is what a later session reads back: an agent that hits the same problem compares the version it is running against the version the report names and learns whether its copy of the skill already contains the fix. That comparison only works if the version named in the report exists as a commit, which is why the closing command insists on a version and why the maintainer’s agent runs it only after the commit is in place.