Workflow

How a physicist runs a research project inside secdev and stays in charge of it: set the project up, write the plan, hand the agent one task at a time, and read what comes back.

The overview promised a loop the agent runs on its own while a person stays in charge of the question. This page shows what the person does, step by step, under the guided contract, in which the agent works on one task at a time and then reports back. The Example page shows what came out of one real project run this way: the agent’s records, its results and the finished paper.

Setting up a project

A project starts with a few steps the person does once; the agent joins at the fourth. From then on the work alternates between the two: the agent runs a task and writes what it found into the draft, the person reads the draft and steers, and at the end the person writes the paper from what the draft holds.

  1. Start a scientific container

    Two secdev images carry the research tooling. newton is the local workbench: numerics, symbolic algebra, LaTeX. einstein is newton plus reach: it can submit jobs to the institute’s compute cluster and run commands on the GPU host. A project that needs the cluster runs in einstein:

    cd ~/projects/example
    secdev einstein
    

    Start the container inside a terminal multiplexer such as tmux, a program that keeps a terminal session alive after its window closes, so the session survives a closed laptop. On start-up the container writes a manifest file, SECDEV.md, that tells the agent what this container can reach and what it must not attempt.

  2. Lay down the scaffold and the contract

    One command, run once in the empty project directory, creates the fixed project structure and asks which contract the agent will work under:

    $ research-init
    research-init: select research contract:
      1) guided-researcher (default)
      2) fully-autonomous-researcher
      3) automatic-optimizer
    choice [1]:
    research-init: is this a multi-agent project
      (adds ROLES.md)? [y/N]:
    research-init: create full research scaffold? [Y/n]:
    

    Guided is the default: one task at a time, and the agent returns when the task is done. The other two contracts are for unattended runs and are described further down . The command creates the following files and directories, each with one job:

    project/
      AGENTS.md      # entry point: points the agent at the files
      CONTRACT.md    # the rules the agent works under
      PLAN.md        # the physicist's goal and tasks; read-only
      TASKS.md       # the tasks, their state and outcome
      PROBLEMS.md    # open and closed conceptual problems
      LOG.md         # the lab book: the agent's memory
      LITERATURE.md  # notes on every paper actually read
      .gitignore     # keeps data and build output out of Git
      code/          # the simulation and analysis code
      data/          # what the runs produced
      literature/    # the downloaded papers
      draft/         # main.tex: findings, written for the
                     # physicist
      paper/         # main.tex: the manuscript; the agent
                     # edits it only on request
    

    The contract, installed as CONTRACT.md, is the file of numbered rules the agent reads at the start of every session, and it determines everything that follows: how the plan works, what the agent must record, what it may not do, and where it may write. It is short enough to read in full, and reading it is the quickest way to see what guidance the agent works under. The guided contract, as the agent receives it:

    # §C. Guided Research Contract
    
    This contract applies when `research-init` has created the full research
    scaffold. It defines project-specific research behavior.
    Container capabilities, installed tools, security constraints, and optional
    remote resources are documented separately in `SECDEV.md`.
    
    ## §C1. Project Scaffold
    
    - MUST read `SECDEV.md`, `CONTRACT.md`, `PLAN.md`, `TASKS.md`, `PROBLEMS.md`,
      and `LOG.md` at session start.
    - MUST NOT create the research layout manually if required scaffold files are
      missing; ask before running `research-init`.
    - `research-init` is opt-in and refuses to overwrite existing files.
    
    ### §C1.1 `PLAN.md`
    
    - Human-owned project contract.
    - MUST read at session start.
    - MUST NOT modify.
    - If edits are requested, tell the human to run `chmod +w PLAN.md`.
    - Front matter:
      - `title`: title for drafts, reports, and summaries.
      - `label`: ASCII slug; use only letters, numbers, dots, underscores,
        hyphens.
      - `mattermost_channel_url`: optional metadata; this guided contract does
        not use Mattermost for reporting.
    - The `Tasks` section contains immutable initial research directions.
    
    ### §C1.2 `TASKS.md`
    
    - Shared, human- and agent-writable list of evolving scientific tasks.
    - MUST contain only high-level scientific objectives, not technical
      micro-steps.
    - MUST use stable task IDs and keep each task `OPEN`, `BLOCKED`, or `CLOSED`.
    - MUST NOT delete tasks. Closed tasks MUST retain their outcome or closure
      reason; blocked tasks MUST retain their blocking reason.
    - MUST update task state and outcome as the research develops.
    
    ### §C1.3 `PROBLEMS.md`
    
    - Shared, human- and agent-readable record of conceptual or structural
      research problems.
    - MUST keep unresolved problems under `OPEN` and resolved problems under
      `CLOSED`, with stable IDs and a resolution or closure reason.
    - MUST NOT delete closed problems.
    - Technical failures, package limitations, and implementation details belong
      in `LOG.md`, not `PROBLEMS.md`.
    
    ### §C1.4 `LOG.md`
    
    - Agent lab book and persistent memory.
    - MUST read at session start.
    - MUST record every experiment, simulation, investigation, failure, and
      dead end.
    - MUST keep newest entries first.
    - Technical micro-steps, implementation problems, and exploratory memory
      belong here.
    - SHOULD use `research-log` to add entries instead of hand-formatting them.
    
    ### §C1.5 `LITERATURE.md`
    
    - Annotated bibliography for `literature/*.pdf`.
    - A paper counts as read only after it has an entry.
    - MUST locate original source before citing when doing online literature
      work.
    - MUST NOT cite from memory.
    - SHOULD use `research-lit` to find PDFs missing entries and create stubs
      before filling verified bibliography details.
    
    ### §C1.6 Folders
    
    - `data/`: store simulation or experimental output; use `research-run` for
      local command runs when practical so outputs and `LOG.md` stay linked.
    - `code/`: store independent codebases or `uv` projects; prefer one
      subfolder per experiment and reference it from `LOG.md`.
    - `draft/`: agent-written research digest for human review; MAY contain
      explanatory prose, plots, caveats, and tested, noteworthy results.
    - `paper/`: final-paper workspace; edit only on explicit human request.
      Prose in final papers MUST be human-written, so agents MUST NOT write
      fully formulated paper paragraphs.
    - `literature/`: store downloaded paper PDFs; every read file needs one
      `LITERATURE.md` entry.
    
    ## §C2. Session Warmup
    
    - At session start, MUST run `research-warmup`.
    - MUST read `SECDEV.md`, `CONTRACT.md`, `PLAN.md`, `TASKS.md`, `PROBLEMS.md`,
      and `LOG.md`; the helper checks their presence but does not replace reading
      them.
    - SHOULD run `research-audit` when starting substantial work or before
      returning final results.
    
    ## §C3. Guided Work Loop
    
    For the user-requested job:
    
    1. Read current project state, tasks, and known problems.
    2. Plan the next concrete step.
    3. Execute.
    4. Evaluate evidence.
    5. Document technical work in `LOG.md`, maintain `TASKS.md` and `PROBLEMS.md`,
       and, when noteworthy, update `draft/`.
    6. Continue until the requested job is fully achieved or genuinely blocked.
    7. Return to the user with the outcome, evidence, and remaining risks.
    
    Loop rules:
    
    - MUST continue after intermediate steps when useful work remains within the
      requested job.
    - MUST return to the user after the requested job is complete.
    - MUST return to the user when progress requires a decision that cannot be
      derived from `PLAN.md`, `TASKS.md`, `PROBLEMS.md`, evidence, or a safe
      default.
    - MUST NOT use Mattermost for progress reporting under this contract.
    
    ## §C4. Planning Before Execution
    
    - For any non-trivial task:
      - MUST plan before editing, running simulations, or modifying drafts.
      - MUST base plan on `PLAN.md`, `TASKS.md`, `PROBLEMS.md`, and `LOG.md`.
    - Trivial tasks are limited to mechanical inspection or status checks that do
      not change project state.
    
    ### §C4.1 Git Discipline
    
    - Unless the user explicitly requests another workflow, SHOULD develop on the
      repository's default branch, normally `master` or `main`, and SHOULD NOT
      create additional branches.
    - SHOULD commit autonomously at reasonable, coherent checkpoints with
      explanatory commit messages.
    - If undoing committed changes is necessary, SHOULD use `git revert` rather
      than rewriting history.
    
    ## §C5. Token Budget
    
    - MUST use `research-warmup` during warmup.
    - SHOULD check `usage` before token-heavy steps, after experiments finish,
      and when nearing limits.
    - If token budget prevents safe completion, document the state in `LOG.md`
      with `research-log` and return to the user with the next concrete action.
    
    ## §C6. Completion And Blocking
    
    - A job is complete only when the requested deliverable is produced, checked,
      and documented.
    - If blocked, MUST document:
      - what was attempted,
      - the evidence for the block,
      - the smallest user decision or external change needed to continue.
    - SHOULD use `research-log --blocked` for blocked work.
    - MUST NOT treat a failed experiment as a scientific conclusion until the
      implementation has been checked.
    
    ## §C7. Integrity
    
    ### §C7.1 Never break a promise.
    
    - If you say "I will do X", do it.
    - Under-promise and over-deliver.
    - If work is deferred, MUST record reason in `LOG.md`, preferably with
      `research-log`.
    
    ### §C7.2 Never manipulate evaluation.
    
    - MUST NOT change metrics, test sets, problem definitions, or fixed
      hyperparameters to make results look better.
    - MUST NOT hard-code results.
    - MUST NOT cherry-pick seeds.
    
    ### §C7.3 Never fabricate citations.
    
    - MUST verify every bibliography entry against the actual source.
    - MUST confirm exact title, full author list, year, venue, and DOI/arXiv or
      other identifier.
    - If the paper cannot be found, do not guess.
    
    ## §C8. Efficiency
    
    ### §C8.1 Make it work before moving on.
    
    - Treat experiment crashes as bugs, not as evidence against the method.
    - MUST NOT discard methods because of implementation failures.
    - If an experiment crashes, investigate, fix, and rerun when still relevant
      to the requested job.
    
    ### §C8.2 Use compute intelligently.
    
    - At session start, SHOULD check local CPU/memory capability.
    - IF `SECDEV.md` documents specialized compute or remote resources suitable
      for the task, SHOULD use them according to that documentation.
    - IF starting a local simulation:
      - MUST estimate expected runtime before launch.
      - SHOULD use `research-run` to capture outputs, timeout metadata, and the
        matching `LOG.md` entry.
      - MUST use a reasonable timeout when the runtime has not been tested
        before.
      - MUST run without a timeout only after prior tests give a good runtime
        estimate.
      - MUST schedule and perform a status check at that expected runtime.
      - MUST investigate simulations that run much longer than expected.
    - Remote jobs that can be actively monitored, such as Slurm jobs, are exempt
      from the local timeout requirement.
    - For embarrassingly parallel work, SHOULD use multiple worker processes when
      it materially reduces wall time and MUST leave CPU headroom.
    
    ## §C9. Scientific Rigor
    
    ### §C9.1 One variable per experiment.
    
    - Change exactly one thing per experiment.
    - If two things change and the metric improves, you cannot know which helped.
    
    ### §C9.2 Evaluate in tiers.
    
    - Tier 1: seconds, does it run?
    - Tier 2: minutes, signal on small subset?
    - Tier 3: full evaluation for reportable claims.
    - MUST NOT draw conclusions from small-scale bug-catching runs.
    
    ### §C9.3 Bound your expectations.
    
    - Before implementing a heuristic, identify theoretical best case and
      estimate maximum possible improvement or correction.
    
    ### §C9.4 Never overstate simplified results.
    
    - MUST state caveats for toy models, approximations, restricted parameter
      regimes, finite-size systems, relaxed assumptions, and surrogate metrics.
    - Restricted contradiction is evidence, not full disproof.
    - Keep caveats visible in `LOG.md`, plots, user reports, and `draft/`.
      If the caveat affects a human-selected paper result, include it in
      `paper/` as a terse note or figure/caption constraint.
    
    ## §C10. Documentation And Reproducibility
    
    ### §C10.1 Record everything.
    
    - MUST log every experiment with goal, method, substeps, outcome, and next
      step, preferably with `research-log` or `research-run`.
    - MUST include failures.
    - MUST document tested, noteworthy results in `draft/main.tex`.
    - SHOULD use plots when available.
    - MUST use tables only when clearly better than plots for the message.
    - MUST keep abbreviation list current.
    - Substeps, dead ends, and later-invalidated results belong in `LOG.md`.
    - If it is not in `LOG.md`, it did not happen.
    
    ### §C10.2 Verify before claiming.
    
    - Assume you are wrong until verified.
    - MUST write verification scripts, not just explanations.
    - MUST actively try to falsify claims.
    - MUST grade claims as verified, partially verified, or unverified.
    
    ### §C10.3 Draft And Paper Roles.
    
    - `draft/` is the agent-human interface. Agents MAY write clear prose,
      explanations, caveats, plots, and result summaries there.
    - `paper/` is the final-paper workspace. Agents MUST edit it only when the
      human explicitly asks.
    - Agents MUST NOT convert `draft/` wholesale into `paper/`.
    - Normal flow is `LOG.md`, `data/`, and `code/` to `draft/`, then
      human-selected content to `paper/`.
    - In `paper/`, agents MAY add paper-grade figures, equations, labels,
      bibliography wiring, provenance notes, and keyword-style placeholders.
    - In `paper/`, agents MUST NOT write fully formulated prose paragraphs; prose
      may only be inserted when supplied by the human.
    
    ### §C10.4 Plot And Figure Quality.
    
    - Figure rules apply to both `draft/` and `paper/`.
    - Every figure MUST support one clear claim and be referenced from nearby text
      or notes.
    - Multi-panel labels such as `(a)`, `(b)`, `(c)` MUST use the same font, size,
      style, and relative position; corresponding labels SHOULD align across
      panels.
    - Text inside figures MUST be legible at final rendered size, with consistent
      font sizes for labels, ticks, legends, annotations, and panel labels.
    - Labels, legends, annotations, arrows, and panel marks MUST NOT overlap data,
      axes, colorbars, or other informative content.
    - Axes and colorbars MUST name quantities and units where applicable.
    - Related figures MUST use consistent variable names, colors, markers, line
      styles, limits, and normalization when comparing the same quantities.
    - Prefer vector formats for plots when practical; avoid decorative styling.
    - Before marking a figure ready, compile or render it and inspect final-size
      output for alignment, readability, overlap, missing labels, and consistency
      with captions or notes.
    
    The guided research contract in full, as research-init installs it into a project. Scroll inside the box to read it. Open the file in a new tab
  3. Write the plan

    PLAN.md is the physicist’s document, and the only place the scientific direction is set. The scaffold installs it read-only, so the physicist unlocks it, writes the goal and the first tasks, and locks it again:

    chmod +w PLAN.md
    $EDITOR PLAN.md
    chmod 0444 PLAN.md
    

    The plan is a scientific brief, not a prompt. It states the goal and the model, and then numbered tasks, each saying what to compute, what counts as a result, where the files go, and how much freedom the agent has to deviate. A first task typically asks for an implementation and a validation against known limits; a later one for a parameter scan with a plot. The opening of a real plan, shortened (its physics is explained on the Example page), shows the shape:

    # Research Plan
    > Owned by the human researcher. The agent must read this
    > file at the start of every session and must not modify it.
    
    ## Goal
    We study the phase diagram of the blockade Hamiltonian
    [...] as function of Ω/Δ [...]. The method is based on
    large scale numerical MPS simulations using TenPy on the
    computer cluster available.
    
    ## Tasks
    Task 1: Read the relevant files and get familiar with the
    problem and the objectives. Then, write a summary of the
    Hamiltonian and the main objective in the draft.tex.
    Task 2: Prepare the steps to perform numerical simulations
    using TenPy. [...]
    Task 3: Perform a fast scan of the phase diagram. [...]
    

    The note at the top of the plan is backed by the contract from the previous step, whose clause on the plan makes it the physicist’s document:

    Commit the scaffold and the plan before the agent starts, so that everything it does afterwards is visible as a change against a clean starting point.

  4. Warm the agent up

    Start the agent from the project root and give it one instruction before any research task:

    Prompt to the agent:

    Warm up by reading AGENTS.md and every file it links. Report back when you understand the project and await my first research task.

    The agent runs research-warmup, which checks that the scaffold is complete, prints how much of its usage allowance is left and what resources the machine has, and shows the Git state; then it reads the manifest, the contract, the plan, the task and problem lists and the lab book. A session that skips the warm-up rediscovers things the lab book already knows. The contract prescribes the warm-up:

  5. Hand over one task and let the agent run

    Give the agent one task from the plan, by number, and leave it alone:

    Prompt to the agent:

    Work on Task 3 as described in PLAN.md.

    The agent plans, calculates, codes, submits jobs, waits, evaluates and documents, for as long as the task takes, without stopping to ask. A task can take an afternoon or run for days on the cluster. The agent keeps two records, for two readers. The lab book is its own memory: every attempt, including the failed ones, in a dense form meant to be read by the agent at the next session start. The LaTeX draft in draft/ is the interface to the physicist: for each finding a short explanation, the plot, the caveats, and pointers to the code and data behind it.

  6. Read the draft, give feedback, repeat

    When the agent returns, the person reads the draft, not the code and not the lab book. The draft is where the agent explains itself to a person, and the plot is usually what decides the next step. The feedback is a scientific conversation in the terminal: which result settles the question, which does not, which idea to drop, what to compute next. Larger corrections go into the plan, which the person unlocks, edits and locks again; then the next task goes to the agent. A project goes round this loop for weeks. The Example page describes the decisions of one project at several of its turns, including a dead end the physicist closed.
  7. Collect the results and write the paper

    At the end the draft holds more than a paper needs: every finding, every dead end, every plot. The person decides which results carry the argument, asks the agent to move those figures, equations and verified references into paper/, and writes the text. The agent may not write a paragraph there. The contract’s clause on the two directories draws the line:

    So the physicist writes the manuscript text, and the agent supplies figures, equations, verified references and, on request, an annotated proof-reading pass with its findings in a separate file. How that division of labour reads in a finished paper is on the Example page.

The guided loop: one task at a time

Under the guided contract the physicist gives one task, the agent works it to completion or to a real blocker, and returns with the outcome, the evidence and the remaining risks. Inside the task the agent does not stop to ask. The loop in full:

The guided loop. The task enters at the top left and the loop ends by returning to the human at the bottom right — the difference between this contract and the autonomous one.

Every pass through the loop leaves an entry in LOG.md, the lab book. An entry has a fixed shape, goal, method, substeps, outcome and next step, and the agent writes it with the research-log helper. The lab book is written for the agent, not for a person: it is the agent’s memory between sessions and the first thing it reads at the next session start, and a physicist opens it only to check a claim:

A light editor view of the top of a LOG.md lab book: two or three complete entries with goal, method, substeps, outcome and next, newest first. (enlarge)
The lab book is the agent’s only memory between sessions, which is why the agent fills it in.

The rule behind the lab book is short, and its last line is the one the agents quote back when asked why they log a failed attempt:

Two further rules matter most when nobody is watching. The agent may not make a result look better by changing what is measured, nor by keeping only the random runs that happened to come out well (choosing the seeds, the numbers that start them), and it may not cite a paper it has not verified:

Rules about record-keeping would be a burden if the agent had to honour them by hand, so the scaffold ships a set of helper commands that make the correct record the convenient one. The agent has already used two of them, research-warmup and research-log. research-run executes a command, captures its output under data/ and writes the matching lab-book entry in one go. research-lit lists downloaded papers that have no entry in the literature file yet, so “I have the PDF” and “I have read it” stay distinguishable. research-audit checks the bookkeeping before a report. The Technical Details page lists them all.

Running jobs on the cluster

In the einstein container the agent submits jobs the same way a person does. The cluster’s job scheduler, the program that queues every user’s computations and starts each one when a machine is free, puts the agent’s jobs into the same queue as everyone else’s, under the same fair-share rules, which give priority to users who have used the cluster least. The contract requires a runtime estimate before a launch and a status check when that time is up. Waiting costs nothing: after submitting, the agent starts a small helper, slurm_wait, as a background process that polls the queue and exits when the job leaves it, and the agent software, Claude Code for example, notices that the helper has exited and resumes the agent. Until then the agent sleeps and spends no tokens. On waking it evaluates the results, fixes what went wrong, resubmits, and makes the plots, so the person coming back in the morning often finds the results already in the draft. The queue itself is nothing special, which is the point:

A dark terminal showing a cluster queue listing with a handful of agent-submitted jobs (running and pending), next to the agent's own monitoring output. (enlarge)
The agent submits to the same queue as everyone else and watches it with the same tools.

The two unattended contracts (experimental)

The fully autonomous contract and the automatic optimizer are experimental: only the guided contract has been through a full project. The fully autonomous contract changes one thing in the guided loop: it no longer returns to the human. The contract’s first clause defines it:

Supervision moves from the terminal to a chat channel named in the plan. At each checkpoint the agent posts a short report there, with the problem, the approach, the result, the next step and a plot, and reads the replies before its next decision. A question to the supervisor must come with a default action and a deadline, so a supervisor who is asleep does not stall the run. The automatic optimizer contract is narrower still: the agent improves one function (for example, a solver whose accuracy the program scores) against a fixed scoring program it may not touch, committing every version before it is scored. Both take their loop and their “never stop” rule from Andrej Karpathy’s autoresearch, an agent that improves a training script overnight against a fixed evaluation [1].

References

  1. autoresearch: AI agents running research on single-GPU nanochat training automaticallyA. Karpathy (2026) · github.com/karpathy/autoresearch