Technical Details

The scaffold is created by one command and governed by a contract the agent reads at every session start. What keeps it honest afterwards is a small set of tools, one file permission and the sandbox.

The Workflow page showed a project from the user’s side. This page describes what runs it, in the order a project meets each piece: the two images, the files the scaffold creates, the three contracts, the helper commands, the route to the compute cluster, and the supervision channel used by the autonomous contract.

Two images: newton and einstein

Both images are built on the secdev image that carries LaTeX, described on the secdev technical page , and share the same research tooling. The following diagram shows what each image adds and what they share:

What each image adds. The rules — contract, plan, lab book, captured outputs, version control — are the same in both.

newton is the local scientific workbench: everything that fits on the machine the container runs on, plus the research scaffold and its helpers. The tools of the newton image:

ToolWhat it is for
numpy, scipy, sympy, pandas, matplotlibThe numerical and symbolic Python stack, and plotting
numba, xarray, dask, JupyterCompiled kernels, labelled arrays, parallel processing, notebooks
Maxima, Octave, gnuplotSymbolic algebra, MATLAB-compatible numerics, plotting from scripts
Fortran, OpenBLAS, LAPACK, GSL, FFTW, HDF5Compiled numerics and large datasets
Lean 4A proof assistant, for checking formal proofs
research-initLays down the project scaffold and installs the chosen contract
research-* helpersSession checks, the lab book, captured runs, the literature file, reports
agent_monitorWatches a detached agent session and nudges it after a budget reset

einstein is newton plus reach. It mounts a dedicated SSH key, not the user’s own, and configures SSH aliases so the agent never needs to type a hostname. A shared remote home is mounted under ~/remote; each project gets one directory there, named after the plan’s label, and the agent may not look into any other. The tools the einstein image adds:

ToolWhat it is for
slurm_execRun a command on the cluster’s master node, typically to submit a job
slurm_waitWait for a cluster job as a tracked background process
gpu_execRun a command on the GPU host, for setup and short smoke tests
mathematica_execRun a command on the computer-algebra host
remount_remoteReconnect the shared remote home if the mount drops

Use newton for anything that fits on the local machine and einstein when the project needs the cluster, the GPU host or Mathematica; both start with secdev newton or secdev einstein.

The scaffold

research-init runs inside either image, in the project root. It asks which contract to install and whether several agents will share the project, then creates the contract, the scaffold files, the fixed directories and a starting LaTeX draft and paper. A multi-agent project also receives a read-only ROLES.md and a collaboration section in the contract. The files research-init leaves behind depend on one another, and all of them are kept under the rules of the contract:

A research project after initialization. The contract governs how every other file is kept; the plan is the only file the human writes alone, and the paper the only one the agent touches only when asked.

Each file has an owner and a timescale of its own, and the two together say what an agent may expect to find where:

FileOwnerPurposeTimescale
SECDEV.mdGeneratedContainer manifestEvery container start
CONTRACT.mdFixed at initThe chosen contract: MUST / MUST NOT clausesStable
PLAN.mdHuman, read-onlyGoal, label, initial tasks, referencesRarely, by hand
ROLES.mdHuman, read-only (multi-agent)One section per agentRarely, by hand
TASKS.md, PROBLEMS.mdSharedTasks and conceptual problems with stable IDs, never deletedDays to weeks
LOG.mdAgentThe lab book, newest firstEvery session
LITERATURE.mdAgentOne entry per PDF; a paper counts as read only when it has an entryContinuous
data/, code/AgentOne subdirectory per run, referenced from the logContinuous
draft/AgentResearch digest for human reviewContinuous
paper/HumanFinal-paper workspace; the agent writes no proseOn request

The files run from stable to fine-grained on purpose: the plan holds the direction, the task and problem lists hold the evolving state, and the log holds the detail; an agent starting a session reads down that gradient. The plan’s protection is one file permission, backed by the contract’s rule not to lift it and by research-audit, which warns whenever the plan is found writable.

The three contracts

The contract is chosen at initialization and copied into CONTRACT.md. The guided and the autonomous contract share the scaffold rules, the integrity clauses and the rigour clauses and differ in when the agent stops; the optimizer has integrity and record rules of its own.

guided-researcher (default). One requested task at a time. The loop (read state, plan, execute, evaluate, document, repeat) continues until the task is complete or genuinely blocked, then returns to the user with outcome, evidence and remaining risks. The agent must return when a decision cannot be derived from the plan, the task list, the evidence or a safe default.

fully-autonomous-researcher (experimental). The same loop with no agent-decided stopping condition. Progress goes to the supervision channel and the scaffold files, never to the terminal conversation; questions carry a proposed default and a timeout; near the end of the token budget the agent waits for the reset instead of stopping.

automatic-optimizer (experimental). One optimizer function, one fixed evaluator, one score. The agent may edit only the optimizer and the records of the current optimisation run; each experiment is committed before evaluation and reverted if it did not improve; the evaluator, metric, seeds and budget are untouchable.

The two unattended contracts have a named source:

The clauses the guided and autonomous contracts share are the ones a careful researcher is trained to follow and an agent will skip by default:

  • change one variable per experiment;
  • evaluate in tiers, from a run of seconds to the full run;
  • assume you are wrong until verified, and try to falsify;
  • grade every claim verified, partially verified or unverified;
  • state caveats for toy models and finite sizes;
  • record every failure.

The contract also sets rules for figures, and they apply to draft/ and paper/ alike.

Helper commands

The contract asks the agent to record every run, every paper read and every dead end in the scaffold files. That record stays complete only if writing it is easier than skipping it, so newton and einstein ship helper commands that write the entries in the required shape:

CommandDoes
research-warmupSession-start checks: required files, token budget, host resources, Git status, supervisor replies
research-auditBookkeeping check: stable IDs, log ordering, unlisted PDFs, unreferenced run directories, plan still read-only
research-logPrepends a contract-shaped entry to LOG.md
research-litLists PDFs without a LITERATURE.md entry
research-runRuns a command with captured output in a new data/ subdirectory and logs it
research-reportChecks supervisor replies, then posts a report to the supervision channel

All of these write or check the same kind of record. The log entry template is the contract’s unit of record, and research-log fills it in:

## YYYY-MM-DD HH:MM UTC  Short title
- **Goal:**     what question this entry tries to answer
- **Method:**   code path / sim params / pointer to
                code/<subdir> and data/<subdir>
- **Substeps:** [ ] todo   [x] done
- **Outcome:**  result, plot reference, or "blocked: …"
- **Next:**     the immediate follow-up (or "none")

Alongside the research helpers, every image provides usage, with which an agent inspects its own usage allowance and can wait across a reset, and host_info, which prints the machine’s resources.

Reaching the compute cluster

The einstein container is not where long computations run. It holds the code and the agent, and the agent sends the heavy work out: simulations go to the compute cluster as jobs, through a wrapper that submits them in the agent’s name, and symbolic calculations go to the computer algebra host through a second wrapper. The results come back into the project directory, and the agent, which has been sleeping, wakes up to evaluate them. From the cluster’s side the agent is an ordinary user with the same queue, the same hosts and the same fair-share rules. The following figure shows the round trip:

The agent uses the same queue, the same fair-share rules and the same hosts as a human user, and spends no tokens while a job runs.

The manifest gives the agent a job template and a wait recipe: submit the job, wait for it in the background, do other useful work, inspect the result on wake-up. GPU work beyond short smoke tests goes through the scheduler and never straight to the host. The contract adds the discipline: estimate the runtime before launching a local simulation, set a timeout when the runtime is untested, and investigate anything that runs far longer than expected.

Supervision channel and agent monitor (experimental)

The supervision channel and the agent monitor exist for the autonomous contract only; the guided contract forbids progress reports outside the terminal.

Supervision channel. The header block of PLAN.md names a channel on Mattermost, the institute’s chat service, and research-report refuses to post before checking for unread replies. The contract requires brief reports at planning and evaluation checkpoints, with plots posted directly, written “like an independent PhD student summarizing findings for a supervisor to decide next steps”.

Agent monitor. agent_monitor watches the detached sessions and injects a continuation prompt after a token reset if a session appears stalled. That is what lets a run cross a budget reset without a person at the keyboard.

Governing a sandboxed command-line agent with written methodological rules is not unique to this project. Zimmer, Pelleriti, Roux and Pokutta describe such a framework in their guide to agentic research [2], and parts of the contracts here are inspired by their work. The contracts differ in emphasis: the research goal lives in a file the human owns and the agent cannot write, and the clauses that matter most are enforced by file modes and the sandbox, not by the prompt alone.

References

  1. autoresearch: AI agents running research on single-GPU nanochat training automaticallyA. Karpathy (2026) · github.com/karpathy/autoresearch
  2. The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine LearningM. Zimmer, N. Pelleriti, C. Roux, and S. Pokutta (2026) · arXiv:2603.15914