# §C. Guided Research Contract

This contract applies when `research-init` has created the full research
scaffold. It defines project-specific research behavior.
Container capabilities, installed tools, security constraints, and optional
remote resources are documented separately in `SECDEV.md`.

## §C1. Project Scaffold

- MUST read `SECDEV.md`, `CONTRACT.md`, `PLAN.md`, `TASKS.md`, `PROBLEMS.md`,
  and `LOG.md` at session start.
- MUST NOT create the research layout manually if required scaffold files are
  missing; ask before running `research-init`.
- `research-init` is opt-in and refuses to overwrite existing files.

### §C1.1 `PLAN.md`

- Human-owned project contract.
- MUST read at session start.
- MUST NOT modify.
- If edits are requested, tell the human to run `chmod +w PLAN.md`.
- Front matter:
  - `title`: title for drafts, reports, and summaries.
  - `label`: ASCII slug; use only letters, numbers, dots, underscores,
    hyphens.
  - `mattermost_channel_url`: optional metadata; this guided contract does
    not use Mattermost for reporting.
- The `Tasks` section contains immutable initial research directions.

### §C1.2 `TASKS.md`

- Shared, human- and agent-writable list of evolving scientific tasks.
- MUST contain only high-level scientific objectives, not technical
  micro-steps.
- MUST use stable task IDs and keep each task `OPEN`, `BLOCKED`, or `CLOSED`.
- MUST NOT delete tasks. Closed tasks MUST retain their outcome or closure
  reason; blocked tasks MUST retain their blocking reason.
- MUST update task state and outcome as the research develops.

### §C1.3 `PROBLEMS.md`

- Shared, human- and agent-readable record of conceptual or structural
  research problems.
- MUST keep unresolved problems under `OPEN` and resolved problems under
  `CLOSED`, with stable IDs and a resolution or closure reason.
- MUST NOT delete closed problems.
- Technical failures, package limitations, and implementation details belong
  in `LOG.md`, not `PROBLEMS.md`.

### §C1.4 `LOG.md`

- Agent lab book and persistent memory.
- MUST read at session start.
- MUST record every experiment, simulation, investigation, failure, and
  dead end.
- MUST keep newest entries first.
- Technical micro-steps, implementation problems, and exploratory memory
  belong here.
- SHOULD use `research-log` to add entries instead of hand-formatting them.

### §C1.5 `LITERATURE.md`

- Annotated bibliography for `literature/*.pdf`.
- A paper counts as read only after it has an entry.
- MUST locate original source before citing when doing online literature
  work.
- MUST NOT cite from memory.
- SHOULD use `research-lit` to find PDFs missing entries and create stubs
  before filling verified bibliography details.

### §C1.6 Folders

- `data/`: store simulation or experimental output; use `research-run` for
  local command runs when practical so outputs and `LOG.md` stay linked.
- `code/`: store independent codebases or `uv` projects; prefer one
  subfolder per experiment and reference it from `LOG.md`.
- `draft/`: agent-written research digest for human review; MAY contain
  explanatory prose, plots, caveats, and tested, noteworthy results.
- `paper/`: final-paper workspace; edit only on explicit human request.
  Prose in final papers MUST be human-written, so agents MUST NOT write
  fully formulated paper paragraphs.
- `literature/`: store downloaded paper PDFs; every read file needs one
  `LITERATURE.md` entry.

## §C2. Session Warmup

- At session start, MUST run `research-warmup`.
- MUST read `SECDEV.md`, `CONTRACT.md`, `PLAN.md`, `TASKS.md`, `PROBLEMS.md`,
  and `LOG.md`; the helper checks their presence but does not replace reading
  them.
- SHOULD run `research-audit` when starting substantial work or before
  returning final results.

## §C3. Guided Work Loop

For the user-requested job:

1. Read current project state, tasks, and known problems.
2. Plan the next concrete step.
3. Execute.
4. Evaluate evidence.
5. Document technical work in `LOG.md`, maintain `TASKS.md` and `PROBLEMS.md`,
   and, when noteworthy, update `draft/`.
6. Continue until the requested job is fully achieved or genuinely blocked.
7. Return to the user with the outcome, evidence, and remaining risks.

Loop rules:

- MUST continue after intermediate steps when useful work remains within the
  requested job.
- MUST return to the user after the requested job is complete.
- MUST return to the user when progress requires a decision that cannot be
  derived from `PLAN.md`, `TASKS.md`, `PROBLEMS.md`, evidence, or a safe
  default.
- MUST NOT use Mattermost for progress reporting under this contract.

## §C4. Planning Before Execution

- For any non-trivial task:
  - MUST plan before editing, running simulations, or modifying drafts.
  - MUST base plan on `PLAN.md`, `TASKS.md`, `PROBLEMS.md`, and `LOG.md`.
- Trivial tasks are limited to mechanical inspection or status checks that do
  not change project state.

### §C4.1 Git Discipline

- Unless the user explicitly requests another workflow, SHOULD develop on the
  repository's default branch, normally `master` or `main`, and SHOULD NOT
  create additional branches.
- SHOULD commit autonomously at reasonable, coherent checkpoints with
  explanatory commit messages.
- If undoing committed changes is necessary, SHOULD use `git revert` rather
  than rewriting history.

## §C5. Token Budget

- MUST use `research-warmup` during warmup.
- SHOULD check `usage` before token-heavy steps, after experiments finish,
  and when nearing limits.
- If token budget prevents safe completion, document the state in `LOG.md`
  with `research-log` and return to the user with the next concrete action.

## §C6. Completion And Blocking

- A job is complete only when the requested deliverable is produced, checked,
  and documented.
- If blocked, MUST document:
  - what was attempted,
  - the evidence for the block,
  - the smallest user decision or external change needed to continue.
- SHOULD use `research-log --blocked` for blocked work.
- MUST NOT treat a failed experiment as a scientific conclusion until the
  implementation has been checked.

## §C7. Integrity

### §C7.1 Never break a promise.

- If you say "I will do X", do it.
- Under-promise and over-deliver.
- If work is deferred, MUST record reason in `LOG.md`, preferably with
  `research-log`.

### §C7.2 Never manipulate evaluation.

- MUST NOT change metrics, test sets, problem definitions, or fixed
  hyperparameters to make results look better.
- MUST NOT hard-code results.
- MUST NOT cherry-pick seeds.

### §C7.3 Never fabricate citations.

- MUST verify every bibliography entry against the actual source.
- MUST confirm exact title, full author list, year, venue, and DOI/arXiv or
  other identifier.
- If the paper cannot be found, do not guess.

## §C8. Efficiency

### §C8.1 Make it work before moving on.

- Treat experiment crashes as bugs, not as evidence against the method.
- MUST NOT discard methods because of implementation failures.
- If an experiment crashes, investigate, fix, and rerun when still relevant
  to the requested job.

### §C8.2 Use compute intelligently.

- At session start, SHOULD check local CPU/memory capability.
- IF `SECDEV.md` documents specialized compute or remote resources suitable
  for the task, SHOULD use them according to that documentation.
- IF starting a local simulation:
  - MUST estimate expected runtime before launch.
  - SHOULD use `research-run` to capture outputs, timeout metadata, and the
    matching `LOG.md` entry.
  - MUST use a reasonable timeout when the runtime has not been tested
    before.
  - MUST run without a timeout only after prior tests give a good runtime
    estimate.
  - MUST schedule and perform a status check at that expected runtime.
  - MUST investigate simulations that run much longer than expected.
- Remote jobs that can be actively monitored, such as Slurm jobs, are exempt
  from the local timeout requirement.
- For embarrassingly parallel work, SHOULD use multiple worker processes when
  it materially reduces wall time and MUST leave CPU headroom.

## §C9. Scientific Rigor

### §C9.1 One variable per experiment.

- Change exactly one thing per experiment.
- If two things change and the metric improves, you cannot know which helped.

### §C9.2 Evaluate in tiers.

- Tier 1: seconds, does it run?
- Tier 2: minutes, signal on small subset?
- Tier 3: full evaluation for reportable claims.
- MUST NOT draw conclusions from small-scale bug-catching runs.

### §C9.3 Bound your expectations.

- Before implementing a heuristic, identify theoretical best case and
  estimate maximum possible improvement or correction.

### §C9.4 Never overstate simplified results.

- MUST state caveats for toy models, approximations, restricted parameter
  regimes, finite-size systems, relaxed assumptions, and surrogate metrics.
- Restricted contradiction is evidence, not full disproof.
- Keep caveats visible in `LOG.md`, plots, user reports, and `draft/`.
  If the caveat affects a human-selected paper result, include it in
  `paper/` as a terse note or figure/caption constraint.

## §C10. Documentation And Reproducibility

### §C10.1 Record everything.

- MUST log every experiment with goal, method, substeps, outcome, and next
  step, preferably with `research-log` or `research-run`.
- MUST include failures.
- MUST document tested, noteworthy results in `draft/main.tex`.
- SHOULD use plots when available.
- MUST use tables only when clearly better than plots for the message.
- MUST keep abbreviation list current.
- Substeps, dead ends, and later-invalidated results belong in `LOG.md`.
- If it is not in `LOG.md`, it did not happen.

### §C10.2 Verify before claiming.

- Assume you are wrong until verified.
- MUST write verification scripts, not just explanations.
- MUST actively try to falsify claims.
- MUST grade claims as verified, partially verified, or unverified.

### §C10.3 Draft And Paper Roles.

- `draft/` is the agent-human interface. Agents MAY write clear prose,
  explanations, caveats, plots, and result summaries there.
- `paper/` is the final-paper workspace. Agents MUST edit it only when the
  human explicitly asks.
- Agents MUST NOT convert `draft/` wholesale into `paper/`.
- Normal flow is `LOG.md`, `data/`, and `code/` to `draft/`, then
  human-selected content to `paper/`.
- In `paper/`, agents MAY add paper-grade figures, equations, labels,
  bibliography wiring, provenance notes, and keyword-style placeholders.
- In `paper/`, agents MUST NOT write fully formulated prose paragraphs; prose
  may only be inserted when supplied by the human.

### §C10.4 Plot And Figure Quality.

- Figure rules apply to both `draft/` and `paper/`.
- Every figure MUST support one clear claim and be referenced from nearby text
  or notes.
- Multi-panel labels such as `(a)`, `(b)`, `(c)` MUST use the same font, size,
  style, and relative position; corresponding labels SHOULD align across
  panels.
- Text inside figures MUST be legible at final rendered size, with consistent
  font sizes for labels, ticks, legends, annotations, and panel labels.
- Labels, legends, annotations, arrows, and panel marks MUST NOT overlap data,
  axes, colorbars, or other informative content.
- Axes and colorbars MUST name quantities and units where applicable.
- Related figures MUST use consistent variable names, colors, markers, line
  styles, limits, and normalization when comparing the same quantities.
- Prefer vector formats for plots when practical; avoid decorative styling.
- Before marking a figure ready, compile or render it and inspect final-size
  output for alignment, readability, overlap, missing labels, and consistency
  with captions or notes.
