Institute for Theoretical Physics III · University of Stuttgart

AI agents at work in a physics institute

Agentic AI is reshaping society. This blog-like website documents several AI-related projects at the Institute for Theoretical Physics III of the University of Stuttgart.

At the Institute for Theoretical Physics III of the University of Stuttgart we put AI agents to work on research, on the institute’s servers, on its web pages and on its paperwork. The goal of all these projects is to explore where agentic AI can be used beneficially in our everyday work. This site documents several of those projects. It is written for readers who are comfortable with technology but do not work here, so each project gets an overview that explains the idea, a page that shows the workflow step by step, and a separate page for the implementation details. The following chart sorts the projects by the part of university life they serve; each chip opens the project’s page:

TeachingsecdevReview skillResearchsecdevResearchIT InfrastructuresecdevWingmanOpenCMSEvolving SkillsAdministrationOfficeEvolving SkillsUniversityAI Agents

One word does a lot of work on this site, so it is worth pinning down first. An agent is a large language model (LLM) that has been given tools and a task and then left to get on with it: it reads files, writes files, runs commands, looks at what happened, and tries again. It is not a chat window. The difference matters because a chat window cannot delete your work, and an agent can. One sentence sums up why that pairing of language and tools matters:

There are two capabilities that made homo sapiens conquer the world: the ability to use tools, and the mastering of language. We equipped computers with both.

If you read one thing, read the research example : a physics paper whose simulations, calculations and figures were produced by agents working under a written contract, a file of rules the agent must obey, with the people in charge writing the plan and the text.

The projects

The idea behind every project

An agent is most useful when it stops asking permission for every step, which is exactly when it is most dangerous. Every project here is a variation on the same resolution: do not constrain the agent — constrain its world. The following figure shows how the projects relate:

How the projects relate. All but Office stand on the same sandbox, and every one of them keeps its credentials out of the agent’s reach. Evolving Skills serves the projects from the other side: it keeps the OpenCMS skill current today (a skill is a set of written instructions that teaches an agent a task, here editing the university website), and the dashed arrows mark where it applies next.

Three rules run through all of the projects, including where they are inconvenient:

  1. Run the model where the data allows, and no further

    For most research code, a model operated by a company elsewhere is fine. For the institute’s administrative material, university policy says otherwise, and Office is the consequence: every part of it, the model included, runs on the institute’s own machines. The data decides, not convenience.

  2. The agent never holds the credential

    The sandbox holds no key to the Git server where the code is stored, so an agent cannot publish its changes. Wingman keeps the key to the institute’s machines outside the agent’s reach entirely. OpenCMS gives an agent a borrowed browser session, never a password. In each case the agent has a way to ask, and holds nothing worth stealing.

  3. Restrictions live in the tools, not the instructions

    A mail service that cannot send mail is a different proposition from a mail service that has been asked not to. Written instructions erode: they get crowded out as a conversation grows, and a model reading an untrusted document can be argued out of them by the document. A capability the tool does not have cannot be talked into existence.

The three rules are shared; the projects that follow them are not alike. One is a sandbox, one a way of doing research, two are agents that operate institute systems, one is an assistant for the office, and one is a procedure for maintaining the others. What they do have in common, besides the rules, is a person who stays in charge, and each project places that person at a different point of the work:

  • secdev , a disposable container in which a coding agent works alone on one project: a person reads the agent’s work before it leaves the machine.
  • Research , a method for letting an agent run calculations and simulations for days under a written contract: a person owns the goal file the agent may not rewrite.
  • Wingman , an agent that administers the institute’s servers: a person approves each command before it reaches a server.
  • Office , an assistant for the administrative staff: a person approves everything that leaves the agent’s workspace, whether a message to a colleague, a file copied into the office archive or a print job, and the agent has no route of its own to a mailbox, a chat or the archive.
  • OpenCMS , an agent that edits the university’s website: a person approves every change before it goes live.
  • Evolving Skills , a procedure for keeping an agent’s written instructions correct: a person reviews every report before it is sent and decides which fix ships.

When to use an agent

An agent will attempt almost any task and return a plausible result, whether or not that result is correct. The question is therefore not what an agent can do, but whether you can check what it did. The two lists that follow turn that question into a test: the first names the conditions under which an agent pays off, the second the cases in which it must not be used:

Use an agent when

  • verifying the result is cheaper than producing it yourself
  • errors would be caught by tests, plots, proofs or review
  • the rules for that kind of data (personal, confidential, public) permit the tool you would use
  • authorship and responsibility stay unambiguous

Do not use one when

  • the task is meant to assess the person’s own ability, as in an exam
  • it would generate text for a thesis or a paper
  • confidentiality or consent is unclear
  • you could not tell whether the output is correct

The first condition in the first list is the whole test: use an agent only when verifying its result is cheaper than producing the result yourself. Everything an agent produces has to be checked by someone who could have produced it themselves. Where checking costs as much as doing, the agent has saved nothing and added a new way to be wrong. An institute seminar put the rule in two lines:

Bad: hidden AI delegation.
Good: AI proposes, researcher verifies, documents, decides.
Institute seminar, August 2026

The sandboxes, approval steps and restricted tools described on this site do not make a task suitable for an agent. They only make running an agent safer once the verification test has said that the task is suitable.

Where the code is

Several of the projects are published for university members on the University of Stuttgart’s GitHub Enterprise: secdev, itp3wingman, the generic OpenCms skill and the report tracker behind Evolving Skills. All of them are licensed under AGPL-3.0-or-later. Office is not published: it is built around this institute’s servers, forms and staff, and would not run anywhere else as it is.