The Workflow page showed the system from the staff member’s chair, and the overview set the system as three pillars: intelligence, tools and guidance. This page takes the pillars apart in that order, and the diagram from the overview is the map of what follows:
The intelligence is the language model and the servers it runs on. The tools are the harness, the per-user agent containers and the tool services, which is where most of the engineering and all of the safety properties live. The guidance is the library of skills and the contract every session receives. Two closing sections list the repositories and hosts, and where each part stands.
Three design rules of the stack
Every design decision below follows from three rules, stated here once:
One small container per capability
Each tool is a separate service with a narrow interface, its own tests and a minimal privilege set, and every part of the stack, tools included, runs in a container described by a Compose file (the text file from which Docker or Podman Compose starts a set of containers), and recurring work runs as scheduled jobs. Replacing one service, or removing one, does not touch the others, and nothing is installed on a host itself, so a service can move from one bare-metal server to another by copying its Compose file and its data. Nothing needs a dedicated operator.
Restrictions live in the service layer
Not in the prompt, not in the model, not in the agent’s good judgement. The agent never holds a credential for anything but the model router, the gateway in front of the model, and the tool services themselves.
Everything the agent works from is a file that staff can edit
The institute’s master data, its address, accounts and contacts, is a plain text file on the network share that the office maintains. The room calendars are a text file too. The skills are text, and the forms are files in a catalogue with a review step. None of this lives in a database that only a developer can change: when a value changes, a staff member changes the file.
The layers
Before the pillars, the stack as it is deployed, from the bottom up. The following diagram shows the layers and the two paths that leave the agent runtime, one down to the model and one sideways to the tools:
The two bottom layers, model serving and routing, are the intelligence pillar. The agent runtime, the tool services and the staff interface (the web page plus the middleware behind it) are the tools pillar. The skills with the form catalogue, and the contract, are the guidance pillar, which the agent runtime reads rather than calls.
Intelligence: the model
The model is the one part of the system that is bought rather than built, and the data-protection rule explained on the overview decides which models qualify: only a model that runs on the institute’s own hardware. The intelligence pillar is therefore two things, the model and the servers that serve it.
The model. A mid-sized open-weight model, one whose trained parameters are published and can be run on one’s own hardware, currently Qwen3.8-27B in 8-bit floating point. It is served with a context window, the amount of text the model can hold at once, of 262,144 tokens, which a long case with several documents needs. Multi-token prediction, a trick in which the model drafts several words at once and checks them, which speeds up short exchanges, is switched off, because it costs speed on the long sessions the office actually runs. The model sees text and produces text; it has no vision, no file access and no network. Everything it knows about a case arrives as text through the harness, and everything it does leaves as a tool call that the harness executes.
Model serving. The model runs on two A100 GPUs with 40 GB of memory each, connected by NVLink, with vLLM as the serving engine. Each of the two cards holds half of every layer of the model, an arrangement called tensor parallelism, and the server is configured for a concurrency of two, so that two agents, one per staff member working at the same time, are served at once. A model router sits in front of the model and issues one virtual key per consumer in place of the model’s own credential, which makes per-user budgets possible and gives one place to change the active model. Changing the model is a change of configuration in the router, not in any agent.
Tools: the harness and the services
The tools pillar is everything that turns text from the model into work on files, mail and forms, and it is where the safety properties live. This section describes its parts in turn:
- the agent runtime, the container that runs the model in a loop;
- the path one request takes through the system;
- the tool services, with what each may and may not do;
- the two features that let a file or a message leave the agent’s workspace, each only on a person’s click, and the printer, which has no button; the skill tells the agent to ask before printing.
The agent runtime
Each staff member has one agent of their own: a rootless container that knows whom it works for, and that can write documents to exactly one place, that person’s workspace. The office archive and the university documents are mounted read-only, and of the mail archive only two mailboxes are mounted, the shared office mailbox and the person’s own. The agent holds no password to anything. It has a key for the model router and tokens for the tool services, and the tokens for mail, search and display are its own and grant only what that person may see. The agent’s built-in web tools are disabled: public research goes through the privacy-gated web-search service.
The office agent does not run in the secdev sandbox of the other projects but in a container of its own, built for this stack. The two share one idea, a document toolchain: the programs a skill may need to read, convert, fill and stamp documents are installed in the container, and a small skill owns each of those jobs, so that an administrative skill calls the right one instead of rebuilding a document by hand. The container’s briefing says so:
A silo is one of the three indexed archives described under Retrieval and ingestion below. The division of labour matters, because it is what keeps a small local model reliable. The toolchain does the mechanical work, and does it the same way every time: stamping a PDF, for example, is a program that finds a free spot on the page, writes the stamp and leaves the original bytes of the document intact. The administrative skill holds the judgement: the invoice-list skill decides which invoice gets which number, and then calls the stamping program. The model never has to produce a document character by character, only to decide what goes where and to call the right program. Nor does it have to check the program’s work: a program that has verified its own output says so in its reply, and the contract forbids the agent to verify the result a second time by rendering it or reading it back, because a model without vision spends longer confirming a deterministic result than producing it.
One request, end to end
What the staff member sees as one conversation is a chain of components. The sequence below follows a single request through all of them:
Three properties of this path are deliberate.
Files do not travel through the model. Documents move between services and into the workspace directly. The model sees what it needs to reason about, not every byte of every attachment.
The staff interface decides who you are, not what you may do. The staff interface checks the login and sends each person to their own container: one login is one agent, one workspace and one set of mailboxes. What that container may touch is settled underneath it, by which folders are mounted writable and which tool services exist at all, so a flaw in the staff interface cannot widen anyone’s reach. The interface adds no capability the agent lacks. The things the interface does on its own, uploading a file, deleting a case, filing a copy in the office archive and sending a message to a colleague, are things the staff member could already do on the network drive or in the chat client. Filing and sending happen only on a click by the case owner, the staff member who started the case, through a route the agent cannot call.
The staff interface is a translator, not the agent. The browser speaks only the words of administration: case, message, question, result, file. One layer between the browser and the agent runtime translates those into the runtime’s own protocol, and that layer is the only part of the system that knows which runtime is in use. Replacing the runtime one day changes that layer and nothing the staff member sees.
The tool services
Each service is a separate container exposing one MCP endpoint, the Model
Context Protocol being the standard interface through which an agent calls a
tool. The services are developed together in one repository, itp3mcp, the
institute’s collection of MCP services, which also holds the search index
and the jobs that feed it. All but the last group of the table below live
there; the tools in the last group are offered by the staff interface
itself. The services, each with its tools and what they do:
| Service | Tool | What it does |
|---|---|---|
| Master data | list_topics, get, search | Read the institute’s reference data, address, accounts, contacts and defaults, from the file the office maintains |
dump | Return the whole reference data set at once | |
list_active_members | List the institute’s current members from the member directory | |
| Search | search, search_office, search_university | Search the three archives by meaning and by words, one archive at a time |
read_indexed_document, get_indexed_document, get_indexed_mail | Return the stored text of an archived document or mail, so the agent reads the extracted text instead of converting the file again | |
recent_files | List the files the person opened or saved last on the network share, from the file server’s access log, for requests such as “the spreadsheet I have open” | |
list_mail_accounts, list_folders, list_mail, read_mail | Browse and read the mailboxes the person may see | |
save_attachment | Save an attachment unchanged into the case folder | |
create_draft, update_draft | Put a mail into the drafts folder, or update one; there is no tool that sends | |
| Forms | search_forms, get_form_schema | Find a form in the catalogue and fetch its current fields |
fill_form, download_form | Fill a form with given values, or fetch the blank original, as a short-lived download | |
| Calendar | list_calendars, get_events, search_events | Read the campus calendars |
check_availability | Say whether a room is free in a given interval | |
| Web search | search_web, read_web_page, research_web | Search the public web and read a public page through a privacy gate that keeps the original question inside |
| Printer | print_file | Print a PDF from a case folder, with copies, sides and colour |
printer_status, list_jobs | Report the printer and its queue | |
| Messages | mitglieder_suchen | Look up the members of the institute’s chat team; sending and reading are reserved for the staff interface |
| Staff interface | dokument_anzeigen, webseite_anzeigen | Open a document, or a page of a listed university domain, beside the conversation |
buero_ablage_anfragen | Ask the case owner to copy a file into the office archive | |
mattermost_nachricht_anfragen | Propose a message to a team member for the case owner to send |
The table shows the security model by what it leaves out: no tool sends, deletes or moves a mail, changes the catalogue, books a room, cancels a print job, writes into an archive or reaches another person’s mailbox or workspace, and none of that is enforced by prompting. The printer, which prints any PDF it is named, is the one tool that the skill gates instead, by telling the agent to ask first. Mail is the sharpest case, and it is send-proof three times over: the service exposes no send tool; it speaks only IMAP, the protocol for reading a mailbox, not the one for sending; and its credentials are restricted to IMAP, so even a compromised service could not send. The one channel out of the office is the chat message to a colleague, which the section on the two doors out of the workspace describes, and that channel reaches only members of the institute’s team. The workspace is mounted at the same path everywhere, so a file the agent names never has to pass through the model.
Three of the services deserve a closer look: the search service with the ingestion behind it, the forms service with its catalogue, and the two services that open a door out of the workspace.
Retrieval and ingestion
The search service indexes three archives: the office archive on the network share, a mirror of the internal university website with its handbook and forms, and the mirror of the configured mailboxes. Every document is split into chunks of a few paragraphs, and every chunk is indexed twice, in a vector database (Qdrant): once as a dense vector from an embedding model, which captures what the chunk means, and once as a sparse vector of term frequencies, which captures the words it contains. A query runs against both and merges the two rankings, so a question phrased in different words from the document still finds it while an exact file reference is not buried. The embedding model runs locally on the file server, in a small inference service of its own, and is currently Harrier , a 0.6-billion-parameter model; a change of model means re-embedding the chunks, not re-extracting the documents. The agent may also ask for a word-only search for a single call, which reports when no document contains all the query terms.
Three scheduled jobs on the file server keep the material current:
| Job | Schedule | Does |
|---|---|---|
| Crawler | Daily | Logs in to the internal university website and mirrors handbook pages and linked documents into the university archive |
| Mail mirror | Hourly | Mirrors the configured IMAP accounts into local, searchable bundles, attachments included |
| Indexer | Hourly | Extracts text from new or changed files in all three archives, with optical character recognition (OCR) for scanned PDFs, and updates the index |
The indexer keeps the extracted text of every file, and the office contract requires the agent to read archive documents from that text instead of converting a PDF again.
The form catalogue
University forms are a heterogeneous mess: Word documents in two generations of format, spreadsheets, PDFs with fillable fields, PDFs with nothing but printed boxes, each with its own idea of where a value goes and how it must be written. Filling one correctly is hard even for a strong model, and a wrongly filled form is worse than no form. The forms service exists to take that difficulty away from the agent at the moment it matters. For each form, a frontier model, one of the large hosted models, writes a connector once: a small program that knows this form’s layout and writes given values into the right places. From then on the local model never touches the file format. It asks the forms service for the form’s fields, hands over the values, and receives the filled document, through the same MCP endpoint as every other tool. The connectors are reviewed, versioned and tested like any other code, in a Git catalogue with a review step. One form in the catalogue looks like this:
forms/<slug>/
metadata.yaml identity, description, category
schema.json fields, types, mandatory values
tests/values.example.json
shared example values
tests/cases/ documents that must fill
tests/raises/ values that must be rejected
originals/<n>-<ext>/ one directory per file format:
<source>.<ext> the original
connector.py writes values into this format
pyproject.toml pinned dependencies (with uv.lock)
A form that exists as both a PDF and a Word document has two connectors, one per file format, but a single list of fields, so the agent fills “the form” and never has to know which format it is in.
From the agent’s side the whole catalogue is three tool calls. The sequence below follows one form from the agent’s search to the filled document in the case folder:
Two doors out of the workspace
The rules above leave the agent with one writable place, its own workspace, and no way to send anything. Two features open one door each, and both are built the same way: the agent can only ask, the case owner has the button, and the copying or sending is done by a service the agent cannot reach.
Filing into the office archive. The archive is read-only for the agent. When a finished document belongs there, the agent asks the staff interface to copy it, naming the file, the target folder and a reason, and the request appears as a box in the conversation. When the case owner clicks Ja (yes), a separate writer service, the only component that can write into the archive, makes the copy, provided the target folder exists and the file has not changed since the agent asked. Every copy and every refusal is logged.
Messages to institute members. The agent can look up the members of the institute’s chat team, and nothing more. When it needs an answer from one of them, it proposes a message; the case owner sees recipient, reason and text, may edit the text, and sends or declines. The message goes out from a bot account, with a footer in English and German that names the staff member and asks for an answer in the thread. A reply is shown to the case owner first and reaches the agent only when the case owner releases it, as a member’s information and not as an instruction. Every send is logged.
The message door has two gates, which guard against two different things. The gate on the way out guards against data leaving the office: a message is the one channel through which text from a case can reach a person, so an agent that had been misled into packing confidential content into a message would still need a staff member to read it and press send. The gate on the way back guards against prompt injection, text that reads like an instruction and that a model may follow although nobody authorised it: an answer from the chat service is written by whoever holds that account, so it is shown to the case owner first, and when released it is framed for the agent as information from a member, not as a command. The round trip of one message:
Guidance: the skills
The guidance pillar is text. The agent reads two kinds of it before it starts: first the office contract, a set of general rules that every session receives as part of its system prompt, and then, if the case was started under a tile, the skill for the task, the step-by-step procedure for exactly this kind of job.
The contract, shown in full in the contract section of the
overview
, fixes what
holds for every task: write only in the case
folder, preserve originals, mark anything missing as Fehlt and anything
doubtful as Zur Prüfung, never invent a value, ask when one answer would
settle a point, and close with the five-section report that the staff interface
sets apart. The clauses that let the agent propose a file for the office
archive and a message to an institute member are part of the contract too,
not of any skill. Examples
quote the contract where the
poster walkthrough meets it.
The skills form a library, one directory per skill. A skill is composed of three kinds of file:
SKILL.md, the procedure itself: what the task is, which tool to call at each step, what to take from where, what to check, when to stop and ask, and how to report;- helper scripts for the deterministic parts, such as deriving a file name or rendering a mail from a template;
- assets the procedure needs, such as the reviewed invitation template and recipient lists of the poster skill.
The library currently holds the procedures the start page offers as tiles:
- looking up a rule, a deadline, an address or a responsibility, with the source;
- fetching a form, a template or a document into the case;
- the procurement documentation, filled from quotations;
- the invoice stamp;
- the seminar poster, with its print job and invitation.
Skills are written the way the rest of the stack is written: by an AI agent, following the guidance of experienced administrative staff, who know the procedure, the exceptions and the forms, and who review the result. A reviewed skill reaches staff as a file copy and a restart of their container; its tile, with its German name, icon and one-line description, comes from the staff interface’s own configuration. Improving the agent at a task therefore means editing a procedure, not retraining a model, and the whole library is reviewed text under version control.
Where the skills could go. Today the office skills are improved by hand: the maintainer reads the closing reports and the cases, sees where a procedure left the agent guessing, and rewrites the step. That does not scale beyond a handful of skills, and it depends on the maintainer noticing. The Evolving skills project, built for the OpenCMS skill, offers a loop that scales: the agent that works through a case keeps a friction log, a note of every place where the written procedure was wrong, unclear or silent; at the end of the case the staff member decides whether to file that log as a report; a maintainer, usually an agent guided by a person, works the queue of reports, reproduces each finding and turns it into the next release of the skill. Applied to the office, the staff would not have to describe what went wrong, because the agent already knows at which step it had to improvise, and the skills would improve from every case that is worked, not only from the ones the maintainer happens to look at. The office skills are plain text with a version, which is all the loop requires; putting it in place is planned, not done.
Repositories and hosts
The whole office framework lives in one master repository, itp3office,
which holds no service of its own. It pins the following independently
deployable repositories as Git submodules, one per part of the stack, and
owns what has to agree across them: the cross-project configuration, the
roadmap and the operations runbook for deploying the parts on the
institute’s servers:
| Repository | Owns |
|---|---|
itp3llm | Model serving on the GPU hosts |
itp3ai | Routing and virtual keys, a test chat interface, the search engine behind the web-search service |
itp3mcp | The tool services, hybrid search, the indexer, the crawler and the mail mirror |
itp3sec | The rootless per-user agent containers and the office document toolchain |
itp3skills | The administrative skills |
itp3forms | The reviewed Git source of the form catalogue |
itp3ui | The staff web interface and the middleware that drives the agent runtime |
The stack runs on three institute servers: a GPU server for the model, an admin server for the router and the search engine, and a file server that holds the document archives and runs the agent containers, the tool services and the scheduled jobs. External systems are the institute’s mail server, the office printer, the campus calendar feeds, the member directory, the chat service, and the internal website of the university administration.
Where each part stands
The stack is deployed, but not every part is finished. The following table records the state of each component:
| Component | State |
|---|---|
| Model serving | Live |
| Routing | Live; per-user virtual keys in use; budgets and request logs not yet configured |
| Tool services | All itp3mcp services live; the display, filing and messaging tools ship with the staff interface |
| Filing and messages | Deployed; messages confirmed in live sessions with every staff member, filing with one |
| Search index | Full index over all three archives; selectable retrieval modes in production |
| Ingestion | Crawler, mail mirror and indexer on scheduled timers |
| Agent runtime | Live, one rootless container per staff member; each prompt names the person |
| Per-user tool tokens | In place for mail, search and display; the printer and the remaining services still take a single token shared by all agents |
| Form catalogue | Reviewed catalogue published; every form on its main branch is live |
| Skills | Information retrieval, document retrieval, procurement documentation, invoice stamp (HÜL), seminar poster |
| Staff web interface | Rolled out to every staff instance; the first pilot of daily use is running |
Next steps
Two things are missing, and both come before any further feature.
Evidence of use. The pilot has to show, through the feedback of the staff, which of the applications above are real use cases that save time and which are demonstrations.
Measurement. The system works only as a whole, as the overview says, and a whole with this many parts, model, prompt, skills, tools, index and interface, has too many parameters to tune by feel: a change anywhere can help one task and hurt another without anyone noticing. What is needed is a full-stack benchmark suite that runs real cases through the complete system and reports how often each succeeds, so that every change can be measured against the same cases. Until that suite exists, nothing on this page says how often a task succeeds.