R Rumus Cloud
2026-08-12

What we learned running an agent platform for a year

Twelve months of scheduled jobs, chat delivery and MCP servers, and the four habits that kept it from becoming a liability.

We started running an always-on agent in August last year, mostly to see whether the scheduled-job pattern held up outside a demonstration. It did, with four corrections we had to make along the way.

Every job needs an owner and a place to fail

A scheduled job that reports to a chat channel looks finished the moment it sends a message. In practice the useful discipline was assigning each job a person who notices when it goes quiet. Jobs without an owner were the ones that silently stopped.

Cost per job, not cost per month

Monthly totals tell you nothing about which job is expensive. We log tokens per run, and the month's invoice becomes an aggregation rather than a surprise. One job that summarised a large document set every hour turned out to be most of the bill and none of the value.

Approvals before actions that cannot be undone

Reading is safe to automate. Writing needs a gate. We keep two classes of job: ones that gather and report, and ones that change something, with the second class requiring a confirmation from a named person.

Prompts lie in wait

A job that reads external text will eventually read something written to mislead it. The defence that held up was boring: allowlists for what the job may call, and treating model output as input rather than instruction.

What we would keep