Self-hosted AI or cloud? Five questions that decide where your agent runs
Self-hosted AI or cloud? Five questions decide it: data, peak loads, volume, upkeep and model updates. Worked through for a freight forwarder.
This article was generated by AI. Labelled in accordance with Article 50 of the EU AI Act. Responsible for publication: Sophera Consulting.
"So does the AI run on our premises or in the cloud?" The question usually carries a worry with two sides. Either customer data ends up with a large provider overseas, or somebody has to buy a server and keep it alive. Neither has to happen. Whether to run AI self-hosted or in the cloud is a decision you make per process, not once for the whole company, and five questions settle it. You can answer all five better than any vendor can.
Self-hosted AI or cloud: there are really three places
For an AI agent, "cloud" means a model provider's API. The agent sends text, gets an answer back, and you pay per call. There is no hardware and no base fee. Whether that works under data protection rules depends on where the provider processes the data and what contract it signs. We use European infrastructure, and the data processing agreement comes with the project.
Self-hosted splits into two options. With a rented server, the model runs on dedicated hardware in a European data centre, used by nobody but you, for a fixed monthly rent. With your own hardware, the machine sits in your server room. The data never leaves your network, and in exchange you pay for the purchase, the electricity and your IT team's time.
The agent itself is identical in all three cases. It reads the same systems, applies the same matching rules, puts the same cases on the exception list and hands the same drafts to a person for approval. The only thing that changes is where the language model does its computing.
Two agents at a freight forwarder
The following example is made up. It is here to show the trade-offs on something concrete.
Say a haulage company with 40 trucks wants to automate two processes. The first is order entry: customers email transport orders, the agent reads them and creates a draft in the transport management system, which a dispatcher approves. The second is checking driver pay. The agent compares tachograph data, expense receipts and trip reports and flags anything that doesn't add up before payroll runs.
Both agents go through the same five questions.
1. What is in the data?
An order email contains company names, loading addresses, pallet counts, weights and usually a contact person's name. That is business data with little personal information in it. An API on European infrastructure, covered by a data processing agreement, is the obvious choice.
Driver pay data contains individual employees' driving and rest times, expenses and absences. Whether that may leave the building is something you decide with your data protection officer and, if you have one, the works council. If the answer is "it stays here", this agent runs self-hosted. We settle that together in the Automation Check before anything is built, so it adds nothing to the build itself.
2. When does the work arrive?
Price comparisons leave this question out. Day-to-day operations notice it first. Transport orders do not arrive evenly. Suppose 150 emails come in on Monday between seven and nine, and another 100 trickle in over the rest of the day.
Through an API, those 150 emails are processed in parallel, and a few minutes later every draft is waiting for dispatch. Your own machine can only do as much as the capacity you bought with it. If it needs 20 seconds per email and handles four at once, the last email of the Monday rush waits a little over twelve minutes. For order intake that is fine. If you want the last order ready within two minutes, you need roughly six times the capacity, and it will sit idle for the rest of the week.
Driver pay, on the other hand, runs overnight before payroll, and nobody is waiting on it. A small machine takes a bit longer and nobody notices.
3. How much work adds up overall?
An API costs per call, a server costs a fixed amount. Below a certain daily volume the API is cheaper, above it the fixed machine wins. We worked through where that point sits, with numbers, in our article on hospital processes and the three places a model can run. On the assumptions used there, 250 order emails a day do not pay for a server on their own.
If driver pay already runs on a dedicated machine, though, order entry can share it, as long as it copes with Monday morning. That is why we look at all of a company's processes together rather than one at a time.
4. Who looks after the machine?
With an API, nobody on your side. With a rented server, the data centre handles the hardware. With your own hardware, it is all yours: the failed power supply, the driver update that is due, the server room that gets too warm in July.
In all three setups, operations are watched by the maintenance agent that comes with every automation we build. It checks the input and output of each interface for changes and sends anything it cannot fix itself to your IT as a ticket. If you have no IT team that can pick up a ticket the same day, an API or a rented server is the better fit.
5. Who decides when the model changes?
Model providers keep improving their models and retire older versions at some point. Through an API you get improvements without lifting a finger, along with changes you never asked for. Self-hosted, the model that was installed keeps running until you replace it. Either way, the maintenance agent tests a new model against past cases and handles the switch.
Then there is choice. Not every model can be brought onto your own hardware, and some of the most capable ones are only available through their provider's API. Whether a given model is good enough for your process is something a test with real cases shows.
What the example company ends up with
A split answer. Order entry runs through an API on European infrastructure: the data has little personal content, Monday brings peaks, and there are no fixed costs. Driver pay runs on a rented server or in-house, because the company wants employee data to stay with it and an overnight job gets by on modest capacity.
Moving later does not mean rebuilding. If order volume grows to the point where a fixed machine pays off, the model moves. Rules, mappings and the exception list stay as they are.
Sophera Consulting builds agents for all three setups, separately for each company, at a fixed price and without a subscription. An agent is set up in one to two days, and testing with real cases happens the same week. If you prefer, the models run locally on your hardware and the data never leaves the building. Ongoing, you pay for model usage or for the machine the model runs on, with no operations fee on top. Which place suits which of your processes is what we work out in the free Automation Check.
Our recommendation
Start with an API when a process handles business data and nobody in-house should be looking after a server. Plan for self-hosting when the data contains something that must not leave the building, or when enough work arrives evenly across the day to keep a fixed machine busy. And decide process by process. A company that puts everything in the cloud, or runs everything in its own server room, ends up with the weaker option for part of its work.
This article was created with the help of AI.