“Can we run this on our own servers?” is one of the first questions we hear about AI automation, and it is almost always asked for a good reason: client contracts, data protection rules, a factory network with no route to the internet, or simply a reluctance to send every invoice and email to a third party.
The honest answer in 2026 is yes: self-hosted AI automation is far more practical than it was two years ago. But self-hosting is usually a decision about control, not cost. This guide sets out the three levels of self-hosting, what the open models can now do, what a GPU actually costs per month, and the work you take on when the model is yours to run.
Three Levels of Self-Hosting
“Self-hosted” gets used for three quite different arrangements. Knowing which one you actually need often removes most of the cost.
| Level | Where the model runs | Who can see the data | Effort |
|---|---|---|---|
| 1. Public model API | The model provider’s infrastructure | The provider, under its terms | Lowest |
| 2. Model in your cloud tenant | A hosted model from a major cloud provider, in a region you choose, under your enterprise agreement | Stays within your cloud account and chosen region | Low to moderate |
| 3. Open-weight model on your infrastructure | GPUs you rent or own, in the cloud or on premises | Nobody outside your organisation | Highest |
Level 2 is the option many teams skip past, and it satisfies a large share of data-residency requirements on its own. Level 3 is the only one that works with no outside connection at all.
One common misunderstanding is worth clearing up. Self-hosting the workflow engine is not the same as self-hosting the model. A self-hosted n8n instance that calls a public model API at every step still sends the data out. Where the orchestration runs is covered in n8n vs Make vs Zapier; this guide is about where the model runs.
Why Companies Choose It
What Open Models Can Now Do
The case for self-hosting changed when capable open-weight models became small enough to run on a single server. OpenAI’s gpt-oss-120b, released in August 2025 under the Apache 2.0 licence, runs on a single 80 GB GPU; its smaller sibling, gpt-oss-20b, fits within 16 GB. Other open families, including Llama, Mistral and Qwen, cover a wide range of sizes.
For the work most automation actually does — extracting fields, classifying and routing, summarising, drafting replies against a template — models of this class are generally good enough. For long, multi-step reasoning across many tools, the leading hosted models still tend to be stronger, though the gap is narrower than it was.
Do not choose a model from a leaderboard. Run your own two or three hundred labelled cases through each candidate and compare the results on the work you actually do.
What It Actually Costs
GPU rental prices have fallen sharply. In early October 2026 the median on-demand price for an NVIDIA H100 was about $3.60 per GPU-hour across roughly 40 providers, with marketplace providers between about $1.50 and $3 and the large hyperscalers between about $7 and $11.
- 730 hours in a month at the median of about $3.63 — about $2,650 a month
- The same GPU at a hyperscaler’s on-demand rate — about $5,000 a month
- A second GPU so the service survives a failure — double either figure
- Engineering time to keep it patched, monitored and evaluated — budget at least two to four days a month
Take a team processing 20,000 documents a month at about 4,000 tokens each, input and output combined: 80 million tokens. At an assumed blended API price of $2 per million tokens, that is about $160 a month. Check current rates; they vary widely by model.
One median-priced GPU at $2,650 a month equals about 1.3 billion tokens at $2 per million — roughly 330,000 documents a month of the size above. Add a second GPU for resilience and some engineering time, and the break-even point more than doubles. Cheaper API models push it higher still.
The conclusion is not that self-hosting is too expensive. It is that for most organisations the GPU bill is the price of control, and it should be justified on control. If the data can leave and the volume is ordinary, an API or a model in your own cloud tenant will almost always cost less. The wider cost picture is set out in what AI automation costs.
| Your situation | Sensible starting point |
|---|---|
| Ordinary business data, normal volume | Public API with a business agreement that excludes training on your data |
| Data must stay in a region or inside your cloud account | Hosted model in your own cloud tenant |
| Data must never reach a third party | Open-weight model on infrastructure you control |
| No internet connection available | Open-weight model on premises, sized for the site |
| Very high, steady volume | Run the break-even calculation; self-hosting may win on cost alone |
The Work You Take On
With an API, someone else keeps the model running. Self-hosting moves that work in-house, and it is easy to underestimate.
A Sensible Path to Self-Hosting
Frequently Asked Questions
Conclusion
Self-hosted AI automation is now a realistic option rather than a research project: capable open models run on a single GPU, and rental prices have fallen. What has not changed is that the GPU bill is mostly the price of control. Classify the data, choose the lowest level of self-hosting that satisfies it, and justify anything beyond that on the break-even arithmetic.
Self-hosted and data-sovereign deployments are a core part of how we build. If your data cannot leave, our AI automation services include designing and running the model inside your own boundary, with the same evaluation and monitoring as any other production system.
