Local LLMs: why companies are building their own AI servers.

Two in three organisations have pulled AI workloads out of the public cloud in the past year. What holds them back is not budget. It is having to document where the data travels before they can switch anything on.

7 min read●Trend●

For years the default answer was "in the cloud". OpenAI, Azure, Bedrock. You pay per token, forget about hardware, it scales on its own. That reasoning has broken in a place nobody had on the spreadsheet: 95% of companies have delayed or cancelled an AI project over data governance, compliance or regulation. Not budget, not talent: knowing what data you hold, where it comes from and who is allowed to touch it.

66%
Have moved AI workloads
off the public cloud
95%
Have stalled a project
over data governance
+53%
Global spend on
AI infrastructure

The first two figures come from The Great AI Re-Architecture, a Cloudera survey run by Wakefield Research among 1,500 architects at companies with over 1,000 employees, published on 11 August 2026. The spending figure is IDC's 2026 forecast: 487 billion dollars on AI infrastructure.

Why now

Legal has learnt to ask

"Where do the American providers process my data?" is the question nobody wanted to ask and everybody now asks, usually the week before an audit.

The underlying problem is not the provider, it is the journey. Moving personal data out of the European Economic Area means holding up the file for an international transfer: assessing the destination country, documenting safeguards, informing the data subject. Processing in your own data centre saves you that file. Mind you, it does not save you the rest of the GDPR, and if the provider of your European hall is a subsidiary of a US parent, the conversation about data access is still open.

The cloud bill does not warn you, it arrives

Paying per token looked cheap until somebody looked at the first month with the pilot already in production. A server has a fixed cost, so past a certain inference volume the maths turns around. Nobody knows in advance where that point sits: it depends on your volume, your model and whether the machine does anything else the rest of the day. Worth working out before you buy, not after.

To be clear, because this industry leans towards absolutes: this does not mean everybody is leaving the cloud. The same survey records organisations that will spend more on cloud and organisations that will spend more on-premise, at the same time. What is dying is the one-size answer.

Latency: you remove the journey, not the thinking

Worth being precise here, because two things get mixed up. Running locally does not make the model reason any faster; it removes the round trip to a data centre that may sit on another continent. If the AI is embedded in a process that is already slow, that journey is the only thing you can cut without touching the model. If your problem is that the model takes its time thinking, moving the machine will not save you.

You no longer need the biggest model

Llama, Mistral, Qwen and Phi publish small versions precisely for this, and for an internal chatbot, classifying documents or summarising minutes, those versions do the job. What is taking hold is the split: the small model takes the bulk of the requests and only the awkward ones escalate to a large one.

What hardware you need

Memory is what rules, and it comes from the model format itself. Quantised to 4 bits it takes around half a gigabyte per billion parameters; in FP16, about two gigabytes. Add the context, which grows with conversation length, and you have your floor.

ModelWeights (4-bit)Recommended total RAMWhat it is for
7B~4 GB16 GB on CPUInternal chatbot, classification, summarising
13B~8 GB32 GB on CPU, or a consumer GPUThe above with more nuance and longer context
34B~20 GB24 GB VRAMReasoning, code, document analysis
70B~40 GB48 GB VRAM or several cardsWhatever a small model cannot solve

The weights come from multiplying parameters by bits. The recommended RAM is higher because it also has to hold the context, the operating system and enough headroom that the server is not running at its limit. And it still does not account for how many concurrent requests you will serve, which is the part almost nobody works out.

And there is the expensive mistake. We have seen a data centre card bought to run a model that summarises meeting minutes: it works, in the same way a lorry works for fetching the bread. Model size tells you whether it fits in the machine. Concurrent requests tell you whether it runs.

What usually gets miscalculated

The same machine that serves a 7B comfortably to three users crawls with thirty, because what saturates is memory bandwidth, not cores. If you are going to size for one thing, size for peak concurrency. The demo with one user always goes well.

Of the remaining parts, the most underestimated is storage. The weights load into memory in full at start-up, so a slow disk is paid for on every restart and every model swap. If you are going to handle several models or large datasets, that is where a lone NVMe stops being enough and a distributed system comes in: we compared them in Storage Scale versus Ceph for inference.

That it can be done without a GPU is not theory. We have built it on IBM Power with vLLM, on AIX with llama.cpp and even on IBM i through PASE, which was the least likely candidate of the three.

Frequently asked questions

How much RAM does a local language model need?

For the weights, half a gigabyte per billion parameters if it is quantised to 4 bits. The rule of thumb is to ask for double what they take: a 7B runs comfortably on 16 GB and a 13B on 32 GB. That difference is the context, the operating system and the headroom that keeps the machine off its limit, which is where the surprises start.

Do I need a GPU to run AI locally?

Not always, but let us be clear about what that means. Up to 13 billion parameters runs on CPU with enough memory, and it is fine for batch processes, overnight jobs or a handful of users. In a chat with people waiting in front of a screen, the CPU shows. A GPU stops being optional when concurrent requests or model size go up. We measured it without a graphics card on vLLM on IBM Power.

Is your own server cheaper than paying per token?

Past a certain volume, yes: the server has a fixed cost and the API grows with usage. Exactly where that point sits depends on the model, the volume and whether the machine does anything else the rest of the day. Below that volume, the API still wins.

Is it legal to process personal data with a cloud LLM?

It can be, but it requires a legal basis, an impact assessment and safeguards on international transfers when the provider processes outside the European Economic Area. Processing at home saves you that file, not the rest of the GDPR. And if the provider of your European hall is a subsidiary of a US parent, the conversation about data access is still open.

Which models can run on a business server?

The open families — Llama, Mistral, Qwen, Gemma, Phi — publish versions in several sizes precisely for this. For internal chatbots, document classification or summarising, the small ones do the job without specialised hardware.

The next step

Before looking at catalogues you have to size it, and that is three questions: which model you will run, how many concurrent requests, and what response time works for you. Memory, CPU, GPU and storage follow from that, in that order and no other.

That is what we built a server configurator for: you pick the use case, the scale and your priorities, and watch the machine take shape in front of you. No part numbers and no imposed brand. When you finish we call you with a firm proposal.


Configure it yourself

How much server does your LLM need?

Pick the use case, the scale and your priorities, and watch the machine take shape. No commitment and no endless forms.