A corporate AI assistant may read contracts, shipment records, customer details and internal reports in a single working day. Before asking how intelligent that assistant is, we should ask a more practical question: where does that information go, and who controls it?
Hosting a large language model (LLM) within infrastructure your organisation controls can make that question easier to answer. It also creates an opportunity to build AI around your business language and share capacity across a group of companies. But ownership brings responsibilities as well as advantages.
What does hosting your own LLM mean?
Self-hosting means running a model on company-managed servers, in a data centre or within a controlled private-cloud environment. Typically, the organisation starts with an existing pretrained model whose licence permits the intended use. It does not need to train a foundation model from scratch.
The company manages the application, inference service that generates answers, supporting data stores and access rules. A rented private-cloud GPU is still a recurring expense; buying on-premises equipment adds an upfront capital investment.
The distinction that matters is operational control. Who administers the systems? Who can access prompts? Where are backups stored? Which services receive data? A private model connected to an external document parser or analytics service may still send information outside the intended boundary.
Data security: control the whole path
Figure 1. A proposed private AI architecture. Keeping data within the boundary depends on the configuration of every component, including document processing, embeddings, logs and backups.
Consider a group with logistics, healthcare and food-service companies. Shared AI infrastructure should not mean shared access to every company's records. A logistics employee should see authorised shipment information; access to patient records should remain restricted to authorised healthcare staff.
For this architecture, I would prioritise the following controls:
- Identity before inference: require corporate sign-in and enforce company, department and user permissions in the application and data layer.
- Permission-aware retrieval: filter documents before they enter the model's context. A prompt saying “do not reveal confidential data” is not an access-control system.
- Protection beyond the model: encrypt traffic and stored records, restrict administrator privileges, minimise sensitive logging and define deletion and backup retention rules.
- Controlled connections: review telemetry, external tools, document conversion, embedding services and outbound network access. Keep sensitive processing local when that is the requirement.
- Bounded actions: give AI tools narrowly scoped permissions and require human approval before consequential actions such as changing a shipment instruction or updating an official record.
Self-hosting reduces dependence on an external inference provider for processing prompts. It does not automatically prevent stolen credentials, misconfigured databases, malicious documents or insider misuse. NIST's Generative AI Risk Management Profile discusses risks including prompt injection, data privacy and data poisoning that remain relevant to private deployments.
External AI services also differ in their contractual terms, retention settings and security controls. Sending a prompt to a business API does not, by itself, prove that the provider trains on it. Compare the actual service terms and architecture rather than assuming every hosted service behaves like a public chatbot.
My perspective: ownership should make accountability clearer
My view is that the strongest reason to self-host is the ability to make the AI system answerable to the business. The organisation can choose its model, control its data connections, schedule upgrades and investigate failures within its own operating processes.
For a group of companies, I would favour a shared platform with separate knowledge collections and permissions for each business. The infrastructure team can manage common capacity while business owners remain responsible for their definitions, data quality and approved uses.
This is also a useful way to define AI security: confidentiality of information, integrity of answers and actions, and availability when employees need the service. Owning a server supports those goals only when someone owns the controls, maintenance and response procedures around it.
One phrase, three businesses: “delivery date”
A model may understand many meanings of a phrase, yet still choose the wrong one when the business context is missing. “What is the delivery date?” is a simple example.
Figure 2. The same words need different records and business definitions. These are fictional examples, not descriptions of any company's deployed AI system.
Logistics: cargo delivery
For a logistics operator, delivery date might mean the planned arrival of cargo at the consignee's warehouse. It may differ from port arrival, customs clearance or the actual proof-of-delivery timestamp.
An assistant given the approved definition and shipment record could answer: “For shipment C-104, planned warehouse delivery is 24 September. Port arrival is 22 September; actual delivery has not yet been recorded.” The useful improvement is that it distinguishes the milestones instead of returning the first date it finds.
Maternity ward: baby delivery
In a maternity ward, the phrase concerns childbirth. The assistant must distinguish an estimated due date, a scheduled delivery and a recorded birth date. These fields are not interchangeable, and an estimated date is not a guarantee.
A suitable administrative response might be: “Do you mean the expected due date or the recorded birth date?” Any subsequent answer should come from the authorised patient record. The example illustrates terminology and record handling, not clinical prediction or medical advice.
Pizza Hut or another pizza restaurant: order delivery
In a pizza restaurant such as Pizza Hut, delivery refers to the food order reaching the customer. A useful answer might be: “Order P-208 has an estimated arrival window of 7:15–7:25 pm,” using current order and dispatch information.
The model should distinguish order time, kitchen completion, driver departure and arrival. It should also say when tracking data is unavailable instead of inventing a confident delivery time.
Does business-specific training become easier?
A focused business task can be easier to adapt and evaluate than a broad “answer everything” assistant. Domain experts can provide a glossary, representative questions and examples of correct outputs. The delivery-date examples become a small test set: does the system choose the correct meaning, source and level of uncertainty?
There are three different ways to supply that context:
- Instructions and examples: define what “delivery” means in the current workflow and show the required answer format.
- Retrieval-augmented generation (RAG): supply relevant, authorised documents or records when the question is asked. This is usually a better starting point for changing facts such as shipment status or order arrival estimates.
- Fine-tuning: update a pretrained model using curated examples when a repeated task needs more consistent terminology, behaviour or structure.
RAG supplies information at answer time; fine-tuning changes model weights. These methods can be combined, as explained in IBM's comparison of RAG and fine-tuning.
Self-hosting gives the organisation control over this adaptation process, but it is not a prerequisite for it. Hosted services can also support business-specific context. Nor does owning the model guarantee higher accuracy: measure correct answers, source selection, employee correction time and refusal to guess against a baseline.
For the three delivery examples, I would begin with business-specific instructions, structured record access and a glossary. Fine-tuning comes later if evaluation shows a persistent behaviour problem. Ordinary staff conversations do not automatically train the model; feedback needs review before it becomes approved training material.
The economics of AI across 50+ offices
Token-based charges can grow substantially when many offices use AI throughout the day. Input tokens include instructions, documents and conversation history; output tokens include generated responses. Long documents, retries and multi-step workflows all affect consumption.
A shared private platform may spread infrastructure costs across the group and make spending more predictable. However, self-hosting is an upfront investment plus recurring operations—not a one-time purchase that makes future AI free.
The calculation below uses deliberately hypothetical USD figures to demonstrate a decision method. These are not vendor prices, hardware quotes or a capacity recommendation.
Assume 50 offices, 100 active users per office, 20 requests per user per working day and 22 working days each month. That produces 2.2 million requests per month. At an average of 3,000 input tokens and 1,000 output tokens per request, consumption is 6.6 billion input tokens and 2.2 billion output tokens monthly.
| Illustrative monthly calculation | Amount |
|---|---|
| API input: 6,600 million tokens × assumed $2 per million | $13,200 |
| API output: 2,200 million tokens × assumed $8 per million | $17,600 |
| Total model API usage | $30,800 |
| Private setup and hardware: assumed $180,000 spread over 36 months | $5,000 |
| Assumed private monthly operations budget | $12,000 |
| Total private monthly equivalent | $17,000 |
Under these assumptions, private hosting is $13,800 lower per month on an amortised basis. A simple cash payback calculation is $180,000 divided by ($30,800 minus $12,000), or approximately 9.6 months, assuming full usage from the start and unchanged operating costs. That calculation excludes financing, taxes and residual value.
But at one-tenth of the request volume, the same API rates produce a monthly bill of only $3,080. A private installation sized and staffed at the assumed level would then be much more expensive. Lower API prices, caching or batch discounts can shift the result again.
The $12,000 operating allowance must be replaced with a real estimate covering engineering support, electricity, cooling, software licences, security, monitoring, backups and repairs. Include spare capacity and disaster recovery. Compare application development and data preparation on both sides; the illustration assumes those shared costs are equal and leaves them out.
Most importantly, the assumed hardware budget has not been proven capable of serving this workload. Benchmark the selected model at peak concurrency, required context length, acceptable latency and comparable answer quality before treating these figures as a business case. Fifty offices is a useful scale example; utilisation and service requirements determine the economics.
The drawback: learning and upgrades become your responsibility
A private deployment does not automatically receive improvements from outside users, researchers or a provider's later training runs. If the organisation freezes a model version and never updates its knowledge sources, the system can become stale.
However, this does not mean the company must teach it everything from the beginning. A pretrained model already contains general capabilities acquired during its original training. Where licences permit, the organisation can adopt newer model releases, evaluate them and deploy them internally. Approved external knowledge can also be added through retrieval.
The real obligation is to maintain company-specific knowledge and decide when to upgrade. Assign owners to refresh policies and glossaries, collect reviewed feedback, test new versions and retain a rollback option. Fine-tuned behaviour may need to be recreated or revalidated when the base model changes.
Other disadvantages deserve equal attention. An internal team must handle outages, patching and capacity shortages. Hardware can become underused or obsolete. A smaller affordable model may struggle with tasks that a stronger hosted model handles well. Licences may restrict usage, and self-hosting alone does not establish legal compliance. Hallucinations remain possible even when every component runs locally.
A practical way to begin
Start with one well-defined workflow, such as interpreting shipment milestones from approved records. Establish a quality baseline and a permission boundary before expanding access.
Pilot across a small number of offices, record peak demand and measure cost per successfully completed task. Include ambiguous questions, missing records and attempts to access another company's data in the evaluation. Give one team operational ownership and one business owner responsibility for the information.
Then decide whether to expand a shared private platform, continue with a managed service or use a hybrid arrangement. In a hybrid design, classify and route data explicitly; do not silently send confidential requests to an external service when the private model is unavailable.
Self-hosted LLMs: pros and cons
| Area | Potential advantage | Limitation or responsibility |
|---|---|---|
| Data security | Direct control over processing, storage and connections | Misconfiguration, insiders and prompt injection still pose risks |
| Business context | Custom definitions, retrieval and model adaptation | Requires clean data, domain experts and measured evaluation |
| Group-wide use | Shared capacity and common governance across offices | Each company still needs enforced data separation |
| Cost | Predictable capacity costs at sustained utilisation | Upfront investment, recurring operations and idle capacity |
| Model improvement | Control over versions, testing and upgrade timing | Outside improvements do not arrive automatically; upgrades need validation |
| Service reliability | Control over deployment and recovery arrangements | Your team owns outages, patching, scaling and disaster recovery |
| Flexibility | Choice of compatible models and deployment tools | Model licences, hardware limits and migration effort still apply |
| Answer quality | Focused workflows can be tuned and evaluated closely | Local hosting does not eliminate hallucinations or guarantee better answers |
Conclusion
Self-hosted LLMs can be a strong fit for organisations that need close control over sensitive information, clear business context and substantial, sustained AI usage. Across 50 or more offices, shared infrastructure may offer attractive economics when utilisation and operating discipline justify the investment.
The delivery-date example captures the opportunity: useful AI understands which business it is serving, consults the right records and respects who is allowed to see them. That requires a well-designed application and well-maintained knowledge as much as a capable model.
My recommendation is to treat private AI as a business platform with named owners, measurable outcomes and a lifecycle budget. Start with a narrow use case, prove security and quality, and scale when the evidence supports it. The value of owning the infrastructure comes from how responsibly and effectively the organisation operates it.
