Private LLM vs public LLM: four ways to ship private AI

May 2, 2026
When the hosted-API answer stops working
Most startups selling into hospitals, banks, insurers, public-sector teams, or pharma research groups follow the same pattern. The product is built on a hosted API from a US provider, the obvious choice, the fast choice. The first deal closes. The second deal closes. The pipeline starts to build. And then a single question from an end-customer's compliance team stops the whole thing cold: where does our data go when your AI processes it?
Two weeks before the Statement of Work was supposed to be signed, the deal moves into limbo. The compliance team is not asking out of curiosity. They are asking because Article 28 of the General Data Protection Regulation makes them legally responsible for every sub-processor in the data path, because the November 2025 designation of 19 Critical ICT Third-Party Providers under the Digital Operational Resilience Act made it harder for any regulated organization to add a US hyperscaler to its register, and because their last security review found that the model provider's data processing agreement does not actually answer the question they need answered.
You, the founder or chief technology officer of a 10- to 30-person startup, are now told you have four to twelve weeks to come back with private AI infrastructure that runs on the customer's hardware, behind their firewall, with no external dependency. The deal is not lost yet. But it is on the clock.
At that point the decision is not whether to ship private AI. The decision is who builds it. There are exactly four options. Three of them look reasonable on the surface. Two of them quietly cost more than they save. One of them is what most well-run startups end up doing, even though they do not call it by that name. The conversation that follows is private LLM versus public LLM: same family of open-weight models, same workload, but a different question about who controls the data path. The four options below are answers to that question. The Euro figures behind each make the comparison honest.
A note on terminology before we start. "Private LLM" in this post means an open-weight model running on infrastructure your end-customer can audit, with no inference traffic crossing your end-customer's network boundary, the architectural opposite of the public LLM API the deal was built on. "Private AI" is the surrounding infrastructure: the serving stack, the retrieval pipeline, the role-based access control, the audit logging, and the operational tooling that turn a private LLM into a deployable system. The two travel together but are not the same thing. "Private AI" is also not "sovereign AI" (which usually means a European-hosted version of a US-controlled service) or "on-premise AI" (which refers to physical location only, not control over the data path). When the procurement conversation turns serious, none of these are interchangeable.
The cost half of the comparison (GPU capex, the engineering quarter, the steady-state burden of running a private LLM at startup volume) is laid out in the real cost of self-hosting an open-weight LLM.
If your end-customer can change, revoke, or monitor your AI, it was never theirs to lose control of.
The four options on the table
Each option below answers the same question: how do you ship private AI infrastructure to your end-customer in the four-to-twelve-week window before the deal slips? The Euro figures are realistic mid-2026 numbers for a 10- to 30-person Dutch startup. Hardware costs are reference figures from the deeplit® hardware compatibility matrix. Engineering costs assume mid-senior Dutch market salaries plus overhead.
Option one: buy GPUs and self-host
You buy or lease the hardware your end-customer's compliance team will accept. You hire or borrow the engineering capacity to deploy an open-weight model on it. You build the serving stack, the retrieval pipeline over private documents, the role-based access control, the audit logging, and the operational tooling to keep it running. The deployment lives on hardware your team owns or controls, in your customer's data center or in a colocation facility they have already accepted.
The cost picture: GPU capex of €50,000 to €150,000 for hardware sized to a small production workload (one or two A100 or H100-class accelerators, server, cooling, networking, and a vendor support contract). One full engineer-quarter to deploy the first version, optimistically. Two to three engineers permanently allocated to maintain it (security patches, model upgrades, incident response, capacity planning). A run-rate of €30,000 to €60,000 per quarter in pure operational engineering once the deployment is live. First-year all-in: €170,000 to €390,000.
The honest part of this option: you own everything. The data path is yours, the keys are yours, the audit boundary is yours. Your end-customer's compliance team can audit any part of the stack and the answer is always "we control it." For a startup whose product is the AI itself (a clinical decision support tool, a fraud detection model, a medical imaging classifier), option one is sometimes the right call because the model is not separable from the product.
The expensive part: the engineering quarter you are spending on the deployment is the engineering quarter you are not spending on the product the customer is paying you for. Most early-stage startups do not have an engineering quarter to spare. The hidden cost is opportunity cost, and at a 10- to 30-person startup it is usually larger than the GPU bill.
Option two: build the deployment story yourself, on someone else's hardware
A variant on option one. You take an open-weight model and deploy it inside your customer's existing cloud account, their Google Cloud project, their AWS account, or their Azure subscription, using Terraform-managed infrastructure-as-code (the GCP pattern is walked through end-to-end in Deploy private AI in your customer's GCP project; the AWS and Azure variants are architecturally identical). The hardware is rented from the customer's existing hyperscaler, billed to their account, sitting inside the network perimeter they already trust. Your code runs there; their data never leaves there.
The cost picture: no GPU capex (the customer pays the hyperscaler bill directly for the GPU instances). The engineering cost is similar to option one. The deployment work does not get easier just because the hardware lives in someone else's cloud. €30,000 to €60,000 of engineering time to build the Terraform installer, the serving stack, the retrieval pipeline, and the audit-logging defaults. Ongoing maintenance is lower than option one (no hardware life-cycle to manage), but you still need at least one engineer keeping the stack current. First-year all-in: €60,000 to €120,000 of engineering, plus the customer-paid hyperscaler bill.
The honest part: this is the strongest commercial option for a startup that has the engineering capacity to do it well. Customer-environment deployment is increasingly what buyer compliance teams ask for by name. It side-steps the GPU-ownership question entirely. It maps cleanly onto the General Data Protection Regulation processor and sub-processor structure. The Critical ICT Third-Party Provider question becomes the customer's existing hyperscaler relationship, not a new dependency on your stack.
The expensive part: you are still spending an engineering quarter on something that is not your product. If you ship it well, you have just built a private AI infrastructure platform for your own use, by accident. If you ship it badly, you have built a security incident waiting to happen, and your end-customer's compliance team will catch it on the next audit. The middle ground (ship something that works in production but is fragile under stress) is the most common outcome and the most dangerous, because it does not announce itself until the third or fourth deployment.
Option three: let the deal slip
The unstated option that founders rarely admit to picking. The deal moves to "indefinite hold." You go back to the pipeline and try to find an end-customer with a less rigorous compliance posture. You hope the next conversation does not hit the same wall.
The cost picture: the booked-but-unsigned deal value, minus the lost engineering hours already invested in the pilot, plus the founder time spent on the deal-stage diligence. For a 10- to 30-person Dutch startup pitching mid-market regulated customers, this is typically €50,000 to €250,000 of contract value walked away from per slipped deal. There is also a less visible cost: the second-order signal to the next customer when they search for case studies and discover that your last named customer in the same vertical is no longer using the product.
The honest part: sometimes the right call. Not every deal is winnable in the four-to-twelve-week window. If the next deal in the pipeline is at a customer with a less rigorous compliance posture, you have not paid the cost of solving the problem yet, and you can defer it.
The expensive part: the next three deals in the pipeline are usually at customers with the same compliance posture. The pattern that put you in this conversation (your end-customer's regulator made the requirement explicit) is structural, not accidental. The deals do not stop coming. They stop closing.
Option four: bring in a private AI infrastructure partner
You buy private AI infrastructure as a managed service. The vendor handles the deployment work (the serving stack, the retrieval pipeline, the role-based access control, the audit logging, the Terraform-managed infrastructure-as-code, the data processing agreement template) and you get back the engineering quarter you were about to spend on it. The deployment runs on either the customer's hardware (the option-one shape) or inside the customer's existing cloud account (the option-two shape), with no external dependency on the vendor for inference or data path.
The cost picture: deeplit® Private is priced on your users, load, and hardware, with the customer providing the hardware separately (either their own GPU server or their existing cloud account). Its platform license stands against a build-it-yourself first-year cost of €60,000 to €390,000 plus the GPU bill. The deployment ships in five to ten working days for the cloud-account variant and two to three weeks for the on-premises variant, against four to twelve weeks for an internal build.
The honest part: this is option two without the engineering quarter. The deployment shape, the audit boundary, the data path, and the compliance posture are identical. The difference is who writes the Terraform.
The expensive part: you are now dependent on a vendor for a critical component of your stack. The mitigations are the ones to negotiate for in the contract: the vendor's management plane runs on customer infrastructure with no back-channel; the Terraform code is escrowed; the data processing agreement makes the boundaries explicit; nothing the vendor does cannot be undone by the customer themselves. Without those mitigations, option four becomes option one with extra steps.

Which option most teams actually pick
Most well-run 10- to 30-person Dutch startups in this position pick option four, even though the conversation usually starts at option one. The math is the reason. The opportunity cost of an engineering quarter at an early-stage startup is somewhere between €60,000 and €150,000, depending on what the engineering quarter would otherwise have built. Option four hands that quarter back: put a quote next to that range and the comparison usually decides itself. The deployment ships in days rather than months. The compliance posture is the same.
The reason the conversation starts at option one is founder instinct: building it yourself feels like the resilient choice, the one that keeps you in control. But the audit boundary, the line between what your end-customer can audit and what they cannot, does not move based on who deploys the stack. If the deployment lives on the customer's hardware (or in their cloud account) and the management plane is operated by the customer, the audit boundary sits in the same place whether you built it or whether deeplit® did. Pick the option that puts the audit boundary in the right place, then optimize the engineering hours.
deeplit® is option four for the Dutch startup market. We deploy open-weight models on your customer's hardware or in their existing cloud account, with the Terraform-managed installer pattern, role-based access control, audit logging, and the data processing agreement template. The platform is designed to support the General Data Protection Regulation and the EU AI Act. The first deployment ships in five to ten working days for the cloud-account variant.
If you are eight to twelve weeks from a deal that just hit a private AI requirement, the next step is a thirty-minute deployment call. Bring the customer's deployment constraints (cloud or on-premises, GPU availability, audit-log retention requirements, data-residency requirements) and we will tell you whether deeplit® is the right call or whether one of the other three options serves you better. We do not win every conversation. We want to win the ones we should.
Book a deployment call
If the procurement conversation has reached your end-customer's compliance team, the architecture-side companion to this post is the audit boundary across hosted-API, your-cloud, and customer-environment AI deployments.
If your end-customer can change, revoke, or monitor your AI, it was never theirs to lose control of.
The math is settled. The remaining question is who writes the Terraform.
Frequently asked questions
How long does it take to deploy private AI in a customer's existing cloud account?
For a Terraform-managed deployment of an open-weight model into a customer's Google Cloud project, AWS account, or Azure subscription, the realistic timeline is five to ten working days when the customer is responsive on access provisioning. On-premises deployments on customer-owned hardware typically take two to three weeks because hardware delivery, cabling, and network configuration sit on the customer's facilities team rather than on the deployment engineer.
Does a startup need ISO 27001 to ship private AI to a regulated end-customer?
No. The end-customer's compliance team is asking for control over the data path, not for a vendor certification. A startup without ISO 27001 can ship private AI by deploying inside the customer's environment with a data processing agreement that makes the audit boundary explicit. ISO 27001 becomes a soft requirement once the startup is selling to large enterprises with formal vendor-security questionnaires; it is not a precondition for the first ten regulated customers.
What is the minimum team size to run private AI in production?
If the deployment is operated by an external partner under option four, the customer's team size for ongoing operations is approximately zero, because the partner handles patching, model updates, and incident response. If the deployment is run internally under option one or two, the minimum is two to three engineers full-time once it is live, with surge capacity for incidents. The most common failure mode at a 10- to 30-person startup is hiring one engineer for the work and burning them out within six months.
Can private AI use the same open-weight models as the hosted-API alternatives?
Yes. The same family of open-weight models that power the hosted-API providers can be deployed privately on customer infrastructure. The architectural differences (where the model runs, who controls the data path, who holds the keys) are independent of which model is loaded. Model selection should be driven by workload requirements (latency, context window, languages supported) rather than by deployment shape.
What is the practical difference between private LLM deployment and using a public LLM API?
The model capabilities are similar; the same families of open-weight models that power many hosted-API offerings are deployable privately. The difference is the data path. A public LLM API sends every inference request out across your end-customer's network boundary to a third-party provider, who is now in the data path and on the customer's sub-processor register under Article 28 of the General Data Protection Regulation. A private LLM deployment runs an open-weight model on infrastructure your end-customer can audit, with no inference traffic crossing the boundary and no third party in the data path. The decision is not about model quality. It is about who controls the audit trail.