Skip to content
[ blog post ]

The real cost of self-hosting an open-weight LLM

Title card reading "The real cost of self-hosting an open-weight LLM" in white on a midnight-blue field with a violet glow rising from the foot, a balance-scale icon above it and the line "notes from deeplit®"

May 2, 2026

H. Kamkar

The GPU bill is the part everyone gets right

Most build-versus-buy comparisons for self-hosting an open-weight large language model get the hardware cost roughly right. GPU prices are well-published. Server, networking, cooling, and support contract costs are quotable. The hardware line item is the part of the comparison that founders feel comfortable with, because it looks like every other infrastructure decision they have made.

The operational engineering cost is the part everyone gets wrong. The published estimates from cost-of-self-hosting writeups across the industry put it at 70 to 80 percent of the total cost of ownership for a self-hosted private AI deployment. Most internal estimates inside startups put it closer to 30 to 40 percent, because the engineering work is invisible until production. The gap between those two numbers is what makes the build path more expensive than the spreadsheet predicts.

This post is the honest math for a 10- to 30-person Dutch startup pricing the build path internally. It is the cost-essay companion to Private LLM vs public LLM: four ways to ship private AI, which covered the four options at a strategic level. This post drills into Option 1 cost: what self-hosting an open-weight large language model actually costs across hardware, engineering, and maintenance, with Euro figures pegged to mid-2026 Dutch market salaries.

The GPU bill is 20% of the cost. The engineers are 80%.

Three line items, in order of how often they are underestimated

Below are the three cost components of running a private AI deployment in production. Each one is sized for a 10- to 30-person Dutch startup running a one- to two-GPU production deployment.

Hardware: the part everyone gets right

GPU capex for a small production private AI deployment is the line item with the least uncertainty. One or two A100 or H100-class accelerators, configured for a low-hundreds-of-concurrent-users workload, with appropriate redundancy: €50,000 to €150,000. Server, cooling, networking, and a vendor support contract: another €10,000 to €30,000. Three-year depreciation is the standard accounting treatment, which means the hardware shows up on the books at €20,000 to €60,000 per year of operating cost.

If the deployment is in a colocation facility (which is common for hospital and bank deployments where the customer requires hardware to live in a specified facility): €5,000 to €15,000 per year of colocation fees. Electricity for a one- to two-GPU production deployment runs €3,000 to €8,000 per year at Dutch grid prices.

Total hardware all-in for the first year: €70,000 to €220,000 if you are buying. €40,000 to €120,000 if you are renting from a hyperscaler at customer expense.

Operational engineering: the part everyone gets wrong

A mid-senior Dutch engineer fully loaded (salary, benefits, employer taxes, equipment, hiring cost amortized) runs €150,000 to €200,000 per year. To run a private AI deployment in production, you need at minimum two engineers with overlapping but distinct skill sets: one engineer with model serving and infrastructure expertise, one engineer with machine learning operations and pipeline experience. With surge capacity for incidents, this is a two-and-a-half full-time-equivalent commitment.

So the steady-state operational engineering cost is €300,000 to €400,000 per year. This is the line item that does not show up in most internal estimates.

What the engineers actually do, in rough proportion of time spent: GPU driver maintenance and CUDA upgrades (10 to 15 percent), model serving framework upgrades (10 to 15 percent), retrieval pipeline maintenance and tuning (15 to 20 percent), audit logging and identity and access management configuration (10 to 15 percent), security patching (5 to 10 percent), capacity planning and incident response (15 to 25 percent), platform hardening and documentation (15 to 20 percent).

For a 10- to 30-person startup, this is two and a half engineers who are not building the product the customer is paying for. The opportunity cost is what the engineering quarter would otherwise have produced. For a startup at this stage, the opportunity cost on €300,000 of engineering spend is somewhere between €600,000 and €1,500,000 of forgone product progress.

Ongoing maintenance: the part nobody plans for

Open-weight model upgrades arrive every 6 to 12 weeks. Each upgrade is a mini-deployment: smoke tests against the new weights, A/B comparison with the production version, gradual rollout, monitoring for regressions. Allocate two to three engineer-days per upgrade. Eight upgrades per year is sixteen to twenty-four engineer-days, or one full engineer-month per year of maintenance just for model upgrades.

Security patching at the operating system, container runtime, and serving framework levels runs monthly minimum. Allocate one to two engineer-days per month. That is twelve to twenty-four engineer-days per year on patching.

Incident response: budget 5 to 10 person-days per quarter even when nothing breaks badly, plus surge capacity for the one or two real incidents per year. Total: 25 to 50 engineer-days per year on incident response.

Sum of maintenance: roughly two to three additional engineer-months per year on top of the steady-state two and a half full-time-equivalent commitment. That pushes the engineering cost to €350,000 to €450,000 per year for a real production deployment.

The same engineering line item shows up regardless of where the hardware sits: your colocation, or your customer's environment. The architectural companion to this cost essay is private AI in your customer's GCP project, which walks the deployment shape behind the customer-environment column of the comparison.

Blueprint of a jet engine in cross-section, white line art on a midnight-blue grid, the combustion chamber at its core glowing violet and a coral point at the intake. Caption: "Most of the cost lives inside the shell, not on it."

The honest comparison

Total year 1. Self-host (build it yourself): €420,000 to €670,000. deeplit® Private: the same GPU bill plus a platform license priced on your users, load, and hardware, with the engineering and maintenance included. The same architectural deployment shape.

Year 3 in steady state. Self-host: €350,000 to €510,000 per year. deeplit® Private: the platform license plus the GPU bill.

What goes into the year-1 number, line by line:

  • GPU hardware (capex or rented). Self-host: €70,000 to €220,000. deeplit®: €40,000 to €220,000, customer-paid.
  • Operational engineering (approximately 2.5 full-time-equivalent). Self-host: €300,000 to €400,000. deeplit®: included in the platform license.
  • Ongoing maintenance and incident response. Self-host: €50,000 to €100,000. deeplit®: included in the platform license.
  • Platform license. Self-host: not applicable. deeplit®: priced on your users, load, and hardware.

Total first-year all-in for self-hosting a one- to two-GPU production private AI deployment, including hardware, engineering, and maintenance: €420,000 to €670,000. Year-2 onwards (no GPU capex if you bought outright; capex amortizes if not): €350,000 to €510,000 per year.

deeplit® Private runs the same architectural deployment shape, with the customer providing the hardware separately (either their own GPU server or their existing cloud account). The GPU bill is the same €40,000 to €220,000, depending on whether you buy or rent. The platform license is priced on your users, load, and hardware, so the comparison to run is that quote against the €350,000 to €450,000 a year of engineering and maintenance it replaces.

The deployment posture, the audit boundary, the data path, and the compliance characteristics are identical. The difference is who writes the Terraform and who runs the on-call rotation.

There are circumstances where the build path is the right call regardless of the cost comparison. When the AI is the product (not infrastructure for the product), owning the stack matters more than the cost comparison suggests. When the team is large enough that two and a half engineers is a small fraction of headcount, the opportunity cost calculation changes. When the customer requires a fully disconnected configuration with no vendor licensing webhook, the build path may be the only option. For most 10- to 30-person Dutch startups in 2026, none of those conditions are true.

If you are pricing the build path internally and want a sanity check on the line items, book a deployment call. Bring your internal estimate; we will price deeplit® Private against it and walk through the numbers honestly. We do not win every conversation; we want to win the ones we should.

If your end-customer's procurement team is asking for the line-by-line cost basis behind the deployment shape, this post is written to be forwarded to them. The engineering-cost argument is theirs to use in their own budget conversations.

The GPU bill is 20% of the cost. The engineers are 80%.

If your engineering quarter is the constraining resource on the roadmap, the comparison decides itself.

Frequently asked questions

Are these numbers for hyperscaler deployment or on-premises?

The hardware line item splits the two: capex of €50,000 to €150,000 for buying hardware, or hyperscaler GPU instance bills (which the customer typically pays directly). The engineering and maintenance line items are roughly the same regardless of where the hardware lives, because the operational work is the same.

What about smaller GPU configurations, like a single A40 or an L40S?

Smaller configurations reduce the hardware line item but not the engineering line item proportionally. A single mid-range GPU configuration runs €15,000 to €35,000 of hardware capex, but the engineering and maintenance work is structurally the same. Against deeplit® Private, the build path looks worse, not better, because the fixed engineering cost dominates.

Does the comparison change if we already have engineers we would reassign?

Reassignment is not free. The engineers you reassign are not building the product they were hired to build. The economic question is the opportunity cost of their reassigned time, not their salary, which they are getting paid either way. For most early-stage startups, the opportunity cost of an engineering quarter is higher than the salary cost.

What is the year-three picture?

Year-three self-hosting cost (hardware fully depreciated, engineering steady-state): €350,000 to €510,000 per year. Year-three deeplit® cost: the platform license plus the customer-paid GPU bill. The build path gets slightly cheaper in year three, but the engineering line item stays the dominant cost across all years.

Does this calculation change if our customer pays for the GPU server directly?

No. The customer-paid GPU shape is how deeplit® Private is deployed: the customer owns the hardware on-premises or in their existing cloud account, and the platform license covers everything else. The line items in the comparison above stay the same; only the question of which entity issues the GPU invoice changes. The architectural pattern is documented in private AI in your customer's GCP project.

Read next

H. Kamkar headshot

H. Kamkar

Building private AI infrastructure for startups selling into regulated industries.