Agentic workflows on internal infrastructure

May 5, 2026
When agentic workflows leave your customer's perimeter on every step
deeplit® builds private AI infrastructure for Dutch technical founders selling AI features into hospitals, banks, insurers, government agencies, and law firms. An agentic workflow on internal infrastructure runs the planner, the open-weight model, the tool registry, and the audit log behind the customer's firewall, on Terraform-managed Kubernetes, fully disconnected from the internet. The deployment pattern below is what gets agentic features past an end-customer's security review in the Netherlands and across the EU.
An agentic workflow is the AI feature most often blocked at the end-customer's security review. A single-prompt assistant sends one question to one model and gets one answer back. An agent plans, calls tools, observes results, replans, calls more tools. Each step crosses a network boundary, and on a hosted-API deployment each crossing leaves the customer's perimeter. You shipped one feature. The chief information security officer counts every tool call as a separate egress event.
What an agentic workflow actually is
An agentic workflow is a software loop in which a language model decides what to do next, calls a tool, observes the result, and decides again. It is distinct from a single-shot prompt, which sends one input and returns one output, and distinct from retrieval-augmented generation, which retrieves once and generates once. The agent retrieves, generates, then chooses to retrieve again, query a different system, write to memory, or stop.
For a security reviewer, four parts matter. The planner is the model deciding the next step. The tool calls are the API surface the agent reaches into. The memory is whatever short-term context window or long-term store carries state between steps. The output is what the agent returns to the user or the downstream system. The tool-use loop, the second of those four, is what makes an agentic workflow architecturally different from a single completion call. It is also the part that decides whether the workflow can run inside your customer's firewall at all.
"The deal does not fail because the model is wrong. It fails because the topology is wrong."
Why hosted-API agentic workflows fail the security review
Every step of a tool-use loop is a separate egress event. A hosted-API agent ships every retrieved document, every intermediate result, every tool-call argument, and every memory write to the vendor. The surface area scales with the number of steps the agent takes.
At least four points where customer data leaves the perimeter on a typical hosted-API agent run:
- Retrieved documents shipped as context on each model call, including patient records, financial statements, contract drafts, or whatever the retrieval tool just returned.
- Tool-call arguments sent to the orchestrator, often containing identifiers, payloads, or downstream system parameters.
- Intermediate plan state and reasoning traces logged on the vendor's infrastructure for observability, debugging, and model improvement.
- Memory writes persisted on storage the customer does not own, with retention defaults set by the vendor rather than the deployer.
This is the multi-step exfiltration surface. A document-classification agent that sorts a clinical record, extracts structured fields, and routes the result is a chain of audit boundaries on a public AI API, not one. The chief information security officer at a Dutch hospital does not read this as one integration. The officer reads it as a chain of places where data covered by the GDPR Article 9 sensitivity bar might leave the controller's environment.
The same pattern blocks fintech agents that route credit decisions, legal-tech agents that draft from privileged documents, and public-sector agents that handle citizen data. The deal does not fail because the model is wrong. It fails because the topology is wrong.
The EU AI Act Article 26 deployer-obligations frame for agentic workflows
Article 26 of the EU AI Act sets the obligations of the deployer of a high-risk AI system, separate from the obligations the regulation places on the provider. When an agentic workflow lands in Annex III territory, the customer of the startup is the deployer. The startup is the provider. The split changes what each party owes the regulator. The compliance date for these standalone high-risk obligations is December 2, 2027, deferred from the original August 2026 date by the EU's Digital Omnibus package. That is not a reprieve. It is a fixed deadline with lead time, and the records the deployer will owe are generated in real time from the first day the agent runs.
The Article 26 obligations the deployer has to discharge include keeping the system's automatically generated logs for a period appropriate to its intended purpose, of at least six months; ensuring human oversight by a competent natural person; monitoring the system's operation and flagging risks to the provider; and reporting serious incidents to the market surveillance authority.
The artifacts are the agent's records: the plan steps the model produced, the tool calls the agent issued, the model outputs the agent observed, the memory writes the agent committed, the decisions the agent made and the inputs to those decisions. Together they form the audit trail Article 26 expects the deployer to produce on request.
When the agent runs on a public AI API, those artifacts live on the vendor's servers, under the vendor's retention policies, behind the vendor's access controls, and in the vendor's jurisdiction. The deployer can request them through a contract clause. The deployer cannot demand them at the moment the regulator asks. The Article 26 record-keeping obligation is real, and the December 2, 2027 compliance date gives every deployer a fixed point to be ready by. The artifacts the regulator will ask for are produced in real time as the agent runs, which means the architecture that decides whether the deployer can produce them is a choice made now, not in 2027.
The architectural alternative is the only one that closes this gap cleanly: the agent runs on infrastructure the deployer controls, the artifacts are written to storage the deployer owns, and the audit log never leaves the customer's perimeter. The deployer's responsibilities are bounded by what the deployer can actually deliver. More on that boundary in EU AI Act and GDPR Article 28.

A reference architecture for an agentic workflow on internal infrastructure
A logistics customer running an agentic workflow over satellite imagery is the cleanest example. The agent decides which satellite passes to query, retrieves the imagery, classifies what it finds, and dispatches results to a downstream operations system. Sovereignty rules out a public AI API; the imagery and the dispatch destinations are not allowed to leave the customer's environment. The architecture reduces to four components, all running inside the customer's perimeter, deployable with Terraform onto Kubernetes, fully disconnected from the internet.
Orchestrator
The orchestrator is the planner-executor loop. It runs as a stateless service inside the customer's Kubernetes cluster, holds no model weights, and keeps no customer data at rest between runs. It talks to the model server and the tool registry over the cluster network only. There is no outbound DNS resolution, no egress route, no health-check ping to a vendor endpoint. When the cluster goes offline, the orchestrator goes offline with it.
Model server
Open-weight models are served on the customer's GPUs by an OpenAI-compatible inference server packaged as a Docker image, deployed and versioned through Terraform. The server has no outbound calls, no telemetry, no automatic update mechanism. New model versions roll out through the same Terraform pipeline as any other infrastructure change, signed and reviewed before they reach production. This is the self-hosting an open-weight LLM pattern, applied inside an agent loop.
Tool registry
The tool registry is the explicit, reviewed list of tools the agent is allowed to call. Each tool is an internal API behind cluster-local authentication: the satellite-imagery store, the internal classifier, the dispatch system. Tool calls never traverse the customer perimeter because the tools themselves are inside it. For the customer-cloud variant, the same pattern applies inside the customer's existing environment, covered in private AI in your customer's GCP project.
Audit log
Every plan step, tool call, model output, and memory write is written to an append-only log inside the customer's storage. Log writes are signed and time-stamped. Retention is configured by the deployer, not the vendor. This is the artifact the deployer's compliance officer reaches for at audit time under Article 26. The log never leaves the customer's perimeter. The deployer produces it for the regulator on the deployer's own timeline.
When this is the wrong choice
This architecture is the right answer when the agent's data crosses a regulatory or contractual boundary. It is overbuilt in three cases. First, low-stakes internal productivity workflows where the data is already public or the agent only touches a knowledge base of marketing materials. Second, low-volume workflows where the GPU economics do not justify dedicated hardware: a handful of runs a week is cheaper on a public AI API and the egress is not a real risk. Third, early-stage prototypes where the tool surface is still moving and the iteration speed matters more than the deployment posture. A simple heuristic: if you cannot name the regulation or the contract clause that the egress would breach, the public AI API is fine.
deeplit® deploys this stack into the customer's environment, with the customer's hardware, behind the customer's firewall. Book a deployment call to walk through the architecture against your specific agent design and your end-customer's compliance posture.
If your customer's compliance lead is asking where the agent's reasoning trace lives at audit time, this post is written to be forwarded to them. The Article 26 framing is theirs to use in their own preparation.
Every tool call on a hosted-API agent is a separate egress event.
Article 26 obligates the deployer; only the architecture decides whether the deployer can comply.
Frequently asked questions
Can an agentic workflow run fully disconnected from the internet?
Yes. The orchestrator, the open-weight model server, the tool registry, and the audit log all run inside the customer's Kubernetes cluster on customer-owned GPUs. There is no outbound network route required for the agent to plan, retrieve, generate, or write its audit log. Software updates ship through the same Terraform pipeline as any other infrastructure change.
How does this differ from retrieval-augmented generation on internal infrastructure?
Retrieval-augmented generation is a single-step pattern: retrieve once, generate once, return the answer. An agentic workflow is a multi-step plan-tool-observe loop where the model decides what to retrieve next, what tool to call, and when to stop. The egress surface scales with the number of steps the agent takes, not with the number of user prompts.
Does this work for voice AI agents?
Yes. A voice agent is an agentic workflow with speech-to-text and text-to-speech as registry tools. A voice receptionist that triages calls, books calendar slots, and queries the customer's knowledge base runs the same orchestrator, model server, tool registry, and audit log inside the customer's environment. The audit log captures the transcript step alongside the agent's plan.
Does deeplit® support the customer-cloud variant?
Yes, as a secondary deployment path. The flagship deployment is on-premises customer-owned hardware, where the architectural promise of fully disconnected from the internet holds cleanly. The customer-cloud variant deploys the same stack into the customer's existing GCP project, AWS account, or Azure subscription when on-premises is not the right fit.
How long does a deployment take?
It depends. The variables are the complexity of the tool registry the agent will call, whether the customer's GPUs are already provisioned, and how the access handoff is structured between deeplit® and the customer's infrastructure team. The typical conversation starts with a short scoping call to map the agent design and the compliance posture before estimating.