Skip to content
Sagar Thakkar
← All writing
Architecture 9 min read
self-hosted ai-architecture llm-gateway observability security agents mlops

Easy to Demo, Hard to Operate: Why AI Pilots Stall in Month Two

AI pilots rarely fail on the model. They stall on memory, security, monitoring and cost. One platform built inside hard limits shows what that takes.

Sagar Thakkar
Sagar Thakkar
AI Systems Architect
Typographic poster reading ANYONE CAN DEMO IT. RUNNING IT IS THE JOB., beside a single slab divided into five bands, one highlighted, above the line ONE BOX. FIVE PLANES. ONE OPERATOR.
TL;DR

AI pilots rarely fail on model quality. They stall on operability: durable memory, a security boundary, monitoring, and per-call cost visibility. This case study builds all four inside hard limits, one host, consumer GPUs and near zero infrastructure spend, and names the cost of every decision a production team would inherit.

For technology leaders with an AI pilot that demos well, and the engineers who will be paged when it does not. What the operability gap actually contains, the five questions that expose it, and one architecture built entirely inside it.


TL;DR

  • Your AI pilot will not stall on the model. It stalls in month two, on memory, security, observability and cost control. None of it shows up in a demo, and all of it shows up in production.
  • Five questions separate a demo from a system, and they are the same at any scale. They are listed below, before the architecture, so they can be asked of any team’s pilot.
  • The interesting constraint here was not capability. No inbound IPv4 because the connection sits behind carrier-grade NAT, consumer GPUs rather than datacentre ones, and an infrastructure budget of roughly nothing.
  • Two decisions carry the design: security by network topology rather than firewall rules, and local inference as a first-class routing tier rather than a fallback.
  • There is no high availability. One host is one fault domain. Recovery means a compose definition plus backups, and the honest objective is hours.
  • What breaks first is VRAM, not money and not the network. Everything else in the cost model is flat until it does.

The constraint was operability, not capability

A prototype that calls a model API is an afternoon of work. One key, one notebook, no memory of yesterday. Nothing watches it. Nothing tells you what it spent. It is impressive in a demo for exactly the reasons it is useless in production.

The gap between that and a platform is not model quality. It is four unglamorous properties. State that survives a restart. A boundary somebody has to cross to reach you. Something that notices when a part stops working. A per-call record of what was spent. Vendors sell the model. The four properties are left as an exercise.

This build treated them as the problem, then made the problem harder by removing the usual escape routes.

Constraints, all real:

ConstraintWhat it rules out
No public IPv4 address, the connection sits behind carrier-grade NATInbound IPv4 port forwarding. Note the qualifier, because IPv6 is a separate question and it is answered below
One workstation, consumer GPUs in the 12 GB class, isolated per cardModel-size headroom for a fully resident model, and any serious concurrency
Roughly zero monthly infrastructure spendManaged databases, managed vector stores, managed observability
Must survive recreation and surface its own healthAnything configured by hand and remembered by nobody

Five questions to ask of any AI pilot

Everything in this case study reduces to five questions. They cost nothing to ask, and a pilot that cannot answer them is a demo, whatever its accuracy numbers say.

  1. Where does state live, and what happens on a restart? If memory is a notebook variable or a single process, the first deploy erases what the system learned.
  2. What is the security boundary, and when did someone last check it? A boundary described in a design document is a claim. A boundary something tests on a schedule is a control. The section below shows how far apart those two can drift.
  3. How would anyone know a part stopped working? Name the failures that alert. Everything not on that list is discovered by a user.
  4. What did yesterday cost, per call and per feature? If the answer arrives with the monthly invoice, cost is not being managed, it is being reported.
  5. Has a backup ever been restored? If not, recovery is a plan, not a capability.

“We assume” is the answer that should worry a leader most, because it sounds like a yes. The rest of this post answers all five for one real platform, with the cost of each answer stated.


Decision one: security by network topology

The rule is that every service binds to loopback and one private mesh address, never to all interfaces. The mesh is a WireGuard overlay, so a laptop or a phone reaches the platform from anywhere, encrypted, with no port opened.

Carrier-grade NAT looks like the limitation that makes this work. Be exact about what it does: it removes inbound IPv4, and that is the whole of its contribution. A container published on a wildcard address stays reachable over IPv6 as soon as the host holds a global address. The usual container runtime publishes on that wildcard by default.

So the binding rule is the control, and a rule is not a control until something checks it. Counting wildcard-published ports and global IPv6 addresses takes one command. Skipping it leaves a design claim standing in for a measurement. The rule was written when the platform ran about two dozen containers. It runs more than forty now, and the newer half never faced the audit.

The fix is not more discipline, it is moving the check off the human. A small reconcile loop now holds a drop rule on the container forwarding chain, scoped to the public interface. A service that publishes on a wildcard is refused at the edge rather than exposed by it. It reapplies every minute, because the chain is recreated whenever the container daemon restarts, and a rule applied once is a rule that disappears quietly. It logs only when it has to re-add, so the log is a drift signal instead of noise. The way to trust it is to delete the rule by hand and watch it come back.

A small number of endpoints are published to the public internet on purpose, per service and behind authentication, because an external client needs them. That is the correct shape: a closed default with narrow, explicit exceptions. It is not the same as zero public attack surface. Once one path is public, topology has stopped being the only control and per-service authentication is carrying real load.

Two costs come with this, and both are permanent:

  • The mesh is now the perimeter. A compromised laptop is inside it. Defence in depth is not optional here, which is why every service also carries its own credentials.
  • A third-party coordination plane is now a dependency. The mesh needs it to establish connections. That is a trust and availability dependency a purely local design does not carry, and it belongs beside the near-zero cost claim.

Blast radius is the other half. One agent runs with real privileges, which is the entire point of it. Its writes are scoped and gated, so automation acts without unbounded reach. An agent that can execute anything is not an architecture, it is an incident waiting for a trigger.


Decision two: local inference as a tier, not a toy

Five planes on one host: operator devices reaching an encrypted mesh, a model plane with gateway and local GPU inference, an agent plane with scheduler and gated executor, a data and knowledge plane, an operations plane with co-located backups, and one narrow authenticated public exception

All model traffic goes through one gateway speaking a single API shape. Behind it sit cloud and local models on the same routing table. Application code never names a provider, so swapping a model is a routing change, and cost and latency are traced per call.

Local inference is not there to save face. It has a job: embeddings and bulk or private work never leave the box, and only hard reasoning is routed out. Embedding a large corpus through a paid API is the line item that quietly dominates a small budget, so that split is what makes the economics work.

Cards are isolated per device and do not pool. A device-selection step prefers the dedicated card and falls back when it is absent. That is how the embedding service survived a card being physically removed. Resilience at this scale means degrading, not failing over.


What the counts actually mean

Numbers published without their unit and their date are decoration. These are the four this platform gets described with, and what each is measuring.

FigureWhat it countsHow to read it
More than forty containersRunning containers on one host in September 2026, up from about two dozen in mid-2026A count of processes, not of microservices. It moves whenever a profile is enabled. The useful claim underneath is that all of them are declared in version-controlled compose definitions, so the count is derived rather than remembered
Fourteen modelsEntries in the gateway routing table, counted mid-2026A configuration count, not a benchmark. What it buys is that no application names a provider
Roughly ninety thousand vectorsScrubbed conversation segments in the memory store, mid-2026, on a counter that only growsNot a capacity figure. It matters for drift: changing the embedding model invalidates every vector, and re-embedding is a linear batch job, so this number is really the price of a model change
Roughly zero per monthMarginal infrastructure cost: no cloud bill for compute, storage or networkingIt excludes electricity, hardware already owned and now depreciating, cloud tokens actually spent, and the mesh coordination plane on a free tier. Zero marginal infrastructure cost is real. Zero total cost of ownership is not a claim anyone should make

The last row is the one that gets quoted wrongly. Free infrastructure and free platform are different sentences.


What breaks first under load

VRAM breaks first. Cards are isolated rather than pooled, so two cards are not one large card. The largest single card sets the ceiling for a model held entirely in VRAM. At the 12 GB class that lands near a fourteen-billion-parameter dense model at four-bit quantisation, once the key value cache is counted.

That ceiling is a performance cliff, not a wall. Past it, layers spill into system RAM and the model still runs, across a much slower bus. A capacity limit becomes a tokens-per-second limit. A thirty-five-billion-parameter model serving a nightly batch is a reasonable trade. The same model in front of a waiting human is not.

Decode speed is bound by memory bandwidth rather than by compute. The upper bound on tokens per second is roughly memory bandwidth divided by the bytes of weights read per token. Quantisation therefore buys speed as well as space. On consumer hardware that bound is low enough that local inference earns its place on cost and privacy, never on latency.

Concurrency is the same constraint wearing a different hat. A local serving process allocates its slots and their key value cache at startup, so requests past the slot count queue rather than fail. The trade on one card is between slot count, context length and model size, and it shows up as waiting.

Cost breaks second, and only indirectly. Infrastructure spend is fixed, so the curve is flat. It stays flat until local capacity saturates, at which point traffic shifts to paid cloud models and the bill starts tracking usage linearly. That step is not gradual. It arrives on the day a workload stops fitting on the box.

The host breaks third, because it is one fault domain, which is the next section.

Notice what is not on this list. Network throughput is irrelevant at this scale, and storage stays cheap until the vector count grows by an order of magnitude. The scarce resource is the one bought years ago that no amount of spending scales this month.


Recovery, without high availability

There is no high availability here. One host, one power supply, one set of disks.

The recovery objectives split in two, and a design that states only one has not been thought through:

  • RPO, 24 hours against logical loss. Corruption, a bad delete, a broken migration. Daily logical backups cover this.
  • RPO, unbounded against physical loss. The object store holding those backups sits on the same machine. Co-located backups are not backups, they are copies. This is the single largest gap in the design, and offsite encrypted replication is the fix.
  • RTO, hours. Restoring data is the fast part. Recreating the environment is the slow part. Container networks and volume permissions do not survive a rebuild cleanly, and secrets have to be re-rendered from the password manager first. It is hours because the steps are known rather than practised.

A backup that has never been restored is a hypothesis. One store here gets the drill. Every six hours its dump is restored into a scratch database and proved, which is what converts a backup job into a recovery point. The rest do not, so their RPO is stated rather than measured. The gap is not the backup schedule. It is the restore coverage.

Detection is worth stating just as plainly. A handful of specific failure modes alert, because somebody predicted them. Everything else surfaces through dashboards and a daily digest. That bounds mean time to repair honestly: predicted faults are caught in minutes, unpredicted ones are caught when the digest lands or when somebody looks. There is no service level objective, because there is no user other than the operator, and writing one would be theatre.


What this platform does not have

The absences define the design more sharply than the inventory does.

  • No general paging, so MTTR is bounded only where a failure mode was anticipated
  • No offsite backups, so no real disaster recovery
  • No restore drill beyond one store, so most recovery points are stated rather than proved
  • No high availability, no redundancy, no failover
  • No multi-tenancy and no identity model beyond one operator

Every one of those is the correct trade for a single-operator platform and the wrong trade for a shared one.


When not to build this

  • A team needs it. A shared platform needs identity, quotas and multi-tenancy. None of that is here, and retrofitting it is a rewrite rather than an addition.
  • Users are waiting on responses. Consumer GPUs will not carry a latency-sensitive product, and the gateway hides that only until the cloud bill arrives.
  • There is no appetite for operations. Managed services win when nobody wants to run anything. Do not build what you will not operate.
  • A compliance boundary is in play. A workstation cannot produce the residency evidence and access attestations a regulated environment asks for. That is a different architecture, and it starts in a cloud region. The Tier-1 bank retrieval walkthrough is the version of this problem where the compliance boundary is the design.

The takeaway

Capability was never the hard part. A platform is the four properties a prototype skips, and the real question is what they cost once the easy answers are removed. Build inside hard limits and the architecture stops being a diagram of components. It becomes a set of decisions with named costs.

For a leader, the practical version is shorter. Take the five questions above into the next pilot review. The pilot that answers them with measurements is ready to run. The one that answers with assumptions is still a demo, and the gap between the two is where the month-two failures come from.


About the Author

I architect production AI systems for regulated enterprises, with delivery experience in banking and logistics. Focus on MLOps maturity, data platforms, and compliance-aware system design.

→ If you are standing something like this up: book a 30-minute architecture review. No pitch. Bring the constraint that is blocking you and we will work out whether this shape fits or whether something else does.

→ More case studies: the blog


Newsletter

New essays, straight to your inbox.

Occasional, in-depth writing on distributed systems, AI agent architecture, and engineering leadership. No spam, unsubscribe anytime.