Immutable golden images across every TeeChat hosted component

TeeChat’s hosted path has three components: the OpenAPI edge, the routing gateway, and the GPU inference engine. For months we have been moving OpenAPI and the gateway into measured Talos confidential VMs that boot from an immutable golden image — guests you can check from Settings or the OpenAPI attest tools. This week the inference engine met the same bar: the production guest no longer offers SSH.

The production inference CVM no longer runs sshd. There is no remote shell into the guest that unscrambles your prompt.

That is the news. The rest of this post is the path that made it honest.

Three components, one rule

When you use hosted chat or OpenAPI, a message still has to be processed somewhere. Our rule has not changed:

Verification is how you check that these live components match published fingerprints — not a hostname alone. See How to verify confidential chat and How to verify TeeChat OpenAPI.

What did change is how much of that path still looks like a normal Linux box an admin can SSH into.

First: OpenAPI and Gateway in Talos

Talos Linux is an immutable OS built for API-driven clusters: no package drift, no “just apt install this,” and no SSH login as the day-to-day control plane. We run those guests as AMD SEV-SNP confidential VMs. Boot measurements go on a public golden list. Clients compare what the hardware attests to what we published.

That is already live for:

ComponentIngressWhat “Talos + CVM” buys
OpenAPI edgeopenapi.teechat.aiTLS ends inside a measured confidential guest. You can challenge the endpoint and pin the measurement (verify guide).
Routing gatewaygateway.teechat.aiThe routing server is a measured SNP guest, not a plain VPS. Settings shows CPU TEE sev-snp and a gateway binary hash checked against the platform manifest.

We hardened these two first because they are TLS termination and session ingress. If TLS or session handling ran on an ordinary host, “the engine is in a TEE” would cover only the last hop.

Then: take SSH off the engine

The inference guest is a different shape — GPU passthrough, a measured OS, a measured app volume. For a long time it still had the familiar ops entry: SSH in if something looks wrong.

That entry is a trust problem. A login that can reach the guest is a login that shares the machine with the only process allowed to see your prompt in plaintext. “We promise we do not peek” is not the same as “there is no shell.”

So the production engine CVM is now built without sshd. Day-to-day there is no remote shell. If a guest is sick, we replace it from the last measured image — we do not patch it by logging in. The launch digest you can verify in Settings is supposed to stay the digest we published, not a digest plus “whatever we wrote after boot.”

This is the same principle as Talos on the edges: immutable, measured, no interactive admin on the live guest. The engine is not a Talos node (GPU and the model stack need a different guest). The operations model is now aligned.

How we stop a golden image from being silently rewritten

An immutable golden image only counts if the measurement matches. We measure the system disk and the application disk separately, and we publish those measurements (hashes) on the public golden list. Hardware attestation reports what the live confidential VM actually booted; the client can check the component you are connected to against the published hashes at any time — Verify attestation in Settings, or an OpenAPI challenge. A mismatch fails verification.

What you should (and should not) conclude

Do conclude

Do not conclude

If you want to check

  1. Desktop or web: Settings → Inference → Hosted chat → Verify attestation.
  2. OpenAPI integrators: Settings → Inference → OpenAPI verify, or teechat-openapi-attest verify https://openapi.teechat.ai.
  3. Allowlist bytes: platform-binaries manifest (and its .sig). Clients on the current train require that signature.

Data sovereignty is still the product promise — your folder, you rule. This post is about shrinking who can reach the confidential VM that holds the unwrap key. That set no longer includes “whoever has SSH to the inference CVM.” More than that, the set is empty today: every hosted component keeps internal confidential data out of operator view.

← All posts