Immutable golden images across every TeeChat hosted component
TeeChat’s hosted path has three components: the OpenAPI edge, the routing gateway, and the GPU inference engine. For months we have been moving OpenAPI and the gateway into measured Talos confidential VMs that boot from an immutable golden image — guests you can check from Settings or the OpenAPI attest tools. This week the inference engine met the same bar: the production guest no longer offers SSH.
The production inference CVM no longer runs sshd. There is no remote shell into the guest that unscrambles your prompt.
That is the news. The rest of this post is the path that made it honest.
Three components, one rule
When you use hosted chat or OpenAPI, a message still has to be processed somewhere. Our rule has not changed:
- The client seals the prompt in an envelope (OPE on confidential chat; TLS into a measured edge on OpenAPI).
- The gateway routes, authenticates, and meters. It should not see chat plaintext.
- The inference engine opens the envelope inside a hardware-isolated confidential VM to run the model, then seals the reply back.
Verification is how you check that these live components match published fingerprints — not a hostname alone. See How to verify confidential chat and How to verify TeeChat OpenAPI.
What did change is how much of that path still looks like a normal Linux box an admin can SSH into.
First: OpenAPI and Gateway in Talos
Talos Linux is an immutable OS built for API-driven clusters: no package drift, no “just apt install this,” and no SSH login as the day-to-day control plane. We run those guests as AMD SEV-SNP confidential VMs. Boot measurements go on a public golden list. Clients compare what the hardware attests to what we published.
That is already live for:
| Component | Ingress | What “Talos + CVM” buys |
|---|---|---|
| OpenAPI edge | openapi.teechat.ai | TLS ends inside a measured confidential guest. You can challenge the endpoint and pin the measurement (verify guide). |
| Routing gateway | gateway.teechat.ai | The routing server is a measured SNP guest, not a plain VPS. Settings shows CPU TEE sev-snp and a gateway binary hash checked against the platform manifest. |
We hardened these two first because they are TLS termination and session ingress. If TLS or session handling ran on an ordinary host, “the engine is in a TEE” would cover only the last hop.
Then: take SSH off the engine
The inference guest is a different shape — GPU passthrough, a measured OS, a measured app volume. For a long time it still had the familiar ops entry: SSH in if something looks wrong.
That entry is a trust problem. A login that can reach the guest is a login that shares the machine with the only process allowed to see your prompt in plaintext. “We promise we do not peek” is not the same as “there is no shell.”
So the production engine CVM is now built without sshd. Day-to-day there is no remote shell. If a guest is sick, we replace it from the last measured image — we do not patch it by logging in. The launch digest you can verify in Settings is supposed to stay the digest we published, not a digest plus “whatever we wrote after boot.”
This is the same principle as Talos on the edges: immutable, measured, no interactive admin on the live guest. The engine is not a Talos node (GPU and the model stack need a different guest). The operations model is now aligned.
How we stop a golden image from being silently rewritten
An immutable golden image only counts if the measurement matches. We measure the system disk and the application disk separately, and we publish those measurements (hashes) on the public golden list. Hardware attestation reports what the live confidential VM actually booted; the client can check the component you are connected to against the published hashes at any time — Verify attestation in Settings, or an OpenAPI challenge. A mismatch fails verification.
What you should (and should not) conclude
Do conclude
- The three hosted components are measured. OpenAPI and Gateway already lived in Talos CVMs. The inference engine no longer offers SSH.
- You can still Verify attestation in Settings before you send, and you can still challenge OpenAPI.
- Operator recovery is “replace the guest,” not “SSH and hotfix.” That is stricter for us. It is better for you.
Do not conclude
- That nobody will ever need to operate the fleet. Hosts, load-balancers, and measured rollouts still exist — outside the guest that holds your prompt.
- That this closes every research finding we track internally. Some checks stay client-side (signed allowlists, quote verification). We will not pretend a no-SSH guest makes those optional.
- That OpenAPI is end-to-end encryption. It is a verifiable TEE proxy with an OpenAI-compatible API. Confidential chat + OPE is the stronger client-enforced path. We said that in the OpenAPI verify post; it is still true.
If you want to check
- Desktop or web: Settings → Inference → Hosted chat → Verify attestation.
- OpenAPI integrators: Settings → Inference → OpenAPI verify, or
teechat-openapi-attest verify https://openapi.teechat.ai. - Allowlist bytes: platform-binaries manifest (and its
.sig). Clients on the current train require that signature.
Data sovereignty is still the product promise — your folder, you rule. This post is about shrinking who can reach the confidential VM that holds the unwrap key. That set no longer includes “whoever has SSH to the inference CVM.” More than that, the set is empty today: every hosted component keeps internal confidential data out of operator view.